Frequently Asked Questions

Last updated: July 28, 2026

This documents compiles the questions most commonly ask about the Brighthive platform and BrightAgent. If you have an additional question not listed in this document or answered in our knowledge base, please reach out to the Brighthive team at support@brighthive.io.

1. Platform Overview & Core Concepts

What is Brighthive, and what is BrightAgent?

Brighthive is a purpose-built harness that automates the upstream work of making data trustworthy — cleaning, validating, contracting, and monitoring data at ingestion and transformation — so every tool downstream (a warehouse, a BI tool, an AI/agent layer) works from clean, well-documented, quality-checked data. BrightAgent is the coordinated set of specialized agents that does this work continuously, functioning like a full data engineering and governance team working around the clock rather than just during business hours. For more information go to 📄 What is Brighthive?.

What are Brighthive Capabilities?

Ingestion & ELT Acceleration

  • Data Ingestion Agent supports 600+ out-of-the-box connectors (via an included, managed Airbyte open-source instance), plus Fivetran, bespoke Python/Lambda ETL, and custom connector builds.

  • Uses the OpenLineage open standard so pipeline observability works consistently whether jobs run in Airbyte, Fivetran, or custom scripts on VMs.

Data Quality, Contracts & Governance

  • Governance Agent automatically generates and maintains a data contract for every data asset and data product, using the Open Data Contract Standard — even in organizations that don't yet have formal contracts, or a formal data team, in place.

  • Runs quality tests at ingestion and transformation time (completeness, validation, row-count checks, schema conformance) and produces an ongoing data quality score per asset.

  • Proactively detects upstream schema drift (e.g., a field changing from date to string) and flags or resolves data contract violations before they break downstream jobs.

Observability & Self-Healing

  • Monitors the full data lifecycle end to end — ingestion, transformation (including jobs versioned in GitHub or run natively in Snowflake), and delivery to BI tools.

  • “Night Shift” runtime lets agents work autonomously overnight to identify and self-heal pipeline breaks, then report results each morning.

  • Data Engineering Agent can commit fixes directly (e.g., updating Snowflake semantic views or dbt-adjacent logic) once an issue is identified.

Metadata, Catalog & Context (MCP)

  • Includes a managed deployment of OpenMetadata, the leading open-source data catalog, layered on top of the customer's existing warehouse or lake — or on top of the warehouse Brighthive provisions.

  • Exposes asset-level context — schema, lineage, quality reports, and policies — to any downstream agent system via MCP (Model Context Protocol), with a smart context-compaction and caching layer that reduces the tokens needed for agents to work with a given data asset.

Orchestration & Access

  • A supervising orchestrator delegates work across the seven BrightAgent agents, headlessly or through the app, Slack, or Microsoft Teams.

  • Brighthive Studio allows customers to build and customize additional agents, and supports the major open-source and private LLMs (bring-your-own-key supported).

  • Multi-cloud, multi-warehouse by design — the same harness governs data across Snowflake, Databricks, Redshift, and others simultaneously when they already exist, or provisions the warehouse of choice as part of a new deployment.

What is the 'workspace context file,' and why does it matter?

It's a layer of captured institutional knowledge — term definitions, data conventions, business rules — that agents draw on so they don't need to be re-prompted with the same context every time. Combined with the managed knowledge/metadata catalog, it's what lets a new analyst, or a new agent, work with the same institutional knowledge a tenured team member would have.

2. Fit With Your Existing Data Stack

Do we need Snowflake, Databricks, dbt, or Horizon already in place to use Brighthive?

No. Organizations can use Brighthive regardless of their data maturity. Brighthive runs in two deployment modes on the same underlying platform:

  1. Layer Mode, for organizations that already have some or all of Snowflake, Databricks, dbt, and Horizon in place. Brighthive deploys upstream of and around these tools, automating the data plumbing they don't natively own. Brighthive is the connective, automated layer spanning the full lifecycle (source, ingestion, transformation, warehouse, and downstream AI/BI consumption) that keeps all of them running on clean, well-documented data.

  2. Managed Stack Mode, for organizations earlier in their data maturity curve with few or none of these tools. Brighthive stands up the modern data stack itself (ingestion, warehouse, transformation, catalog and governance) using the same open-source foundation and agent harness, so quality and governance are built in from day one rather than retrofitted later. The result is that an organization without a modern stack today can reach the same state of upstream data quality and observability as a mature-stack one

If we already have a mature data stack, what's left for Brighthive to do?

Brighthive is deliberately complementary, not competitive, with these tools: it occupies the part of the data lifecycle they either don't cover natively or only cover once data has already arrived and a person has initiated the work.

Concretely:

  • Snowflake owns the warehouse, compute, Cortex AI, and native quality checks; Brighthive cleans, validates, and contracts data upstream of Snowflake, and feeds Cortex and semantic views the documentation, lineage, and quality state they need but don't generate themselves.

  • Databricks owns lakehouse compute and ML/AI workloads; Brighthive applies the same ingestion, quality, and observability layer across Databricks as it does Snowflake, giving a single governance layer across multiple warehouses instead of separate tooling per platform.

  • dbt owns transformation/modeling; Brighthive wraps dbt jobs with automated data contracts, upstream quality testing, and lineage — reading dbt-managed queries (e.g., via GitHub) to monitor for breaking changes rather than duplicating dbt's transformation logic.

  • Horizon/Cortex own native cataloging, governance, and AI querying initiated by a person; Brighthive provides the proactive, always-on, cross-platform layer upstream of and around them — continuous monitoring and remediation rather than reactive, human-initiated checks scoped only to what's already inside Snowflake.

Snowflake, Databricks, dbt, and Horizon remain the systems of record, transformation, and native governance. Brighthive is the connective, automated layer spanning the full lifecycle — source, ingestion, transformation, warehouse, and downstream AI/BI consumption — that keeps all of them running on clean, well-documented data.

We don't have a modern data stack yet — is Brighthive still relevant, or is it too advanced for where we are?

It's built for this case too, via Managed Stack Mode:

  • Ingestion: the Data Ingestion Agent connects to source systems from day one using the bundled Airbyte instance (600+ connectors), replacing the need to separately evaluate and procure an ELT tool.

  • Warehouse: Brighthive deploys and configures a warehouse (e.g., Snowflake, Databricks, or another supported target) sized to the customer's needs, rather than requiring the customer to select and stand one up independently first.

  • Transformation: baseline transformation logic and semantic modeling are established with quality checks and documentation built in from the first pipeline, rather than added later as a retrofit.

  • Catalog & governance: the OpenMetadata-based catalog and Open Data Contract Standard contracts are in place from day one, so the organization never accumulates the undocumented, unowned data assets that mature-stack customers are often trying to clean up years later.

The result is that an organization without a modern stack can reach the same state of upstream data quality and observability as a mature-stack customer, without first spending a budget cycle assembling and integrating an ELT tool, a warehouse, a transformation framework, and a catalog separately.

Can Brighthive govern data across more than one warehouse at the same time (e.g., Snowflake and Databricks together)?

Yes. The platform is multi-cloud and multi-warehouse by design. The same harness governs data across Snowflake, Databricks, Redshift, and others simultaneously when they already exist, or provisions the warehouse of choice as part of a new deployment, giving a single governance and quality layer instead of duplicated tooling per platform.

If we start in Managed Stack Mode without a modern stack, can we later bring in Snowflake, Databricks, dbt, or Horizon directly?

Yes. This is a designed glide path, not a re-platform. Because Brighthive already maintains contracts, lineage, and quality history from day one, that governance foundation can hand off cleanly to a newly adopted Snowflake, Databricks, dbt, or Horizon environment rather than starting governance from scratch.

3. Data Ingestion, Deployment & Infrastructure

Can Brighthive deploy inside Azure instead of AWS?

Yes. Brighthive is cloud-agnostic and can deploy into a customer's AWS or Azure environment. There is a formal AWS partnership that brings additional resources and streamlined deployment (including running inference through AWS Bedrock), but that partnership is not a prerequisite — customers running Snowflake on Azure, for example, connect natively without needing to migrate cloud providers.

Do we need a VPC or VNet before we can start?

No. Unlike most ELT tools, Brighthive does not require a dedicated VNet or VPC to get started. You can connect existing sources and run BrightAgent today. The platform scales into a private VNet/VPC when ready, and VM-deployable scripts are available for on-prem ELT jobs in the meantime.

What connectors and ingestion methods does Brighthive support?

The Data Ingestion Agent supports 600+ out-of-the-box connectors via an included, managed Airbyte open-source instance, plus Fivetran, bespoke Python/Lambda ETL, and custom connector builds. For more information go to 📄 What integrations does Brighthive offer?.

Since Brighthive doesn't store our data, what happens if table or column names change? Does the system remap identifiers?

Brighthive's metadata and lineage layer tracks entities and relationships independently of any single table or column name, so when naming changes, it can re-infer relationships rather than requiring a manual rebuild. Fields that 'look similar but need validation' are surfaced for a human to confirm rather than silently guessed. Identifier mapping is agent-assisted, with a human checkpoint before it's trusted downstream.

Does this require a database or vector store to embed (vectorize) our data?

Not for the platform's core structured-data workflows (ingestion, quality, governance, transformation). Brighthive's context layer uses embeddings and a metadata/catalog layer to support retrieval over mixed structured and unstructured content and to power natural-language querying, but you don't need to stand up or manage your own vector database to use the platform.

Do you have security certifications relevant to regulated industries?

Brighthive is built for regulated environments and maintains SOC 2 Type II, HIPAA compliance, GDPR alignment, and ISO 42001 (AI management systems). Inference and data processing run inside the customer's own cloud environment, so data does not need to leave your perimeter for the platform to operate.

Can Brighthive guarantee our data stays in a specific country or region (data sovereignty)?

Yes. Brighthive deploys single-tenant, bring-your-own-storage (BYOS), with each customer on their own cloud account, which supports region-specific deployment (for example, Canadian data residency) for customers in regulated or public-sector environments.

4. Data Quality, Cleaning & Monitoring

How do you check "data quality"?

Quality checks run at both ingestion and transformation time, covering completeness, validation, row-count checks, and schema conformance, producing an ongoing data quality score per asset. Checks are written and run as reproducible, versionable code (using the Great Expectations framework) rather than one-off manual review, and findings are prioritized with a severity score plus a plain-English explanation of business impact.

What kind of data cleaning does the platform actually do?

The platform proactively detects upstream schema drift (for example, a field changing from date to string) and flags or resolves data contract violations before they break downstream jobs. It identifies issues such as invalid categorical values, inconsistent date/timestamp formats, and referential-integrity problems, and proposes — or, with approval, executes — remediation, from a targeted transformation fix to a broader ingestion rule change.

How do you build active data-quality monitoring for things like schema changes, with automated alerts?

The platform monitors the full data lifecycle end to end — ingestion, transformation, and delivery to BI tools — and proactively detects schema drift and data-contract violations rather than waiting for a person to notice. A 'Night Shift' runtime lets agents work autonomously overnight to identify and self-heal pipeline breaks, then report results each morning, and the Data Engineering Agent can commit fixes directly (for example, updating semantic views or transformation logic) once an issue is identified.

How would BrightAgent identify good data versus anomalous data — and will it recommend how to correct it?

Yes. Quality checks profile each dataset, flag schema drift and contract violations without requiring an explicit prompt, and put issues in context with a severity score. The platform then proposes remediating actions — up to and including self-healing overnight via the Night Shift runtime — for a human to review or approve, depending on how the workflow is configured.

Has data quality measurably improved for real customers using the platform?

Yes. Brighthive baselines an organization's overall data quality score at the start of an engagement and tracks improvement from there. Customers have seen baseline scores move from roughly 60% to 83%, and from 83% to 99%, within the first six weeks of engagement.

Can we check the confidence level behind a chatbot / natural-language response?

The platform is designed so that low-confidence natural-language answers are flagged rather than presented as fact. BrightAgent runs a structured self-check after each response (did it have to assume anything, was a key term ambiguous, did it rely on stale metadata, etc.) and can surface a context-gap report when confidence is low, rather than only returning an answer with no indication of reliability.

5. Agent Behavior, Models & the MCP Context Layer

How much maintenance or customization does the platform's context layer need? Will it make the same decision every time given the same input?

Quality and transformation logic is written as deterministic, reproducible code (for example, Great Expectations tests, versioned transformation jobs) — so given the same data and the same configured rules, you get the same result every run rather than a re-guess from scratch each time. What does improve over time is context: the workspace/catalog layer and accumulated policies and contracts mean answers to natural-language questions get more accurate and more tailored to your organization's conventions the longer the platform runs.

How do you keep Brighthive's agents from doing destructive things to the data?

Governance policies and data contracts are enforced deterministically, not just advised by a language model, and all agent-authored changes to pipelines and schemas are version-controlled (for example, committed to GitHub), so any change is visible, reviewable, and reversible. The Night Shift runtime is scoped to identifying and self-healing pipeline breaks — not open-ended data modification — and higher-impact actions are surfaced for review rather than executed silently.

Does Brighthive support MCP (Model Context Protocol)?

Yes. Brighthive exposes asset-level context (schema, lineage, quality reports, and policies) to any downstream agent system via MCP, with a smart context-compaction and caching layer. This means tools like Snowflake Cortex, Databricks AI, or a custom agent don't have to spend tokens re-deriving context about a data asset every time they query it, which directly reduces AI/agent token spend on tools you already run.

Can we access or interact with BrightAgent outside of the web app?

Yes. The supervising orchestrator that delegates work across BrightAgent's agents can be accessed through the app, Slack, or Microsoft Teams via chat.

6. Governance, PII & Compliance

How can I validate that PHI/PII is not accessed inappropriately by an AI agent?

PII (including PHI in healthcare contexts) is automatically detected and tagged, and governance policies enforce anonymization, hashing, or access restriction on tagged fields. For more information in how to create these policies within your Brighthive Workspace go to: 📄 Add new Policy to Workspace

How do you handle governance policies that today live only in documents, with no real enforcement?

The Governance Agent automatically generates and maintains a data contract for every data asset and data product — even in organizations that don't yet have formal contracts or a formal data team in place — using the Open Data Contract Standard (ODCS), so governance rules become something actively monitored and enforced rather than documentation that agents can't read. You can create additional custom policies in your Brighthive Workspace that are automatically enforced by the Governance Agent. For more information in how to create these policies within your Brighthive Workspace go to: 📄 Add new Policy to Workspace