Introducing DataHub Cloud v2.2

Build the agents that automate your data work

An analytics agent can run SQL against your warehouse in seconds. Whether it runs the right SQL against the right table using the right metric definition is a different problem, and it’s the one most teams hit after the first week of production. DataHub Cloud v2.2 addresses that problem directly: the Context Platform is now in public beta, available to every DataHub Cloud customer, and a new Agents capability lets teams build custom AI assistants that automate documentation, governance, and discovery work on top of it. 

Let’s take a look at what’s new at DataHub.

Context Platform now in Public Beta

The Context Platform is available now in Public Beta to every DataHub Cloud customer. It sits between your analytics agents and your enterprise data ingesting metadata from 150+ sources, extracting semantic meaning from existing query history, routing AI-generated context to domain experts for review via Slack, and serving the validated result to any agent via MCP

Miro layered it in and took their analytics agent from roughly 40% accuracy to over 90% on benchmark questions. The lift came from query history, cross-source context, and business meaning the agent could not infer from schema alone.

This release also ships a redesigned Evals experience. A pass/fail score tells you something is broken. The new experience tells you which context document caused the failure, so your team fixes the right thing and can verify the fix without leaving the platform. Schedule evals on a cadence, catch regressions before they reach a business user, and watch accuracy improve as context compounds.

  • Set up Context Intelligence in minutes with a guided conversational flow. No configuration expertise required! 
  • Trace a failing eval to the exact context document that caused it and fix it without leaving the platform

Learn more about the Context Platform.

Custom agents that run your data workflows

Agents extend Ask DataHub from a conversational assistant into an automation platform. Work that used to require a human, such as documenting new tables, generating governance reports, or debugging a failing query, now runs on a schedule, triggered by an event, or on demand. The platform is built on three concepts:

  • Agents: purpose-built AI assistants with their own instructions, tools, and a scoped view of your data
  • Tasks: repeatable units of work assigned to an agent, triggered on a schedule, manually, or by events
  • Decisions: checkpoints where an agent pauses and waits for human input before continuing

Two agents ship ready to use today, no setup required:

  1. Context Curator Agent continuously generates semantic context from warehouse query logs and BI dashboards, keeps it current as your data changes, and routes proposals to the right domain experts for review. 
  2. Ingestion Agent connects new sources, troubleshoots failed ingestion runs, reads connector source code to diagnose the problem, and tells you exactly what to change without needing to know which specialist to ask.

Beyond the built-in agents, the platform lets you build agents tailored to your data estate, your domains, and your workflows. Unlike fixed-purpose tools, each agent you create has its own instructions, tools, and scope built for your specific data landscape, not a generic template someone else designed.

Three use cases to start with:

  1. Data discovery. A “Snowflake Production” agent scoped to that warehouse, or a “Marketing Domain” agent that guides analysts to certified tables. Each answers questions only within its lane.
  2. Governance on autopilot. A nightly task drafts descriptions and assigns owners for new tables. A weekly task sends your data steward a coverage report. When an agent hits something it isn’t sure about—a sensitive tag, an ambiguous ownership call—it raises a Decision, waits, and picks up where it left off.
  3. Cross-stack debugging. Connect agents to Snowflake, GitHub, and more via AI Plugins. An agent can inspect a failing query in Snowflake, cross-reference its lineage in DataHub, and pinpoint the upstream change that broke it.

The Agents capability is available in Private Beta as part of the Context Platform add-on. See the Agents guide for full details.

Watch the DataHub Cloud Agents Demo.

Ontology Explorer

A glossary of hundreds of definitions doesn’t tell you how “net revenue” connects to “gross revenue,” or what breaks downstream when you change a term. The Ontology Explorer does. It makes business concepts navigable the same way data lineage makes pipelines navigable—visually, interactively, and from any glossary term page. Agents retrieving a definition get its dependencies alongside it. Teams can define their own relationship types beyond the built-in vocabulary.

Data Product Marketplace

Data products finally have their own home. Browse everything your organization publishes at Marketplace organized by domain, searchable by name, nested the way your business is actually structured. Parent-child taxonomies mean a Finance data product can nest Payments, Revenue, and Refunds underneath it, so teams browse the way the business is organized rather than the way the warehouse is structured. Stewards propose new domains and data products through the same accept/reject workflow already used for glossary terms, so structural additions go through review before they land.

You can now:

  • Browse all data products in a dedicated Marketplace with a hierarchical sidebar, domain filtering, and in-sidebar autocomplete search
  • Model nested data product taxonomies with parent-child relationships that mirror your org structure
  • Propose new domains and data products through the governed accept/reject workflow, giving governance leads a full audit trail for every structural addition

Semantic Model container lineage

Semantic Models now render as containers in the lineage graph rather than as nodes in the middle of the path. Metrics sit inside their model, and lineage flows cleanly to the physical datasets they read from: Metric → Semantic Model Dataset → Physical Dataset. When a metric returns the wrong number, you can now see exactly which table to look at.

Note: Existing Semantic Model metadata needs a re-ingest or migration run for metric lineage to populate in the new shape.

Learn more about semantic model lineage in our docs.

More in DataHub Cloud v2.2

One-click Slack App install. App Config and Refresh tokens are no longer required. Click “Connect to Slack” on the Integrations page and you’re done.

Domain and Data Product proposals. Stewards can propose creating Domains and Data Products through the existing accept/reject workflow, the same flow used for glossary terms, giving governance leads an audit trail for structural catalog additions.

Volume assertion accuracy. Anomaly detection predictions align better with sub-daily seasonality, hold more stable over time, and use tighter sensitivity bounds to reduce false positives.

New ingestion connectors: 

  • Confluent Cloud Stream Catalog metadata support via Kafka Connect
  • Cross-account Lake Formation lineage for Glue
  • Contract import from S3, GCS, HTTP, and Git via ODCS
  • Snowflake semantic views now preserve logical-table casing

Summary of what’s new in DataHub Cloud v2.2

CapabilityWhat it doesWho will benefit the most
Context Platform (now in Public Beta)Ingests metadata from 150+ sources, extracts semantic meaning from query history, and serves validated context to analytics agents — improving text-to-SQL accuracy from ~50% to ~90%Data platform engineers, analytics engineers, domain experts
Custom agents that run your data workflowsBuild and deploy custom AI assistants that answer questions in chat, run scheduled documentation and governance tasks, and pause for human review when judgment is neededData platform admins, data stewards
Agentic Ingestion ConfigurationDiagnoses failed ingestion runs by reading connector source code and routing questions to the right specialist automaticallyData platform engineers
Ontology ExplorerVisualizes relationships between business glossary terms in an interactive graph so governance teams can trace how concepts connect before changing a definitionData governance leads, data stewards
Data Product MarketplaceDedicated browse experience for data products with parent-child taxonomies, domain filtering, and governed proposals for structural additionsData consumers, analysts, governance leads
Semantic Model Container LineageRenders semantic models as grouping containers in the lineage graph so metric lineage traces cleanly from metric to physical datasetAnalytics engineers, data platform teams

Let’s build together

We’re building DataHub Cloud in close partnership with our customers and community. Your feedback helps shape every release. Thank you for continuing to share it with us.