Governance at Scale: August 2026 Town Hall Highlights
A data catalog can sit in production for a year and still not add up to governance. Governance is what happens when people actually use the catalog, when definitions are agreed on rather than reinvented, and when both humans and agents can trust what they read. That gap, between having a catalog and having a foundation people rely on, ran through the August town hall.
The theme was governance at scale. Trustpilot walked through governing a 19-year data estate across two clouds, and were candid that the hard part was cultural, not technical. Founding Product Manager, Maggie Hays, introduced the DataHub Metrics Catalog, which makes business metrics first-class, governed entities instead of definitions scattered across tools. We introduced Otto, a community assistant grounded in DataHub’s own code and docs. And DataHub Core 1.7.0 shipped with 11 new sources.
The connective tissue is trust. An estate people reach for by reflex, a metric an agent can read without guessing, an assistant that cites its source before it answers. Different surfaces, one requirement.
From the community: a hackathon wrap-up, CONTEXT speakers, and meetups
A quick lap around the community before the sessions.
- The Agent Hackathon has wrapped. Build with DataHub: The Agent Hackathon ran fully online and drew more than 3,000 participants and over 600 submissions, with 100+ pull requests and upstream contributions opened into the DataHub project and supporting repos along the way. Check out the winning projects in our Agent Hackathon winners announcement blog.
- CONTEXT 2026 is set for November 4. The half-day virtual summit is built for practitioners putting context into production. This month we announced three community speakers: Alexandre Miyazaki (iFood), returning after the May town hall to go a level deeper on how iFood wired DataHub into its stack; Jishanahmed “JARS” Shaikh, on building a production-ready DataHub connector in under an hour; and Uttam Nayak (Nike), on making AI agents’ queries audit-provable. Register for free at datahub.com/context.
- Meetups, home and away. The DataHub India meetup in Bengaluru (August 14) brought practitioners together in person. Next up is the DataHub User Meetup DACH, online on October 13.

Meet Otto, a community assistant grounded in DataHub’s context
Shirshanka Das, co-founder and CTO of DataHub, introduced the community’s newest member. Otto (full name Otto the Otter) is DataHub’s AI community assistant, now in public beta and living in the DataHub Slack community. It first appeared a while back as DataHub’s mascot. This release turns it into something you can actually delegate questions to.
The design choice that matters is where Otto gets its answers. It is grounded in real DataHub sources, the official docs and the codebase, and it reads the relevant code before answering rather than paraphrasing from memory, with source links so you can verify and dig deeper. That makes it a fit for the practical questions the community asks every day: how to configure a connector, why lineage is not showing, whether a recipe is correct. To use it, tag @Otto the Otter in the community channels, the same way you would any other agent.
Otto is a public beta and improving quickly, so feedback to the community team is welcome. Before it shipped, it went through four rounds of evaluation against DataHub’s own repo. Herald, the community’s earlier assistant, continues to run. The investment from here goes into Otto.

Trustpilot: governing a 19-year data estate across two clouds
David Walker, Staff Data Engineer, and Hugo Hobson, Senior Data Engineer, on Trustpilot‘s data enablement team gave the month’s deepest session: what governance looks like when you inherit almost two decades of accumulated data, and the five-pillar framework they used to roll DataHub out to the entire business.
Who Trustpilot is, and what the black box cost
Trustpilot runs the world’s leading open review platform and recently passed 400 million reviews. It employs more than 1,000 people, with product and engineering teams in London, Copenhagen, and Edinburgh. Trustpilot runs multi-cloud: AWS carries operational workloads, GCP carries analytics and ML, data replicates from AWS across to GCP, and third-party sources such as Salesforce land in the data lake. Underneath sits a familiar modern stack, with DataHub Cloud cataloging assets from both clouds.
The numbers David put up frame the scale. DataHub now gives Trustpilot visibility into more than 500,000 data assets across the two clouds, 568 documented glossary definitions, and a 19-year estate that engineers can finally navigate.
Before DataHub, it was a bit of a black box.
David WalkerStaff Data Engineer, Trustpilot
The starting point was hard to work with. Answering “where is that data, who owns it, and what depends on it” meant a manual hunt across GitHub, Slack threads, and whoever happened to hold the context in their head. There was no central catalog, no consistent ownership model, and no real lineage. Downstream dependencies could be reconstructed only partially, and only by hand.
The cost was concrete. Two people on the data operations team spent a large share of their capacity answering discovery questions a catalog would have answered for them, capacity that came straight out of the infrastructure work they were meant to be doing.

Why DataHub Cloud, and contributing back
Trustpilot evaluated DataHub against other vendors and adopted it in 2024. David put the decision down to three things:
- Coverage of the platforms Trustpilot already ran was materially better and has kept improving, strong on the AWS side (DynamoDB, MongoDB, Kafka, Kafka Connect) as well as BigQuery on GCP.
- The DataHub project is open source, which fits Trustpilot’s “we win together” operating principle.
- The Trustpilot team could get close to new ingestion work early.
That last point is more than sentiment. Trustpilot has contributed real work upstream: DocumentDB modeled as a separate platform from MongoDB, ClickHouse and Iceberg sink support for Kafka Connect, a run of GCP Vertex AI ingestion fixes, and close collaboration on GCP Knowledge Catalog (formerly Dataplex) ingestion. BigQuery support for Analytics Hub linked datasets is on the way.
Five pillars, framed as why, what, and how
Adopting the tool was the easy part. Hugo was candid that a year after DataHub was in place, it was still a niche tool for data practitioners rather than something the business used. People were stopping ingestions without realizing what sat downstream, the kind of avoidable break a catalog exists to prevent. Around 18 months ago the team set out to change that, and treated it as a cultural problem rather than a technical one.
We saw this more as a cultural challenge than a technical one.
Hugo HobsonSenior Data Engineer, Trustpilot
Instead of handing DataHub to everyone and hoping, the team wrote a user-facing strategy document that defined what data governance meant at Trustpilot. It set out five pillars, and broke each one into the same three questions: why should I care, what does it mean, and how is it done here.
1. Ownership
Ownership comes first, because nothing else holds until you know who is responsible. Trustpilot owns almost everything at the team level and splits ownership into two types, a data (business) owner and a technical owner, ideally held by the same team. Hugo was clear that the framework sets expectations but leaves the collaboration to the owners themselves. As he put it, the team is not there to be parents.
2. Classification
Classification is the set of standards applied across assets, starting with criticality and personal data, with teams free to add their own.
3. Metadata
Metadata is deliberately broad, and Hugo cheerfully admitted the pillar is vague on purpose, because everything is metadata. Its real focus is descriptions, table and column, of which Trustpilot had almost none.
4. Lineage
Lineage is the biggest day-to-day value driver and mostly arrives out of the box, so the pillar is really about making people aware of what they already have.
5. Data quality
Data quality is the least mature pillar. Trustpilot splits it into operational quality (freshness, volume, availability) and integrity quality (correctness, validation, referential integrity), on the reasoning that monitoring whether data arrived is not the same as trusting that it is right.

Rolling it out: top of the tree, early adopters, build on top
The rollout came down to three moves. The team started at the top of the tree, the transactional source databases, so ownership and classification could propagate downstream. Tag a column as PII at the source, and it stays PII as it flows. They gave keen early adopters white-glove support, turned them into champions, and learned the edge cases the framework had missed in the wild. And they encouraged teams to build on top of DataHub.
That third move reached the most people. Trustpilot’s analytics team had nearly 500 Looker dashboards and a familiar problem: no one knew which was the right one. They created an internal quality standard they call the Kite Mark, attached glossary terms to the dashboards that earned it, and surfaced that metadata inside Looker through the DataHub Chrome plugin. Commercial users got DataHub context without leaving the tool they already worked in, which pulled in a set of people the data team would not otherwise have reached.
Where it landed, and where it goes next
The clearest outcome is cultural. “Have you looked in DataHub?” is now the reflex before anyone pings a team, and Hugo pointed to that shift from asking to self-serving as the result the team is proudest of. The catalog has carried a live migration from self-hosted MongoDB to Amazon DocumentDB, where lineage traversal shows the downstream blast radius before the work starts. Out of nearly 500 Looker dashboards, business users now find the governed source of truth without technical help. And DataHub assertions add an earlier layer of quality signal across a number of sources.
It’s nice to see teams starting a DataHub-first approach.
Hugo HobsonSenior Data Engineer, Trustpilot

A first look at the DataHub Metrics Catalog
Maggie Hays, Founding Product Manager at DataHub, introduced the DataHub Metrics Catalog, which makes business metrics first-class entities in the catalog rather than notes buried in a glossary. It is the shipped version of the Metrics Catalog teased in earlier town halls, and Maggie noted the team has been talking about metrics in the DataHub community since at least 2021.
“Which revenue number is right?”
The problem she opened with is a familiar one. The same metric gets defined again and again, within and across tools, with no shared record of which version is authoritative. Revenue might be a SUM in a dbt model, a slightly different SUM in a Looker measure, and a third definition in a Snowflake semantic view that excludes voided transactions. Each is reasonable on its face. Together they produce metric drift: no owner of record, no lineage back to a source column, and agents left to guess which one to use.

One standard underneath: Apache Ossie
Rather than invent a format, DataHub built on the standard the industry is rallying around. Apache Ossie (incubating), formerly Open Semantic Interchange (OSI), is a vendor-neutral spec for exchanging semantic metadata across analytics, AI, and BI platforms, expressed as declarative YAML. The payoff is three-fold: you define a metric once and it moves between platforms instead of being retyped; finance, marketing, and sales read the same number because they read the same definition; and agents reason from declared business logic instead of guessing at column names. DataHub modeled its new entities on the Ossie core classes, so what you catalog stays portable.
You define it once, and it applies everywhere.
Maggie HaysFounding Product Manager, DataHub
Two new entity types
The Metrics Catalog introduces two entity types. A Semantic Model groups logical datasets, defines how they join (with cardinality), and exposes the dimensions and measures that can be derived from them. It preserves the source DDL or YAML verbatim, and it backs one or more metrics. A Metric is the named, governed SQL calculation of a business measurement, such as total_revenue or daily_active_users. It carries an expression in one or more SQL dialects, derived-from and related metrics, full governance (owners, domains, tags, glossary terms), and AI context.
The lineage runs from the KPI down to the source column. A metric sits on a semantic model, which draws on logical datasets, which resolve to physical datasets. The semantic model acts as a container of its members. Only the cataloged definition becomes a Metric, so the three scattered revenue definitions collapse into one governed entity with a traceable path back to raw.orders. Metrics can also nest, so a downstream metric can reference a canonical one rather than redefining it.

Built to be read by agents
The part Maggie tied back to DataHub’s agent work is AI context, an ai_context block that Apache Ossie recommends and that can be set on a semantic model, a single metric, or an individual field. It captures synonyms (the terms a business actually uses, so an agent can match “total sales” to gross revenue), agent instructions (how to calculate the metric, for example recognizing revenue by ship date rather than order date), example questions the model can answer, and golden SQL an agent can pattern-match against. The goal is to keep an agent from inventing its own business logic at query time.
How metrics and semantic models get into the catalog
Snowflake Semantic Views are the first ingestion source. Enable semantic_views in the Snowflake recipe, turn on column lineage, and Semantic Model and Metric entities emit directly from the native views. For everything else, the Python SDK ships the same builders you already use for datasets, datahub.sdk.SemanticModel and .Metric, so any source can emit these entities programmatically. Teams that modeled semantic views as datasets under an earlier stand-in have a migration script to move them to the new entity type.
The Metrics Catalog is available in DataHub Core v1.7.0+, gated behind a METRICS_ENABLED flag while the feature is being finished, and on DataHub Cloud 2.1.0+. Once enabled, a Metrics area appears in the left navigation with a Beta badge. On the roadmap: surfacing metrics in universal search and browse and making them retrievable via DataHub’s Model Context Protocol (MCP); more ingestion sources, including the dbt Semantic Layer, Databricks metric views, Sigma, and BI-tool measures; column-level lineage from source columns straight through to metrics; and Ask DataHub resolving natural-language questions to cataloged metric definitions.

What shipped in DataHub Core 1.7.0
I closed the product block with the 1.7.0 release. The headline is 11 new sources, ten shipping as Alpha and ThoughtSpot as Beta. Alpha means experimental and still subject to change; Beta means ready to use but not yet tested across every edge case.
- BI and analytics: Cube, ThoughtSpot, MicroStrategy, and QuickSight.
- Database and streaming: SAP Datasphere, Amazon Kinesis, TimescaleDB, and TiDB.
- Governance and docs: ODCS (Open Data Contract Standard), BigID, and GitHub.
ODCS was the single most requested item in the release. TiDB came from the community, built by PingCAP, the team behind the database, and contributed upstream.
Beyond connectors, the release added a multi-language UI that is on by default and follows your browser locale, Kafka profiling for topic-level stats, Apache Spark 4.x support, and S3 and ABS profiling without a PySpark dependency. Across the Airflow, Dagster, Prefect, and Great Expectations plugins, emit is now async by default.
Two fixes are worth calling out for the coverage they recover:
- Redshift multi-line SQL is no longer dropped from lineage, so formatted or dbt-generated queries now contribute the edges they should.
- Power BI and Mode column-level lineage now resolves on non-lowercase columns.
Full detail on what we shipped is in the release notes for DataHub Core 1.7.0.

Thank you, contributors
The release landed on the back of 25 external contributors, all named on the following slide. The wider numbers behind the project: 3.6 million monthly downloads, 3,600+ GitHub forks, and 3,000+ organizations using DataHub. If DataHub is useful to you, starring the repo is one of the simplest ways to support it.

What it all adds up to
Governance at scale meant different things across the hour. Trustpilot showed it is as much a cultural achievement as a technical one, a foundation people now reach for by reflex. The DataHub Metrics Catalog made business metrics first-class, one governed definition an agent can read instead of three that drift. And the community showed up in force: a new assistant in Otto, and a 1.7.0 release with 11 new sources built alongside 25 outside contributors.
Where to go next:
- Watch the full town hall recording for the live demos.
- Read the Trustpilot customer story for the full five-pillar build.
- Explore the DataHub Metrics Catalog feature guide and try it in DataHub Core.
- Register for CONTEXT 2026 on November 4.
- Browse the DataHub integrations index for the 1.7.0 sources.
- Meet Otto and ask your next question in the DataHub Slack.
- Register for the next DataHub town hall to catch the next set of stories live.


