← back to the archiveCover illustration for “Data gravity is becoming agent governance”
ESSAYday 42·2 weeks ago·by Andy Padia

Data gravity is becoming agent governance

Databricks' Genie agents inherit Unity Catalog permissions, semantics, and lineage — the model is swappable, the governed access layer is not. You can win the agent layer by owning the data control plane, not a model.

Databricks has been extending its agent tooling around governed data — Genie agents that must have their data registered in Unity Catalog, draw on structured retrieval, and inherit the catalog's permissions, business definitions, and lineage. The documentation is refreshingly unsexy: an agent gets up to 30 tables or views, conversations are capped, everything registered and audited. (Revenue figures floating around the coverage — a $1.7B AI run rate — I couldn't tie to agent adoption, so leave those out.) The unsexy part is the whole strategic point, and it cuts against the loudest assumption in enterprise AI right now.

That assumption: the frontier model vendor owns the agent stack. Whoever has the best model wins the enterprise, top to bottom. Databricks is making the opposite bet, and it is a good one — you can win the agent layer without owning the model, by owning the governed control plane the agent has to pass through.

Why the sandbox demo lies

Every enterprise agent pilot looks great in a sandbox, and then a specific, predictable thing kills it at production review. The demo agent answered questions beautifully — because in the sandbox it ran on a service account that could see everything, against tables someone hand-picked, with "revenue" meaning whatever the builder assumed. Production is where the questions arrive that the sandbox never asked. When this agent answers for a regional manager, does it see only that region's rows? Whose definition of "active customer" did it use — and is that the one finance signs? When it took an action, who is accountable, and where is the audit line?

Those are not model questions. A better model answers the sandbox questions faster and the production questions no better, because they are questions about access, semantics, and accountability — properties of the data platform, not the reasoning engine. An agent that cannot inherit row-level permissions, trusted business definitions, and auditable ownership is a liability that demos well, and that is exactly the agent legal shuts down at scale.

What actually accretes gravity

Here is the swappability asymmetry that makes this a strategy rather than a feature. The model layer is becoming genuinely fungible — new frontier release every few months, prices falling, and (as the Economist's worldview map incidentally showed) even same-lab models differ enough that you re-evaluate anyway. Swapping models is a Tuesday.

Now try to swap the other layer. Your row-level and column-level permissions, encoding years of who-can-see-what decisions. Your lineage graph, showing which numbers derive from which sources. Your business semantics — the negotiated, politically-load-bearing definition of every metric that matters. Your sharing and credential rules. None of that is portable, because it is your organization's accumulated agreements about its own data, written down and enforced. That is data gravity, and in the agent era it has become agent governance.

rendering diagram…

The platform that already holds that governed layer can offer agents that pass production review by inheriting it — the agent sees exactly what the asking user may see, uses the metric definitions finance already blessed, and writes an audit line by construction. A best-of-breed model bolted onto ungoverned data can't match that at any capability level, because the gap isn't intelligence. It's provenance.

The counterpoint, and the buyer's move

Be fair to the other side: model quality still matters, and a governed platform with a mediocre model loses to a governed platform with a great one — which is precisely why Databricks-style plays emphasize broad model choice on top of the governed layer. The bet is not "the model is irrelevant." It is "the model is the swappable part, the governance is the sticky part, so own the sticky part and rent the rest." That is a more durable position than owning a model that a competitor leapfrogs next quarter.

At work, the buyer implication I now flag: enterprises are shopping for agents as if they were shopping for models — comparing reasoning benchmarks — when the question that determines production success is whether the agent inherits their existing governance or forces them to rebuild it. A team that has invested years in a data catalog, permissions, and semantic definitions should weight "does this agent layer sit on my governed data" far above "whose model is 3% better on some eval." The first is where deployments live or die; the second is a number that changes next month.

Steal this evaluation reframe: for any enterprise agent platform, score it first on governance inheritance — does it enforce your row/column permissions, use your blessed metric definitions, and produce an audit trail without a rebuild — and only second on model quality (which you can swap anyway). If a vendor's demo runs on a god-mode service account and hand-picked tables, you are watching the sandbox that lies. Ask to see it run as a restricted user, against your semantics. The ones that can are selling agent governance. The rest are selling a chatbot that will fail the review you haven't scheduled yet.

Models are the tenants; governed data is the landlord — own the control plane the agents must pass through, and you own the agent layer without ever shipping a model.

#agents#data-platforms#governance#databricks#strategy
← older drop
On-prem AI is an architecture decision
newer drop →
Agent harnesses are release platforms

related drops

explore all 78 drops →
← back to the archiveday 59