This time last year, I released a Deeplearning.ai course with Andrew Ng on Governing AI agents with the goal of educating developers on the basics of data protection for effective and responsible agent deployment. The motivation for the course came from IBM’s 2025 breach study, which found 97% of the organizations that suffered an AI-related breach lacked proper AI access controls, and 63% had no AI governance policy at all.
A year is an eternity in AI development, and we have come a long way from agents with zero governance or grappling with governing a single agent. Teams are now confronting a new governance challenge: agent sprawl (the uncontrolled growth of autonomous AI agents across an organization without centralized tracking, ownership, or governance). According to Gartner, by 2028 the average Fortune 500 enterprise will use over 150,000 AI agents. However, according to the firm, only 13% of organizations believe that they have the right AI agent governance in place. Since last year, the ability to deploy has gotten easier than ever. The onset of coding agents like claude code and codex, as well as a variety of low-code/no-code options have lowered the barrier to agent deployment substantially. One result is teams are now faced with the possibility of their agent fleet engaging in hugely wasteful token usage and incurring unforeseen costs. Another serious challenge is the additional pathways for sensitive data to leak out. Unsurprisingly, governance challenges have evolved.
Currently, I’m watching customers build agents faster than ever, and the shape of the problem has shifted. A year ago we focused on adding the four pillars of governance to a single agent: lifecycle management, risk management, security and observability.
We knew even then that building an agent wasn’t the hard part. You can stand up an agent in an afternoon, wire in observability, point it at a copied-over slice of data, and it looks as if it’s production-ready. Then you try to run it for real, against live systems and at scale, and you hit the wall that actually matters: infrastructure at scale. Now, this hasn’t changed, but what’s the shift? The difference is this is happening with dozens of agents/sub-agents at the same time. Agents are multiplying faster than a lot of teams are able to govern them. Teams now need to be equipped with a data platform that scales centralized tracking, permissions, and governance with the number of agents being built. Let’s examine how the pillars of governance have evolved in 2026.
Where we were: four pillars and one agent
Last year I taught this free course through DeepLearning.AI on governing AI agents, where we built a governed HR analytics agent on Unity Catalog and MLflow. The whole course distilled into four pillars:
-
Lifecycle Management (Separation of Duties): version, deploy, and retire agents with full lineage across dev, staging, and prod.
-
Risk Management (Defense in Depth): overlapping defenses such as PII detection, guardrails, compliance controls, and monitoring, from data ingestion through model performance.
-
Security (Least Privilege Access): agents and users get only the minimum permissions their role requires, enforced through authentication, encryption, and granular access controls.
-
Observability (Audit Everything): log every input, output, and decision for complete traceability and compliance.
The pillars sound abstract until you sit in the room with legal, auditors and leadership. Then they collapse into three very concrete questions.
-
What can the agent reach? In our build, the answer was: no direct table access, ever. Everything ran through layers, from masking to views to groups to functions, with data classification enforced at each one and aggregation-only access on sensitive tables. In practice that means the agent can answer “What’s attrition in engineering this quarter?” while being structurally incapable of surfacing any one person’s salary.
-
What changed, and can I undo it? The agent was registered as a versioned Unity Catalog model and deployed from that version, sitting on top of version-controlled functions and views. So when answer quality drops on a Tuesday, you know exactly what shipped on Monday, and you can revert it instead of debugging a black box in production.
-
Can I reconstruct what happened? Every function call was logged, MLflow traced every run, and there was an audit trail from the query all the way down to the underlying data. When an auditor asks what the agent touched on March 3rd, that’s a query, not a three-week investigation.
But notice the scope of all this. The entire game, a year ago, was getting one agent safely into production.
Where we are now: governance has to scale
While the four pillars still hold, now every one of them has to apply across a whole fleet of agents at once. That takes two things: policies that apply globally, to every agent, and the infrastructure to enforce it.
Every agent needs governed access to live data, not a copied-over sample that looked fine in the prototype. Also, every agent generates its own record: traces, spending, and access logs; and all of that has to land somewhere you can query. Now, hand-configuring lineage and least-privilege for one carefully-built agent simply doesn’t survive contact with a hundred of them running on a data platform that was never wired for it. The hard work moves down a layer, from the agent to the infrastructure it runs on.
The good news is that the infrastructure to support governing agents at scale is becoming a top priority of companies that are building agents and want to mitigate risk– as well as a priority for companies who have previously experienced unauthorized acts and breaches of duty from their agents. The answer to sprawl is a control plane: one governed layer that every agent runs through, sitting on top of the same data platform that already holds your tables, permissions, and lineage. On Databricks, that layer is Unity Gateway, a system that governs how developers reach AI coding agents, models and tools. Admins configure and govern centrally; developers just run a command, ug claude or ug codex, and get an approved agent with the right settings already baked in. Four capabilities do the heavy lifting.
-
Agent Configuration. Admins define the authorized set, which models, MCP servers, skills, and budgets a team can use, then publish it. Developers install the Unity Gateway CLI once and launch approved tools with ug claude or ug codex; every launch checks for local drift and enforces the published rules. Governance stops being something you wire into each agent and becomes an org-wide default developers inherit automatically.
-
Smart Routing. Instead of sending every request to the biggest available model, the gateway matches task complexity to model capability: cheap models for simple work, powerful ones for the hard problems. On Databricks’ own internal coding benchmark, smart routing alone produced a 35% cost saving. Governance now quietly decides which model runs on which task.
-
Smart Budgets. Spend becomes first-class. You set monthly budgets, shared across a team or per user, and decide what happens at each threshold: send an alert, block further requests, or both. The gateway can also nudge toward cheaper options as spend climbs, recommending a smaller model or a low-cost open model once you cross, say, 80% of budget. Cost, which used to surface in a finance spreadsheet weeks later, becomes something you govern in near real time.
-
Unified Tracing. Every tool call is traced automatically: its name, arguments, errors, token counts, and latency land in a unified table you can query. That turns cost control into a lookup instead of an investigation. In one case, Databricks traced roughly $499,000 a year in wasted tokens to seven small bugs in tool servers, and fixed them in about an hour.
Governance scales the same way, through policy. Policies use attribute-based access control (ABAC): you write one rule against governed tags and it applies everywhere, so an MCP server can be permissioned down to individual tools, and access keys off the attributes of the user or agent making the call. Service policies add guardrails on every request and response, blocking or masking sensitive data and catching prompt injection, unsafe content, and hallucinations before they reach a user. A new agent inherits the right permissions and guardrails from who it belongs to, not from a bespoke grant.
This isn’t a theory. One customer, Concurrence, described routing all their traffic “through a single governed path while maintaining identity-level attribution and access to approved models and MCP tools,” a deployment that ran 61 billion input tokens across roughly 360,000 requests. That’s the difference between governing an agent and governing an organization’s entire agent footprint.
The shift: same pillars, bigger surface
The four pillars didn’t get replaced, they scaled to cover a fleet and we added a 5th pillar: cost.
|
Pillar |
Then: one HR agent |
Now: a fleet of coding agents |
|
Lifecycle Management |
Version and deploy one agent via MLflow + UC |
Central agent configuration published to the fleet, with permissions and lineage on every model, MCP, and skills |
|
Security |
Least-privilege and masking on one dataset |
GRANT/DENY, ABAC, and contextual policies provide dynamic access controls to models, services, and tooling. Granular MCP tool prevents unapproved tool usage without removing all MCP functionality. |
|
Observability |
MLflow traces and sessions |
Unified trace tables and dashboards across the whole fleet |
|
Risk Management |
Catch failure modes before production |
Service-policy guardrails on every request and response: sensitive data, prompt injection, unsafe content, hallucinations |
|
Cost |
Over-engineered agents and lack of cost insights led to surprise billing, often stifling development. |
Smart routing for model efficiency, cost visibility, rate limits and budget caps to prevent tokenmaxxing while promoting valuemaxxing. |

In practice, the pillars now map to concrete capabilities delivered through Unity Gateway, Unity Catalog, and MLflow: central agent configuration, policies (guardrails and ABAC grants), unified traces, and smart routing and budgets. The one-line version: governance went from gatekeeper to control plane. A year ago the question was “can this agent see this row?” Now it’s “which model runs, on which task, at what cost, under whose identity, across every agent in the company?”
What’s next
The course closed on a roadmap: blue-green deployments for agents, a centralized gateway for endpoint management, and anomaly detection on agent behavior. It’s satisfying to watch those move from “coming soon” to “shipping.” The gateway layer, in particular, is now real infrastructure rather than a slide.
If you’re building agents, the fundamentals haven’t changed. The four pillars are still the on-ramp, and the free 75-minute course still walks you through them end to end. What’s changed is the ceiling. Start by auditing your own agents against the pillars. Then ask the bigger question: not just whether each agent is governed, but whether you can steer all of them at once. Because if your agent deployment is stuck waiting on a compliance review, the blocker probably isn’t the model.

