Pre-Calculated Metrics as the Unit of Agentic Analytics
Pre-calculated metrics let agents reason safely without defining data on the fly.

Agentic analytics has moved out of pilot projects and into production systems that monitor data, reason through anomalies, and act without a human typing a query first. The Model Context Protocol (MCP) has made it trivial to wire an agent to nearly any data source. The question that decides whether any of this works safely is no longer about connectivity. It's about what the agent finds once it gets there.
Why agentic analytics has reached a structural decision point
A traditional BI tool, no matter how fast, waits. Someone asks a question, the tool returns an answer, and then the loop ends until the next question arrives. Agentic analytics breaks that pattern entirely: the system monitors metrics on its own schedule, notices when something looks wrong, investigates the cause, and either recommends a fix or acts on it, cycling continuously through that monitoring-reasoning-acting-learning sequence. That is a different orientation toward data than any previous generation of business intelligence software has had, and it changes what the underlying infrastructure needs to guarantee.
The distinction matters because the industry has started drawing a sharp line between automation and agency. Rules-driven systems execute predefined logic. Goal-driven systems pursue an objective across multiple steps, they call other tools and systems along the way, and no human specifies each move. Tell an agent to watch customer churn, and it doesn't just report a number when asked. It decides when the number looks wrong, pulls related data to figure out why, and may trigger a downstream action, and no person initiates any single step. Turning a corporate objective like "reduce churn" into that kind of multi-step analytical process requires an architecture underneath the agent that can support it reliably.
That's where MCP changes the calculus. Before protocols like MCP standardized how agents connect to data sources, a meaningful share of engineering effort went into building and maintaining those connections. MCP removed most of that friction. An agent can now reach a data warehouse, a catalog, a query engine, or a metadata store through a common interface, with little custom integration work required. That's a genuine gain, but it also means the old bottleneck, connectivity, no longer absorbs any of the risk. Once access is easy, everything that used to depend on access being hard now depends entirely on what sits behind it: a governed semantic layer, or nothing.
What breaks when an agent reaches raw tables directly
Given a connection string and a production database, an agent will answer every question put to it. It will also do so with total confidence whether the answer is right or wrong, and nothing in that confident tone will tell you which one you got. An agent without a governed semantic layer has no standardized definition of "revenue," "active user," or "churn" to draw from, so it improvises one, often differently depending on which tables it happens to query and how it chooses to join them. A financial summary and an operations report can come from the same underlying data and still show two different revenue figures, both delivered with the same unhedged certainty, and nothing in either output flags that they disagree.
This creates a trust problem before it creates a technical one. When two executives compare numbers from the same analytics system and get contradictions, they question whether the system can be trusted at all, and that doubt tends to spread to every report the system has ever produced, not just the one that broke. Unrestricted agent access to a cloud data warehouse leads to exactly this kind of breakdown, and the fix isn't a smarter model. Role-based permissions, audit logging, and runtime tracing have to be built into the architecture from the start, because they are the mechanism that keeps an agent's confident answer tethered to a defined, checkable calculation.
Framing the issue as a hallucination problem misses what's actually happening. A model that invents a number isn't malfunctioning: it's filling a gap nobody gave it the definition to fill correctly. Without governed metadata, lineage, and access rules, an agent doesn't fail to answer. It answers anyway, guessing at definitions and aggregations in ways that are difficult to audit after the fact. That failure pattern appears starkly in regulated industries. When logic lives in ad-hoc scripts with no lineage attached, and a regulator eventually asks how a number was calculated, the honest answer becomes a search for the analyst who built the spreadsheet months or years earlier, not a query against documented logic.
There's a second category of damage that has nothing to do with correctness: production risk. Agents that call a database like Postgres directly bypass whatever filtering logic lives in the application layer. Row-level security enforced at the database itself is the only control standing between the agent and data it shouldn't see. And agents aren't naturally cautious about cost the way a human analyst is, pacing against a dashboard's query budget. An agent running analytical workloads straight against a production transactional database can degrade performance for the live application depending on that same database, simply because nothing in its design tells it to hold back.
The Pre-Calculated Metric
The fix is a different unit of consumption: the pre-calculated, governed metric, a value defined once, modeled centrally, and exposed to agents as a contract. An agent given only raw tables will guess at definitions, as shown above. An agent given semantic context but no way to query anything can't produce an answer. The governed metric solves both problems at once, because it hands the agent a defined output and the boundary it operated within to get there.
Without a trustworthy semantic layer, no other part of an agentic system can be trusted. Data teams need a centrally aligned semantic layer with standardized metric definitions, so that the same question returns the same answer no matter which team, tool, or agent asks it. A well-built MCP server for analytics reflects this directly: it exposes semantic resources and query tools side by side, so an agent retrieves a metric's definition, asks what dimensions it's allowed to slice by, runs a governed query against that contract, and gets back a result with its provenance attached. None of that resembles a raw table scan, and that's the point.
Restricting agents to metric names instead of table names is least-privilege access applied at the analytics layer. A governed query tool takes a metric name, a set of filters, and requested dimensions, not arbitrary SQL, and the query service sitting behind it enforces policy before returning anything, attaching lineage to the result automatically. That's a meaningful difference from an agent that simply knows how to write a SQL query: one agent operates inside a defined contract, the other operates wherever its own reasoning happens to take it. Governed agentic analytics in practice means every answer an agent returns carries its own lineage, already knows which fields are sensitive, and respects access controls defined once in a catalog, with no gap between what the agent can technically see and what governance says it should see.
None of this requires stripping agents of autonomy in how they reason. It requires drawing a boundary around what they're allowed to request. An agent can have real latitude in how it investigates an anomaly, which metrics it checks, which dimensions it slices by, which hypotheses it tests. What it shouldn't have is the ability to define a new metric on the fly or decide for itself which data is in scope. The intelligence in a well-built agentic system lies in how the agent navigates inside those boundaries, not in whether it can step outside them.
How agentic report generation works when metrics are pre-modeled
Once metrics are pre-modeled and governed, report generation is no longer a gamble on what an agent decides to query; it becomes a predictable, checkable loop. The architecture settles into three layers running in sequence. A semantic layer defines each KPI once and makes it available for any agent to call. A reporting agent reads from that layer and executes governed queries against it. An LLM narrative layer then writes the prose that accompanies the numbers, grounded in cited metric values.
The agentic behavior in this setup lives in an orchestration layer that breaks a high-level goal into smaller tasks, sequences the tool calls needed to complete them, evaluates whether the results it's getting make sense, and decides whether to refine its approach or escalate to a human. All of that reasoning stays bounded by the metric contracts that govern it, so the agent has room to work but no room to redefine what it's measuring. The output of that process is a structured analytical artifact, a dataset, a calculation, a chart, something that updates automatically when the underlying data changes, rather than a one-off narrated answer nobody can go back and verify.
In regulated contexts, the lineage this produces does real work. Every query an agent runs gets logged, every transformation gets tracked, and every data source gets cited, all built into the output as it's generated rather than reconstructed afterward by someone combing through logs. When a regulator asks how a number was calculated, the answer is a machine-readable trail, not a hunt for the person who remembers building the underlying spreadsheet. The same lineage that satisfies a regulator's question also gives executives a reason to trust the report sitting in their inbox each morning, because the number didn't arrive unexplained.
The clearest illustration of what this architecture makes possible at speed comes from post-merger integration. A travel services company, described in Promethium's enterprise case analysis, deployed an agentic analytics platform after a large-scale merger and went from kickoff to first production insights in under four weeks, and leadership had self-service access throughout, with zero data migration required. The speed wasn't a function of the agents being clever. It came from agents mapping both companies' schemas onto a shared semantic model up front, so that every query run afterward executed across both legacy systems as though they had always been one governed source. That's the compounding payoff of doing the definitional work once: each subsequent agent call gets faster and cheaper, because nobody has to re-derive what "revenue" or "active customer" means every time a new question comes in.
The governance gap the July 2026 MCP specification exposed
The MCP specification released on July 28, 2026 reworked the protocol to be stateless at its core: it strips out transport-level session management, so it scales across ordinary HTTP load-balanced infrastructure instead of needing the kind of persistent connections stateful protocols depend on. For analytics specifically, that change fits the workload well. An agent investigating an anomaly might fire off dozens of short-lived calls in sequence, pulling metadata, checking semantic definitions, running queries, checking a catalog for lineage, and a stateless protocol core handles that volume of brief, independent calls cleanly, without the overhead of maintaining state across each one.
The same July revision also sharpened the protocol's security guidance, bearing directly on whether agents should have raw access. MCP's updated guidance describes "confused deputy" risks that can arise in MCP proxy servers, and it states directly that token passthrough is forbidden. That prohibition closes a specific hole: an agent can't take a user's own credentials and pass them through to a raw database connection to escalate what it's permitted to see.
This is where the metric-as-unit pattern turns out to be a security pattern as much as a semantic one. If an agent calls a metric name through an MCP server, then it operates inside a governed boundary by construction. An agent that calls arbitrary SQL through that same server does not, no matter how carefully the server is configured otherwise. A properly governed analytics MCP server accepts a metric name, filters, and dimensions, never freeform SQL, and it enforces policy before it returns anything, attaching provenance to the result. Combined with the ban on token passthrough, this makes the governed metric API the actual enforcement boundary, the one place where an agent's request either fits inside policy or gets rejected.
That requirement lines up with where enterprise AI governance is heading more broadly. PwC's 2026 AI predictions identify responsible AI governance as moving from talk to traction: organizations are now expected to show that their AI systems operate inside defined, auditable boundaries, not simply assert that they do. Pre-calculated metrics served through a governed MCP server are what that expectation looks like at the analytics layer specifically. Every call an agent makes is bounded by a metric contract, logged as it happens, and attributable to a specific request, which is a different standard than hoping an agent's SQL stays reasonable.
The real objection: does governing metrics reintroduce the analyst bottleneck?
There's a real objection to all of this, and it deserves a direct answer. If an agent can only act on metrics a data team has already modeled, then the whole promise of agentic analytics, speed and autonomy, now depends on how fast that data team can define new metrics. That's not obviously faster than the BI workflows agentic systems were supposed to replace, and if every new question requires a human to build a new metric definition first, the agent has simply moved the bottleneck upstream.
The objection holds up only if it conflates two different kinds of work that don't scale the same way. Defining a metric is a one-time cost. Re-deriving a definition from scratch on every single query is a cost that recurs forever and grows worse as more agents start asking questions. Once a metric is pre-calculated and governed, every subsequent call against that definition is fast and cheap, and it returns the same correct answer regardless of which agent or team is asking. The ad-hoc alternative looks cheap at the start of each query and then gets expensive fast, because someone still has to check whether that query's logic was right, and that validation cost multiplies as more agents make similar requests.
A centrally aligned semantic layer is valuable precisely because the same question gets the same answer across every team, tool, and agent that asks it, and that value compounds as more agents come online. The governance work here is front-loaded, not a recurring tax on every new question. Once core metrics like CAC, LTV, ARR, and churn are defined and governed a single time, every agent that ever asks about them going forward inherits the correct definition automatically, with no analyst needed in the loop for that particular question ever again. The bottleneck the objection fears is real for the first version of each metric. It disappears for every use of that metric after.


