A deep dive on the Trust Fabric. Why attribution is the foundation, why enforcement lives at the gateway, and why evidence is the output that makes autonomous software defensible.
This is the seventh and final piece in a series on enterprise multi-agent architecture. The flagship laid out five planes and a Trust Fabric. Five deep dives followed, one for each plane: the Agent plane and harness engineering, the Model plane and the AI gateway, the Memory and Knowledge plane and the discipline of curating before storage, the Tool and Action plane and the MCP gateway, and the Orchestration and Experience plane and the three properties that turn a demo into a product. This piece closes the series on the Trust Fabric, the cross-cutting layer that turns the five planes into a system your business can put its name on.
The five questions that ended the meeting
A board committee convenes to review the enterprise’s AI program in late 2026. The CIO opens with the achievements. Fourteen agents in production. A meaningful reduction in average handling time on the workflows they support. Three business units running commitments through agents rather than through humans. Two more units in pilot. Real productivity moving through the system. The board approves the next tranche of investment.
Then one director asks a set of questions.
Which agent made the recommendation to terminate the Vendor X contract last quarter? Who at our company authorized that action? What data did the agent read to make the recommendation? Under what policy was the agent operating? How would we prove any of this to a regulator?
The CIO does not have answers.
The engineering team has traces. They have logs. They have model call records and tool call records and memory reads. But the traces live in one platform, the logs in another, the model records in a gateway, the memory reads in a data warehouse, the audit trail in the ticketing system. The correlation from a user’s intent, through the agent’s identity, through the model calls, through the tool calls, through the memory reads, through the downstream effect on the vendor is a manual investigation that would take an engineer four hours, assuming they can even reach every system.
The board meeting ends without the incident being an incident. The CIO leaves the room knowing that the next question of that kind, from a regulator rather than a director, will not go the same way. And the deadlines for exactly that kind of question are already on the calendar. The EU AI Act’s high-risk provisions entered enforcement on August 2, 2026. DORA has been in enforcement since January 17, 2025. Colorado’s AI Act took effect on June 30, 2026. The window for architecture that produces intelligence without producing accountability has closed.
This is the failure mode the Trust Fabric prevents. Not a crash. Not a breach. A quiet realization that the system produces answers but not evidence, and that the difference matters more with every quarter that passes.
The thesis
The Trust Fabric is the cross-cutting layer that turns the five-plane architecture into a system that is accountable, auditable, and defensible.
Its core purpose is threefold. Attribution, because every action an agent takes has to trace back to an identifiable agent, a delegating user, a governing policy, and a business intent, without which no other control functions. Enforcement, because policies live at the point of action rather than in slide decks, executing at runtime through the same gateways that carry model calls, tool calls, and memory operations. Evidence, because the system produces the traces, logs, evals, cost attribution, and human review outcomes that a regulator, auditor, board, or investigator needs, as first-class outputs rather than as by-products.
Security lives here as one of the disciplines these three properties enable. So do compliance, reliability, financial predictability, and operational resilience. The Trust Fabric is the architecture that makes the five planes below it operable at enterprise scale, in front of executives who need to explain them, in front of regulators who need to audit them, and in front of customers who need to trust them.
The rest of this piece defends each property and shows what they mean in practice.
What changed in 2025 and 2026
Seven inflection points reshaped what serious looks like on this layer. The Trust Fabric was a research idea eighteen months ago. It is now a regulatory expectation, a market category, and an executive-level accountability.
The EU AI Act’s high-risk provisions entered enforcement on August 2, 2026. Regulation (EU) 2024/1689 does not name agents. It applies to AI systems, and every enterprise agent operating in the EU or serving EU customers is an AI system inside the regulation’s scope. Article 10 requires documented data governance for training, validation, and testing datasets. Article 12 requires automatic event logging throughout the system’s lifecycle. Article 14 places the duty of human oversight on the deployer, not the model provider or the framework vendor. Article 15 requires accuracy, robustness, and cybersecurity. Article 26 lists the deployer’s specific duties. Annex III classifies the high-risk cases. If your agent touches creditworthiness, insurance risk and pricing, employment, essential services, or certain customer-facing decisions, you have deployer obligations that are now in force.
DORA has been in force since January 17, 2025. For any EU financial entity, the Digital Operational Resilience Act treats AI agents as ICT systems and applies Chapter II ICT risk management obligations to them. It does not need to name AI to reach it. The clean summary from a Peliqan analysis I read earlier this year: a financial entity using an AI agent on personal data through an MCP server must simultaneously satisfy GDPR Articles 5, 32, and 48, EU AI Act Article 26 deployer duties, and DORA Articles 5 through 30. DORA is lex specialis only against NIS2. It does not displace GDPR or the AI Act.
Colorado’s AI Act took effect on June 30, 2026. For US-headquartered enterprises that operate high-risk AI systems affecting Colorado consumers, the Act requires impact assessments, consumer notifications, and reasonable care against algorithmic discrimination. It is the first US state law of its scope. It will not be the last.
Singapore’s IMDA published the Model AI Governance Framework for Agentic AI in January 2026. The framework is the first comprehensive governance document written explicitly for autonomous agents rather than for AI systems generally. It requires each agent to carry a verifiable digital identity and an audit trail of which agent acted under whose authorization. The framework is voluntary today. It is the direction the ISO working groups are heading.
The NIST AI Agent Standards Initiative launched in February 2026. The associated NCCoE concept paper names the gap directly: agents are commonly treated as generic service accounts without dedicated identity, authorization, or accountability controls. The initiative organizes the US federal effort to close that gap. NIST SP 800-207 Zero Trust Architecture explicitly addresses non-person entities in its Section 5.7, and the companion SP 800-207A references SPIFFE as a service-mesh identity primitive. The scaffolding for federal expectations on agent identity is now visible.
Anthropic published Trustworthy Agents in Practice in April 2026 and Zero Trust for AI Agents in May 2026. The two documents together frame the industry position that no single vendor can secure agentic AI, that shared infrastructure and open standards are the path forward, and that a serious enterprise deployment requires capability tiers across seven concerns: identity, access, observability, monitoring, I/O controls, integrity, and governance. Google’s SAIF 2.0 Agent Risk Map, Microsoft’s Entra Agent ID and Agent Governance Toolkit, Forrester’s AEGIS framework from April 2026, and the CAGE-1 framework from July 2026 fill out the picture. Different vocabularies, converging content.
The non-human identity crisis became a board-level topic. By mid-2026, industry reports converged on numbers that had been dismissed as alarmist a year earlier. Non-human identities outnumber human identities by ratios reported between 50 to 1 and 144 to 1 depending on the survey. Ninety-seven percent of them have excessive privileges. Seventy percent of 2026 identity-related security incidents were linked to autonomous AI agents. A Cloud Security Alliance analysis found that over sixteen percent of organizations do not track the creation of AI-related identities at all. Only twenty-two percent of security practitioners assess AI agent identities as independent identities. Verizon’s DBIR reports twenty-two percent of breaches start with credential abuse. The 2025 EchoLeak vulnerability (CVE-2025-32711) demonstrated a zero-click AI compromise pulling data out of Microsoft 365 Copilot through pure text embedded in ordinary business documents. The attribution gap is not a theoretical concern.
FinOps for AI became a discipline. The FinOps Foundation’s State of FinOps 2026 report identified AI cost management as the number one forward-looking priority for the year. The share of FinOps practitioners with responsibility for AI spend rose from thirty-one percent in 2025 to ninety-eight percent in 2026. EY analysis reported that customer service AI cost per interaction moved from $0.04 in 2023 to $1.20 in 2026, roughly thirty times higher, driven by orchestration, tool calls, and reasoning loops. AWS launched a FinOps agent at FinOps X 2026 in June. The unit economics of agent systems, once an engineering concern, became a CFO concern.
Seven shifts. One direction. The Trust Fabric moved from an idea to a set of obligations, backed by real regulation, real vendors, real budgets, and real board-level scrutiny.
Attribution: the foundation
Attribution is the property that lets any other control function. If you cannot say which agent did what, on whose behalf, under what policy, at what moment, you cannot enforce access controls, you cannot investigate incidents, you cannot produce audit evidence, and you cannot calibrate autonomy over time. Every other property in the Trust Fabric depends on identity as the anchor.
The 2026 architectural answer is workload identity for agents. The emerging standard is SPIFFE (Secure Production Identity Framework For Everyone) with SPIRE as its runtime. SPIFFE issues each workload a short-lived, cryptographically verifiable identity document (an SVID) based on the workload’s runtime attributes rather than on a static secret in a config file. The identity is tied to what the agent is, not to a string someone dropped into a repository three sprints ago. When an incident happens, you can answer the two questions that matter immediately: which agent caused the event, and what else did that agent have access to. Without unique, attributable identity, those questions are unanswerable in the moment they matter.
Delegation is the next layer. Enterprise agents rarely act on their own authority. They act on behalf of a user, an application, or another agent. Preserving the delegation chain matters both operationally (so you can prove who authorized what) and legally (so the Article 14 human oversight duty attaches to the right person). The 2026 standard is OAuth 2.0 Token Exchange (RFC 8693) with the act claim, which preserves the chain of principals across handoffs. When a Renewal Supervisor delegates to a Risk Reassessment agent that delegates to a Financial Signal Analyzer, the audit trail carries the full chain, not just the last-mile identity.
Non-human identity governance is the harder problem. Non-human identities outnumber human identities dramatically across every 2026 industry report I have read, with the ratio depending on the survey methodology. What is uniform is the direction: agent proliferation makes the problem worse, not better, and the tooling most enterprises inherited from the cloud era does not scale to agent-density workloads. Static credentials, over-permissive service accounts, and orphaned tokens are the initial access vector in most 2026 breach reports on agent systems. The mitigations are known and increasingly available: SPIFFE-issued workload identity, short-lived credentials minted at the moment of use, delegation preserved through OAuth 2.0 Token Exchange, and centralized inventory that treats the identity itself as governed infrastructure. What is missing in most enterprises is the discipline to actually implement them before the incident that forces the conversation.
The practical implication for a CTO is that identity is not a security team concern to inherit. It is an architectural investment your team owns, and every quarter you defer it makes the eventual retrofit more expensive. Start here.
Enforcement: policy at the point of action
Policy that lives in a document does not enforce anything. Policy that runs at the point of action does.
The good news is that the enforcement points already exist across the five planes below the Trust Fabric. The AI gateway from the Model plane enforces model-layer policy. The MCP gateway from the Tool and Action plane enforces tool-layer policy. The memory plane’s write-path governance enforces memory-layer policy. The Orchestration and Experience plane’s approval flows enforce human-oversight policy. The Trust Fabric’s job is not to build new enforcement points. It is to unify the policies that flow through the existing ones.
Three architectural patterns anchor the discipline.
Policy as code. Every rule that matters, expressed in machine-readable form, versioned, tested, and deployed alongside the code it governs. Open Policy Agent (OPA), Cedar, and similar policy engines are the substrate. Rules of the form “the Risk Reassessment agent may read SOC 2 reports but not vendor commercial terms unless delegated by a user with commercial authority” become executable predicates the gateways enforce on every call. The alternative, policy that lives in an internal wiki and gets ignored under deadline pressure, does not survive contact with production.
Just-in-time credentials. Every credential minted at the moment of use, scoped to the specific action, expiring in minutes. The credential never enters the agent’s context window. The agent knows it can perform the action; it does not carry the secret that authorizes it. This one discipline collapses the largest category of credential exposure risk, and it lives at the gateways.
Deny by default, allow by contract. New tools, new memory scopes, and new model routes are unavailable to agents until an explicit contract is added to the policy set. The contract names the tool, the scope, the agent identity, the approving user, and the risk tier. The default is closed. This is how you avoid the failure mode where an agent gets access to something it should not have because it was easier to grant a broad scope than to enumerate specific actions.
The deeper version of enforcement is that policy has to be continuous rather than periodic. Pre-deployment reviews, quarterly audits, and annual attestations were how enterprises governed conventional software. Agents change the operating tempo. An agent that behaves inside policy in the morning can drift out of policy by the afternoon as its context, its tools, or its models shift. Enforcement at the gateway means every call gets the current policy applied, not the policy that applied at the last audit.
Evidence: the output the business consumes
The board committee scene at the top of this piece is a story about missing evidence. Not missing controls. Not missing intelligence. Missing evidence, produced in a form the business can consume.
Evidence is the observable, exportable, auditable output of the Trust Fabric. It is what a regulator asks for, what an auditor tests, what a board member reads, what a security officer investigates, and what a CFO uses to attribute cost. If the architecture does not produce evidence as a first-class output, every one of those consumers has to reconstruct it from raw system state, which is expensive when they have time and impossible when they do not.
Four categories of evidence matter most.
Trace evidence. Every action, every model call, every tool call, every memory read and write, correlated end to end through W3C Trace Context propagation. The OpenTelemetry GenAI semantic conventions stabilized through 2025 and 2026 and are now the standard vocabulary. The instrumentation is neutral; the backend is a choice. LangSmith has the deepest LangChain and LangGraph integration and shipped LangGraph Studio v2 with time-travel debugging in early 2026. Arize Phoenix ships fifty-plus research-backed metrics and is open source. Braintrust closed an $80 million Series B at an $800 million valuation in early 2026 on an evaluation-first philosophy. Langfuse is the strong open-source, self-hosted option. Datadog and Honeycomb keep LLM spans alongside the rest of the application telemetry. Traceloop and OpenLLMetry are the OTel-native instrumentation options. Helicone provides a drop-in proxy for teams that want the fastest install. My working advice for CTOs I talk to: instrument with OpenTelemetry from day one, treat the backend as swappable, and pick the platform that fits the workflow rather than the vendor with the biggest marketing budget.
Eval evidence. The golden set, versioned and mined from production traces, gated on every deploy. The eval suite is a change-control artifact, not a research artifact. It runs in CI, blocks releases that regress, tracks quality per workflow over time, and produces the numeric evidence a compliance officer can point to when a regulator asks whether the system’s performance has been monitored. Regenerate the golden set from recent production traces on a rolling cadence so it stays honest about current traffic.
Cost evidence. Token-level accounting, attributed by agent, by workflow, by customer, by tenant, by team. Not per-model-call, which is the metric most teams start with and which tells them almost nothing useful. Cost per outcome. Cost per completed vendor risk reassessment. Cost per resolved support case. Cost per approved renewal. These metrics let a CFO compare the unit economics of an agent to the unit economics of a human doing the same work, which is the comparison the board wants and the one most engineering dashboards cannot produce. The FinOps tooling for this now exists. Amnic and Cloudchipr ship native token tracking. IBM Cloudability offers policy-based chargeback. Snowflake and AWS both ship first-party FinOps agents. The tooling is not the blocker. The discipline of measuring cost per outcome from day one is.
Human oversight evidence. Every approval, every intervention, every escalation, logged as structured data with the reviewer’s identity, the reasoning provided, and the outcome. The evidence has to survive the reviewer moving to another team or leaving the company. Microsoft Copilot Studio’s advanced approvals feature is one visible implementation of the pattern. The underlying discipline is that oversight is only oversight if it produces evidence a regulator can verify.
The architectural claim is that all four categories should flow through a common evidence bus, correlated by trace ID, filterable by agent identity and user delegation, and exportable in formats the compliance function already uses. Building four parallel evidence systems is expensive. Building one is where the leverage lives.
The seven concerns of the Trust Fabric
Attribution, enforcement, and evidence are the three properties. Underneath them, the Trust Fabric addresses seven concerns, each of which touches every plane below it.

Identity. Every agent has a verifiable identity that appears in every log line and access request. Users delegating to agents have their own identities preserved through the delegation chain. Non-human identity governance is a first-class discipline.
Access and policy. Rules expressed as code, enforced at the gateways, denied by default, allowed by explicit contract. Continuous rather than periodic. The gateway is the enforcement point across all five planes.
Observability. Traces, spans, session capture, correlated end to end through W3C Trace Context. Instrumented with OpenTelemetry, exported to a backend the compliance function can query.
Evals. Golden sets mined from production traces, versioned, gated in CI, regenerated on a rolling cadence. Every deploy passes the eval suite or does not ship.
FinOps. Token economics measured as cost per outcome, attributed by agent and workflow and customer, budgeted with hard caps and soft alerts, reported to finance in the same cadence as any other operational cost.
Human oversight. Approval by risk tier rather than on every action. Structured logging of every intervention. Human-on-the-loop as the default pattern, with the specific approval points selected by risk rather than by convention.
Compliance. EU AI Act deployer duties, DORA ICT risk management, GDPR Article 22, ISO/IEC 42001, and whatever local regulations apply to your jurisdiction and your industry. Compliance evidence is a scheduled export, not a project.
Each concern touches every plane. Identity flows through the Agent plane’s harness, the Model plane’s gateway, the Memory plane’s write path, the Tool plane’s MCP gateway, and the Orchestration plane’s approval flows. Observability instruments each layer with the same semantic conventions. FinOps measures token spend across the model, the memory, and the tool. The Trust Fabric is not a separate layer bolted on top. It is the discipline that runs through all five planes.
A worked example: the Risk Reassessment agent’s Trust Fabric
Recall the Risk Reassessment agent from the prior six pieces. Its job is to assemble a current view of a vendor’s risk profile and produce a structured risk score. Here is what its Trust Fabric looks like end to end.
The agent has a SPIFFE-issued workload identity issued by the SPIRE server when the agent container starts. The identity is bound to the agent’s runtime attributes and rotates every hour. The identity appears in every log line the agent emits, every gateway call the agent makes, and every trace span the agent produces.
When a procurement leader delegates a reassessment to the agent, her identity is captured through the enterprise IdP, and an OAuth 2.0 Token Exchange with the act claim binds her identity as the delegating principal for the agent’s session. Every downstream call from the agent, into the AI gateway, into the MCP gateway, into the memory plane, carries both identities. When the Renewal Supervisor delegates to the Risk Reassessment agent that delegates to the Financial Signal Analyzer, the delegation chain is preserved end to end.
Access and policy are enforced at three gateways. The AI gateway enforces model-layer policy: which models this agent may call, which providers are allowed for which jurisdictions, what the token budget for the session is. The MCP gateway enforces tool-layer policy: which tools this agent may reach, which actions are read-only versus write, whether dry-run mode is required. The memory plane’s write-path governance enforces memory-layer policy: which authority tiers are visible to this agent, which tenants are scoped in, which content is filtered out by staleness. All three sets of rules are expressed as OPA policies, versioned in git, deployed alongside the agent code, and tested in CI.
Observability instruments every layer with OpenTelemetry. A single trace ID follows the reassessment from the procurement leader’s click, through the Renewal Supervisor’s delegation, through the Risk Reassessment agent’s work, through every model call and tool call and memory read, to the final structured output. The trace lives in the enterprise APM platform alongside every other application trace. When something goes wrong, the investigation starts from the trace ID and reaches every relevant span in under a minute.
Evals run on every deploy. The golden set was mined from six months of production reassessments where a human reviewer disagreed with the agent’s recommendation or where the recommendation later proved wrong. Every deploy runs the golden set and blocks release on regression. The set is regenerated quarterly from the most recent quarter’s production traces so it stays honest about current workload.
FinOps runs continuously. The token cost of each reassessment is captured at the gateway, attributed by agent, workflow, and customer, and rolled up to a dashboard the CFO checks weekly. Cost per completed reassessment is the primary metric. Cost per model call is a subordinate metric used for debugging. Budgets are set per team, per customer, and per action class, with hard caps that stop the agent before a runaway loop can produce a five-figure invoice.
Human oversight runs by risk tier. Low-risk reassessments proceed autonomously. Medium-risk cases route to the procurement leader’s queue. High-risk cases require explicit approval and a second review from the head of procurement. Every approval is captured as structured evidence with the reviewer’s identity, the reasoning provided, and the outcome.
Compliance evidence exports on a scheduled cadence. The Article 12 record-keeping requirement of the EU AI Act is satisfied by the trace export. The Article 14 human oversight duty is satisfied by the approval evidence. The DORA ICT risk requirements are satisfied by the identity, access, and observability outputs the compliance team already consumes. The evidence lives in the compliance data warehouse alongside every other regulatory export the firm produces.
When the board asks the CIO which agent made the recommendation to terminate the Vendor X contract, the answer takes ninety seconds to produce, includes the agent identity, the delegating user, the policy that applied, the data the agent read, and the human approval trail. The next question of that kind from a regulator produces the same answer in the same ninety seconds.
That is what one agent’s Trust Fabric looks like when it works. The value is that the business, the regulator, the auditor, the board, and the customer each get the specific evidence they need, on demand, without an engineer having to reconstruct it from raw system state.
The 90-day move
If you are reading this and wondering where to begin, here is what I would do this quarter.
- Stand up workload identity for every agent. Pick SPIFFE/SPIRE or the equivalent, issue a unique identity to every agent instance, wire it into every gateway and log line. This is the keystone investment for the Trust Fabric. Nothing else works without it.
- Preserve the delegation chain. Wire OAuth 2.0 Token Exchange with the
actclaim into every agent-to-agent handoff and every user-to-agent delegation. When an incident happens, the chain is your answer to “on whose authority.” - Instrument end to end with OpenTelemetry. W3C Trace Context propagation across every gateway. GenAI semantic conventions on every span. Backend that the compliance function can query. This is the evidence bus everything else feeds.
- Codify policy at the gateway. Every access rule expressed in OPA or the equivalent, versioned in git, tested in CI, enforced at the AI gateway and the MCP gateway. Deny by default. Allow by explicit contract.
- Mine the eval golden set from production. Sample from real traffic, weighted toward failures and escalations. Gate every deploy on the eval suite. Regenerate quarterly.
- Measure cost per outcome. Cost per completed reassessment. Cost per resolved case. Cost per approved renewal. Attribute by agent, workflow, and customer. Report to finance on the same cadence as any other operating cost.
- Calibrate human oversight by risk tier and capture the evidence. Low, medium, high. Different approval regimes. Every approval captured as structured evidence with the reviewer’s reasoning.
That is roughly a quarter of focused work for a small team. It sounds like a lot because it is. It is also the difference between an architecture that scales into regulated production and an architecture that runs into the wall the first time a director or a regulator asks the questions the board asked at the top of this piece.
Closing the series
This is the last piece. It is worth stepping back one level and naming what the seven pieces together are trying to say.
The claim I have been defending across the series is that the architecture is the product. Not the model. Not the framework. Not the interface. The architecture. Every plane refined that claim from a different angle. The Agent plane made the case that harness engineering is where reliability lives. The Model plane made the case that a substitutable, tiered portfolio is what survives the model market’s commoditization. The Memory and Knowledge plane made the case that curating before storage is the discipline that keeps the system’s long-term character honest. The Tool and Action plane made the case that standardization, composability, and controllable autonomy are what let autonomous software do useful work at scale. The Orchestration and Experience plane made the case that topology fits the task and delegation is the product. The Trust Fabric makes the case that attribution, enforcement, and evidence are what turn all of it into a system a business can put its name on.
None of the seven pieces is a complete architecture on its own. Each one hides the mess so the reasoning above it can compose. Together they form the reference architecture I wish I had been handed when I started building agent systems three years ago.
The claim about durability matters most now, at the end. The frameworks will change. The models will change. The vendors will consolidate and re-fragment. The regulations will evolve. What will not change is that agents that operate in the real world need to be reliable, substitutable, memorable, capable of action, orchestrated into outcomes, and accountable. The seven-piece decomposition is a way of separating the concerns that let a team build for that reality rather than for whatever the current model release cycle happens to be optimizing for.
The next chapter of the field is already visible from where we are. Agent-to-agent protocols will mature into cross-organizational interoperability. Non-human identity will become as governed as human identity is today. Evidence will become a scheduled export rather than a manual reconstruction. Cost per outcome will become as ordinary a metric as unit cost of goods sold. The regulatory expectation will be that enterprises operating agents can answer every question the board asked at the top of this piece, and the tooling to answer them will be commodity infrastructure. The teams that get there first will have a durable advantage. The teams that arrive late will pay the retrofit tax.
If you take one thing from the series, take this: the model is not the product; the architecture is the product. Build for durability. Design the planes deliberately. Thread the Trust Fabric through all of them. The intelligence is the easy part. The discipline is the work.
Thank you for reading. This is the last piece in this series. The next one, whenever it comes, will start from what this series established rather than repeat it. If any of the seven pieces changed how you approach a problem, that is what I hoped for when I started. If you disagree with a claim I made, write about it. The field will be better for the argument.
Further reading
- Regulation (EU) 2024/1689. EU Artificial Intelligence Act. High-risk provisions in force August 2, 2026. Articles 10, 12, 14, 15, 26, and Annex III.
- Regulation (EU) 2022/2554. Digital Operational Resilience Act. In force January 17, 2025.
- Colorado. AI Act (SB24-205). Effective June 30, 2026.
- Singapore IMDA. Model AI Governance Framework for Agentic AI. January 2026.
- NIST. AI Agent Standards Initiative. February 2026. NIST SP 800-207 Zero Trust Architecture, Section 5.7 on non-person entities. NIST SP 800-207A.
- ISO/IEC 42001:2023. Information technology, Artificial intelligence, Management system.
- Anthropic. Trustworthy Agents in Practice. April 2026. Zero Trust for AI Agents. May 2026.
- Google. SAIF 2.0 Agent Risk Map. 2026.
- Microsoft. Entra Agent ID and Agent Governance Toolkit. 2026.
- Forrester. AEGIS Framework. April 2026.
- Sure, R. W. CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI. July 2026.
- OWASP. Top 10 for Agentic Applications. December 2025. Peer-reviewed by NIST, Microsoft AI Red Team, AWS.
- SPIFFE. Secure Production Identity Framework For Everyone. SPIRE runtime. https://spiffe.io
- IETF. RFC 8693, OAuth 2.0 Token Exchange. The
actclaim for delegation. - Cloud Security Alliance. Non-Human Identity Governance Vacuum. 2026 whitepaper on the NHI crisis in agent systems.
- FinOps Foundation. State of FinOps 2026. AI cost management as the top forward-looking priority.
- EY. Agentic AI Enterprise Token Cost. June 2026.
- Independent evaluations of LangSmith, Arize Phoenix, Braintrust, Langfuse, Datadog LLM Observability, Honeycomb LLM Observability, Traceloop, Latitude, Galileo, Maxim, Helicone, AgentOps, Confident AI (all current 2026).