Table of Contents
Executive Summary
Agentic AI assurance needs to keep pace with the deployment of autonomous AI agents.
While AI products can be certified and management systems audited, a specific AI agent deployment – a particular agent, with its permissions, data, and human oversight – receives guidance and self-assessment but little independent, evidence-based assurance.
This gap creates three problems: institutional over-trust, an inability to demonstrate control adequacy when an agent causes harm, and unpriceable risk that slows adoption.
Drawing on regulation, established assurance practice and the agentic risk surface (broadly defined), this paper proposes twelve criteria any solution should meet.
The next step is to compare existing standards, frameworks and assurance schemes against these criteria and engage standards bodies to close the gap.

The Agentic AI Assurance Gap at the Point of Deployment
When a significant system is supplied by one party but deployed and operated by another, some form of independent assurance or inspection is common in regulated and high-trust domains.
However, while AI management systems and some AI products or services can be certified or otherwise assessed, independent, evidence-based assurance of an organisation’s own specific agent deployment remains less mature.
Despite available deployment guidance, self-assessments and a growing AI-assurance market, our initial review has not identified a widely adopted, agentic-specific standard or scheme that enables independent conclusions about whether the controls of a particular deployment are adequately designed, implemented and operating effectively.
The Agentic AI Assurance Gap Creates Three Problems
Three scenarios illustrate the potential problem:
1. Over-Trust at Institutional Scale – in the current environment, a board may feel assured when it reads “our platform is certified and we hold ISO 42001.” Yet the organisation’s specific deployment may still not have been independently examined: the permissions granted, the tools made available, the data connected, the monitoring and governance. This can create false comfort because supplier or management system certification does not by itself demonstrate that deployment-specific controls are adequate or discharge the board’s own oversight responsibilities.
2. An Unanswerable Question When Something Goes Wrong.
A distinct risk scenario involves a well-functioning, non-compromised agent acting within its granted permissions but taking a harmful path the user did not intend to authorise; Anthropic’s March 2026 submission to NIST describes examples of this scenario.
This type of event may arise from the deployment context rather than a conventional product defect and, depending on the applicable obligations and circumstances, the deployer may need to show that its controls were adequate, implemented and working.
If the controls are found wanting, regulatory scrutiny, client challenge, litigation or other consequences may follow.[1]
Yet our initial review has not identified a generally accepted, agentic-specific benchmark that currently defines what evidence would demonstrate, at deployment level, that such controls are adequately designed, implemented and operating effectively.
This can leave boards with clear responsibilities but no widely accepted, agentic-specific assurance benchmark against which to test them.
3. Unpriceable Risk Slows Adoption – consistent assurance can help risks be compared and may support pricing. Without comparable evidence, insurers may find an agentic deployment harder to underwrite, while management may find it harder to distinguish a well-governed agent from a poorly governed one. Good practice could therefore go unrewarded, reducing incentives to aim for high standards and potentially slowing adoption in regulated or high-trust organisations.
[1] The EU AI Act Article 26, DORA’s operational-resilience obligations, and SS1/23 assign relevant obligations to deployers or using organisations within their respective scopes.

We Expect Assurance in Other Domains. Should Agentic AI Be Different?
When a company prepares its accounts under an applicable financial reporting framework (the guidance), management takes responsibility for them (the assertion), and an independent auditor expresses an opinion on whether the financial statements give a true and fair view (the assurance).
Similarly, when an organisation relies on a service provider, an independent practitioner may report on the provider’s controls. A SOC 2 Type 2 report can address whether relevant controls operated effectively over a period; an ISAE 3000 engagement can cover other suitable non-financial assurance subject matters.
And in everyday life, a car built by a manufacturer to type-approval standards must then pass ongoing and regular independent roadworthiness tests.
Our initial review has not identified a broadly adopted equivalent focused on the controls of a specific deployment: certification of a provider, product or management system does not necessarily inspect how an agent is configured and governed in use. AI-assurance services already exist, and the UK market is growing, but the agentic AI deployment assurance category remains immature.
Other domains provide further, though not exact, analogies:
- Organisations handling payment-card data must validate PCI DSS compliance in the manner required by the relevant payment brands or acquirer; some assessments are performed by a Qualified Security Assessor, while eligible organisations may self-assess.
- Under DORA, European financial entities must maintain digital operational resilience testing programmes; for entities other than microenterprises, testing follows a risk-based approach and must be undertaken by independent internal or external parties, and may include vulnerability assessments and penetration testing.
- Building-control regimes also commonly inspect completed work rather than relying only on approved designs and materials; for example, higher-risk buildings in England require a completion certificate before occupation.

The Criteria for an Effective Agentic Assurance Solution
To address this agentic AI assurance gap, deployers of their own agentic AI need a solution for assuring their boards, auditors and regulators.
As a first step, we propose 12 agentic AI assurance criteria that any solution should meet to be effective. Each answers one or more of the three problems above:
- Independence and evidence-based verification (criteria 1, 7, and 9) address false comfort.
- Three-level depth and tested oversight quality (6 and 11) produce the evidence a board would need when challenged.
- A common, traceable, current method (10 and 12) supports more consistent comparison and, potentially, risk pricing.
The criteria derive from three features of the assurance landscape – regulation, established assurance practice, and the agentic risk surface – set out in the Methodology below.
| Criterion | Test to be applied to each solution |
|---|---|
| 1. Deployer-anchored subject | Assessment covers the deployer’s own control environment. |
| 2. Workflow-in-context unit | Conclusion varies with the permissions, data, and oversight of the deployment. |
| 3. Broad risk-surface coverage | Coverage across defined agentic risk categories, cross-checked against multiple external taxonomies and frameworks. |
| 4. Proportional applicability | Documented, risk-justified statement of applicability for each control. |
| 5. Shared-responsibility allocation | Per-control provider / builder / deployer split; role migration handled. |
| 6. Three-level depth | Distinction between design, implementation, and operating effectiveness. |
| 7. Evidence-based verification | Artefacts and testing required; attestation sampled and cross-checked. |
| 8. Tempo matched to change | Re-assessment triggers detect deployment-side change between cycles. |
| 9. Independence | Assessor independence mandated; designer-verifier conflicts managed. |
| 10. Regulatory traceability | Maintained mappings to ISO/IEC 42001, NIST AI RMF, EU AI Act, sectoral regimes. |
| 11. Oversight quality | Effectiveness of human oversight tested, not just its presence. |
| 12. Currency | Governed, version-controlled updates at the pace of agentic change. |

The Criteria in Detail
Each criterion states a requirement that derives from one of three features of the landscape – regulation, established assurance practices, or the identified agentic risk surface – which are set out in the Methodology below. The derivation noted under each criterion maps it to one of these features.
1. Deployer-focused
Criterion. The subject of the assurance is the deploying organisation – its decisions, permissions, oversight, and governance of agentic workflows – not solely the provider’s product or the provider’s organisation.
Derivation. EU AI Act Article 26 obligations for deployers of high-risk AI systems; DORA Articles 5–6 and 28, under which financial entities retain responsibility for ICT risk, including third-party risk; and SS1/23, under which firms within scope retain responsibility for models they use, including externally developed models.
How it will be tested. Does the solution’s assessment subject include the deployer’s own control environment, or only the provider’s product and practices?
2. Workflow-in-context unit of assessment
Criterion. The unit assessed is the agentic workflow in its operating context: the agent plus its granted permissions, data access, tool catalogue, orchestration with other agents, and human hand-offs – not the agent product in isolation.
Derivation. Anthropic’s March 2026 submission to NIST shows why outcomes can depend on permissions and operating context, not solely on the underlying product; EU AI Act Article 26(1), (4) and (5) impose context-specific duties on deployers of high-risk AI systems.
How it will be tested. Can the solution’s assessment reach a different conclusion for two deployments of the same product with different permissions, data, and oversight?
3. Broad risk-surface coverage
Criterion. The solution should cover a broad, defined agentic risk surface: individual agent behaviour, multi-agent interaction, security, governance and human factors – including deployment-specific issues such as shadow agents, over-trust, oversight quality and organisational change.
Derivation. Agentic Risks’ categories A–E were cross-checked against the NIST AI Risk Management Framework functions, the Cloud Security Alliance AI Controls Matrix domains, and the OWASP Top 10 for Agentic Applications and MIT AI Risk Repository to test breadth and reduce dependence on a single taxonomy.
How it will be tested. Mapping the solution’s control set against all risk categories: which does it address, and at what depth?
4. Proportionality with justified applicability
Criterion. Control expectations are selected in proportion to the assessed risk of the specific workflow, with every inclusion and exclusion justified and recorded, so that ‘considered and excluded’ is distinguishable from ‘never considered’.
Derivation. ISO/IEC 27001 clause 6.1.3(d), which requires a Statement of Applicability identifying necessary controls, justifying their inclusion and explaining exclusions from Annex A; ISO/IEC 42001’s risk-based management-system approach; and ISO 31000’s principle that risk management should be customised to the organisation’s context.
How it will be tested. Does the solution require a documented, risk-justified applicability decision for each control, or apply a fixed scope?

5. Shared-responsibility allocation
Criterion. Each control is explicitly allocated across the value chain – platform provider, builder and deployer – including recognition that a party’s role and duties may change when applicable legal conditions are met, such as substantially modifying a high-risk AI system or changing an AI system’s intended purpose so that it becomes high-risk.
Derivation. EU AI Act Article 25 value-chain provisions; the shared-responsibility model established in cloud computing; and Agentic Risks’ controls on accountability chains and substantial modification.
How it will be tested. Does the solution provide a per-control responsibility split, and does it address role migration for organisations that build agents on third-party platforms?
6. Three-level assurance depth
Criterion. The assessment distinguishes design adequacy, implementation, and operating effectiveness, and reports findings at the level at which the control failed, since each level implies a different remediation.
Derivation. ISAE 3000’s evidence-based assurance framework; the SOC 2 Type 1 / Type 2 distinction; and established internal-control practice under COSO.
How it will be tested. Does the solution assess and report at all three levels, or attest to presence and design only?
7. Evidence-based verification
Criterion. Conclusions rest on verifiable artefacts – registries, logs, drill records, signed inventories – machine-testable where possible, with self-attestation reserved for matters no artefact can evidence and subject to sampling and independent cross-checks.
Derivation. ISAE 3000 evidence requirements; ISO 19011 audit-evidence guidance; the anti-gaming logic of Agentic Risks’ controls on anti-gaming indicators and metric-integrity audits.
How it will be tested. What evidence does the solution require – artefacts and testing, or questionnaire and attestation?
8. Assurance tempo matched to change tempo
Criterion. The cadence of assurance matches the speed at which the assured environment changes: drift, tool-catalogue changes, permission creep, and memory accumulation are detected within a period proportionate to the workflow’s autonomy and impact, rather than at fixed annual or quarterly intervals alone.
Derivation. EU AI Act Article 26(5) deployer monitoring and Article 72 provider post-market monitoring; DORA’s digital operational resilience testing; and Agentic Risks’ behavioural-drift risk, which can move at configuration speed rather than certification-cycle speed.
How it will be tested. What is the solution’s re-assessment trigger and cadence, and can it detect deployment-side change between cycles?

9. Independence and conflict management
Criterion. The party providing the assurance is independent of the parties that designed, built, or operate the controls, with any residual conflict disclosed and managed – including where the assurer also offers design or remediation services.
Derivation. ISO/IEC 17021-1 and ISO/IEC 42006 requirements on competence, consistency and impartiality for AI management-system certification bodies; the Institute of Internal Auditors’ Three Lines Model.
How it will be tested. Does the solution mandate assessor independence, and how are designer-verifier conflicts handled?
10. Regulatory traceability
Criterion. Findings can be mapped to the legal obligations, standards and voluntary frameworks the deployer is required or has chosen to address – such as the EU AI Act, sectoral regimes, ISO/IEC 42001 and the NIST AI Risk Management Framework – so that a gap can be described as a potential regulatory or governance exposure and remediation can contribute to, rather than by itself establish, compliance evidence.
Derivation. Deployers’ need to demonstrate compliance and governance across multiple applicable regimes and frameworks.
How it will be tested. Does the solution publish maintained mappings from its controls to the major regimes?
11. Human accountability and oversight quality
Criterion. A named, competent human owner exists for every agentic workflow, and the quality of human oversight – not merely its presence – is tested, including for rubber-stamping and automation bias.
Derivation. EU AI Act Article 14, which requires high-risk AI systems to be designed for effective human oversight, and Article 26(2), which requires deployers of high-risk systems to assign oversight to natural persons with the necessary competence, training and authority; SS1/23 accountability expectations for firms within scope; and Agentic Risks’ human-oversight and staff-over-trust risks.
How it will be tested. Does the solution test oversight effectiveness – review capacity, override rates, approval quality – or only require that oversight exist?
12. Currency under governed change control
Criterion. The criteria and control content of the solution are maintained under version control by a governed process, at a cadence matching the evolution of agentic risk, so that assurance performed this year is not performed against last year’s threat model.
Derivation. The observable pace of agentic AI evolution; standards-body revision practice; the good practice of version-control.
How it will be tested. What is the solution’s update mechanism, and what is its observed revision cadence relative to the pace of agentic change?

Next steps for Agentic Risks
- Publish for external comment, debate, and input into an updated version of the criteria. Comments are invited here until 30 September, after which we will publish the updated version of the criteria.
- On completion, we will engage the standards bodies to understand what in-flight initiatives they have in this area that could help solve this problem. We will then perform a gap analysis between the updated criteria and selected existing standards, frameworks, schemes, legal requirements and assurance methods.
- At the time of publishing the draft criteria, these include AIUC-1, ISO/IEC 42001 together with ISO/IEC 42006, CSA AICM/STAR, NIST AI RMF, the EU AI Act, BSI AIC4, Singapore’s Model AI Governance Framework for Agentic AI and AI Verify, and ISAE 3000.
Methodology
To derive effective criteria, we examined three features of the assurance landscape – regulation, established assurance practices and the identified agentic risk surface.
Anchoring each criterion is deliberate: it makes the reasoning transparent and allows readers to challenge the source, our interpretation of it, or the criterion derived from it.
1. Regulation
The applicable obligations that regulation places on relevant organisations and use cases, since an assurance solution should help organisations discharge their legal obligations.
In the regimes considered here, deployers or using organisations retain duties that generally cannot be discharged solely by pointing to a supplier’s certificate. For example:
- The EU AI Act places distinct obligations on deployers of high-risk AI systems, including use in accordance with instructions, competent human oversight, input-data relevance where the deployer controls those data, monitoring and log retention (Article 26). Article 25 governs specified circumstances in which roles shift along the AI value chain, while Article 72 requires providers of high-risk AI systems to establish post-market monitoring systems.
- The EU Digital Operational Resilience Act (DORA) requires financial entities to remain fully responsible for compliance when they use third-party ICT services.
- In the UK, the Prudential Regulation Authority’s supervisory statement SS1/23 applies to specified banks, building societies and PRA-designated investment firms and requires them to manage model risk for models they use, including externally developed models.
2. Established assurance practices
An effective solution should draw on established disciplines used in audit and certification to support reliable assurance conclusions:
- ISAE 3000 requires practitioners to obtain sufficient appropriate evidence for non-financial assurance engagements; the AICPA’s SOC 2 regime distinguishes reports on control design at a point in time from reports that also address operating effectiveness over a period.
- ISO/IEC 27001’s Statement of Applicability must identify the necessary controls, justify their inclusion, state whether they are implemented and justify exclusions from Annex A.
- ISO/IEC 17021-1 and ISO/IEC 42006 set competence, consistency, impartiality and conflict-management requirements for bodies that certify AI management systems.
- ISO 19011 provides guidance on evidence-based auditing and treats audit evidence as verifiable information.
- The Institute of Internal Auditors’ Three Lines Model distinguishes management roles from the independent assurance provided by internal audit.
3. The agentic risk surface
One important deployment-side scenario, articulated in Anthropic’s March 2026 submission to the US National Institute of Standards and Technology, involves a well-functioning, non-compromised agent acting within its granted permissions but following a path the deployer never intended to authorise.
Much of this risk surface – permission scoping, oversight quality, unregistered ‘shadow’ agents, staff over-trust, board governance and organisational change – may be missed by an assessment limited to the AI product alone.
Therefore, an effective solution should address risks that arise from deployment context as well as product design. We cross-checked the proposed coverage against external sources – the OWASP Top 10 for Agentic Applications, the MIT Risk Repository, the NIST AI Risk Management Framework, and the Cloud Security Alliance AI Controls Matrix – alongside Agentic Risks’ Enterprise-Wide Agentic AI Risk Control Framework.
Each contributes a different perspective:
- The OWASP Top 10 for Agentic Applications is a globally peer-reviewed, security-focused framework for autonomous and agentic systems, cataloguing risks that include goal hijacking, tool misuse, identity and privilege abuse, memory and context poisoning, and rogue agents.
- The MIT AI Risk Repository is a living database of more than 1,700 risks extracted from 74 frameworks and classifications, organised through causal and domain taxonomies.
- The Enterprise-Wide Agentic AI Risk Control Framework pairs its risk categories with specific controls and extends beyond security into governance, human factors and deployment operations, including board oversight, oversight quality, staff over-trust, shadow agents and organisational change.
About Agentic Risks
Agentic Risks IP publishes its agentic AI risk and control frameworks as a not-for-profit initiative and provides free agentic risk management webinars through the Institute of Risk Management. Agentic Risks Ltd, the commercial business, provides independent risk management services to organisations deploying agentic AI.
Agentic AI introduces a new form of software-based delegated work that may reshape how organisations allocate tasks and accountability: the human-agent organisation. Enter it with confidence.
Frequently Asked Questions
Agentic AI assurance is independent, evidence-based confirmation that a specific AI agent deployment’s controls are adequately designed, implemented, and operating effectively. It differs from certifying an AI product or auditing a management system: it examines the deployment itself – a particular agent, with its granted permissions, connected data, tools, and human oversight – in its real operating context. It asks not whether an agent can be safe, but whether this organisation’s use of it demonstrably is.
Guidance, self-assessment, and assurance are three distinct things. Guidance tells an organisation what it should do; self-assessment is the organisation rating its own compliance; assurance is an independent party concluding, on evidence, that the controls are adequate, implemented, and working. For agentic AI deployment, guidance and self-assessment are increasingly available, but independent AI assurance is far less mature – which is the gap this work addresses. The distinction matters because a self-declared score is not the same as independent evidence a board or regulator can rely on.
No. ISO/IEC 42001 certification of an organisation’s AI management system, or certification of an AI product, does not by itself assure a specific AI agent deployment. A certificate can attest only to what the certifier examined; unless its scope includes the deployment, it cannot address the permissions granted, data connected, tools made available or human oversight applied. The same certified agent can be deployed safely or recklessly, and the certificate reads identically in both cases.
An organisation can demonstrate its agentic AI controls at three levels: design (are the controls capable of mitigating the risk?), implementation (do they actually exist in the environment?), and operating effectiveness (have they worked over a period, not just on the day of assessment?). Evidence should be verifiable artefacts – agent registries, permission configurations, logs, kill-switch drill records – rather than assertion alone. Reporting findings at the level where a control fails matters, because a design gap, a build gap, and an operating gap each require a different remedy.
Effective agentic AI assurance should assess the deployer’s control environment and the AI agent deployment in context, including its permissions, data, tools, orchestration and human hand-offs. It should also test broad risk coverage, proportionality, shared responsibility, control design, implementation, operating effectiveness, evidence, reassessment triggers, independence, regulatory traceability, human oversight and currency.
Comments Section
Comment on the Agentic AI Assurance at the Point of Deployment
We invite you to leave your thoughts below. Please leave your name and email address, so we can get in touch, and to minimize spam.


