Six dimensions, one per discipline, and five levels from Ad hoc to Optimised. Find where you are and what to do next.
Levels describe how AI use is organised, not how much AI is used. A team with modest AI use and strong controls is more mature than one with heavy, untracked use.
| Level | What it looks like |
|---|---|
| 1. Ad hoc | Individuals use AI tools on their own initiative; no inventory, no policy; AI output untracked. |
| 2. Experimenting | Sanctioned pilots with named tools; a basic usage policy; practice shared informally; value anecdotal. |
| 3. Managed | Approved tool catalogue; risk tiers defined and applied in review; provenance recorded; baseline metrics. |
| 4. Governed | Controls enforced in pipelines as policy-as-code; audit trail queryable; the second line assures; roles staffed. |
| 5. Optimised | Delegation adjusted from evidence; context managed as a product; value and cost optimised continuously. |
Aim for the level your risk exposure needs, not for level 5 everywhere. A bank running critical-tier work with AI needs Governed controls; a startup doing low-tier work may be well served by Managed.
Each row is a discipline; each cell describes that level for it.
| Dimension | 1. Ad hoc | 2. Experimenting | 3. Managed | 4. Governed | 5. Optimised |
|---|---|---|---|---|---|
| AI-Augmented Delivery | engineers use AI tools privately; no one can say which changes were AI-assisted. | named teams pilot approved tools in development; results shared informally. | tiers applied at planning; AI-assisted changes marked; review depth follows the tier. | the pipeline enforces provenance, licence, eval and approval gates as policy-as-code across all stages. | delegation per stage adjusted from change-failure and rework data; retirement automated from the element register. |
| Human–AI Workflow Design | individuals decide for themselves what to hand to AI; approvals and escalations are informal or missing. | pilot teams agree simple delegation rules; agents run at low autonomy with a person watching. | process maps show human, AI and tier for each step; approval gates and escalation queues defined and used. | autonomy caps and gates enforced by tooling; approval and escalation records queryable; second line reviews high and critical workflows. | autonomy levels raised or lowered from override, escalation and incident data; review load balanced continuously. |
| Tooling & Context Engineering | individuals choose their own tools and paste context by hand; integrations run on personal tokens. | a shortlist of sanctioned tools; shared prompts circulate in wikis and chat without owners. | approved catalogue with tiers of use; context assets in repositories with owners; agents have their own identities. | integrations only through the allow-list or gateway; evals and context hygiene enforced in the pipeline; non-human identities recertified. | context managed as a product with measured effect on output quality; tools swapped using the exit plan without disruption. |
| Governance & Guardrails | AI tools used on personal accounts; no policy, no inventory, no view of which providers hold company data. | a basic usage policy; sanctioned pilots with named tools; provider terms checked case by case. | approved tool catalogue and tier-based policy; inventory drawn from the element register; AI providers in vendor management. | guardrails enforced as policy-as-code; kill switches tested; second line assures from pipeline data; internal audit covers AI-assisted work. | controls and delegation boundaries tuned from monitoring and audit evidence; revocation is fast and routinely exercised. |
| Skills & Roles | people teach themselves; nobody knows who is competent to review AI-assisted work. | pilot teams share practice informally; some introductory training exists. | baseline AI literacy for all tool users; roles assigned as responsibilities; approval authority defined per tier. | training records and competence checks enforced at the approval gate; roles staffed; second line assures. | enablement adjusted from review quality and incident data; career paths reward review and specification. |
| Value & Measurement | value claimed from anecdote and licence counts; no baseline exists. | pilots report usage and satisfaction; outcomes are estimated, not measured. | baselines taken before adoption; value hypotheses written and tested; AI-assisted vs not compared at team level. | total cost reported per value stream; portfolio decisions to expand or stop rest on reviewed evidence. | delegation and spend adjusted continuously from flow, quality and cost data. |
Pick the level that best describes your organisation today for each dimension. The descriptor for your choice appears below it. Nothing is stored or sent: the assessment runs entirely in your browser.
Run it with delivery teams and the second line together; the disagreements are often more useful than the score. See also how self-assessment fits the control pillars.