Section 4
4UCSMM as the Yardstick
Cognitive XOps uses UCSMM to answer one question every six months: are we actually becoming more capable, or just busier?
4.1 Four dimensions and five levels
| Dimension | What it covers in a software team |
|---|---|
| D1 · AI governance | Executive oversight, a register of every model in use, risk-based classification, and the authority to stop an AI system. |
| D2 · Cyber resilience | The ability to absorb model failures, keep working when an AI provider is down, and recover cleanly. |
| D3 · Threat detection | Behavioural baselines, prompt-injection testing, context-drift monitoring and checks on where data comes from. |
| D4 · Adaptive management | How quickly lessons from incidents and retrospectives turn into updated pipeline policy. |
Each dimension is placed on one of five levels: Level 1 (Initial), Level 2 (Managed), Level 3 (Defined), Level 4 (Quantitatively managed) and Level 5 (Cognitive-optimised).

Text in this figure
Initial · Managed · Defined · Quantitatively managed · Cognitive-optimised · Above Level 3: live dashboards and automated metrics · D1 · AI governance · D2 · Cyber resilience · D3 · Threat detection · D4 · Adaptive management · Overall maturity = the weakest dimension
How the level is set
Each item in the UCSMM instrument is rated from 1 to 5, and each dimension has five items. A dimension’s level is its average rounded down, so an average of 2.6 means Level 2. No dimension rises above Level 3 without live dashboards and metrics pulled automatically from systems. The assessment is repeated every six months.
In this playbook we also report overall maturity as the level of the weakest dimension. A strong policy cannot make up for a team that cannot recover from an outage.
4.2 The XOps Profile
The official UCSMM level comes from the study’s registered twenty-item instrument. Alongside it, Cognitive XOps tracks twenty operational indicators of its own, mapped to the same four dimensions. They do not replace the UCSMM instrument. They show where the day-to-day work is moving between two formal assessments.
| Dim. | Indicator | Looks like Level 1 | Looks like Level 3 | Looks like Level 5 |
|---|---|---|---|---|
| D1 | 1. Model register | No list of models in use | A standard register, kept by hand | Continuous, automated inventory |
| D1 | 2. Risk classification | None | A formal five-dimension taxonomy | Real-time routing by risk |
| D1 | 3. Stop authority | Nobody can stop an AI system | A written delegation matrix | Automated circuit breakers |
| D1 | 4. Executive sponsorship | Ad hoc | A dedicated AI steering board | Board-level cognitive audits |
| D1 | 5. Policy enforcement | Paper policy | Policy gates in CI/CD | Compliance-as-code |
| D2 | 6. Degraded-mode playbooks | Chaos during outages | Defined failover to small local models | Self-healing local routing |
| D2 | 7. State-Save and rest | Ignored | Mandatory 15-minute State-Save | Rest prompts driven by telemetry |
| D2 | 8. Rollback speed | Hours | Rollback within one sprint | Automated zero-downtime rollback |
| D2 | 9. Incident culture | Blame | Blameless retrospectives | Regular Failure Swap sessions |
| D2 | 10. Financial circuit breakers | None | Monthly token limits | Real-time halts per commit |
| D3 | 11. Prompt-injection testing | None | Periodic manual tests | Continuous automated fuzzing in CI/CD |
| D3 | 12. Provenance checks | Unverified | SBOM and licence scans | Automated training-data provenance |
| D3 | 13. Context-drift telemetry | None | Weekly drift reviews | Real-time distribution alerts |
| D3 | 14. Shadow AI discovery | Blind | Network proxy logs | In-overlay SASE telemetry |
| D3 | 15. Crypto-agility scanning | Legacy RSA and ECC only | Annual crypto audit | Automated scanning for FIPS 203, 204 and 205 readiness |
| D4 | 16. Feedback-loop speed | Months | Updates after each sprint retrospective | Automatic policy updates after incidents |
| D4 | 17. Training | None | Wise AI Ambassador programme | Practical gates delivered continuously |
| D4 | 18. Verification-gate audits | Occasional sampling | The four-part verification gate | Automated hallucination scoring |
| D4 | 19. Dual-dashboard tuning | Delivery metrics only | DORA and CLI reviewed together by hand | Throttling proposals generated automatically |
| D4 | 20. Continuous improvement | Static | Annual framework review | Pipelines that improve from their own data |
Tip: use ← → to move between sections.
