Prompts-as-Code Kit
Version, review and ship prompts like any other code.
Why
Same prompt, different answer
For decades, software behaved predictably. The same input and the same code gave the same output, and a unit test could prove it. AI changes that. A model can answer the same prompt differently twice, and the flaws it introduces are often subtle enough to slip past tests that were written for predictable code.
How it works
From draft to ready to ship
- 1
Write the spec
Name the prompt, give it an owner and a model, and keep the system instructions, template and retrieval settings together, because the playbook versions all three.
- 2
Add test cases
Write inputs with the behaviour you expect. They are what the evaluation gates and reviewers run against.
- 3
Pass the gates
Tick each control only when it is true: evaluation, red-teaming, the four-part verification gate, and, for agents, the five LCAC constraints.
- 4
Version the change
Bump the version and write what changed. A new version starts with an empty checklist, because it has not been through the gates yet.
- 5
Export and commit
Download prompt.yaml and commit it next to the code that calls it; share the Markdown in the pull request.
The kit
Build a prompt spec
Everything stays in this browser. Keep several prompts in the library, and export each one when it is ready.
Review checklist
Tick an item only when it is true for this version.
Export
# prompt.yaml · Prompts-as-Code Kit
# Cognitive XOps, module 01: prompts, system instructions and retrieval settings,
# versioned in Git, reviewed, and tested. www.fawzooz.ai
id: ""
name: ""
version: "0.1.0"
owner: ""
purpose: ""
model:
primary: ""
degraded_mode_fallback: ""
prompt:
system: ""
template: ""
inputs: []
retrieval: ""
constraints: ""
agent:
enabled: false
finops:
token_budget: ""
circuit_breaker: false
risk:
reversibility: easy
approver: ""
evaluation:
# minimum scores for the automated evaluation gates
faithfulness: null
answer_relevance: null
context_grounding: null
red_team: [prompt_injection, jailbreaks, context_leakage, hallucinated_dependencies]
tests: []
standards: []
review:
reviewer: ""
checks:
code_git: false
code_reviewed: false
code_tested: false
code_owner: false
eval_faithful: false
eval_relevant: false
eval_grounded: false
red_injection: false
red_jailbreak: false
red_leakage: false
red_deps: false
gate_facts: false
gate_logic: false
gate_missing: false
gate_impact: false
run_breaker: false
run_fallback: false
approve_delegation: false
approve_mapped: false
ready_to_ship: false
changelog:
- version: "0.1.0"
date: ""
change: initial
note: ""
Your prompts are stored only in this browser (localStorage). Nothing is sent anywhere.
Where the kit adds structure: the playbook prescribes the controls, not a file format. The spec fields, semantic versioning, the three reversibility bands and the YAML layout are this kit’s scaffolding around those controls.
Reference
The modules this kit puts to work
Every checklist item below comes from one of these modules, in the playbook’s own words.
- Module 01
Prompts as code
We treat prompts, system instructions and retrieval settings exactly like source code: versioned in Git, reviewed, and tested. Before anything merges, automated evaluation gates built on tools such as DeepEval and RAGAS check that answers stay faithful to their sources, relevant to the question and grounded in the right context.
- Module 03
Zero Trust in the reasoning layer
Read moreShow less
Agents get the least context they need for the task in front of them, an approach we call Least-Context Access Control (LCAC). Five constraints apply to every agent: API keys with the least privilege possible; a hard ceiling on reasoning steps (for example, ten loops); predefined points where a human must approve before the agent continues; signed execution logs; and a kill switch that stops the agent instantly.
- Module 04
Continuous red-teaming
Adversarial testing runs inside the build pipeline. Before a pull request is approved, automated scripts try prompt injection, jailbreaks, context leakage and hallucinated dependencies against the change. If a guardrail breaks, the build stops.
- Module 07
AI FinOps
We track token use, inference cost and the return on semantic caching for each deployment. A financial circuit breaker halts the pipeline when a change causes a sudden spike in model calls, which protects against the Unbounded Consumption risk in the OWASP Top 10 for LLM Applications (2025).
- Module 08
Degraded mode
Every team has a written plan for the day its main AI provider is down, slow or rate-limited. Work falls back to smaller local models, or moves to a human review queue, so delivery slows down rather than stops.
- Module 09
The developer’s daily playbook
Developers register any new AI tool through an open intake register, which sorts it against a five-dimension risk taxonomy.
Read moreShow less
A delegation matrix decides who may approve what, based on how easily an action can be reversed. And every AI-generated change passes a four-part verification gate before commit: check the facts, validate the logic and security, look for what is missing, and assess the impact.
- Module 11
Standards alignment
Each control maps to recognised standards: ISO/IEC 27001:2022 for information security, ISO/IEC 42001:2023 for AI management systems, ISO 45003:2021 as guidance on psychosocial risk, the NIST AI Risk Management Framework (AI RMF 1.0, 2023), and the EU Artificial Intelligence Act.
Source
Built on the applied playbook “Cognitive XOps” (2026): module 01, Prompts as code, with the controls of modules 03, 04, 07, 08, 09 and 11 and the Wave 1 roadmap rule that every model has an owner.
Elgendi, M. F. (2026). Cognitive XOps: An applied playbook. www.fawzooz.ai
