Prompts-as-Code Kit

Version, review and ship prompts like any other code.

Why

Same prompt, different answer

For decades, software behaved predictably. The same input and the same code gave the same output, and a unit test could prove it. AI changes that. A model can answer the same prompt differently twice, and the flaws it introduces are often subtle enough to slip past tests that were written for predictable code.

How it works

From draft to ready to ship

  • 1

    Write the spec

    Name the prompt, give it an owner and a model, and keep the system instructions, template and retrieval settings together, because the playbook versions all three.

  • 2

    Add test cases

    Write inputs with the behaviour you expect. They are what the evaluation gates and reviewers run against.

  • 3

    Pass the gates

    Tick each control only when it is true: evaluation, red-teaming, the four-part verification gate, and, for agents, the five LCAC constraints.

  • 4

    Version the change

    Bump the version and write what changed. A new version starts with an empty checklist, because it has not been through the gates yet.

  • 5

    Export and commit

    Download prompt.yaml and commit it next to the code that calls it; share the Markdown in the pull request.

The kit

Build a prompt spec

Everything stays in this browser. Keep several prompts in the library, and export each one when it is ready.

1Spec

Fields marked * are required before the prompt can ship.

2Prompt
Inputs found in the templateNo {{inputs}} in the template yet.
3Model and runtime
4Risk and approval
How easily can the change be reversed?
Standards mapped
5Evaluation gates

The playbook names what is checked; set the minimum scores your evaluation tool must return.

6Test cases

Expected behaviour for known inputs. At least one complete case is required.

#1

Review checklist

Tick an item only when it is true for this version.

01Versioned, reviewed and tested

Module 01 · Prompts as code; Wave 1 · assign an owner to every model.

01Automated evaluation gates

Module 01 · built on tools such as DeepEval and RAGAS, before anything merges.

04Continuous red-teaming

Module 04 · tried against the change before the pull request is approved. If a guardrail breaks, the build stops.

09The four-part verification gate

Module 09 · every AI-generated change passes it before commit.

03Least-Context Access Control

Not an agent, so the LCAC constraints do not apply.

07 · 08Cost and outage

Module 07 · AI FinOps; module 08 · Degraded mode.

09 · 11Approval and standards

Module 09 · the delegation matrix; module 11 · Standards alignment.

Export

# prompt.yaml · Prompts-as-Code Kit
# Cognitive XOps, module 01: prompts, system instructions and retrieval settings,
# versioned in Git, reviewed, and tested. www.fawzooz.ai
id: ""
name: ""
version: "0.1.0"
owner: ""
purpose: ""
model:
  primary: ""
  degraded_mode_fallback: ""
prompt:
  system: ""
  template: ""
  inputs: []
retrieval: ""
constraints: ""
agent:
  enabled: false
finops:
  token_budget: ""
  circuit_breaker: false
risk:
  reversibility: easy
  approver: ""
evaluation:
  # minimum scores for the automated evaluation gates
  faithfulness: null
  answer_relevance: null
  context_grounding: null
red_team: [prompt_injection, jailbreaks, context_leakage, hallucinated_dependencies]
tests: []
standards: []
review:
  reviewer: ""
  checks:
    code_git: false
    code_reviewed: false
    code_tested: false
    code_owner: false
    eval_faithful: false
    eval_relevant: false
    eval_grounded: false
    red_injection: false
    red_jailbreak: false
    red_leakage: false
    red_deps: false
    gate_facts: false
    gate_logic: false
    gate_missing: false
    gate_impact: false
    run_breaker: false
    run_fallback: false
    approve_delegation: false
    approve_mapped: false
  ready_to_ship: false
changelog:
  - version: "0.1.0"
    date: ""
    change: initial
    note: ""

Your prompts are stored only in this browser (localStorage). Nothing is sent anywhere.

Where the kit adds structure: the playbook prescribes the controls, not a file format. The spec fields, semantic versioning, the three reversibility bands and the YAML layout are this kit’s scaffolding around those controls.

Reference

The modules this kit puts to work

Every checklist item below comes from one of these modules, in the playbook’s own words.

  • Module 01

    Prompts as code

    We treat prompts, system instructions and retrieval settings exactly like source code: versioned in Git, reviewed, and tested. Before anything merges, automated evaluation gates built on tools such as DeepEval and RAGAS check that answers stay faithful to their sources, relevant to the question and grounded in the right context.

  • Module 03

    Zero Trust in the reasoning layer

    Read moreShow less

    Agents get the least context they need for the task in front of them, an approach we call Least-Context Access Control (LCAC). Five constraints apply to every agent: API keys with the least privilege possible; a hard ceiling on reasoning steps (for example, ten loops); predefined points where a human must approve before the agent continues; signed execution logs; and a kill switch that stops the agent instantly.

  • Module 04

    Continuous red-teaming

    Adversarial testing runs inside the build pipeline. Before a pull request is approved, automated scripts try prompt injection, jailbreaks, context leakage and hallucinated dependencies against the change. If a guardrail breaks, the build stops.

  • Module 07

    AI FinOps

    We track token use, inference cost and the return on semantic caching for each deployment. A financial circuit breaker halts the pipeline when a change causes a sudden spike in model calls, which protects against the Unbounded Consumption risk in the OWASP Top 10 for LLM Applications (2025).

  • Module 08

    Degraded mode

    Every team has a written plan for the day its main AI provider is down, slow or rate-limited. Work falls back to smaller local models, or moves to a human review queue, so delivery slows down rather than stops.

  • Module 09

    The developer’s daily playbook

    Developers register any new AI tool through an open intake register, which sorts it against a five-dimension risk taxonomy.

    Read moreShow less

    A delegation matrix decides who may approve what, based on how easily an action can be reversed. And every AI-generated change passes a four-part verification gate before commit: check the facts, validate the logic and security, look for what is missing, and assess the impact.

  • Module 11

    Standards alignment

    Each control maps to recognised standards: ISO/IEC 27001:2022 for information security, ISO/IEC 42001:2023 for AI management systems, ISO 45003:2021 as guidance on psychosocial risk, the NIST AI Risk Management Framework (AI RMF 1.0, 2023), and the EU Artificial Intelligence Act.

Source

Built on the applied playbook “Cognitive XOps” (2026): module 01, Prompts as code, with the controls of modules 03, 04, 07, 08, 09 and 11 and the Wave 1 roadmap rule that every model has an owner.

How to cite

Elgendi, M. F. (2026). Cognitive XOps: An applied playbook. www.fawzooz.ai