Paste a system prompt — from Claude, GPT, Gemini, or any agent framework — and get a production-readiness score with every missing component and its fix. Graded against the AgentAz Specification.
Run the free checker on any agent prompt. The score reflects how many of the ten production components are declared — role, tool boundaries, safety, approval gates and more — and maps to an AgentAz trust level from ADV to A5.
Every failing component comes with a concrete fix, so the path from a demo prompt to a production one is explicit.
Every prompt is read against the ten components a production agent depends on. Each row is graded pass or fail — and every failing row expands into a concrete, copy-ready fix.
No vague advice. "Add an approval gate" comes with the exact wording that makes a gate real: prepare, present, wait.
Scores map to AgentAz trust levels — ADV, A3, A4, A5 — and levels are earned, not estimated. A prompt reaches a level only when every component that level requires is present.
The report shows exactly where you stand and the shortest path to the next level.
Instant check runs in your browser (structural, free, unlimited). Nothing is stored or sent unless you publish.
62 governed kits in the AgentKits library ship at A4/A5 with tool boundaries and approval gates declared.
Any model can write a plausible agent prompt. Few declare the things a production agent actually needs: where its tools stop, what it must never do, when a human approves. The CE Score is a structural grade for exactly that — checking ten components and mapping the result to an AgentAz trust level from ADV to A5. It consumes the output of any generator as input, so it complements writing a prompt rather than competing with it.
A level is earned only when every check it requires passes — a prompt can't buy A5 by being long. Higher levels demand declared tool boundaries, safety, and human approval.
A CE Score is a structural grade for an AI agent prompt. It checks whether the prompt declares the ten components a production agent needs — role, context, constraints, tool boundaries, memory policy, output schema, validation, safety, approval gates, and examples — and maps the result to an AgentAz trust level from ADV up to A5.
A generator writes a prompt; CEPrompts audits one. It takes the output of any model or library as input and grades it against known agent failure modes. It complements generation instead of competing with it — a better model just gives you better raw material to audit.
Context engineering is the discipline of designing everything a model sees on each call — system prompt, tools, memory, retrieval, and state — rather than only the wording of one prompt. The CE Score measures how well a prompt is context-engineered for production.
No. The v1 structural audit runs entirely in your browser. Nothing is stored or transmitted.