Ten categories, ten points each, scored against what the submitted repository, the submission document and a live probe actually show. Every score carries a written justification and a note on what would make it better.
| Category | Pod of Conviction | Ratel Works Pod | Prove-It Pod | Max |
|---|---|---|---|---|
| Acceptance coverage Do items A–F actually work, on demand, without the pod driving? | 6 | 7 | 4 | 9 |
| Architecture & technology Are the technology choices ones you would defend in six months? | 7 | 7 | 6 | 8 |
| Code quality Could a new engineer change this safely next week? | 6 | 6 | 6 | 8 |
| Testing & verification Do the tests prove the acceptance items, or just exercise code? | 6 | 7 | 5 | 8 |
| Deployability & operations Clean clone to running system, following only the written steps. | 7 | 5 | 6 | 8 |
| Documentation & handover Could HAUD take this over without the pod in the room? | 7 | 8 | 7 | 8 |
| AI usage & prompt craft The A/B payload — how the pod actually worked with models. | 8 | 8 | 7 | 9 |
| Process & decision evidence Decision log and daily journal — kept as work, not written on the 30th. | 7 | 6 | 7 | 8 |
| Candour & self-assessment The requested LLM assessment, and whether it tells the truth. | 7 | 8 | 7 | 9 |
| Model choice & sovereignty Which models, were they the right ones, and did they respect the boundary? | 8 | 7 | 7 | 8 |
| Token & spend efficiency What the budget bought — EUR 1,500 per pod, hard, every euro logged. | 8 | 6 | 7 | 9 |
| Security & data handling Consent, opt-out, secrets and the sovereignty boundary. | 4 | 7 | 2 | 8 |
| Total | 81 | 82 | 71 | 100 |
| Pod | Tokens reported | Spend reported | Cap |
|---|---|---|---|
| The Pod of Conviction | 1,009,989,029 | $1,446 ≈ €1,242 | €1,500 |
| The Ratel Works Pod | 13,620,898,737 | not reported | €1,500 |
| The Prove-It Pod | 1,114,540,355 | $1,158 | €1,500 |
Each pod's own figures, from its own records — not measured by this assessment, and not comparable like for like: the pods counted different things. Per-pod pages give the source and the caveats.
Read the column, not the ranking. One pod counts cache reads, which dominate an agent workload and inflate a total by roughly an order of magnitude against a pod counting input and output alone. One reports GLM-5.2 only. One excludes its DevOps AI usage by its own note. These are three different measurements wearing the same unit.
| Pod | Hours stated | Commits | Active days | In the 3-h window |
|---|---|---|---|---|
| The Pod of Conviction | not reported | 458 | 14 | 21.0% |
| The Ratel Works Pod | not reported | 926 | 7 | 33.2% |
| The Prove-It Pod | 299 h | no history | — | — |
Charter §2: pod work is confined to a shared three-hour window per working day — 09:00-12:00 CEST — with automation excepted and humans not. 'The cap is the experiment's control: up to forty-five pod-hours per person, identical for both pods by construction.' Work outside the window was to be visible, and time lost to business-as-usual logged as an impediment.
One pod reported hours; two did not. The pod that reported them is the only one whose figures can be questioned — which is a consequence of disclosure, not of working longer. A commit is not an hour, and no score is derived from any number in this table.
| Pod | Score | Assessment |
|---|---|---|
| The Pod of Conviction | 5 / 10 | The strongest individual prompt documents sit next to the weakest evidence trail. The in-repo artifacts are genuinely good: prompt.md and APP_GENERATION_PROMPT.md carry hard constraints (exact stack, 'Do not us… |
| The Ratel Works Pod | 8 / 10 | Few prompts, but the best-crafted ones in the trial. The four briefs in docs/submission/prompts/ are prompts-as-work-orders with a specificity level no other pod approaches: GLM-BRIEF cites defects by file and … |
| The Prove-It Pod | 8 / 10 | The most complete prompt record of the trial, and most of it is verbatim. Five workstreams each left artifacts in their own working style: the FE series is day-numbered (Day 1-12) with real API schemas, polling… |
Scored separately from the twelve categories. The trial's question was how each pod worked with models, and the prompts are the record of that. Volume was not rewarded — the largest library did not score highest.