What this assessment is, how each score was reached, and the things it deliberately does not do.
Ten categories, ten points each. Every score picks a described band rather than a bare number, because a described band is far more consistent between subjects and between sittings.
| Score | Band | Meaning |
|---|---|---|
| 0–0 | absent | Nothing in the submission addresses this. |
| 1–3 | token | Gestured at, but not really done. A stub, a heading, a promise. |
| 4–6 | partial | Real work, with real gaps. Usable but not finished. |
| 7–8 | solid | Does the job properly. What a competent team ships. |
| 9–10 | exemplary | Better than the bar asked for. Worth copying into the playbook. |
| Stage | What happens |
|---|---|
| Intake | The repository is checked for the seven requested artifact classes and the submission document for its seven requested sections. Absences are reported with paths; nothing is inferred. |
| Deterministic probes | Clean clone, documented start command, test run. The result is what happened, not what was claimed. |
| Live probe | An independent Telegram account drives the POC as a first-time user would. Acceptance items C and F are refusal properties, so the probe actively tries to bypass approval and opt-out rather than asking whether they are supported. |
| Judgement | A human reads the evidence and writes the score, the justification and the feedback together. |
Two categories cover the AI story, and they ask different questions.
Model choice & sovereignty reads which models were named, where in the tree they appear, and whether the charter boundary held: product code and trial data touch self-hosted open-weight models only, with third-party APIs confined to the platform layer. Newest is not best. A pod that ran a mid-tier open model well scores above one that reached for a frontier API it was not permitted to use.
Token & spend efficiency reads what the EUR 1,500 bought. Underspending is not virtue and overspending is not vice — the charter calls an exhausted budget a workflow finding, not a failure. What is scored is whether the pod knew where the money and the tokens went and can show it.
Model tiers are a snapshot of judgement at the assessment date, not a measurement, and they date quickly. Where a model is named in a product path that the charter reserves for open weights, the tool records a question — a mention is not proof of use, and a human rules on each one.
The charter names two pods and defines the audit exchange as a two-way swap between them. A third entrant is scored on the same ten categories, but it was not run by a counterpart pod. That difference is marked on the submission rather than adjusted for, because an unaudited build and an audited one are not the same evidence, and quietly averaging them would hide it.