PODJUDGE

Method

What this assessment is, how each score was reached, and the things it deliberately does not do.

Standing separately from the panel. The Pod Trials charter, signed 10 August 2026 before any code existed, locks a six-axis rubric scored by a three-person panel. That rubric is not re-weighted by anyone, including this tool. What you are reading is a separate, artifact-facing assessment: it scores what a repository, a submission document and an independent live probe can be made to show.

Scoring bands

Ten categories, ten points each. Every score picks a described band rather than a bare number, because a described band is far more consistent between subjects and between sittings.

ScoreBandMeaning
0–0absentNothing in the submission addresses this.
1–3tokenGestured at, but not really done. A stub, a heading, a promise.
4–6partialReal work, with real gaps. Usable but not finished.
7–8solidDoes the job properly. What a competent team ships.
9–10exemplaryBetter than the bar asked for. Worth copying into the playbook.

How evidence is gathered

StageWhat happens
IntakeThe repository is checked for the seven requested artifact classes and the submission document for its seven requested sections. Absences are reported with paths; nothing is inferred.
Deterministic probesClean clone, documented start command, test run. The result is what happened, not what was claimed.
Live probeAn independent Telegram account drives the POC as a first-time user would. Acceptance items C and F are refusal properties, so the probe actively tries to bypass approval and opt-out rather than asking whether they are supported.
JudgementA human reads the evidence and writes the score, the justification and the feedback together.

What this does not do

Models, tokens and spend

Two categories cover the AI story, and they ask different questions.

Model choice & sovereignty reads which models were named, where in the tree they appear, and whether the charter boundary held: product code and trial data touch self-hosted open-weight models only, with third-party APIs confined to the platform layer. Newest is not best. A pod that ran a mid-tier open model well scores above one that reached for a frontier API it was not permitted to use.

Token & spend efficiency reads what the EUR 1,500 bought. Underspending is not virtue and overspending is not vice — the charter calls an exhausted budget a workflow finding, not a failure. What is scored is whether the pod knew where the money and the tokens went and can show it.

Model tiers are a snapshot of judgement at the assessment date, not a measurement, and they date quickly. Where a model is named in a product path that the charter reserves for open weights, the tool records a question — a mention is not proof of use, and a human rules on each one.

Entrants outside the charter

The charter names two pods and defines the audit exchange as a two-way swap between them. A third entrant is scored on the same ten categories, but it was not run by a counterpart pod. That difference is marked on the submission rather than adjusted for, because an unaudited build and an audited one are not the same evidence, and quietly averaging them would hide it.