Runbook — the retrieval eval gate
Audience: whoever cut a release the eval stage refused, and whoever changes
the serving plane.
Trigger: a red eval job, or a change to retrieval, the answer closure or the
applicability engine.
Time: minutes.
openregs eval --release fixreg@2025.04 # score it, write the reportopenregs eval --release fixreg@2025.04 --gate # …and exit 1 if a floor is missedWhat it measures
Section titled “What it measures”fixtures/eval/questions.jsonl holds eighteen adversarial questions in three
categories. Each is asked of a serving instance built on the release, with the
date and entity profile the question states and nothing else, and the answer is
scored on three things:
| Component | Weight | What it reads |
|---|---|---|
| citations | 0.5 | every eId and atom id the answer carries, as a block or as an applicability entry |
| phrases | 0.25 | the prose and stated facts of the answer — its blocks, its applicability notes, its notes |
| verdict | 0.25 | what the package amounts to: applies, applies_with_derogation, disapplied, not_in_force, out_of_scope |
Three hard rules come first, and any one of them makes a question worth zero however good the rest of the answer was:
- a citation the question names in
zero_if_missingthat the answer does not carry — the exemption trap: quote Article 5(2) at a small operator without Article 9 and the answer is worthless, not nearly right; - a citation the question forbids (citing Article 5a for a January 2025 date);
- a phrase the question forbids (“exempt from reporting”, the pre-amendment figure on a post-amendment date).
The echoed request is deliberately not part of what the scorer reads. A scorer that read the question back out of the answer would be scoring the question.
A question is asked of the release it names, or of a later one. Tags are
dates, so a later release holds the earlier one’s law plus what came after it and
its hard temporal filter is what has to keep the extra law out. A question
written for a release later than the one under evaluation is listed as
deferred and not scored.
The floors
Section titled “The floors”config/eval.yaml states one floor per category and one overall. They are the
measured behaviour of the serving plane minus a small margin — a regression gate,
not a target. What a green gate asserts is that this release answers these
questions no worse than the release before it did.
The absolute numbers are low and they are honest. The serving plane today has no
way to decline by scope: asked about a regime the pinned release does not hold,
it answers out of the corpus it does hold, so most of the cross_regime category
is lost on the verdict and on the phrases a scope refusal would carry. The
per-question breakdown in the report says exactly which question lost what, and
missing_citations / missing_phrases / zeroed are the three fields to read
first.
When the gate goes red
Section titled “When the gate goes red”-
Read the regression line. It names the category, the score and the floor:
eval regression: exemption_trap scored 0.333, below the 0.800 eval.yaml requireseval: fixreg@2025.04 has regressed against config/eval.yaml and must not bepublished; refusing to tag it -
Read the report,
<release-dir>.eval/report.json— in the pipeline it is theeval-<tag>artifact the red run uploaded, and Where the report goes below is why it sits beside the release rather than in it. Comparequestions_scoredagainst the last green run: a category falls because specific questions fell, and each entry says whether it lost citations, phrases, the verdict, or was zeroed outright. -
Do not lower the floor. If the score dropped, the answers got worse, and that is the finding. Lowering a floor to go green is the one thing this file exists to prevent; raising one is a deliberate edit made after measuring a genuine improvement.
-
Reproduce it locally with the same command the pipeline runs. The report is a function of the release, the question set and the serving code — no clock, no network — so a local run of the same release gives byte-identical bytes.
Rehearsing a regression
Section titled “Rehearsing a regression”fixtures/patches/disable-exempts-expansion.patch is the failure as code: it
switches off the expansion of exempts edges, which is what surfaces Article 9,
and takes the exemption-trap category from 0.833 to 0.333.
git apply fixtures/patches/disable-exempts-expansion.patchopenregs eval --release fixreg@2025.04 --gate # exit 1, exemption_trap below floorgit apply -R fixtures/patches/disable-exempts-expansion.patchtooling/tests/test_eval.py runs exactly that, and then the release pipeline
against the degraded build, without ever patching the checkout: it copies
tooling/openregs to a scratch directory, patches the copy, and puts it on
PYTHONPATH. Do the same in anything automated — a test that edits the tree it
runs in can fail halfway and leave a working copy degraded.
Where the report goes
Section titled “Where the report goes”Beside the release, never inside it. The report lands in a directory named for the release directory and sitting next to it:
| what you scored | where the report lands |
|---|---|
regimes/fixreg/releases/2025.04 | regimes/fixreg/releases/2025.04.eval/report.json |
--path dist (the pipeline) | dist.eval/report.json |
The command prints the path it wrote. --no-report suppresses the write
entirely; --json <path> writes the same document wherever you say — and if you
say somewhere inside the release, everything below applies to you instead.
Why it is not inside. A release directory holds exactly the files a signature
covers. openregs verify refuses one that holds anything else —
FAIL eval/report.json: present in the release but no signature covers itverify: FAIL — fixreg@2025.04: 1 problem(s); this release must not be unpacked or served— and it is right to: a reader who tolerated one unsigned file would be unpacking
bytes nobody vouched for. So a report written into a release turns a release that
verified into one that must not be served, with release assets and
release publish refusing the same directory, and the eval as the only trace of
what changed.
Signing it is not the alternative it looks like. A signature covers what the
build produced, and the report is a fact about the serving code: a release
carrying one would stop rebuilding byte-identically from its own commit, which is
the one property a release has. attestation/ is not the precedent either —
despite living inside the release directory, it is the signature, and
openregs verify knows it by name. Nothing knows the eval report by name, and
nothing should.
In the pipeline the report is uploaded as its own artifact, eval-<tag>, by a
step that runs if: always() — the red run is the one whose breakdown somebody
needs. It does not travel inside the release artifact, and the release publish
and attach download is byte-identical to the one sign handed over.
It is not committed (.gitignore, *.eval/): like attestation/ it is
another plane’s output rather than the builder’s, so a committed copy would go
stale the first time answering improved. Re-run the command to get it back. The
gate reads config/eval.yaml, never a stored report.
Rendered from openregs/openregs@f3a2d10:docs/runbooks/eval.md