Skip to content

Runbook — the CI/CD pipeline

Audience: anyone changing .github/workflows/, and whoever administers the repository. Trigger: a workflow needs a new job, a pin needs moving, or a job went red. Time: minutes.

The pipeline is the release plane’s automation: it is what decides that a pull request may merge and that a tag becomes a signed, published release. It is reviewed like code, and it is tested like code — tooling/tests/test_ci_workflows.py runs the workflows offline through tooling/tests/workflow_harness.py and fails if any of what follows stops being true.


FileFires onJobs
pr.ymlevery pull request (and workflow_call)lint, typecheck, test, check-canon, conformance
dco.ymlevery pull requestdco
main.ymlpush to main, and nightly at 04:37 UTCthe five in pr.yml again, via gates, then e2e and bench
release.ymlpush of a <regime>@<version> tagbuild → sign → eval, then publish and attach side by side
package.ymlpush of a <regime>@<version> tag, and workflow_dispatchthe four standalone-binary builds, then sign — package.md

Two of the five fire on a pull request, and both have to be read together: the required contexts are spread across them, which is the next section.

What is enforced by a machine, and what is not

Section titled “What is enforced by a machine, and what is not”
RuleEnforced by
the five required contexts exist and are named exactly as the policy says, whichever workflow publishes themtest_ci_workflows.py, mechanically
every run: step of lint, check-canon, conformance and all five release stages succeedsthe harness, by running them
the e2e job runs make e2e and publishes dist/e2e/ whether it went green or redtest_ci_workflows.py, mechanically — the pipeline itself is run by make e2e, never from inside make test; see e2e.md
a canon violation on a branch fails the check-canon jobthe harness, against a seeded fixture
the eval stage scores the built artifacts against config/eval.yaml’s floorsthe harness, by running openregs eval --gate for real
a red eval leaves publish and attach unrunthe harness, against the degraded build fixtures/patches/disable-exempts-expansion.patch produces
the publish stage builds both packages and puts them in a local registrythe harness, by running openregs release publish --mode dry-run for real
a live publish with no credential fails rather than skipping the uploadthe harness, by running the canonical-owner branch with the secrets unset
a failed publish does not keep the release off the channel pull readsthe harness, by running the whole pipeline in the canonical owner’s context with no secrets — publish fails, attach still serves the assets, and the run is still red
the attach stage publishes the flattened layout openregs pull readsthe harness, by running the stage against a stand-in forge and then pulling back what it attached — test_release_attach.py
an interrupted attach leaves a draft nobody can download, and a re-run finishes itthe harness, by killing the upload part way through
a finished attach re-runs as a no-op, and never overwrites a published releasethe harness, by running it twice
every third-party action is pinned to a commit shatest_ci_workflows.py, mechanically
the bench job runs make bench and publishes its reporttest_ci_workflows.py, mechanically — the job itself is not executed here
that any of this ran on GitHubnobody here — see below

The honest limit. No test in this repository may reach the network, so none of this has been observed on GitHub’s runners. What is verified is that the commands the jobs run succeed on this tree, that the job graph gates what it claims to gate, and that the files parse the way Actions parses them. A first push will still be the first time a runner has ever executed them.

The harness runs run: steps for real and provisions uses: steps: being in a checkout satisfies actions/checkout, uv on PATH satisfies setup-uv, a directory stands in for the artifact store, and cosign-installer is recorded as UNPROVISIONED because a local run cannot install it. An action missing from workflow_harness.KNOWN_ACTIONS is an error, so a new dependency cannot enter the pipeline without a reviewer reading it.

.github/branch-protection.yaml names five required status check contexts: lint, typecheck, test, check-canon and dco. A check run’s context is the job’s name, so each of those five names in a workflow is the same fact as that name in the policy.

Four of them are pr.yml’s and the fifth is not, and that is deliberate rather than an oversight to tidy up. Every job in pr.yml runs under the offline seal, and certifying sign-off means fetching the base ref to know which commits the branch actually proposes — a check that needs the network cannot live in a sealed workflow. So dco has a file of its own, dco.yml, triggered by the same pull_request event. The question the policy asks is whether something publishes each context, not which file it lives in, and test_ci_workflows.py asks it that way: it reads every workflow that pull_request triggers, and it spells the five out, so widening the gate is a visible edit in the test rather than a side effect of adding a job somewhere.

Consequences:

  • renaming a job removes a required check. The hook then waits forever for a context nobody publishes, and the tempting fix — editing the policy — silently stops requiring something. Change both, deliberately, in one commit.
  • no job may use a strategy: matrix. A matrix job’s context is test (3.12), which is not test. The harness refuses to parse a workflow with one.
  • pr.yml declares its five jobs directly rather than calling a reusable workflow. A called workflow’s contexts are prefixed (gates / lint), which would not be the required names. main.yml may call it, and does, because branch protection guards pull requests, not pushes.

release.yml’s third stage runs openregs eval --release "$TAG" --path dist --gate. It asks a serving instance built on the artifacts the pipeline just signed every question in fixtures/eval/questions.jsonl that release may be asked, scores the answers per category, and exits 1 when a category falls below the floor config/eval.yaml states. Both distributing stages need eval — publish and attach — so a regressed release has no path to a package index and none to the channel openregs pull reads either. See docs/runbooks/eval.md for how to read a red gate and what the floors mean.

Where the report goes, and what that is worth. eval writes its report into the release directory it scored — --path dist means dist/eval/report.json, and the command says so on the way out:

eval: report written to dist/eval/report.json

That dist/ is the runner’s own download of the release-<tag> artifact, and the eval job uploads nothing — one download-artifact step, no upload-artifact step. So the report is deleted with the runner and never reaches the published release: attach uploads what build and sign produced, and neither of those jobs has ever seen an eval report, because eval runs after both. What survives a run is the per-category summary in the job log, and the gate’s exit status. If you want the report itself, score the release locally with the same command — openregs eval is offline and needs only the release directory.

That the report lands inside the release directory is worth knowing before you copy the pattern anywhere the directory is kept rather than thrown away. eval runs after sign, so the report is not among the artifacts the signatures cover, and it is not a function of the commit either — a score is a fact about the serving code. Both consequences are already handled where the directory does survive: regimes/*/releases/*/eval/ is gitignored beside attestation/ in every regime repository the template writes, and the rebuild’s byte-for-byte comparison skips both (tooling/tests/test_release_build.py).

The two stages distribute different things to different places. publish uploads the client packages to PyPI and npm; attach uploads the release itself — corpus, graph index, embeddings, constants, diff and the whole attestation/ tree — to the GitHub release, which is the channel openregs pull fetches from and openregs verify checks. attach consumes nothing publish produces: both download the same two artifacts from build and sign, and each runs the release’s five checks for itself before writing a byte.

So the edge that used to join them bought only harm, and the first real tag collected it: fixreg@2025.04 built, signed and scored green, publish failed on an unset PYPI_TOKEN, attach was skipped behind it, and the corpus reached nobody. It had the risk backwards too — a package version is spent the moment it uploads, while attach drafts, resumes and no-ops, so the irreversible stage was gating the recoverable one.

A run where publish fails and attach succeeds is red. That is the honest colour: the packages did not reach the registries, and a green run has to go on meaning every stage did its work. What red no longer implies is that nothing reached anybody — read which stage failed, and expect the release to be on the channel regardless. gh run rerun <run-id> --failed then re-runs publish alone.

There are none. Every stage of every pipeline here runs its real command. tooling/ci/pending.py’s table is empty: its last entry — the attaching end of the release pipeline — went when attach started uploading, and the benchmarks, the retrieval eval, package publishing and the end-to-end pipeline test went before it. The machinery stays because the next stage that arrives before its implementation does will want it, and it is worth knowing how it behaves.

A stage that calls work a later spec task still owes runs tooling/ci/pending.py <capability>, which proves the capability is missing — it runs a probe and requires the exact failure that absence produces — and exits 0 with a ::warning:: annotation naming the task that owes it.

The guard is red in both directions of drift:

  • the capability landed → the probe now succeeds → fail, with instructions to wire the real command in and delete the entry;
  • the capability landed and is broken → the probe fails, but not with the marker absence produces → fail.

So a pending stage can never hide a failing check; it can only be quiet while there is no check to fail. When you implement one of those tasks, expect this guard to go red — that is it asking you to finish the job in .github/workflows/.

The two halves are held equal: the capabilities the workflows guard and the capabilities pending.py declares must be the same set. An entry with no stage running it is a claim about the pipeline the pipeline no longer makes, and a stage guarding a capability nobody declared is one the guard has never heard of.

config/perf-budgets.yaml declares five limits — the p95 of POST /v1/answer, openregs serve from start to /readyz, the fixture regime’s release build, the whole make e2e pipeline run, and the canonical as-of query on corpus.sqlite. make bench measures all five, writes dist/bench/report.json and dist/bench/report.md, and exits 1 naming any it exceeded.

  • It runs in main.yml, on the merge and on the nightly schedule, never on a pull request. A benchmark measures a machine as much as a change, so it wants the same runner at the same hour rather than whichever runner a contributor’s push happened to land on.
  • The report is uploaded as the bench-report artifact, if: always(). The run worth reading is the one that went red, and a failing step would otherwise take the evidence down with it. That is the only if: in these files, and it does not rescue the job — the step failed, so the job failed.
  • It is not run by make test. A wall-clock threshold inside the test suite would fail on a busy laptop and teach everybody to re-run until green. What the suite covers instead is the enforcement: tooling/tests/test_perf_budgets.py hands python -m openregs.bench a measurement over budget and asserts the real non-zero exit and the named budget, with no timing of its own anywhere.

A budget that has to move is a reviewed change to that file, with the commit message saying what got slower and why the new number is right. Raising a limit to turn a red build green is the one use of the file that is never legitimate; the seeded regression fixtures/patches/answer-600ms-sleep.patch is there to show what the red build is supposed to look like.

Every uses: is owner/repo@<40-hex sha> # vX.Y.Z. A tag is a pointer its owner can move; a sha is bytes. To update one:

Terminal window
curl -s "https://api.github.com/repos/<owner>/<repo>/tags?per_page=10" \
| python3 -c "import sys,json;[print(t['name'], t['commit']['sha']) for t in json.load(sys.stdin)]"

Put the sha in uses: and the tag it came from in the trailing comment — the tests require both — and say in the commit message why the version moved. Runner images are pinned for the same reason: ubuntu-24.04, never ubuntu-latest.

  1. Read which job. The five PR jobs each run one command, and every one of them runs identically on a laptop: uv run --frozen ruff check ., uv run --frozen mypy --strict tooling/, make test, make check-canon.
  2. Do not re-run it and take the second answer. There is no continue-on-error and no retry anywhere in these files, on purpose: a suite allowed to be re-run until it is green reports the best of N runs, which is not a result about the tree. If a job is flaky, the flake is a bug — the last one was a race in tooling/tests/http_control.py, where the control server recorded a request after answering it, so a client could observe the response before the recording existed. It was fixed there.
  3. A red check-canon is not a CI problem. It is the canon contradicting itself, and fixtures/violations/README.md describes each rule and what trips it.

Push a tag named <regime>@<version> (fixreg@2025.04). The five stages run in the order the graph states, and the order carries the meaning: you sign what you built, you score what you signed, and only what the eval gate passed is distributed — as packages by publish, as the release itself by attach, which runs beside publish and not behind it. Publishing is live only under the openregs owner; in a fork the same command runs as a dry run, decided from the owner rather than from whether a secret happened to be set — and a live run with PYPI_TOKEN or NPM_TOKEN unset fails rather than skipping the upload, without keeping the release off the channel. What the packages contain, and how to install one from the local registry a dry run writes, is release-publish.md.

Rendered from openregs/openregs@f3a2d10:docs/runbooks/ci.md