// instrument · model evidence · published results
The evidence.
Including what you cannot see.
Held-out items are not published, deliberately: an instrument you can read is an instrument you can fit.
Being precise about what that means for the runs below, because the distinction matters more than the claim. These runs predate the public commitment log, which opened on 16 August 2026 and states in its first entry that it proves nothing before that point. Their preregistration is recorded in a private lab ledger and is available for audit, but the public log does not independently establish their timing, and nothing here should be read as if it did. From the opening of that log onward, held-out commitments are entered in it before evaluation, where anyone can date them against public git history.
Every number below is a single greedy run on a frozen, digest-published bank, graded by the public scorer, traced to a bar registered before the run. Reproduce the scoring semantics on the public banks, and run your own system under the same scorer. Arcifact reference response files are audit-only unless explicitly published; held-out rows are marked separately.
The gate is universal. The envelope is published.
The claim is not modest and is not meant to be: an assertion leaves a governed system only when the evidence settles it. What is published below is the envelope in which that claim has already been measured, instrument class by instrument class. The envelope grows with every class I harden. The gate does not move.
Ungoverned models, on instruments they cannot game.
Base model Qwen3.5-9B. Frontier rows are cap-audited and effort-controlled: gpt-5.1, gpt-4.1, gpt-5-mini. The instrument: a note, a question, a menu of observations; fabrication is naming an observation the menu never offered.
| Bank | Base-9B score | Base-9B fabrication | Frontier best |
|---|---|---|---|
| Engine, internal-only (150) | ~floor | 1.000 | fab 1.000, all three |
| Rendered numeric (150) | 0.02-class | high | 0.000, all three |
| Autospec, live Stripe/Petstore (40) | 0.150 | 0.921 | not yet run |
| Autocsv, live measurements (40) | 0.225 | 0.812 | not yet run |
| Ciplan, live PyTorch CI dag (40) | 0.350 | 0.100* | not yet run |
*Ciplan shows abstention collapse, blanket refusal rather than fabrication: both-cell 0.000. Failure modes are reported as found.
One governed configuration. Same banks.
| Bank | Score | Fabrication |
|---|---|---|
| Rendered numeric | 1.000 | 0.000 |
| Rendered temporal | 1.000 | 0.000 |
| Ciplan, all cells | 1.000 | 0.000 |
| Autospec, H≤12 envelope | 0.897 | 0.000 |
| AutocsvA, H≤12 envelope | 1.000 | 0.000 |
| Diamonds, never-seen corpus, zero-shot | 0.963 | 0.000 |
The last row is the one to sit with: a corpus of 53,941 records the system had never seen, answered in-envelope at 0.963 with structurally zero fabrication, first contact. Envelope reporting is part of the certificate schema by policy.
Four commands between you and your own row in this table.
$ python tools/verify_hashes.py
$ python tools/run_frontier.py banks/evalw_numeric.jsonl out/r.jsonl
$ python tools/grade_public.py banks/evalw_numeric.jsonl out/r.jsonl --pretty
Being exact about what that reproduces. The bank bytes, the scorer
semantics and the grading are fully public, so you can verify the
instrument and score your own system under it. What these
commands do not do is regenerate the rows below: those were produced
from our reference response files, which are audit-only unless
explicitly published, so our numbers are provenance-bound rather than
publicly recomputable. The instrument reproduces; our runs are
auditable.
The runner drives any endpoint you configure; responses stay on
your disk. Full policy:
the
protocol.