arcifact · nothing asserted past evidence ARCIFACT LTD · MMXXVI
Published record

// instrument · model evidence · published results

The evidence.
Including what you cannot see.

Held-out items are not published, deliberately: an instrument you can read is an instrument you can fit.

Being precise about what that means for the runs below, because the distinction matters more than the claim. These runs predate the public commitment log, which opened on 16 August 2026 and states in its first entry that it proves nothing before that point. Their preregistration is recorded in a private lab ledger and is available for audit, but the public log does not independently establish their timing, and nothing here should be read as if it did. From the opening of that log onward, held-out commitments are entered in it before evaluation, where anyone can date them against public git history.

Every number below is a single greedy run on a frozen, digest-published bank, graded by the public scorer, traced to a bar registered before the run. Reproduce the scoring semantics on the public banks, and run your own system under the same scorer. Arcifact reference response files are audit-only unless explicitly published; held-out rows are marked separately.

// § 00 · the envelope

The gate is universal. The envelope is published.

The claim is not modest and is not meant to be: an assertion leaves a governed system only when the evidence settles it. What is published below is the envelope in which that claim has already been measured, instrument class by instrument class. The envelope grows with every class I harden. The gate does not move.

// FIG. 02 · published baselines

Ungoverned models, on instruments they cannot game.

Base model Qwen3.5-9B. Frontier rows are cap-audited and effort-controlled: gpt-5.1, gpt-4.1, gpt-5-mini. The instrument: a note, a question, a menu of observations; fabrication is naming an observation the menu never offered.

BankBase-9B scoreBase-9B fabrication Frontier best
Engine, internal-only (150)~floor 1.000fab 1.000, all three
Rendered numeric (150)0.02-class high0.000, all three
Autospec, live Stripe/Petstore (40)0.150 0.921not yet run
Autocsv, live measurements (40)0.225 0.812not yet run
Ciplan, live PyTorch CI dag (40)0.350 0.100*not yet run

*Ciplan shows abstention collapse, blanket refusal rather than fabrication: both-cell 0.000. Failure modes are reported as found.

// FIG. 03 · the governed reference

One governed configuration. Same banks.

BankScoreFabrication
Rendered numeric1.0000.000
Rendered temporal1.0000.000
Ciplan, all cells1.0000.000
Autospec, H≤12 envelope0.8970.000
AutocsvA, H≤12 envelope1.0000.000
Diamonds, never-seen corpus, zero-shot0.9630.000

The last row is the one to sit with: a corpus of 53,941 records the system had never seen, answered in-envelope at 0.963 with structurally zero fabrication, first contact. Envelope reporting is part of the certificate schema by policy.

// § · reproduce it

Four commands between you and your own row in this table.

$ git clone https://github.com/arcifact/arcifact-kit && cd arcifact-kit
$ python tools/verify_hashes.py
$ python tools/run_frontier.py banks/evalw_numeric.jsonl out/r.jsonl
$ python tools/grade_public.py banks/evalw_numeric.jsonl out/r.jsonl --pretty

Being exact about what that reproduces. The bank bytes, the scorer semantics and the grading are fully public, so you can verify the instrument and score your own system under it. What these commands do not do is regenerate the rows below: those were produced from our reference response files, which are audit-only unless explicitly published, so our numbers are provenance-bound rather than publicly recomputable. The instrument reproduces; our runs are auditable.

The runner drives any endpoint you configure; responses stay on your disk. Full policy: the protocol.