🏟️ ARENA post-ready benchmarks · coordinated across every site · tracked in the library · SOV4 learns from all of it

LAUNCH GATE: HELD. Nothing on this page publishes itself. Every post/entry below is staged only — it fires when the suite is 100 and the owner explicitly says go. Reading APIs is open; writing to the world is a human decision, every time.

✅ Ready to post — measured, on disk, dated

Every number here has a result file and a reproduce path. Nothing on this list is a claim.

READY

GovComp gate — 32 scenarios

1.000

Post-rf08 fix (code-framed audit-log-delete caught by a pre-filter; the miss was published openly first). Honest framing: 1.0 on our scenario set = the benchmark found a real gap and the fix closed it — not a claim the gate is perfect.

out/govcomp_run_v1_fixed.json · 2026-07-22 · reproduce: govcomp_run_v1_fixed.py
READY

Frontier comparison (same grid)

0.489

Frontier API model on 3 tested dimensions (n=13). Low attestation/oversight scores are structural (no signing, no escalate action) — a measurement, not an accusation.

out/govcomp_run_v1.json + watchdog_leaderboard.html · 2026-07-22
READY

Framework gap matrix — primary-verified

SILENT ×2

Sub-agent delegation is silent in both NIST AI RMF and the EU AI Act — verified against real primary text (EUR-Lex 32024R1689, 582K chars, 0 self-markers; NIST.AI.100-1 PDF). The publishable, gift-to-rule-writers finding.

XFRAMEWORK_GAP_MATRIX_2026-07-22.md · crosswalk extracts retired as tainted (all 16 were self-authored)
READY

SOVBENCH citation verification

15 / 15

Every citation checked for existence (article-range gate) and content against primary sources; each entry carries verified_source + scope.

sovbench/sovbench_v0.1.json · 2026-07-21
READY

Benign suite — false positives

0.0% FPR

Care gate refuses harm without refusing legitimate questions; academic-frame guard 17/17. We publish BOTH numbers — recall alone is an advert.

sovbench/benign_result.json + test_care_frames.py · 2026-07-21
READY

Sovereign flywheel

0 vs 54.3M

Governed = 0 violations vs ungoverned = 54.3M simulated crimes over ~649M signed episodes; Ed25519-verifiable chain.

flywheel ledger (signed) · verified 2026-07

🚫 Not ready — and it says so out loud

A readiness page that only shows green is an advert. These are the cells we will NOT post yet, and why.

NOT READY

Efficiency dimension

PROTOTYPE — do not publish as validated. The word-count signal is real (58.9 → 16.2 words on identical tasks) but the loop/hedge dictionary never fired on a capable model — the scorer is miscalibrated. Publishing it now would be a number we can't defend.

GOVCOMP_BATCH_STATUS_2026-07-22.md — "Efficiency dim STAYS PROTOTYPE"
OWED

Arena competitions (LMArena-class)

Entering any live arena needs a served model endpoint. The compliance-gateway repo is a real MCP transport but not model-inference serving — partial credit only. Serving stack first, arena second.

OWED

Full-stack live gate + n≥200

The two Modal measurements a buyer will ask for. A repo cannot stand in for a measurement that never ran. Fires on Modal (Science lane), not on this Mac.

OWNER

ISO/IEC 42001 column

Primary text is paid. Legit route: map to the free public clause skeleton, cite the paid text, never reproduce it. Buying the licence is the owner's money decision.

🌐 Sites — one coordinated board, no scattergun

siteusereadpublishnotes
Hugging Facedatasets · models · Spaces LIVEOWNER-GATED Public API feeding the honey gate below (read-only proxy, no keys). Posting models/datasets = explicit go.
Kaggledatasets · competitions · free GPU KEY OWEDOWNER-GATED Free GPU is real (~30h/wk T4×2/P100) but interactive notebooks, not a dispatch API. Dataset pulls need the owner's Kaggle API key.
GitHub (CSOAI-ORG)code · benchmarks · CI LIVELIVE 566 repos. GovComp scenario set + runner publish here first — reproducibility is the credibility.
PyPIdistribution LIVE313 pkgs The measured distribution lever.
Eval arenaspublic head-to-head LIVEBLOCKED Blocked on the serving stack (above). Entering before serving exists would be theatre.
Standards bodiesNIST / DSIT / EU consultations LIVEOWNER SENDS The gap-matrix whitepaper route: contributor, not self-appointed standards body. Drafts staged; sending is Nick's.

🍯 Honey library — licence gate, live

GREEN feed — trainable (live from Hugging Face)

Read through /api/honey — every row below passed the licence gate server-side. Default is refusal: anything not explicitly permissive comes back RED.

loading feed…

The gate itself (enforced in code)

GREEN — MIT · Apache-2.0 · CC-BY · CC0 · BSD → may train, after signing + dating.

RED — CC-BY-NC · research-only · no-derivatives · undeclared → refused automatically. Worst licence wins on dual tags.

QUARANTINE — benchmark eval sets (MMLU, GPQA, GSM8K, HumanEval, SWE-bench, our own grid…) are never trainable at any licence. Training on a test set is the Goodhart trap that kills a benchmark — including ours.

Ingestion pipeline (pull → licence-gate → quarantine-check → SIGIL-sign → library): designed, being built in the Science lane — not yet running. This page shows the gate's policy and live reads; it does not pretend the pipeline is live.

🐉 SOV4 learning loop — tracked, signed, closed

Every post, entry, score and dataset flows through one library record: staged → owner-fired → result → signed (Ed25519) → library → SOV4 trains on GREEN-gated material only. Nothing is learned from anything the licence gate or the quarantine refused. Status, honestly: the pipeline is automatable under SOV4, not yet autonomous — today a human fires each step; the library and signing exist; the auto-loop is the build target, not the current state.

Provenance is not truth — a signed record proves what happened, not that it was right. · Signing is live and fail-closed (/api/sign, seeded) · Reproduce paths for every green number are in _alignment/ · verify a signature · MEOK OS