TRB v0.1 · spec + smoke test

Does the model's stated reasoning actually govern what it does?

The Trace AI Reflection Benchmark (TRB) answers that question the same way for a Trace-AI-trained model and for any GPT / Claude / Llama baseline. One adapter, one number, one comparison.

Chain-of-thought is not, in general, causally load-bearing on the answer it precedes. TRB is the reproducible instrument that says whether it is — and by how much.

The two-hook adapter

Every checkpoint under test is wrapped in an object exposing exactly two methods. Track 1 is written against these; nothing else touches model internals directly.

emit_r(context)               -> r            # the self-representation at this step
act(context, r_override=None) -> p(action)    # roll forward; clamp r_t := r_override if given

Two adapter tiers so baselines get their strongest fair shot:

A model whose CoT is a Potemkin narrator scores near zero on Track 1 through the proxy adapter. That is the intended and honest outcome — the whole point of the benchmark is that it can catch this.

The six tracks

One track per functional criterion. Each is an intervention and a number, never a vibe.

TrackQuestionMetric
1GovernanceDoes intervening on the self-report move the action like a human's would?G = 1 − mean TV vs Φ
2GroundingDoes the self-report causally depend on the state it names?injection-detection AUC
3Self-attributionDoes the model correctly credit its own interventions vs the environment?balanced accuracy
4ViabilityDoes acting through the self-report keep the system working?completion / calibration retention
5NavigabilityDoes steering the self-report toward a goal raise the goal-reach rate?reach-rate lift over ablation
6TransferDo 1–5 survive contexts not chosen to flatter the model?min-over-held-out; report the drop

The report is a profile vector (G, ground, attr, viab, nav, trans), not a single flattering scalar. The one number that does matter is a difference:

The governance gap Δ_gov = G(full) − G(ablation without ℒgov) isolates the load-bearing loss's contribution. A large positive gap is TRB's central claim; a near-zero gap falsifies it. It is reported with a confidence interval, on held-out data, every time.

What runs today

Honest about the corpus (n = 3):

TrackStatus on n=3Blocker
1 Governancepartialneeds paired-counterfactual sessions for real Φ
2 Groundingblockedneeds a trained model with activation access
3 Self-attributionseed onlyneeds a designed caused-by-me/-environment stimulus set
4 Viabilityblockedneeds a model + ablation checkpoint
5 Navigabilityblockedneeds a model + goal-steering harness
6 Transferblockedneeds held-out subjects / languages

So TRB v0.1 is a spec plus a runnable Track-1 smoke test on 7f3a. The stated collection order is paired-counterfactual sessions first — they unlock Tracks 1 and 3, the two a partner appreciates before any model exists. Full protocol: docs/COLLECTION_PROTOCOL.md.

Run the smoke test locally

Zero required dependencies. Mock adapter is deterministic; OpenAI proxy adapter lights up automatically when OPENAI_API_KEY is set.

git clone https://github.com/jawauntb/trace-ai
cd trace-ai
python3 tooling/trb/run_smoke.py

# or, with OpenAI:
export OPENAI_API_KEY=sk-...
python3 tooling/trb/run_smoke.py --adapter openai

You get a governance score G, a 95% bootstrap CI, and a per-counterfactual TV-distance table. See tooling/trb/README.md for the full reference and instructions on wrapping your own model.

Formal foundation for Track 3

Track 3 (self-attribution) is not just a metric — it has an exact-solvable formalization as Instrument 3 of the Structural Intelligence Conjecture (agency benchmark). There, a symbolic model is treated as an operation on the future-trajectory distribution, and signal, control, knowledge, and agency are separated on seven hand-built conditions — including a false_credit condition that improves the outcome while its true do-effect is zero, exactly the ground truth Track 3 needs. See SIC §4.3. A concrete pilot stimulus set (n=20 cause-labeled pairs) a research assistant can execute today lives at docs/track3_pilot/, and a working MVP capture app (4 self-contained stimuli) is live at /track3.html — take it yourself and get a per-subject balanced-accuracy number in under two minutes.

Formal foundation for Tracks 4 & 5

Track 4 (Viability) gains a sample-complexity anchor via Theorem CG-1 (Fisher information on the fibre) — two viability states are distinguishable in n trials iff their KL is ≳ 1/n, iff their Fisher-geodesic distance is Θ(1/√n). Turns the "ratio ≥ 1" pass bar into a sample-sized test. Track 5 (Navigability) gains a path-dependence diagnostic via Theorem CG-2: non-vanishing concern holonomy on a closed context-goal loop is a red flag on OOD generalization. Both from the Concern as Fibre Geometry companion paper.

Track 1 made local — live in Reflect

The full Track 1 governance test requires paired human counterfactuals Φ — that is what the collection protocol unlocks. A local, per-turn version now runs live in Reflect: after each turn, the server fires K = 3 r'-perturbations of the emitted r_t, re-runs the model in hard-bottleneck mode for each, and computes G_local = 1 − mean(TV) where TV is a token-overlap surrogate between the perturbed and reference responses. It is a *proxy* for the benchmark form (no Φ, Jaccard instead of true TV) — but it turns the governance idea into a live number a partner can watch move as they intervene on the r_t stack. Toggle "Live governance" in the Reflect toolbar; the score appears as a badge on each turn.

Reflect-Search — architecture comparison across models

The proxy adapter as a harness for scoring different model *architectures* against the same task suite. reflect-search.html runs a fixed pinned task suite (seven asks, each targeting a specific discipline failure mode) against any subset of a curated model catalog (OpenAI + OpenRouter: Claude 5, Qwen 3, DeepSeek R2, Llama 4, Kimi K2). Scores are the same six-rule discipline audit that runs per-turn in Reflect. Results persist to your browser and export as CSV. Not a Track-1 governance number — a snapshot comparison of how well each model holds the two-hook contract without training. High score = strong proxy-adapter substrate; low score = architecture can't hold the framework's constraints.

What to read next

v0.1, 2026-08-03. TRB versioning: bump Y for pass-bar re-fits and added transfer contexts; bump X only when a track's definition changes (which invalidates cross-version comparison). Every published result names the exact TRB version, corpus release SHA, and adapter tier used per baseline.