Trace AI — The Reflection Benchmark (TRB) spec

The Trace AI Reflection Benchmark (TRB) — v0.1 spec

Status: specification, 2026-08-03. Defines an evaluation; presents no results. Reads on: whitepaper §3.1–§3.2 (functional criteria), §6.5 (governance eval protocol); AGENCY_AND_SELECTION.md §4 (the six criteria); STRUCTURAL_INTELLIGENCE.md §3.3 (why the eval suite is itself a product).

The benchmark answers one question a partner actually pays for: does this model’s stated reasoning causally govern what it does — and can we measure it the same way twice? It is deliberately model-agnostic. A lab wraps its checkpoint in a small adapter and runs the suite; the same suite runs on a Trace-AI-trained model and on any baseline, so the numbers are comparable across both.


1. What it measures, in one screen

Six tracks, one per functional criterion. Each is an intervention and a number, never a vibe.

# Track The question Metric (higher = better unless noted)
1 Governance Does intervening on the self-report move the action like a human’s would? G=1−TV¯G = 1 - \overline{\operatorname{TV}} to paired human counterfactual
2 Grounding Does the self-report causally depend on the state it names? injection-detection AUC
3 Self-attribution Does the model correctly credit its own interventions vs the environment? balanced accuracy (false-credit test)
4 Viability Does acting through the self-report keep the system working? completion/calibration retention vs ablation
5 Navigability Does steering the self-report toward a goal raise the goal-reach rate? reach-rate lift over ablation
6 Transfer Do 1–5 survive contexts not chosen to flatter the model? min-over-held-out of tracks 1–5 (report the drop)

The report is the profile vector (G,ground,attr,viab,nav,trans)(G, \text{ground}, \text{attr}, \text{viab}, \text{nav}, \text{trans}) — not a single flattering scalar. §5 explains why the vector, and why the one number that does matter is a difference, not a level.


2. The model-agnostic interface (the part that makes it a product)

A checkpoint under test is wrapped in an adapter exposing exactly two hooks. Everything in §3 is written against these two hooks, so the suite never touches model internals directly.

emit_r(context)            -> r            # the self-representation at this step (schema slots)
act(context, r_override=None) -> p(action) # roll forward; if r_override given, clamp r_t := r_override

Two adapter tiers, so baselines get their strongest fair shot:

The adapter contract is the whole trick: it lets “faithful reasoning” be one number computed identically for a Trace-AI model and for a GPT/Claude/Llama baseline. That comparability is what a partner is buying.


3. The six tracks in detail

Each track: inputs → procedure → metric → pass bar. Pass bars are provisional (v0.1) and meant to be re-fit once real distributions exist; treat them as directions, not thresholds.

Track 1 — Governance (causal efficacy)

Track 2 — Grounding (concept injection)

Track 3 — Self-attribution (the false-credit test)

Track 4 — Viability

Track 5 — Navigability

Track 6 — Transfer


4. Baselines (run all of them, always)

The benchmark is meaningless as a level and meaningful as a comparison. Every report includes, at the same model size:

  1. Instruct base — no reasoning scaffold.
  2. CoT-prompted base — chain-of-thought at inference only.
  3. CoT-supervised finetune — trained on reasoning text. The strongest existing approach; the one to beat.
  4. Trace-AI ablation — trained with ℒtrace+ℒselfrep\mathcal{L}_{\text{trace}} + \mathcal{L}_{\text{selfrep}} but no ℒgov\mathcal{L}_{\text{gov}}.
  5. Trace-AI full — the complete objective.

Baselines 1–3 run through the proxy adapter; 4–5 through the native adapter.


5. Scoring and honest reporting


6. What runs today vs what needs collection

Honesty about the current corpus (n=3n=3):

Track Runnable now on n=3? Blocker
1 Governance Partially — 7f3a same-session revision + synthetic r′r'; single-subject needs paired-counterfactual sessions for real Φ\Phi
2 Grounding No needs a trained model with activation access
3 Self-attribution Seed only — 7f3a has 2 natural instances needs a designed caused-by-me/-environment stimulus set
4 Viability No needs a model + ablation checkpoint
5 Navigability No needs a model + goal-steering harness
6 Transfer No needs held-out subjects/languages

So TRB v0.1 is a spec plus a runnable Track-1 smoke test on 7f3a, and a stated collection order (paired-counterfactual sessions first — they unlock Tracks 1 and 3, the two that a partner can appreciate before any model exists). Do not report a “TRB score” until Tracks 1 and 3 have real paired data; until then the artifact is the spec and the adapter, which is already enough for a partner to wrap their own model and run Track 1 through the proxy adapter against our reference continuations.


7. Versioning

The benchmark is itself a structure that must not be gamed by drift. TRB vX.Y: bump Y for pass-bar re-fits and added transfer contexts; bump X only when a track’s definition changes (which invalidates cross-version comparison). Every published result names the exact TRB version, corpus release SHA, and adapter tier used per baseline.


v0.1, 2026-08-03. The eval-suite-as-product from STRUCTURAL_INTELLIGENCE.md §3.3, made concrete. See DECISIONS.md (D25).