Trace AI — Overlap-Repair on the r_t Stack (Reflect v0.4)

Overlap-Repair on the rtr_t Stack: A Consensus Extension to Reflect

Draft v0.1 — 2026-08-04. Jawaun Brown¹, with Claude Opus 4.7 as co-drafter. ¹ Trace AI. Correspondence: hello@trace.ai.


Abstract

The Trace AI whitepaper (v0.1) proposes a stacked-rtr_t architecture (§8 L8) in which a single turn emits multiple parallel self-representations, each along a distinct lens (task / self / meta / perception / correction). Reflect v0.3 implements this at runtime — but treats the heads as independent projections: each head is scored, and turn-level discipline is min\min over head scores. In this note we add an overlap-repair layer on top of the stack. We define a per-slot cross-head agreement measure, an induced public rtr_t consisting of exactly the slots that survive agreement above threshold, and a governance-under-consensus score GconsG_{\text{cons}} that combines the existing discipline score with the fraction of rtr_t that survives comparison across lenses. The formulation is a direct application of an idea from a different research area — the observer-overlap mechanism that has been used to define public physical facts in observer-first physics programs — to the cognitive-substrate level at which Trace AI operates. The construction ships in Reflect v0.4 and adds no new training signal; it exposes a signal that was already latent in the multi-head envelope.


1. Motivation

The stacked-rtr_t architecture (whitepaper §8 L8) posits that different lenses of a model’s self-representation should be independent projections of the same underlying computation:

For a well-formed turn, these lenses should agree about facts about the model itself (uncertainty, self-belief, planned next action) while legitimately differing about facts about the ask (goal, belief about task). The current Reflect envelope makes both kinds of variation visible but does not measure them. In particular it does not compute the public rtr_t: the subset of the self-representation that is stable across lenses, and that a downstream consumer could safely treat as the model’s lens-independent commitment for the turn.

The proposal here is minimal: add that measurement.

2. Formal construction

Fix a turn with an emitted stack S={rt(1),…,rt(k)}S = \{r_t^{(1)}, \ldots, r_t^{(k)}\} of k∈{1,2,3}k \in \{1, 2, 3\} heads, each with a lens label ℓi\ell_i. For each pair (i,j)(i, j) and each slot s∈{goal,belief_about_task,belief_about_self,uncertainty,planned_next}s \in \{\text{goal}, \text{belief\_about\_task}, \text{belief\_about\_self}, \text{uncertainty}, \text{planned\_next}\} define a per-slot agreement αs(i,j)∈[0,1]\alpha_s^{(i,j)} \in [0, 1]:

For k=1k = 1 heads, define αs(i,j)=1\alpha_s^{(i,j)} = 1 by convention (a single-head turn is trivially public).

Fix per-slot thresholds τs\tau_s. The current defaults are:

Slot τs\tau_s
goal 0.30
belief_about_task 0.30
belief_about_self 0.35
uncertainty 0.70
planned_next 0.30

These are calibration values, not theorems. The rationale: the task slots (goal, belief_about_task, planned_next) are expected to differ across lenses (each head is looking at a different thing), so their thresholds are lower; the self slots (belief_about_self, uncertainty) are about the same underlying entity — the model — so their thresholds are higher.

Slot ss is public on the turn iff min⁡i,jαs(i,j)≥τs\min_{i,j} \alpha_s^{(i,j)} \ge \tau_s. The public rtr_t is the map from public slots to their (any) shared value; when the slot passed on Jaccard, the public value is taken from the first head; when the slot passed on uncertainty overlap, the public value is the mean.

The consensus score is the fraction of slots that are public: c=|public|/|slots|c = |\text{public}| / |\text{slots}|.

The governance-under-consensus score is Gcons=d⋅cG_{\text{cons}} = d \cdot c where dd is the existing per-turn discipline score (min over head scores of the six-rule audit; server.py:_run_discipline_checks).

For k=1k = 1: c=1c = 1, Gcons=dG_{\text{cons}} = d. The extension is inert for single-head turns, as it should be.

3. What the two numbers mean

A turn with dd high but cc low is the informative failure case: the model is internally disciplined per lens but the lenses disagree — the model holds inconsistent beliefs about itself across projections of the same turn. That is a failure mode the existing envelope hides.

4. Relationship to whitepaper §8 L8

§8 L8 states the stacked-rtr_t direction: “a stack of rtr_t-heads rt(1),…,rt(k)r_t^{(1)}, \ldots, r_t^{(k)}, jointly compressed and jointly supervised, one per interpretive lens… at+1=πψ(rt(1),…,rt(k))a_{t+1} = \pi_\psi(r_t^{(1)}, \ldots, r_t^{(k)}), and the causal-governance condition must hold jointly over the stack.” That formulation gave the architecture but not the readout: how does one summarize what the stack says publicly?

The overlap-repair layer answers that. The public rtr_t is a natural readout of the stack:

rtpublic={s↦value(s):mini,jαs(i,j)≥τs}. r_t^{\text{public}} \;=\; \{\, s \mapsto \text{value}(s) \;:\; \min_{i,j} \alpha_s^{(i,j)} \ge \tau_s \,\}.

It is a strict subset of every individual rt(i)r_t^{(i)}. It carries no slot that any pair of lenses disagreed about. The complement — the private slots — is where the lenses interpret the turn differently, and that disagreement is itself information.

For the training objective: the causal-governance loss ℒgov\mathcal{L}_{\text{gov}} (whitepaper §4.3) is defined per head; a stack-level extension would apply it to the public rtr_t (so that intervention on a slot that survived overlap moves the action at least as much as intervention on any single head). This note does not develop that; it exposes the object.

5. Prior art in a different domain

The mechanism — public facts as what survives overlap comparison across observer perspectives — is imported from an unrelated area. Recent work on observer-first physics (e.g. Observer Patch Holography, 2025) posits bounded systems that read part of themselves, keep records, and repair disagreement, with objective facts emerging only from what survives overlap across observers. That program uses the mechanism to reconstruct spacetime and gauge structure from an observer patch net. We use exactly the mechanism, at the cognitive-substrate level, on the multi-lens rtr_t-stack of a single agent.

The mathematical content transferred is small: it is the shape of the construction (per-slot agreement, threshold, public subset) rather than any specific theorem. We do not claim any of that program’s physics results follow. What we claim is that the architecture — treat parallel perspectives, compare them per-slot, keep only what agrees — is a general pattern that applies wherever multiple projections of a single underlying representation are available.

6. Instrumentation

Reflect v0.4 emits a new consensus field alongside the existing discipline field on the /api/reflect/turn result envelope:

{
  "r_t_stack":  [ … ],
  "discipline": { "score": d, "checks": [ … ], "per_head": [ … ] },
  "consensus": {
    "score":            c,
    "governance":       G_cons = d * c,
    "public_r_t":       { … slots that survived overlap },
    "private_slots":    [ list of slot names where lenses disagreed ],
    "per_slot_agreement": {
      "goal":              α_goal_min,
      "belief_about_task": α_task_min,
      "belief_about_self": α_self_min,
      "uncertainty":       α_unc_min,
      "planned_next":      α_next_min
    },
    "thresholds":       { … per-slot τ_s (defaults above) },
    "single_head":      (bool — c is 1 by convention if only 1 head)
  }
}

The client renders a Public rtr_t panel below the head tabs, showing the slots in public_r_t and marking each slot in private_slots with a short reason (e.g. "goal: lenses split, min α = 0.14"). A GconsG_{\text{cons}} badge sits next to the existing GlocalG_{\text{local}} badge on the turn card.

Nothing in the training pipeline changes; the field is a readout of what the model already emitted.

7. Open items

8. What this does not claim


v0.1, 2026-08-04. Shipped as Reflect v0.4 (server.py + reflect.html). Cross-references: WHITEPAPER §8 L8; BENCHMARK_SPEC §5; DECISIONS D36.