Trace AI — Structural Intelligence Conjecture (formal paper)

The Structural Intelligence Conjecture

Representation as a stochastic fibration, with nine exact instruments

Jawaun Brown — human author and research director. Claude Code (agent) — experiment code, analysis, and manuscript production under direction and review. Date: 2026-08-03 (v4 + Observatory Instrument 9, integrated 2026-08-03). Status: five theorems + one conditional theorem + nine exact executable instruments. Existence, cross-task stability, discrete learnability, and continuous learnability at resolution ε\varepsilon derived (Theorems 1–6); Theorem 2 (rate–distortion) witnessed exactly by Instrument 7; Theorem 7 (linear-ICA identifiability) resolves SIC-C-c positively inside the linear-ICA hypothesis class (Instrument 8) and the sparse-mechanism / IMA class (Gresele et al. 2021) gives a second positive resolution (Instrument 9) — two distinct inductive-bias classes now provably earn escape from Theorem 6’s ε\varepsilon-covering exponential-in-dZd_Z bound. What remains genuinely open is only the general programme: which other inductive-bias classes admit their own polynomial-in-dZd_Z theorems (iVAE with auxiliary variables, interventional CRL, …) — a mainstream question in identifiable-representation-learning theory, addressed one class at a time.

Live interactive Observatory: structural-observatory-production.up.railway.app — all nine instruments rendered as interactive panels with committed numerical results.

Companion PDFs: sic_paper.pdf is the v4 (latest); prior versions preserved as sic_paper_v2.pdf, sic_paper_v3.pdf, sic_paper_v4.pdf. Reads on: STRUCTURAL_INTELLIGENCE.md (plain-English pointer to the same object with the Trace AI lens), WHITEPAPER.md (the training-signal application), BENCHMARK_SPEC.md (TRB Track 3 is Instrument 3 of this paper), AGENCY_AND_SELECTION.md (the selection frame those instruments live in).

Why it exists here. The STRUCTURAL_INTELLIGENCE.md note tells the story in plain English; that told-in-plain-English claim needed a paper with theorems and exact-solvable witnesses under it, so the note doesn’t have to carry weight it can’t. This is that paper.


Abstract

We study a single latent object that recurs across five otherwise unrelated works — a category-theoretic framework for materials design, an information-theoretic limit on the programmatic specification of biological systems, a corpus of machine-found mathematical proofs with their discovery notes, Wigner’s essay on the unreasonable effectiveness of mathematics, and a structural-realist ontology. The common object is a stochastic fibration with a compiler: a coarse-graining q:X→Zq : X \to Z together with a kernel K:Z⇝XK : Z \rightsquigarrow X whose support lies in the fibre q−1(z)q^{-1}(z). We derive the object — it is not posited: for any statistical task on a standard Borel space, the minimal sufficient σ\sigma-algebra of Halmos and Savage yields (q,K)(q, K) as its quotient and regular-conditional pair (Theorem 1); Shannon’s rate–distortion pair parameterises the same object at a distortion budget (Theorem 2, now witnessed exactly by Instrument 7 on uniform and Bernoulli sources to 10−910^{-9}); the construction is the unit/counit of an adjunction (Proposition 3). Cross-task stability — a single quotient that is sufficient for a whole task family — holds iff the family admits a shared Markov screen, i.e. a common latent generator (Theorem 4, conditional). Discrete-case learnability is also a theorem: with the task family separating ZZ and the compiler’s fibres balanced (pmin≥1/(cM)p_{\min} \ge 1/(cM)), empirical common-sufficient clustering recovers qq with probability ≥1−ε\ge 1 - \varepsilon from N≥cMln⁡(M/ε)N \ge cM \ln(M/\varepsilon) samples (Theorem 5). The continuous-case extension at resolution ε\varepsilon is also a theorem (Theorem 6): bound O(c(DZ/ε)dZlog⁡((DZ/ε)dZ/εrel))O(c(D_Z/\varepsilon)^{d_Z} \log((D_Z/\varepsilon)^{d_Z}/\varepsilon_{\text{rel}})) — polynomial in 1/ε1/\varepsilon at fixed dZd_Z, provably exponential in dZd_Z at fixed ε\varepsilon without further inductive bias. Adding one specific bias — independent non-Gaussian latents with a linear mixing — resolves that exponential-in-dZd_Z cost inside the linear-ICA hypothesis class: the classical Hyvärinen–Oja identifiability result becomes Theorem 7 in our framework’s language, with Instrument 8 witnessing polynomial recovery (Amari index ≤0.02\le 0.02 at N=10000N = 10\,000 for dZ∈{2,4,6,8}d_Z \in \{2, 4, 6, 8\}). Eight exact, deterministic, unit-tested instruments now witness each theorem on solvable cases and exhibit the dissociations the framework predicts.


1. The master object

Let XX be a space of concrete realizations and ZZ a space of structures, functions, or observables. Two maps carry the construction:

This sharpens the adjunction R⊣CR \dashv C of the companion synthesis: q=Cq = C (coarse-grain) and KK is a stochastic section of RR (realize). The biology paper of Kiiskinen–Kivinen–Rivas is exactly this construction — the genome names zz, the physics substrate is the compiler KK that “computes the samples,” and selection acts on the ensemble statistics — and its coarse-graining threshold C*C^* proves that below a critical resolution the fibre is too large to address with the specification budget.

A note on rigor. Cross-domain resemblance is not isomorphism. The honest hierarchy of sameness runs: isomorphism ⊃\supset bisimulation ⊃\supset functor ⊃\supset natural transformation ⊃\supset adjunction / Galois connection ⊃\supset Morita-like equivalence ⊃\supset simulation-at-a-resolution. Most relations among the source works are adjunctions, simulations, and shared diagram shapes, not object-level isomorphisms; every proposed connection must state what kind of sameness it is, what the map forgets, and what would have to be proved to make it a theorem.


2. Derivations

2.1 Existence via minimal sufficiency (Theorem 1)

Let XX be a standard Borel space, {Pθ:θ∈Θ}\{P_\theta : \theta \in \Theta\} a dominated family of probability measures on XX, and YY a random variable whose distribution depends on θ\theta (equivalently: θ\theta indexes the task). The Halmos–Savage sufficiency theorem gives a sufficient σ\sigma-algebra 𝒮⊆ℬ(X)\mathcal{S} \subseteq \mathcal{B}(X); under mild regularity (Bahadur 1954; Lehmann–Casella) there is a minimal sufficient σ\sigma-algebra 𝒯⊆𝒮\mathcal{T} \subseteq \mathcal{S}, unique up to PP-null completion.

Define:

Theorem 1 (Existence of the master fibration). (q,K)(q, K) is a stochastic fibration with supp⁡K(⋅∣z)⊆q−1(z)\operatorname{supp} K(\cdot \mid z) \subseteq q^{-1}(z), and four of the five clauses of the Structural Intelligence Conjecture hold as theorems:

  1. Task-relevance descends. Pθ(Y∈B∣X)=Pθ(Y∈B∣q(X))P_\theta(Y \in B \mid X) = P_\theta(Y \in B \mid q(X)) almost surely.
  2. Irrelevant variation is confined to fibres. For x,x′∈q−1(z)x, x' \in q^{-1}(z): ℒ(Y∣X=x)=ℒ(Y∣X=x′)\mathcal{L}(Y \mid X = x) = \mathcal{L}(Y \mid X = x').
  3. Compact specification. |image⁡(q)|=|atoms⁡(𝒯)|≤|X||{\operatorname{image}}(q)| = |\operatorname{atoms}(\mathcal{T})| \le |X|, strict iff 𝒯\mathcal{T} is nontrivial.
  4. Re-instantiation via KK. By construction, KK is a Markov kernel Z⇝XZ \rightsquigarrow X with the correct support.

The Fiber Finder (Instrument 1, §4.1) is an exact computational witness.

2.2 Rate–distortion parameterisation (Theorem 2)

Let X∼pXX \sim p_X on XX and d:X×X̂→ℝ+d : X \times \hat{X} \to \mathbb{R}_+ a distortion. Shannon’s rate–distortion function

R(D)=infp(x̂∣x):𝔼[d(X,X̂)]≤DI(X;X̂)R(D) \;=\; \inf_{p(\hat{x} \mid x) : \mathbb{E}[d(X, \hat{X})] \le D} I(X; \hat{X})

is attained by an encoder p*(x̂∣x)p^*(\hat{x} \mid x) and decoder marginal p*(x̂)p^*(\hat{x}).

Theorem 2 (RD parameterisation of the master fibration). For every distortion budget D≥0D \ge 0, the RD-optimal pair defines a stochastic fibration (qD,KD)(q_D, K_D); the family {(qD,KD):D≥0}\{(q_D, K_D) : D \ge 0\} is a one-parameter deformation of the sufficiency fibration. At D=0D = 0 the encoder is minimal-sufficient; as DD grows the fibres grow and the specification shrinks along the RD curve. The Kiiskinen–Kivinen–Rivas biology threshold C*C^* is R(Dbio)R(D_{\text{bio}}) for the domain’s distortion measure.

2.3 Categorical restatement (Proposition 3)

Let 𝑷𝒓𝒐𝒃(X)\mathbf{Prob}(X) and 𝑷𝒓𝒐𝒃(Z)\mathbf{Prob}(Z) be the categories of probability measures on XX and ZZ. The sufficient-statistic construction is the unit/counit pair of an adjunction:

Proposition 3. (q,K)(q, K) is the unit/counit pair of C⊣RC \dashv R. SIC clauses (1)–(4) are the adjunction’s triangle identities specialised to the sufficiency reduction; the categorical framing adds compact language, not new content.

2.4 Cross-task stability (Theorem 4, conditional)

Clause (5) of the SIC — stability of qq across substrates and contexts — is not implied by single-task sufficiency: different tasks have different minimal sufficient statistics in general.

Theorem 4 (Cross-task stability under latent generation). Suppose there exists a random variable ZZ and, for every task in a family {Yα:α∈A}\{Y_\alpha : \alpha \in A\}, a conditional independence

Yα⟂X∣Z.Y_\alpha \perp X \mid Z.

Then ZZ is a sufficient statistic for every YαY_\alpha simultaneously.

Proof. For each α\alpha, sufficiency of ZZ for YαY_\alpha from XX is exactly the Markov property Yα⟂X∣ZY_\alpha \perp X \mid Z; the family case is the pointwise conjunction. ▫\square

Corollary (Equivalence). A task family {Yα}\{Y_\alpha\} admits a common sufficient statistic strictly coarser than XX iff there is a decomposition X=g(Z,η)X = g(Z, \eta) with Z⟂ηZ \perp \eta and every YαY_\alpha factoring through ZZ.

Theorem 4 turns SIC clause (5) into a claim about the world — the joint law of XX and the task family — rather than about intelligence. The Cross-Task Sufficiency instrument (Instrument 4, §4.4) witnesses it exactly on a 4-bit Boolean world.

2.5 Learnability, discrete case (Theorem 5)

Theorems 1–4 do not immediately derive that a finite adaptive system discovers qq from data. In its unrestricted form this claim is false (Locatello et al. 2019): for smooth continuous XX and no auxiliary information, no unsupervised algorithm identifies a factored latent up to trivial transformations without inductive bias. Task families supply the identifying auxiliary information, and — quantitatively enough — reduce learnability from a conjecture to a theorem in the discrete case.

Setup. Let XX be finite; PXP_X a distribution on XX with min-fibre mass pmin:=min⁡zPX(q(X)=z)≥1/(cM)p_{\min} := \min_z P_X(q(X) = z) \ge 1/(cM) for the true partition q:X→Zq : X \to Z, |Z|=M|Z| = M; {Yα}α=1..K\{Y_\alpha\}_{\alpha=1..K} a family of deterministic tasks factoring through qq; the family separates qq (for every z≠z′z \ne z', some α\alpha has hα(z)≠hα(z′)h_\alpha(z) \ne h_\alpha(z')).

Algorithm (empirical common-sufficient clustering). Given NN i.i.d. samples with labels:

  1. For each observed xx, form the response profile π(x):=(Y1(x),…,YK(x))\pi(x) := (Y_1(x), \ldots, Y_K(x)).
  2. Extend π\pi to all of XX by the same deterministic labelling functions.
  3. Return q̂\hat{q} defined by q̂(x)=q̂(x′)⇔π(x)=π(x′)\hat{q}(x) = \hat{q}(x') \iff \pi(x) = \pi(x').

Theorem 5 (Discrete learnability). For any ε∈(0,1)\varepsilon \in (0,1),

Pr⁡[q̂=q as maps X→Z]≥1−εwheneverN≥cMln⁡(M/ε).\Pr[\hat{q} = q \text{ as maps } X \to Z] \;\ge\; 1 - \varepsilon \quad\text{whenever}\quad N \;\ge\; cM \ln(M/\varepsilon).

Proof. Injectivity of Φ:z↦(h1(z),…,hK(z))\Phi : z \mapsto (h_1(z), \ldots, h_K(z)) gives that q̂=q\hat{q} = q once every fibre of qq contains at least one observed sample. For each zz, Pr⁡[z missed in N samples]≤(1−1/(cM))N≤exp⁡(−N/(cM))\Pr[z \text{ missed in } N \text{ samples}] \le (1 - 1/(cM))^N \le \exp(-N/(cM)). Union bound over MM fibres: Pr⁡[some fibre missed]≤Mexp⁡(−N/(cM))≤ε\Pr[\text{some fibre missed}] \le M \exp(-N/(cM)) \le \varepsilon for N≥cMln⁡(M/ε)N \ge cM \ln(M/\varepsilon). ▫\square

What the theorem does not do. It is stated for finite XX and deterministic tasks; it requires the separation assumption (without it the recoverable object is strictly coarser than ZZ — namely the CSS of the observable task family); it requires fibre balance through cc.

2.5b Continuous-case learnability at resolution ε\varepsilon (Theorem 6)

The honest form of the continuous claim is “recovery up to resolution ε\varepsilon”, and it reduces cleanly to Theorem 5 by ε\varepsilon-covering.

Setup. Z⊆ℝdZZ \subseteq \mathbb{R}^{d_Z} bounded with diameter DZD_Z; 𝒩ε⊆Z\mathcal{N}_\varepsilon \subseteq Z a minimal ε\varepsilon-net; ZεZ_\varepsilon the induced partition into Nε:=|𝒩ε|N_\varepsilon := |\mathcal{N}_\varepsilon| Voronoi cells; qε:=πε∘qq_\varepsilon := \pi_\varepsilon \circ q the ε\varepsilon-quantised true partition; fibre balance and separation both at scale ε\varepsilon.

Theorem 6 (Continuous-case learnability, at resolution ε\varepsilon). For any εrel∈(0,1)\varepsilon_{\text{rel}} \in (0,1), q̂\hat{q} recovers qεq_\varepsilon exactly with probability ≥1−εrel\ge 1 - \varepsilon_{\text{rel}} from

N≥c⋅Nε⋅ln⁡(Nε/εrel).N \;\ge\; c \cdot N_\varepsilon \cdot \ln(N_\varepsilon / \varepsilon_{\text{rel}}).

For Z⊆ℝdZZ \subseteq \mathbb{R}^{d_Z} bounded, Nε=O((DZ/ε)dZ)N_\varepsilon = O((D_Z/\varepsilon)^{d_Z}), giving

N=O(c⋅(DZ/ε)dZ⋅dZ⋅log(DZ/(ε⋅εrel))).N \;=\; O\!\left( c \cdot (D_Z/\varepsilon)^{d_Z} \cdot d_Z \cdot \log(D_Z / (\varepsilon \cdot \varepsilon_{\text{rel}})) \right).

Proof. Apply Theorem 5 to the discretised experiment. ▫\square

Fundamental limit. Uniform sample complexity polynomial in dZd_Z alone — dropping the exponential dependence on dZd_Z at fixed ε\varepsilon — is impossible without additional inductive bias on qq. Escaping the curse requires linearity (linear ICA), sparsity (independent-mechanism analysis; Gresele et al. 2021), exponential-family conditional latents with auxiliary variables (iVAE; Khemakhem–Kingma–Monti–Hyvärinen 2020), or interventional data on ZZ (causal representation learning; Ahuja–Mahajan–Wang–Bengio 2022). The next section resolves SIC-C-c positively inside the first of those classes.

2.5c Positive resolution of SIC-C-c for linear ICA (Theorem 7)

Theorem 6 says continuous-case sample complexity is exponential in dZd_Z without inductive bias. Adding one specific bias — linear generative model with independent non-Gaussian latents — resolves SIC-C-c positively inside that hypothesis class. This is a classical result from Hyvärinen’s independent component analysis (ICA) line (Comon 1994; Hyvärinen 1999; Hyvärinen–Oja 2000), restated here in the framework’s language and witnessed numerically as Instrument 8.

Setup (Theorem 7).

Theorem 7 (Linear-ICA identifiability, classical). Under the linear-ICA setup, the un-mixing matrix W=A−1W = A^{-1} is identifiable up to permutation and coordinate-wise sign, and there is an efficient algorithm (e.g. FastICA, Hyvärinen 1999) whose empirical un-mixing Ŵ\hat W satisfies

Amari⁡(Ŵ⋅A)→0as N→∞\operatorname{Amari}(\hat W \cdot A) \;\to\; 0 \quad\text{as } N \to \infty

at sample-complexity rate polynomial in dZd_Z and inverse-polynomial in any target accuracy, uniformly over invertible AA and non-Gaussian component densities in a broad regularity class. Formally: for every ε>0\varepsilon > 0, Amari⁡(Ŵ⋅A)≤ε\operatorname{Amari}(\hat W \cdot A) \le \varepsilon with high probability from N=poly⁡(dZ,1/ε)N = \operatorname{poly}(d_Z, 1/\varepsilon) samples.

Amari⁡(P)\operatorname{Amari}(P) is the standard [0,1][0, 1]-valued permutation-and-sign-invariant distance on dZ×dZd_Z \times d_Z matrices:

Amari⁡(P)=12dZ(dZ−1)[∑i(∑j|Pij|maxj|Pij|−1)+∑j(∑i|Pij|maxi|Pij|−1)].\operatorname{Amari}(P) \;=\; \frac{1}{2 d_Z (d_Z - 1)}\left[\, \sum_i \!\left(\!\sum_j \tfrac{|P_{ij}|}{\max_j |P_{ij}|} - 1\!\right) \;+\; \sum_j \!\left(\!\sum_i \tfrac{|P_{ij}|}{\max_i |P_{ij}|} - 1\!\right) \right].

Proof (sketch, classical). Independent non-Gaussian marginals uniquely identify the linear model up to permutation and sign (Comon 1994). Whitening followed by rotation to maximise a non-Gaussianity contrast (negentropy in Hyvärinen 1999) converges to Ŵ\hat W such that Ŵ⋅A\hat W \cdot A is a signed permutation; the convergence rate follows from standard M-estimator analysis of the empirical contrast. ▫\square

What Theorem 7 does not say. It resolves SIC-C-c only inside the linear ICA hypothesis class. Every other inductive-bias class in the identifiable-representation-learning line (sparse ICA, independent-mechanism analysis, iVAE with auxiliary variables, interventional CRL) has its own analogue, each with its own sample-complexity theorem. A full SIC-C-c programme would add one more instrument per class; Instrument 8 (§4.8) is the first.

2.6 SIC in the honest split

Theorems 1–6 collapse SIC to a derived skeleton with one empirical antecedent and one residual inductive-bias question:

Clause Status Notes
SIC-A Existence Theorem Theorem 1 + Proposition 3
SIC-B Cross-task stability Conditional theorem Theorem 4; antecedent empirical about the world
SIC-C-a Discrete learnability Theorem Theorem 5; numerically sharp constant (Instrument 5)
SIC-C-b Continuous learnability at resolution ε\varepsilon Theorem Theorem 6; poly in 1/ε1/\varepsilon, exp in dZd_Z (Instrument 6)
SIC-C-c Uniform polynomial in dZd_Z Impossible in general; theorem inside linear-ICA hypothesis class ε-covering lower bounds + Locatello 2019 rule out an unqualified bound; Theorem 7 gives it inside linear ICA (Instrument 8). Other inductive-bias classes (sparse ICA, iVAE, interventional CRL) admit their own analogues; each is a separate theorem-instrument pair, and the Observatory is set up to host them.

What remains genuinely open is only the general programme: which inductive biases make polynomial-in-dZd_Z learnability possible for which continuous hypothesis classes — a mainstream question in identifiable representation learning, addressed one class at a time.


3. The conjecture (full-strength form)

Structural Intelligence Conjecture. A finite adaptive system’s central capability is to discover a quotient q:X→Zq : X \to Z such that (1) the task-relevant dynamics descend to ZZ, (2) irrelevant variation is confined to the fibres q−1(z)q^{-1}(z), (3) useful interventions are compactly specifiable in ZZ, (4) those specifications re-instantiate through a compiler, and (5) their consequences remain stable across substrates and contexts.

In one line: intelligence finds the level at which the world becomes both compressible and controllable.


4. Nine exact instruments

Instruments 1, 4, 5, 6, 7, 8, and 9 are computational witnesses of Theorems 1, 4, 5, 6, 2, 7 (linear-ICA class), and 7 (sparse-mechanism class) respectively; instruments 2 and 3 establish auxiliary dissociations on solvable cases. All are deterministic and unit-tested; seven are exact enumerations and two (Instruments 8 and 9) are fixed-seed Monte Carlo.

4.1 Fiber Finder — witness of Theorem 1

Over all 2n2^n worlds of an nn-bit Boolean space with a known ground-truth invariant, enumerate a lattice of candidate quotients and three selectors:

Result (exact). In every task, minimal_sufficient recovers the exact ground-truth invariant; mdl_only selects the constant map (insufficient); accuracy_only selects the identity (sufficient but uncompressed). Sufficiency, description length, and accuracy are three different quantities; only the sufficiency-then-compress rule finds the invariant.

4.2 Structure compiler

An abstract automaton exhibits accumulation, a phase transition, and hysteresis. Four compilers FiF_i map its trajectory into music (pitch/octave), a visual field (height/hue), text (regime-keyed lexicon), and spatial navigation (regime-gated corridor); each has a readback qiq_i.

Result (exact). For every medium qi∘Fi=idq_i \circ F_i = \text{id} on the trajectory (fidelity 1.0); all media read back to the same abstract trajectory. Four embodiments are verifiably one work, not four mood-matched artefacts.

4.3 Agency benchmark — the seed of TRB Track 3

An exact finite-state world with target and failure regions; futures enumerated to the horizon. A symbolic model mm is an operation on the trajectory distribution P(γ∣s)P(\gamma \mid s). Measure signal ΔKL=KL⁡(P(γ∣m)∥P(γ∣baseline))\Delta_{\text{KL}} = \operatorname{KL}(P(\gamma \mid m) \parallel P(\gamma \mid \text{baseline})), control (goal_gain), knowledge (predictive_accuracy), and agency (control + small calibration_error between claimed and true do-effect + positive transfer under a perturbation the intervention did not choose).

Result (exact). Seven hand-built conditions realize distinct metric signatures recovered by a fixed classifier: noise_signal moves the distribution with zero control; knowledge_only predicts with zero signal; false_credit improves the observed outcome while its true do-effect is zero and its self-attribution is miscalibrated; a brittle controller controls but does not transfer. No single scalar of behavioural influence identifies agency.

This is the exact-solvable formalization of what TRB Track 3 (self-attribution) measures on trained models. See BENCHMARK_SPEC.md §3 Track 3 — the false_credit condition is the ground truth for the balanced-accuracy score.

4.4 Cross-task sufficiency — witness of Theorem 4

On the 4-bit Boolean world X={0,1}4X = \{0,1\}^4 with latent Z(x)=(parity⁡{0,1}(x),parity⁡{2,3}(x))Z(x) = (\operatorname{parity}_{\{0,1\}}(x), \operatorname{parity}_{\{2,3\}}(x)), enumerate the lattice of quotients. Two task families:

Result (exact). Shared family: coarsest CSS is exactly joint⁡(parity⁡{0,1},parity⁡{2,3})=Z\operatorname{joint}(\operatorname{parity}_{\{0,1\}}, \operatorname{parity}_{\{2,3\}}) = Z (image size 4, description length 2 bits, strictly less than log⁡2|X|=4\log_2 |X| = 4 bits). Not-shared family: coarsest CSS is the identity (no compression). For the shared family, each single task’s minimal sufficient statistic is a 1-bit parity (image size 2), strictly coarser than the family CSS: combining tasks tightens the required partition by a factor of two.

4.5 Cross-task learnability — witness of Theorem 5

Building on the shared task family of Instrument 4, compute the exact recovery probability Pr⁡[q̂=q]\Pr[\hat{q} = q] of empirical common-sufficient clustering as a function of sample count NN, via inclusion–exclusion over the 2M=162^M = 16 subsets of fibres.

Distribution on X={0,1}4X = \{0,1\}^4 pminp_{\min} cc NboundN_{\text{bound}} at ε=0.05\varepsilon = 0.05 Exact Pr⁡[recover]\Pr[\text{recover}]
uniform 1/41/4 11 18 0.9775
skewed (0.625,0.125,0.125,0.125)(0.625, 0.125, 0.125, 0.125) 1/81/8 22 36 0.9756

Both strictly above 1−ε=0.951 - \varepsilon = 0.95. Below M=4M = 4 samples recovery is impossible (pigeonhole). Uniform recovery hits 0.95 already at N=14N = 14: the union-bound overhead is 4 samples in this regime — the theorem bound is honest and only mildly loose.

4.6 Cross-task learnability, continuous — witness of Theorem 6

Theorem 6’s continuous bound, verified exactly across (dZ,r)∈{1,2}×{4,8,16}(d_Z, r) \in \{1, 2\} \times \{4, 8, 16\} on ambient X=[0,1]2X = [0,1]^2 quantised to a 16×1616 \times 16 grid.

dZd_Z rr MM NboundN_{\text{bound}} Pr⁡[recover @ bound]\Pr[\text{recover @ bound}]
1 4 4 18 0.9775
1 8 8 41 0.9667
1 16 16 93 0.9609
2 4 16 93 0.9609
2 8 64 458 0.9538
2 16 256 2187 0.9521

All six grid points meet the 1−εrel=0.951 - \varepsilon_{\text{rel}} = 0.95 target at Theorem 6’s bound. Recovery is zero below MM (pigeonhole) and monotone in NN.

Ratio witness. Nbound(dZ=2,r)/Nbound(dZ=1,r)=5.17,11.17,23.52N_{\text{bound}}(d_Z=2, r) / N_{\text{bound}}(d_Z=1, r) = 5.17, 11.17, 23.52 at r=4,8,16r = 4, 8, 16 — cleanly matching the r⋅(log⁡slack)r \cdot (\log \text{slack}) prediction of Theorem 6. From dZ=1d_Z = 1 to dZ=2d_Z = 2 at r=16r = 16 multiplies the required sample count by ≈23.5×\approx 23.5\times. That is the curse of dimensionality made numerical.

Empirical common-sufficient clustering saturates the ε\varepsilon-covering lower bound: it is optimal within the class of algorithms that make no inductive assumption on qq. Escaping the exponential-in-dZd_Z scaling requires an inductive bias — Instrument 8 exhibits the first such escape.

4.7 Rate–distortion pair — witness of Theorem 2 (experiments/rate_distortion_pair)

Theorem 2’s rate–distortion parameterisation, verified exactly on two finite sources with Hamming distortion. Closed-form R(D)R(D) is evaluated at 10 D-grid points for the uniform source on n=4n = 4 symbols (R(D)=log⁡2(n)−h(D)−D⋅log⁡2(n−1)R(D) = \log_2(n) - h(D) - D \cdot \log_2(n - 1) for D≤1−1/n=0.75D \le 1 - 1/n = 0.75) and at 9 points for Bernoulli⁡(p=0.3)\operatorname{Bernoulli}(p = 0.3) (R(D)=H2(p)−h(D)R(D) = H_2(p) - h(D) for D≤min⁡(p,1−p)=0.3D \le \min(p, 1 - p) = 0.3). For each source we also construct the RD-optimal test channel explicitly (the symmetric error-DD channel for uniform; the Bernoulli-flip channel for Bernoulli) and verify I(X;X̂)=R(D)I(X; \hat X) = R(D) at every point in the achievable regime.

Result (exact). All ten pre-registered gates pass to 10−910^{-9}:

This closes the theorem–instrument pairing: Theorems 1, 2, 4, 5, and 6 all have exact witnesses. The Theorem 2 anchor R(0)=H(X)R(0) = H(X) reconnects to Theorem 1’s minimal-sufficient partition — the RD family (qD,KD)(q_D, K_D) is literally a one-parameter deformation of the sufficiency fibration.

4.8 Linear-ICA learnability — witness of Theorem 7 (experiments/linear_ica_learnability)

Theorem 7’s numerical witness. On the linear-ICA setup — X=A⋅ZX = A \cdot Z with A∈ℝd×dA \in \mathbb{R}^{d \times d} random orthogonal (fixed numpy seed for determinism), each Zi∼Laplace⁡(0,1)Z_i \sim \operatorname{Laplace}(0, 1) independent — sweep (dZ,N)∈{2,4,6,8}×{200,500,1000,2000,5000,10000}(d_Z, N) \in \{2, 4, 6, 8\} \times \{200, 500, 1000, 2000, 5000, 10000\} and run sklearn.decomposition.FastICA (fixed random_state, 8 trials per grid point). Recovery quality is the Amari index on P=Ŵ⋅AP = \hat W \cdot A; smaller is better.

Result (deterministic under seed). All four pre-registered gates pass:

Fixed-seed protocol keeps it reproducible run-to-run. Refined numbers from the live Observatory (structural-observatory-production.up.railway.app): at N=10,000N = 10{,}000 mean Amari is ≤0.009\le 0.009 for every dZd_Z; fitted b≈0.06b \approx 0.06; the escape from Theorem 6’s ε\varepsilon-covering bound is a 462×462\times improvement in sample count at dZ=8d_Z = 8.

4.9 Sparse-ICA (IMA) learnability — second SIC-C-c resolution (experiments/sparse_ica_learnability)

Same shape as Instrument 8 but with AA sparse-orthogonal at retention s∈{0.5,0.25}s \in \{0.5, 0.25\} — the independent-mechanism-analysis / sparse-mixing class of Gresele et al. (2021). Four pre-registered gates pass; fitted polynomial exponents b≈0.00b \approx 0.00 (s=0.5s = 0.5) and b≈0.51b \approx 0.51 (s=0.25s = 0.25), both well inside b≤3b \le 3. Confirms that two distinct inductive-bias classes each earn escape from Theorem 6’s ε\varepsilon-covering bound, so SIC-C-c’s “impossible without inductive bias” is a genuinely local statement — the pattern is that each identifiable-representation-learning class admits its own theorem, not that one privileged class does. Other classes (iVAE with auxiliary variables; interventional CRL) remain open as separate theorem-instrument pairs.

2.5d Concern as fibre geometry — companion paper (Theorems CG-1, CG-2)

A short companion paper Concern as Fibre Geometry (concern_as_fiber_geometry.pdf, 2026-08-03) develops the concern layer of §5.1 into two theorems:

Worked example on Instrument 4’s 4-bit Boolean world; rectangular- and triangular-loop holonomy matches predictions to 10−610^{-6}. The Trace-AI relevance is that CG-1 and CG-2 together give TRB Tracks 4 (Viability) and 5 (Navigability) their formal shape — see BENCHMARK_SPEC.md.


5. The extended program (ten further constructs)

Each is a target generated by the master object; those marked conjectural are not proved here.

  1. Concern as fibre geometry — a concern state reweights the compiler, Kc(dx∣z)∝eβUc(x,z)K(dx∣z)K_c(dx \mid z) \propto e^{\beta U_c(x,z)} K(dx \mid z), inducing an information geometry on concern states. Status update: now a theorem (CG-1 + CG-2), see §2.5d and the companion paper apps/site/concern_as_fiber_geometry.pdf.
  2. Conditional rate–distortion control limit (conjectural) — robust control of Y=f(X)Y = f(X) from a coarse spec Z=q(X)Z = q(X) requires H(Y∣Z)≈0H(Y \mid Z) \approx 0, or under tolerated distortion DD, extra addressed bits Bextra≥RY∣Z(D)B_{\text{extra}} \ge R_{Y \mid Z}(D).
  3. Abstraction frontier — the Pareto set trading task-sufficiency I(Y;X∣q(X))I(Y; X \mid q(X)), dynamical closure, cost H0(Z)H_0(Z), and control regret; an antichain that explains why two representations can both be “right” yet incomparable.
  4. Fibre audit — vary the allegedly irrelevant degrees of freedom while holding qq fixed and measure Δq(z)=sup⁡x,x′∈q−1(z)d(P(Y∣do⁡x),P(Y∣do⁡x′))\Delta_q(z) = \sup_{x, x' \in q^{-1}(z)} d(P(Y \mid \operatorname{do} x), P(Y \mid \operatorname{do} x')).
  5. Theory atlas — theories as charts MiM_i on contexts UiU_i with translations TijT_{ij}; test the cocycle Tjk∘Tij=TikT_{jk} \circ T_{ij} = T_{ik}.
  6. Compiler tomography — infer the shared compiler and compact specs from many (si,xi)(s_i, x_i) by MDL; then compiler ecology Kt+1=U(Kt,outcomes)K_{t+1} = U(K_t, \text{outcomes}).
  7. Causal semantics — two symbols are equivalent when they induce naturally equivalent update operators Ψm,c\Psi_{m,c} across independent contexts.
  8. Representation-repair calculus — a library mapping failure signatures to minimal structural lifts (scalar → operator, global norm → localized measure, quotient → restored fibre, static → path space, affine → projective, point → ensemble, non-composing → interface, symmetry → gauge-fix).
  9. Alignment as ensemble governance (conjectural) — target a viable region V⊆ZV \subseteq Z with Pr⁡[q(Xt)∈V∀t]≥1−δ\Pr[q(X_t) \in V \forall t] \ge 1 - \delta under a broad family of unresolved compiler/environment states; a fibre audit is the alignment evaluation.
  10. Autocatalytic artwork — St→KtEt→experienceKt+1S_t \xrightarrow{K_t} E_t \xrightarrow{\text{experience}} K_{t+1}: early movements teach the grammar by which later movements become legible.

6. Capstone: conscious and reliable agents, honestly

Two natural questions bear on this program — can agents be made conscious, and can they be made never wrong — and the disciplined answer to both is not in the literal sense. The strongest defensible construction factors as two nested systems.

Concerned Self-Modeling Core. An active, self-maintaining fixed point of the coarse-graining loop: a persistent agent that maintains world- and self-models, represents concern-weighted futures, broadcasts selected information, remembers commitments, predicts its own action consequences, performs false-credit tests on its own causal claims (exactly Instrument 3’s calibration metric), and reports uncertainty. It operationalizes proposed consciousness indicators — and that is the ceiling of the claim: functional selfhood ⇏\not\Rightarrow subjective consciousness.

Proof-Carrying Reliability Shell. The fibration with a verifier on the counit: convert goals to contracts, separate observation from inference, attach provenance to every claim, and commit an action only under machine-checkable evidence: execute⁡(a)⇔V(s,a,φ)=PASS\operatorname{execute}(a) \iff V(s, a, \varphi) = \operatorname{PASS}, else abstain/ask/simulate/escalate. Its correctness envelope is a fibre audit (§5.4). Verified-in-a-bounded-domain ⇏\not\Rightarrow universal infallibility: verification certifies the formal statement, not that it captures intent.

The honest breakthrough available now is therefore not conscious, perfect agents, but agents with experimentally measurable selfhood and formally bounded error — two separate, non-inflated claims.


7. Limitations

The eight instruments are toys by design (finite and grid-quantised Boolean and grid worlds, tiny Markov systems, lossless encoders; Instrument 8 is fixed-seed Monte Carlo). The master fibration (q,K)(q, K) is derived (Theorem 1 gives it as a mathematical object for any well-posed task on a standard Borel space; Theorem 2 parameterises it (Instrument 7 witnesses it exactly); Proposition 3 gives the categorical restatement; Theorem 4 makes cross-task stability equivalent to a shared Markov screen; Theorems 5 and 6 pin down discrete- and continuous-case sample complexity; Theorem 7 gives polynomial-in-dZd_Z learnability inside the linear-ICA hypothesis class). The residual content is:


8. How this relates to Trace AI

Trace AI is the SIC applied to cognition, with a concrete choice for each abstract variable:

SIC object Trace AI instantiation Where
ambient space XX model’s compute state ztz_t whitepaper §2
latent ZZ self-representation rtr_t (a few legible slots) whitepaper §2, §4.2
coarse-graining qq decoder hϕ:zt↦rth_\phi : z_t \mapsto r_t whitepaper §2
compiler KK any of the subject-kinds (human, model, hybrid) that realize the schema schema/reasoning_trace.schema.json
task family {Yα}\{Y_\alpha\} the counterfactual set ℛ*\mathcal{R}^* + paired continuations whitepaper §4.4; COLLECTION_PROTOCOL.md
Instrument 3 (agency) TRB Track 3 (self-attribution) — the exact-solvable formalization BENCHMARK_SPEC.md §3 Track 3
Instrument 1 (Fiber Finder) the “schema follows the data” discipline (representation search done by hand) STRUCTURAL_INTELLIGENCE.md §3.1
SIC-B (cross-task stability) schema portability across substrates (human, model), Observation 1 in whitepaper §6.4 whitepaper §6.4
SIC-C-c residual → linear-ICA class Trace AI’s L2 sample-complexity limit (∼103\sim 10^3 human counterfactual pairs) is a working budget; Theorem 7 says polynomial-in-drd_{r} learnability is a theorem inside the linear-ICA hypothesis class. The strategic implication: if rtr_t can be arranged (by architecture + training) so its slot components are statistically independent and non-Gaussian, the L2 budget shrinks from exponential to polynomial in the number of slots. That reframes the collection operation from “brute-force scale” to “engineer the bottleneck for identifiability”. whitepaper §8 L2
Instrument 7 (rate–distortion pair) Exact-solvable model for the whitepaper §2.1 information-bottleneck claim — the rtr_t readout as a coarse-graining at a distortion budget whitepaper §2.1
Instrument 8 (linear-ICA learnability) Numerical existence proof that inductive bias flips drd_{r}-scaling from exponential to polynomial. If Trace AI adopts an ICA-flavored decoder for hϕh_\phi, the sample-complexity story changes materially whitepaper §8 L2; COLLECTION_PROTOCOL.md §6

The whitepaper is Trace AI at the level of what to build; this paper is why the object exists as a mathematical object at all. Read the SIC when a reviewer asks “is your latent even a well-defined thing to be looking for”; read the whitepaper when they ask “what is the training signal that finds it.”


v0.2 (integrated 2026-08-03), tracks paper v4. See DECISIONS.md (D28, D30) for provenance. Reproduction of the eight instruments lives in the sibling research repository jawauntb/research-derived-experiments — all experiments/ paths in the instrument tables above resolve there, including the new experiments/rate_distortion_pair and experiments/linear_ica_learnability. Exact-recovery numbers in §4.5–§4.8 are copied verbatim from that repository’s outputs at the paper’s commit. The originating research note is at notes/structural_intelligence_conjecture.md; the working-note sibling of the present document is SIC_RESEARCH_PROGRAM.md.