Reflect-Search · architecture leaderboard

One task suite × many models × the six discipline rules.

Reflect-Search is the framework's answer to "which architecture serves reflective reasoning best" — treated as an empirical question a partner can run today. Same fixed task suite, same JSON contract, same six-rule discipline audit. Score reveals which model architectures already satisfy the framework's constraints without training and which don't. The RESEARCH_DIRECTIONS §1 full form.

Honest scope. This is a snapshot comparison, not a training result. What it measures: how well each model follows the two-hook contract, emits a legible r_t, and satisfies the six discipline rules without any ℒ_gov training. A high score here means the model is a strong proxy-adapter substrate; a low score means it either can't hold the JSON envelope or falls into helplessness / hedging / paraphrase-under-override. Both are useful facts about the architecture. Neither is a Track-1 governance score (which needs paired human Φ).
API key Not saved. Server accepts your key per-request and never stores it.

Task suite

Models

Aggregate rankings

Each cell fires /api/reflect/turn in soft-bottleneck mode with the task's prompt and the selected model. Scores are the server-side six-rule discipline audit (see how-reflect-works). OpenRouter models require OPENROUTER_API_KEY on the server (already set on this deployment). Results persist to your browser's localStorage so a page reload preserves them; use Clear results to start over. This is not the TRB benchmark — see Benchmark for the definitive governance-gap score, which requires paired human counterfactuals.