Latent reasoning: research directions

Each node is a paper or a hub concept. Colours are literatures; solid edges are real connections; dashed red edges are connections that do not yet exist: searched for and not found as of 16 September 2026, which is to say the research agenda. Dotted purple edges are direct rebuttals, which this field has an unusual number of. Dashed edges are dated, scoped search claims rather than proofs of absence, and are withdrawn when someone finds the thing. Click a node to open the work; drag to rearrange; hover for the one-line story.

builds on / descends from same object, two readings endpoint of the dial rebuts / contradicts GAP: does not exist

Corrections welcome: file an issue on microprediction/latentreasoning.

The gaps, spelled out

Dashed edges are claims about the literature, so they are stated here in full with a confidence. An absence is only interesting if someone actually looked, and only honest if it is withdrawn when someone finds the thing.

Five claimed gaps have been withdrawn or narrowed since 16 September 2026 after a reader supplied counterexamples, all of which checked out. A joint scaling law for looped models, a joint treatment of padding and looping, reinforcement learning over latent diffusion with verifiable rewards, and state-space operators inside a latent recursive scaffold all exist and now appear on the map as real connections. What follows is a dated, scoped search claim: not a proof of absence.

  1. Interpretability results have not reached the deployment question. The finding that latent traces decode 65–93% of the time and the controversy over a production model whose card records reduced monitorability do not cite each other. Nobody has applied the decoding methods to the deployed case. Confidence: high. The most policy-relevant gap here.
  2. No independent measurement of monitorability loss. OpenAI has published its own finding. It is self-reported, not run against a shared public suite, and unreproduced. Open recurrent-depth checkpoints exist and such a suite exists. Nobody has run one on the other. Confidence: high.
  3. No shared benchmark for latent computation. Narrowed. A purpose-built instrument does exist: a 4,000-item benchmark that forces internal inference by requiring the answer to be signalled through the language of the first response token, so it cannot be externalised as text. What is still missing is community convergence: it measures internal reasoning in stock models rather than discriminating between latent-reasoning architectures, its authors cannot rule out heuristic exploitation, and critique papers continue to build bespoke diagnostics instead of adopting a common one. Confidence: moderate, for the narrowed claim.
  4. Which auxiliary objective, and why. Revised. This previously claimed that policy-gradient objectives assume single-step decisions and so cannot train latent recurrence. That is false: for a differentiable recurrence with gradients retained, a terminal reward's total derivative already includes every unrolled update. What is true is narrower: standard objectives have empirically underperformed on looped models, and the causes on offer (long Jacobian products, sparse reward, truncation or detaching in implementations) are optimisation problems, not invalidity. The open question is which auxiliary objectives and truncation/halting choices actually improve gradient quality and compute efficiency under controlled comparison. Confidence: moderate, and now a question rather than an absence.
  5. The superposition arguments still stand apart. Narrowed. Padding and looping are now analysed together with threshold-circuit characterisations, so the broad claim that the toolkits never meet was wrong. What has not been connected is the superposition-and-parallel-trace line of argument to either of them. Confidence: moderate.
  6. Weights-as-memory and cache-as-scratchpad have never been compared head to head. Test-time training compresses context into weights; KV-cache distillation compresses it into the cache. No paper evaluates them against each other on reasoning tasks. Confidence: medium.
  7. State-space recursion is untested beyond tiny scale. Narrowed. Mamba-2 hybrid operators have been placed inside a latent recursive scaffold, so the blanket claim was false. But that work is ~7M parameters on abstract reasoning tasks, and nothing establishes a linear-state model as the reasoning medium at language-model scale. Confidence: medium, for the narrowed claim only.

Found a counterexample to any of these? That is the most useful possible contribution: open an issue. The four withdrawn above arrived that way.