Latent reasoning: research directions
Each node is a paper or a hub concept. Colours are literatures; solid edges are real
connections; dashed red edges are connections that do not yet exist: searched for and not found as of 16 September 2026, which is to say the research agenda. Dotted purple
edges are direct rebuttals, which this field has an unusual number of. Dashed edges are dated,
scoped search claims rather than proofs of absence, and are withdrawn when someone finds the
thing. Click a node to open the work; drag to rearrange; hover for the one-line story.
builds on / descends from
same object, two readings
endpoint of the dial
rebuts / contradicts
GAP: does not exist
Corrections welcome: file an issue on
microprediction/latentreasoning.
The gaps, spelled out
Dashed edges are claims about the literature, so they are stated here in full with a
confidence. An absence is only interesting if someone actually looked, and only honest if
it is withdrawn when someone finds the thing.
Five claimed gaps have been withdrawn or narrowed since 16 September 2026 after a reader supplied
counterexamples, all of which checked out. A joint scaling law for looped models, a joint
treatment of padding and looping, reinforcement learning over latent diffusion with verifiable
rewards, and state-space operators inside a latent recursive scaffold all exist and now appear
on the map as real connections. What follows is a dated, scoped search claim: not a
proof of absence.
-
Interpretability results have not reached the deployment question.
The finding that latent traces decode 65–93% of the time and the controversy over a
production model whose card records reduced monitorability do not cite each other. Nobody has
applied the decoding methods to the deployed case. Confidence: high. The most
policy-relevant gap here.
-
No independent measurement of monitorability loss.
OpenAI has published its own finding. It is self-reported, not run against a shared public
suite, and unreproduced. Open recurrent-depth checkpoints exist and such a suite exists. Nobody has run one on the other. Confidence: high.
-
No shared benchmark for latent computation.
Narrowed. A purpose-built instrument does exist: a 4,000-item benchmark that forces
internal inference by requiring the answer to be signalled through the language of the first
response token, so it cannot be externalised as text. What is still missing is community
convergence: it measures internal reasoning in stock models rather than discriminating
between latent-reasoning architectures, its authors cannot rule out heuristic exploitation,
and critique papers continue to build bespoke diagnostics instead of adopting a common one.
Confidence: moderate, for the narrowed claim.
-
Which auxiliary objective, and why.
Revised. This previously claimed that policy-gradient objectives assume single-step
decisions and so cannot train latent recurrence. That is false: for a differentiable
recurrence with gradients retained, a terminal reward's total derivative already includes
every unrolled update. What is true is narrower: standard objectives have
empirically underperformed on looped models, and the causes on offer (long Jacobian
products, sparse reward, truncation or detaching in implementations) are optimisation
problems, not invalidity. The open question is which auxiliary objectives and
truncation/halting choices actually improve gradient quality and compute efficiency under
controlled comparison. Confidence: moderate, and now a question rather than an
absence.
-
The superposition arguments still stand apart.
Narrowed. Padding and looping are now analysed together with threshold-circuit
characterisations, so the broad claim that the toolkits never meet was wrong. What has not
been connected is the superposition-and-parallel-trace line of argument to either of them.
Confidence: moderate.
-
Weights-as-memory and cache-as-scratchpad have never been compared head to head.
Test-time training compresses context into weights; KV-cache distillation compresses it into
the cache. No paper evaluates them against each other on reasoning tasks.
Confidence: medium.
-
State-space recursion is untested beyond tiny scale.
Narrowed. Mamba-2 hybrid operators have been placed inside a latent recursive
scaffold, so the blanket claim was false. But that work is ~7M parameters on abstract
reasoning tasks, and nothing establishes a linear-state model as the reasoning medium at
language-model scale. Confidence: medium, for the narrowed claim only.
Found a counterexample to any of these? That is the most useful possible contribution: open an issue. The four
withdrawn above arrived that way.