Recognition of insufficiency during generative reasoning
The broader phenomenon is awareness of unawareness: a reasoning system may become sensitive that its current basis is inadequate before it knows what is missing or how to repair it.
How does a reasoning system recognize that its current basis is inadequate, and how does that recognition control what it does next?
This means that the system detects a problem before it knows the repair, which we call recognition before identification.
Why would we need such a thing?
Solving the object-level problem is only half of what an autonomous reasoner has to do. It also has to decide whether to commit, continue thinking, retrieve information, ask for help, try another approach, use a tool, or stop.
Something has to control that allocation.
Confidence already does some of this. Recent work from Google DeepMind
and Princeton showed that internal confidence is not merely something
models report: manipulating confidence representations changes whether
a model answers or abstains. Google DeepMind · Princeton · Nature
Now, confidence answers questions like how likely is my current answer to be correct? But there are other questions, such as is the reasoning process I am currently using capable of resolving this yet?
A system can have substantial compute available and still repeatedly pursue the same bad interpretation. More computation is not necessarily what it needs. It may need some signal that the current reasoning is failing to capture something necessary.
Recognition before identification
Suppose a language model is solving a semantic problem and repeatedly uses an obvious meaning of a word.
The relevant alternative meaning may already exist in the model's parameters but has not become available in the current reasoning.
The interesting case is not simply that the model eventually generates another meaning. It is whether, before identifying that meaning, there is evidence that the model has become sensitive to the inadequacy of what it currently has.
The model does not yet know that BANK also means X. It may only have something weaker: something about what I am doing is failing to account for the problem. That is the phenomenon I want to study.
Humans have related metacognitive phenomena. Koriat's work on feeling
of knowing shows that people can make judgments about unavailable
memories before retrieving the target, potentially from partial or
accessible information. PubMed Ackerman & Thompson's work on
metareasoning studies how monitoring ongoing reasoning affects
cognitive control.
Trends in Cognitive Sciences
And there is a real gap around that question
Internal confidence causally influences whether a model answers or
abstains. Google DeepMind · Princeton · Nature
There is now work on thought sufficiency: deciding
when an LLM has reasoned enough and can stop. arXiv
There is work on inferential boundary awareness:
detecting that necessary premises are missing instead of fabricating
them. arXiv
There is work explicitly called Knowing What's Missing,
but their method asks the model to first hypothesize what specific
information is missing and then verify its absence. arXiv
There is internal metacognitive monitoring for whether tools
are necessary. arXiv
Decision-theoretic work asks whether agents should act or request more
information. arXiv
And reasoning behaviors such as uncertainty, backtracking, and
hypothesis testing correspond to steerable directions inside reasoning
models. arXiv
So this is not an isolated idea. The field is converging on the broader problem of knowing when reasoning is adequate.
These are close, but the question I care about comes earlier: can there be a useful monitoring signal before the system knows what is missing? I have not found work that cleanly answers it for ongoing generative reasoning.
Why it could matter
If such monitoring exists, it could support several forms of adaptive control.
A coding agent may recognize that its current hypothesis does not explain the observations before knowing the bug. A medical reasoner may recognize that its current explanation leaves findings unexplained before identifying the omitted diagnosis. A research agent may recognize that its current evidence cannot support a conclusion before knowing which source it needs. A planning agent may recognize that its plan has failed to account for something before knowing exactly what that thing is.
The same monitoring process could therefore matter for adaptive computation, information seeking, reframing, tool use, and abstention, rather than treating each as an unrelated heuristic.
What I am testing now
Clusterwords gives me a controlled environment for creating situations where a competent reasoner becomes trapped by a plausible interpretation while another relevant meaning or structure has not yet been identified.
The current experiments have already produced such cases. On engineered 32-word boards, Codex can repeatedly commit to the same wrong residual split despite having substantial resources remaining. Fable currently solves these boards almost perfectly, while Haiku often fails for broader competence reasons.
The interesting question is therefore not merely why a model makes an error, but what makes a competent model decide not to act, and why.
A model can refrain because it is uncertain, because it is following a conservative policy, because it is stuck, or because it has actually detected something inadequate about its reasoning. Those possibilities need to be separated experimentally.
The first behavioral experiments can therefore reduce the control decision to ACT versus NOT ACT, while manipulating the reasoning situation around that decision. Later experiments can examine the reasoning trajectory, hidden representations, and whether recognition can be induced or suppressed without providing the missing solution.
What uncertainty is doing here
Uncertainty is not the main research question. Confidence, semantic uncertainty, self-consistency, internal confidence representations, and related measures are useful controls because they provide simpler explanations for behavior.
The goal is not to prove that "insufficiency is different from uncertainty."
The goal is to determine what allows a reasoner to detect the limits of its ongoing reasoning before it has found the missing solution, and whether that detection changes subsequent control.
What could turn out to be true
There may be no special representation of insufficiency. What looks like recognition could be completely explained by confidence, conflict, error signals, progress monitoring, search heuristics, or some combination of them.
Models may also never recognize inadequacy before identification; they may simply continue generating possibilities until something works.
Or recognition may exist behaviorally without corresponding to one clean internal variable.
Those are all informative outcomes.
The project therefore does not assume that I have discovered a new cognitive variable and need to prove that it exists.
The questions
- Can monitoring detect inadequacy before the missing content becomes available?
- What would make a competent model decide not to act when its current reasoning is inadequate?
- Why did it decide not to act?
- What allows a reasoning system to detect that its ongoing reasoning is inadequate before it has found the missing content, and how, if at all, is that information used to control what it does next?
That is the computational problem I am currently studying.
- How do humans recognize that the possibilities currently available during reasoning may be insufficient?
- What information in the dynamics of ongoing reasoning provides evidence that the currently available possibilities may not be enough?
- How does a reasoner distinguish "I have not solved this yet" from "what I currently have may itself be insufficient"?
- Can recognition of insufficiency occur before a reasoner can identify what possibility is missing?
- Does recognition of insufficiency depend on a single diagnostic signal, or on the integration of multiple signals over time?
- How do failure, confidence, progress, retrieval, and changes in candidate possibilities contribute to recognition of insufficiency?
- Does the history of reasoning matter (repeated failure, repeated return to the same possibilities), or is the current reasoning state sufficient?
- How does the partial development of a possibility affect judgments of insufficiency? Does a vague but promising possibility feel different from no promising possibility at all?
- When a useful possibility has failed to enter consideration, are there detectable signatures in the reasoning process before that possibility is eventually discovered?
- Can a computational model predict when a person will conclude that their currently available possibilities may be insufficient?
- Can that model distinguish recognition of insufficiency from low confidence, difficulty, retrieval failure, conflict, or simply needing more time?
- What causes recognition of insufficiency to change subsequent reasoning? Does it lead to more search, different search, refinement of partial possibilities, retrieval, or representational change?
- When does additional search reflect recognition of insufficiency rather than ordinary persistence?
- Do people differ systematically in their ability to recognize insufficiency, and do those differences predict successful discovery of previously unconsidered possibilities?
- Does experience with a task change recognition of insufficiency by changing which possibilities become available, by changing how their adequacy is monitored, or both?
- Do contemporary machine reasoners recognize when their currently available possibilities may be insufficient, or do they primarily continue evaluating and elaborating what they have already generated?
- When humans and machines fail because a useful possibility never entered consideration, do they show different signatures before the failure?
- Do machines use signals analogous to those that predict recognition of insufficiency in humans?
- Can a computational mechanism derived from human recognition of insufficiency predict machine failures that confidence alone cannot?
- Can introducing such a mechanism improve a machine reasoner's ability to change its search when its currently available possibilities are inadequate?
A possible mechanism: generative reasoning
Gt and Ct are the proposed generative mechanism through which Ît may arise. They are one possible mechanistic account of what happens, especially in sequential reasoning; they are no longer the definition of the phenomenon itself. The question they address:
What signals during ongoing reasoning lead an agent to recognize that the possibilities currently available may be insufficient?
Three working objects (abstractions, not brain modules):
Gt (generative reasoning): the processes that determine which candidate possibilities become available and how they change during reasoning.
Ct (currently available possibilities): what is on the table right now. A possibility can be fully specified, partial, vague, or weakly activated. Ct is not assumed to be a literal discrete set.
At (recognition of insufficiency): the metacognitive sensitivity that what is currently available may not be enough. In the notation above this is Ît:
At ⟶ Ît
Î is a working target rather than an assumed cognitive module or established distinct state. Whether recognition of insufficiency can be distinguished from uncertainty, low confidence, impasse, and related states is the first empirical question, not an assumption.
At ⇏ c* ∈ Ct
Recognition can happen before the missing possibility is known.
Uncertainty vs insufficiency, intuitively
Uncertainty: in the earlier framing, "I have plausible alternatives, but I do not know which is correct; the uncertainty concerns alternatives that are already represented." That is one form of uncertainty, uncertainty over represented candidate solutions. It is a special case of 𝒰, not the general definition.
Recognition of insufficiency: I have reason to suspect that what is currently available may not contain what I need. The intuition still holds. Experimentally it is measured against I, which is known by construction.
Recognition vs control
What Î is not
Recognition of insufficiency is different from simply making an error, low confidence, being stuck, already knowing the missing alternative, and an automatic strategy switch. Those can be signals, correlates, causes, or consequences of Î, but they are not the thing itself. Whether it is different from uncertainty between represented alternatives is kept, but as the empirical question above rather than an assertion.
Recognition vs control response
Î is not defined by what it triggers. Given recognition, the control response can vary:
Î ⟶ continue · retrieve · generate alternatives · reframe · search externally · nothing
Which response wins (including committing anyway) is a downstream question about costs and task demands, not the definition of Î. Monitoring and control stay separate.
The reasoning trajectory (Clusterwords)
Clusterwords (the game on the home page) lets us observe something richer than commit-versus-don't-commit:
C0 ⟶ C1 ⟶ C2 ⟶ ⋯
Possibilities are generated, partially formed, refined, abandoned, recombined, and sometimes recognized as inadequate. Suppose the board hides EAGLE, BIRDIE, BOGEY, PAR:
"these feel related" ⟶ "these four belong together somehow" ⟶ "golf-related" ⟶ "golf scoring terms"
The question is not where exactly this becomes a hypothesis. It is: how does this possibility evolve inside Ct, and what does its evolution tell the reasoner about whether what is currently available is adequate? The guesses give observable anchors in that trajectory. The trajectory is still, perhaps eventually, the most interesting part, but it creates too many confounds for a first experiment. So it comes second:
first: one-shot / closed information ⟶ later: sequential
If the first experiment suggests that insufficiency is meaningful beyond uncertainty, I becomes It, Î becomes Ît, and the system can take an action:
at ∈ {seek information, act}
Then the original question can be tested: does recognizing insufficiency provide a better signal for deciding when to seek additional information than uncertainty alone? That is where explore ⟷ exploit enters.
Bounded agent
A bounded agent cannot consider every possible explanation, action, or future consequence due to limited resources.
How the framing evolved history
Version A: a literal hypothesis set.
ℋt = {h₁, h₂, h₃}, P(h* ∉ ℋt) > 0
"How do I know h* isn't in my set?"
Version B: but what even counts as h? Can hypotheses be partial?
h = these four share r, r = ?
Version C: don't require a literal hypothesis set.
Gt ⟷ Ct ⟶ At
Possibilities of different degrees of specification become available during generative reasoning. The project studies what signals tell the reasoner that what is currently available may be insufficient, without first solving where a hypothesis begins.
Current: separate actual insufficiency from its recognition, and both from uncertainty. Version C was missing the distinction between insufficiency and recognition: At is Ît, and I was not there at all.
I ⟶ Î phenomenon: can the system recognize objective insufficiency?
𝒰 ⟷? I, Î does uncertainty explain recognition, or can they dissociate?
Gt ⟷ Ct ⟶ Ît later mechanism: what dynamics of candidate generation produce recognition?
Not a rejection of the earlier framework; a reorganization. The construct-validity layer that had been missing underneath it.
Questions we still need to answer (Ct: what's on the table)
- What is the representational object? Is it a rule, causal model, gist, category, explanation, structured relation?
- Can it be incomplete or coarse? Can the person have "something about these belongs together" without the full explanation?
- Can that partial representation guide a deliberate decision?
- Can it later be refined while preserving earlier content?
- When does the literature treat it as the same representation becoming more specific versus a new hypothesis replacing the old one?
- What behavioral evidence shows that the representation is actually available to the person?
- What counts as a currently available possibility in a language model? Must a possibility have been explicitly generated into context by the time of a decision, or can an unexpressed internal representation count as available?