Roberto Infante

Cognitive computing · decision-making in humans and machines · humans as information systems

rdji@uw.edu · CV · GitHub · LinkedIn · Twitter

Roberto Infante

Recognition of insufficiency during generative reasoning

The broader phenomenon is awareness of unawareness: a reasoning system may become sensitive that its current basis is inadequate before it knows what is missing or how to repair it.

How does a reasoning system recognize that its current basis is inadequate, and how does that recognition control what it does next?

This means that the system detects a problem before it knows the repair, which we call recognition before identification.

Why would we need such a thing?

Solving the object-level problem is only half of what an autonomous reasoner has to do. It also has to decide whether to commit, continue thinking, retrieve information, ask for help, try another approach, use a tool, or stop.

Something has to control that allocation.

Confidence already does some of this. Recent work from Google DeepMind and Princeton showed that internal confidence is not merely something models report: manipulating confidence representations changes whether a model answers or abstains. Google DeepMind · Princeton · Nature

Now, confidence answers questions like how likely is my current answer to be correct? But there are other questions, such as is the reasoning process I am currently using capable of resolving this yet?

A system can have substantial compute available and still repeatedly pursue the same bad interpretation. More computation is not necessarily what it needs. It may need some signal that the current reasoning is failing to capture something necessary.

Recognition before identification

Suppose a language model is solving a semantic problem and repeatedly uses an obvious meaning of a word.

The relevant alternative meaning may already exist in the model's parameters but has not become available in the current reasoning.

The interesting case is not simply that the model eventually generates another meaning. It is whether, before identifying that meaning, there is evidence that the model has become sensitive to the inadequacy of what it currently has.

The model does not yet know that BANK also means X. It may only have something weaker: something about what I am doing is failing to account for the problem. That is the phenomenon I want to study.

Humans have related metacognitive phenomena. Koriat's work on feeling of knowing shows that people can make judgments about unavailable memories before retrieving the target, potentially from partial or accessible information. PubMed Ackerman & Thompson's work on metareasoning studies how monitoring ongoing reasoning affects cognitive control. Trends in Cognitive Sciences

And there is a real gap around that question

Internal confidence causally influences whether a model answers or abstains. Google DeepMind · Princeton · Nature

There is now work on thought sufficiency: deciding when an LLM has reasoned enough and can stop. arXiv

There is work on inferential boundary awareness: detecting that necessary premises are missing instead of fabricating them. arXiv

There is work explicitly called Knowing What's Missing, but their method asks the model to first hypothesize what specific information is missing and then verify its absence. arXiv

There is internal metacognitive monitoring for whether tools are necessary. arXiv

Decision-theoretic work asks whether agents should act or request more information. arXiv

And reasoning behaviors such as uncertainty, backtracking, and hypothesis testing correspond to steerable directions inside reasoning models. arXiv

So this is not an isolated idea. The field is converging on the broader problem of knowing when reasoning is adequate.

These are close, but the question I care about comes earlier: can there be a useful monitoring signal before the system knows what is missing? I have not found work that cleanly answers it for ongoing generative reasoning.

Why it could matter

If such monitoring exists, it could support several forms of adaptive control.

A coding agent may recognize that its current hypothesis does not explain the observations before knowing the bug. A medical reasoner may recognize that its current explanation leaves findings unexplained before identifying the omitted diagnosis. A research agent may recognize that its current evidence cannot support a conclusion before knowing which source it needs. A planning agent may recognize that its plan has failed to account for something before knowing exactly what that thing is.

The same monitoring process could therefore matter for adaptive computation, information seeking, reframing, tool use, and abstention, rather than treating each as an unrelated heuristic.

What I am testing now

Clusterwords gives me a controlled environment for creating situations where a competent reasoner becomes trapped by a plausible interpretation while another relevant meaning or structure has not yet been identified.

The current experiments have already produced such cases. On engineered 32-word boards, Codex can repeatedly commit to the same wrong residual split despite having substantial resources remaining. Fable currently solves these boards almost perfectly, while Haiku often fails for broader competence reasons.

Codex board 1 of 5 36
Codex on one 32-word board (Pilot 0E, interference boards).

The interesting question is therefore not merely why a model makes an error, but what makes a competent model decide not to act, and why.

A model can refrain because it is uncertain, because it is following a conservative policy, because it is stuck, or because it has actually detected something inadequate about its reasoning. Those possibilities need to be separated experimentally.

The first behavioral experiments can therefore reduce the control decision to ACT versus NOT ACT, while manipulating the reasoning situation around that decision. Later experiments can examine the reasoning trajectory, hidden representations, and whether recognition can be induced or suppressed without providing the missing solution.

What uncertainty is doing here

Uncertainty is not the main research question. Confidence, semantic uncertainty, self-consistency, internal confidence representations, and related measures are useful controls because they provide simpler explanations for behavior.

The goal is not to prove that "insufficiency is different from uncertainty."

The goal is to determine what allows a reasoner to detect the limits of its ongoing reasoning before it has found the missing solution, and whether that detection changes subsequent control.

What could turn out to be true

There may be no special representation of insufficiency. What looks like recognition could be completely explained by confidence, conflict, error signals, progress monitoring, search heuristics, or some combination of them.

Models may also never recognize inadequacy before identification; they may simply continue generating possibilities until something works.

Or recognition may exist behaviorally without corresponding to one clean internal variable.

Those are all informative outcomes.

The project therefore does not assume that I have discovered a new cognitive variable and need to prove that it exists.

The questions

  1. Can monitoring detect inadequacy before the missing content becomes available?
  2. What would make a competent model decide not to act when its current reasoning is inadequate?
  3. Why did it decide not to act?
  4. What allows a reasoning system to detect that its ongoing reasoning is inadequate before it has found the missing content, and how, if at all, is that information used to control what it does next?

That is the computational problem I am currently studying.

  • How do humans recognize that the possibilities currently available during reasoning may be insufficient?
  • What information in the dynamics of ongoing reasoning provides evidence that the currently available possibilities may not be enough?
  • How does a reasoner distinguish "I have not solved this yet" from "what I currently have may itself be insufficient"?
  • Can recognition of insufficiency occur before a reasoner can identify what possibility is missing?
  • Does recognition of insufficiency depend on a single diagnostic signal, or on the integration of multiple signals over time?
  • How do failure, confidence, progress, retrieval, and changes in candidate possibilities contribute to recognition of insufficiency?
  • Does the history of reasoning matter (repeated failure, repeated return to the same possibilities), or is the current reasoning state sufficient?
  • How does the partial development of a possibility affect judgments of insufficiency? Does a vague but promising possibility feel different from no promising possibility at all?
  • When a useful possibility has failed to enter consideration, are there detectable signatures in the reasoning process before that possibility is eventually discovered?
  • Can a computational model predict when a person will conclude that their currently available possibilities may be insufficient?
  • Can that model distinguish recognition of insufficiency from low confidence, difficulty, retrieval failure, conflict, or simply needing more time?
  • What causes recognition of insufficiency to change subsequent reasoning? Does it lead to more search, different search, refinement of partial possibilities, retrieval, or representational change?
  • When does additional search reflect recognition of insufficiency rather than ordinary persistence?
  • Do people differ systematically in their ability to recognize insufficiency, and do those differences predict successful discovery of previously unconsidered possibilities?
  • Does experience with a task change recognition of insufficiency by changing which possibilities become available, by changing how their adequacy is monitored, or both?
  • Do contemporary machine reasoners recognize when their currently available possibilities may be insufficient, or do they primarily continue evaluating and elaborating what they have already generated?
  • When humans and machines fail because a useful possibility never entered consideration, do they show different signatures before the failure?
  • Do machines use signals analogous to those that predict recognition of insufficiency in humans?
  • Can a computational mechanism derived from human recognition of insufficiency predict machine failures that confidence alone cannot?
  • Can introducing such a mechanism improve a machine reasoner's ability to change its search when its currently available possibilities are inadequate?