Moonshot AI released Kimi K3 on July 16: a 2.8 trillion parameter mixture-of-experts model with a 1 million token context window and the company’s Kimi Delta Attention architecture, per Reuters the largest model ever slated for an open-weight release. Semiconductor stocks fell on the news, and by July 20 Moonshot was pausing new subscriptions because it could not serve the demand.
K3 tops several coding and agentic leaderboards, undercuts US frontier pricing, and promises full weights by July 27. Markets read it as a second DeepSeek moment: proof that the efficiency frontier keeps moving, and a question mark over how much compute the frontier actually requires.
What the benchmarks actually say
The claims hold up, with a caveat: they are benchmark-specific, not universal dominance. K3 ranks first on frontend-coding arenas and performs at or above US frontier models on several agentic and long-horizon coding evaluations, including Terminal-Bench 2.1 and SWE-style suites, according to Moonshot’s materials and third-party evaluations from Artificial Analysis and Arena. On broad knowledge and reasoning it is competitive rather than leading. API pricing lands around $3 per million input tokens and $15 per million output tokens through providers such as OpenRouter, well under closed-frontier list prices.
We profiled Moonshot in January, when Kimi was quietly becoming the default Chinese model inside Western coding tools. K3 is that trajectory reaching the frontier. It is also the sharpest data point yet in the pattern we tracked when China’s AI stack shipped a 2.7T model, a robotics brain and a new chip in one week.
Why chips sold off
Bloomberg’s read is that K3 may be more about memory than compute: its architecture leans on efficient attention over long contexts, shifting the bottleneck toward high-bandwidth memory rather than raw FLOPs. That rhymes with our own reporting that memory and megawatts, not GPUs, now cap the AI buildout. If frontier capability keeps getting cheaper to serve, the market’s question is not whether AI demand is real, it is which layer of the stack captures it.
The signal
Two things are true at once. Demand for K3 was strong enough that Moonshot paused paid signups within four days. And the selloff says investors no longer assume every capability gain translates into proportional silicon spend. The weights are not out yet: until July 27 this is an API launch with an open-weight promise, and the technical report is still pending. But if Moonshot delivers, the open-weight frontier will have jumped from the 975B-parameter class to nearly 3 trillion in under two weeks, and the gap between open and closed will be measured in weeks, not years.
AI and business tech news, verified by a physician who reads the filings. One email a week, no noise.
