IMAGE CREDITS: COREWIRE

Moonshot’s Kimi K3 is the biggest open-weight bet ever, and chip stocks fell on it

Moonshot AI’s Kimi K3: 2.8 trillion parameters, 1M context, top coding-arena scores, $3/$15 pricing, weights promised July 27. Semis sold off, and Moonshot paused signups on demand.

Moonshot AI released Kimi K3 on July 16: a 2.8 trillion parameter mixture-of-experts model with a 1 million token context window and the company’s Kimi Delta Attention architecture, per Reuters the largest model ever slated for an open-weight release. Semiconductor stocks fell on the news, and by July 20 Moonshot was pausing new subscriptions because it could not serve the demand.

K3 tops several coding and agentic leaderboards, undercuts US frontier pricing, and promises full weights by July 27. Markets read it as a second DeepSeek moment: proof that the efficiency frontier keeps moving, and a question mark over how much compute the frontier actually requires.

What the benchmarks actually say

The claims hold up, with a caveat: they are benchmark-specific, not universal dominance. K3 ranks first on frontend-coding arenas and performs at or above US frontier models on several agentic and long-horizon coding evaluations, including Terminal-Bench 2.1 and SWE-style suites, according to Moonshot’s materials and third-party evaluations from Artificial Analysis and Arena. On broad knowledge and reasoning it is competitive rather than leading. API pricing lands around $3 per million input tokens and $15 per million output tokens through providers such as OpenRouter, well under closed-frontier list prices.

We profiled Moonshot in January, when Kimi was quietly becoming the default Chinese model inside Western coding tools. K3 is that trajectory reaching the frontier. It is also the sharpest data point yet in the pattern we tracked when China’s AI stack shipped a 2.7T model, a robotics brain and a new chip in one week.

Why chips sold off

Bloomberg’s read is that K3 may be more about memory than compute: its architecture leans on efficient attention over long contexts, shifting the bottleneck toward high-bandwidth memory rather than raw FLOPs. That rhymes with our own reporting that memory and megawatts, not GPUs, now cap the AI buildout. If frontier capability keeps getting cheaper to serve, the market’s question is not whether AI demand is real, it is which layer of the stack captures it.

Open weights are not free compute

For enterprise buyers, a 2.8 trillion parameter open model deserves a cold look at the serving math. Open weights eliminate the license fee, not the infrastructure: a model this size runs across multiple nodes of high-end accelerators with enormous memory footprints, and the operational expertise to serve it reliably is scarcer than the hardware. In practice, most organizations will consume K3 the same way they consume closed models, through hosted API providers, which is why the serving layer keeps attracting capital on the thesis that someone has to run all these free weights.

What open weights do change is negotiating leverage and exit rights. A team building on K3 through one provider can move to another, or eventually in-house, without rewriting its product, and that portability disciplines pricing across the whole market. Closed-model vendors have already responded by making pricing itself the product. The strategic question for buyers is no longer just which model is smartest, it is which combination of weights, provider and price leaves the most options open a year from now.

The signal

Two things are true at once. Demand for K3 was strong enough that Moonshot paused paid signups within four days. And the selloff says investors no longer assume every capability gain translates into proportional silicon spend. The weights are not out yet: until July 27 this is an API launch with an open-weight promise, and the technical report is still pending. But if Moonshot delivers, the open-weight frontier will have jumped from the 975B-parameter class to nearly 3 trillion in under two weeks, and the gap between open and closed will be measured in weeks, not years.

Get the Signal

AI and business tech news, verified by a physician who reads the filings. One email a week, no noise.

Dr. Joseph Joshua

Dr. Joseph Joshua is the founder and editor of Corewire. A medical doctor by training, he brings the evidence-first discipline of clinical medicine to technology journalism: claims get checked against primary sources before they get published. He has produced technology and B2B content for companies across…

View Bio

Keep Reading