DeepSeek, the Chinese lab that repeatedly embarrassed bigger rivals on efficiency, may be taking the next logical step. The company is developing its own AI inference chip, Reuters reports, citing three people familiar with the matter. DeepSeek has not confirmed the effort, and no timeline or specifications have been reported; treat the details accordingly.
The reported focus is inference, serving models, rather than training them, and the strategic aim is reducing dependence on both Nvidia and Huawei silicon.
Why inference, and why it matters
Inference is where AI economics live once models ship: every user request burns serving cost, which is why margins in AI software hinge on it. A lab famous for wringing performance from constrained hardware designing silicon for its own workloads is the full-stack playbook Google wrote with its TPUs, attempted under export-control constraints that make dependence on any single supplier a strategic risk.
If the report holds, it also reframes the memory and foundry demand picture: another serious buyer designing for the same constrained supply chain everyone else is bidding on.
The foundry question is the whole question
Designing a chip is the part export controls cannot stop; fabricating one is the part they were built to stop. A DeepSeek inference part would need a foundry, and the options define the ceiling. TSMC and Samsung cannot legally fabricate advanced AI silicon for a Chinese lab under current US rules, which leaves SMIC, whose most advanced processes run generations behind the leading edge and without access to the EUV lithography tools that define it. An inference chip on a trailing node can still make economic sense, inference tolerates older processes better than training does, but it concedes the efficiency frontier the chip exists to chase.
That constraint explains why the effort is rational anyway. For a Chinese lab, the alternative to imperfect domestic silicon is not perfect foreign silicon, it is silicon that can be revoked by the next round of export rules. Sovereignty over a worse chip beats dependence on a better one, which is the same logic driving the broader domestic stack buildout. DeepSeek’s specific talent is the wildcard: a lab that made constrained hardware overperform in software may be the one design team for whom a constrained foundry is a familiar problem rather than a disqualifying one.
What to watch
Watch for any DeepSeek confirmation or tape-out reporting, which foundry would fabricate the part, a question export controls make existential, and whether other Chinese labs follow. An unconfirmed report from a wire service with three sources is worth covering; it is not yet worth treating as fact, and we will label developments accordingly. For the wider pattern, see our analysis of the inference silicon race, and for how it fits a broader national buildout, see MiniMax’s reported 2.7 trillion-parameter model and MetaX’s chip production ramp.
AI and business tech news, verified by a physician who reads the filings. One email a week, no noise.
