Two announcements a week apart tell one story. OpenAI and Broadcom unveiled Jalapeño, OpenAI’s first custom inference processor, a reticle-sized ASIC taped out in a nine-month cycle and targeting deployment from late 2026, per Tom’s Hardware’s analysis. Days later, Reuters reported that DeepSeek is developing its own inference chip. The labs are done renting the layer where their economics live.
Why everyone suddenly wants inference silicon
Training happens once; inference happens every time a user asks. As we argued in our margins analysis, serving cost is the gravity of the whole AI economy, it decides whether products at the application layer keep 25 percent or 70 percent of their revenue. A chip tuned to one lab’s serving patterns attacks that cost directly: OpenAI says engineering samples already run production workloads at target frequency and power.
The strategic logic mirrors what Google did with TPUs a decade ago, but the timeline is new: nine months from design start to tape-out, with OpenAI using its own models to accelerate chip design. Silicon iteration is starting to move at software speed, which also reshapes demand on the memory supply chain that feeds every accelerator.
The two races inside the race
The American version is about margin and Nvidia leverage: every hyperscaler and lab wants a second source it controls. The Chinese version, DeepSeek’s, is about sovereignty under export controls, where dependence on any foreign supplier is an existential risk. Same instinct, different stakes: control the layer your economics depend on.
What to watch
Watch whether Jalapeño’s late-2026 deployment holds, which foundry DeepSeek can actually use, and Nvidia’s pricing response, the clearest tell of how seriously it takes the custom-silicon exodus. The same sovereignty logic is now visible one layer up the stack in China’s open-weight model strategy.
AI and business tech news, verified by a physician who reads the filings. One email a week, no noise.