IMAGE CREDITS: COREWIRE

Thinking Machines ships Inkling, a 975B open model that refuses to play the benchmark game

Mira Murati’s Thinking Machines released Inkling: 975B total parameters, 41B active, multimodal, Apache 2.0, weights on Hugging Face. The pitch is fine-tuning on Tinker, not benchmark wins.

Thinking Machines Lab, the startup former OpenAI CTO Mira Murati founded, released its first open-weights foundation model on July 15. It is called Inkling: a mixture-of-experts transformer with 975 billion total parameters, 41 billion active per token, native multimodality across text, images and audio, a context window of up to 1 million tokens, and an Apache 2.0 license with full weights on Hugging Face.

The specs are frontier-class, but the positioning is the story: Inkling is explicitly not chasing the leaderboard. Thinking Machines describes it as a balanced generalist built to be fine-tuned on Tinker, its customization platform. The bet is that the durable business in open models is adaptation, not raw benchmark wins.

What shipped

Per the model card, Inkling is a 66-layer decoder-only transformer with a sparse MoE feed-forward backbone: each token is routed to 6 of 256 experts plus 2 shared experts. It was pretrained on 45 trillion tokens of text, image, audio and video data. Multimodality is encoder-free, with images handled through hierarchical patch encoding and audio through discrete tokens, all projected into one shared space. Weights ship in the original BF16 checkpoint and a quantized NVFP4 version.

A lighter sibling, Inkling-Small, was previewed alongside the release with 12 billion active parameters and, per the company, strong early results on reasoning and agentic tasks at lower cost and latency. Its full weights are promised once testing completes. Both models expose a controllable thinking-effort dial for trading performance against token spend.

Customization as the product

Thinking Machines is not claiming state of the art. Its announcement leans on calibration, efficiency and adaptability, and routes developers toward fine-tuning Inkling on Tinker, with an Inkling Playground added to the console for testing. That makes the open weights an acquisition funnel for a paid platform, a pattern the open-model economy has been converging on: Together AI raised $800 million as open-model demand tripled on essentially the same thesis, that someone has to serve and tune all these free weights.

What owning the weights actually buys

The fine-tuning-first pitch lands on a real enterprise pain point: deprecation risk. Teams that built on closed APIs have repeatedly watched models they tuned prompts against get retired, repriced or silently changed. Weights under an Apache 2.0 license cannot be taken away, cannot be repriced, and can be pinned to a version forever. For regulated industries that must reproduce a model’s behavior years later, that permanence is worth more than a few benchmark points.

The license choice matters as much as the weights. Apache 2.0 carries none of the usage restrictions and revenue thresholds attached to some rival open releases, which means no legal review cycle before commercial deployment and no clause that changes the deal at scale. It is the same lock-in calculus we mapped in the agentic AI buyer’s guide and in the free-compute land grab, inverted: Thinking Machines is betting that giving up control of the base model is precisely what earns the customization business on top of it. The unanswered question is whether adaptation revenue can carry a frontier lab’s training costs, and Inkling is the first serious test of that model at this scale.

The signal

The open-weight frontier is suddenly crowded at the top. Moonshot’s Kimi models came from nowhere into Western coding tools, and Inkling now puts a 975B-parameter American open model on the table days before Moonshot promises weights for a model nearly three times that size. For enterprises, the calculus is shifting from which closed API to buy toward which open base to own and tune. That is exactly the fight Murati’s company was built for, and it chose to enter it without pretending to a benchmark crown it does not hold. In a market where pricing is increasingly the product, an Apache 2.0 license is the most aggressive price there is.

Get the Signal

AI and business tech news, verified by a physician who reads the filings. One email a week, no noise.

Dr. Joseph Joshua

Dr. Joseph Joshua is the founder and editor of Corewire. A medical doctor by training, he brings the evidence-first discipline of clinical medicine to technology journalism: claims get checked against primary sources before they get published. He has produced technology and B2B content for companies across…

View Bio

Keep Reading