IMAGE CREDITS: COREWIRE

Thinking Machines ships Inkling, a 975B open model that refuses to play the benchmark game

Mira Murati’s Thinking Machines released Inkling: 975B total parameters, 41B active, multimodal, Apache 2.0, weights on Hugging Face. The pitch is fine-tuning on Tinker, not benchmark wins.

Thinking Machines Lab, the startup former OpenAI CTO Mira Murati founded, released its first open-weights foundation model on July 15. It is called Inkling: a mixture-of-experts transformer with 975 billion total parameters, 41 billion active per token, native multimodality across text, images and audio, a context window of up to 1 million tokens, and an Apache 2.0 license with full weights on Hugging Face.

The specs are frontier-class, but the positioning is the story: Inkling is explicitly not chasing the leaderboard. Thinking Machines describes it as a balanced generalist built to be fine-tuned on Tinker, its customization platform. The bet is that the durable business in open models is adaptation, not raw benchmark wins.

What shipped

Per the model card, Inkling is a 66-layer decoder-only transformer with a sparse MoE feed-forward backbone: each token is routed to 6 of 256 experts plus 2 shared experts. It was pretrained on 45 trillion tokens of text, image, audio and video data. Multimodality is encoder-free, with images handled through hierarchical patch encoding and audio through discrete tokens, all projected into one shared space. Weights ship in the original BF16 checkpoint and a quantized NVFP4 version.

A lighter sibling, Inkling-Small, was previewed alongside the release with 12 billion active parameters and, per the company, strong early results on reasoning and agentic tasks at lower cost and latency. Its full weights are promised once testing completes. Both models expose a controllable thinking-effort dial for trading performance against token spend.

Customization as the product

Thinking Machines is not claiming state of the art. Its announcement leans on calibration, efficiency and adaptability, and routes developers toward fine-tuning Inkling on Tinker, with an Inkling Playground added to the console for testing. That makes the open weights an acquisition funnel for a paid platform, a pattern the open-model economy has been converging on: Together AI raised $800 million as open-model demand tripled on essentially the same thesis, that someone has to serve and tune all these free weights.

The signal

The open-weight frontier is suddenly crowded at the top. Moonshot’s Kimi models came from nowhere into Western coding tools, and Inkling now puts a 975B-parameter American open model on the table days before Moonshot promises weights for a model nearly three times that size. For enterprises, the calculus is shifting from which closed API to buy toward which open base to own and tune. That is exactly the fight Murati’s company was built for, and it chose to enter it without pretending to a benchmark crown it does not hold. In a market where pricing is increasingly the product, an Apache 2.0 license is the most aggressive price there is.

Get the Signal

AI and business tech news, verified by a physician who reads the filings. One email a week, no noise.

Dr. Joseph Joshua

Dr. Joseph Joshua is the founder and editor of Corewire. A medical doctor by training, he brings the evidence-first discipline of clinical medicine to technology journalism: claims get checked against primary sources before they get published. He has produced technology and B2B content for companies across…

View Bio

Keep Reading