Thinking Machines Lab, the startup former OpenAI CTO Mira Murati founded, released its first open-weights foundation model on July 15. It is called Inkling: a mixture-of-experts transformer with 975 billion total parameters, 41 billion active per token, native multimodality across text, images and audio, a context window of up to 1 million tokens, and an Apache 2.0 license with full weights on Hugging Face.
The specs are frontier-class, but the positioning is the story: Inkling is explicitly not chasing the leaderboard. Thinking Machines describes it as a balanced generalist built to be fine-tuned on Tinker, its customization platform. The bet is that the durable business in open models is adaptation, not raw benchmark wins.
What shipped
Per the model card, Inkling is a 66-layer decoder-only transformer with a sparse MoE feed-forward backbone: each token is routed to 6 of 256 experts plus 2 shared experts. It was pretrained on 45 trillion tokens of text, image, audio and video data. Multimodality is encoder-free, with images handled through hierarchical patch encoding and audio through discrete tokens, all projected into one shared space. Weights ship in the original BF16 checkpoint and a quantized NVFP4 version.
A lighter sibling, Inkling-Small, was previewed alongside the release with 12 billion active parameters and, per the company, strong early results on reasoning and agentic tasks at lower cost and latency. Its full weights are promised once testing completes. Both models expose a controllable thinking-effort dial for trading performance against token spend.
Customization as the product
Thinking Machines is not claiming state of the art. Its announcement leans on calibration, efficiency and adaptability, and routes developers toward fine-tuning Inkling on Tinker, with an Inkling Playground added to the console for testing. That makes the open weights an acquisition funnel for a paid platform, a pattern the open-model economy has been converging on: Together AI raised $800 million as open-model demand tripled on essentially the same thesis, that someone has to serve and tune all these free weights.
The signal
The open-weight frontier is suddenly crowded at the top. Moonshot’s Kimi models came from nowhere into Western coding tools, and Inkling now puts a 975B-parameter American open model on the table days before Moonshot promises weights for a model nearly three times that size. For enterprises, the calculus is shifting from which closed API to buy toward which open base to own and tune. That is exactly the fight Murati’s company was built for, and it chose to enter it without pretending to a benchmark crown it does not hold. In a market where pricing is increasingly the product, an Apache 2.0 license is the most aggressive price there is.
AI and business tech news, verified by a physician who reads the filings. One email a week, no noise.
