Thinking Machines Lab has released Inkling, its first open-weights model and the clearest statement yet of the company’s commercial thesis: a capable foundation model matters most when developers can adapt it to their own work. The startup, founded by former OpenAI CTO Mira Murati, is not claiming to have produced the industry’s outright strongest system. Instead, it is pairing a large multimodal model with Tinker, its platform for fine-tuning.

Inkling is a mixture-of-experts transformer with 975 billion total parameters and 41 billion active at a time. It accepts text, images and audio, supports a context window of up to one million tokens, and was pretrained on 45 trillion tokens spanning those modalities. The weights are available publicly, while Tinker offers a managed route for adapting the model. In its launch announcement, Thinking Machines describes that combination—not raw peak performance—as Inkling’s central appeal.

That positioning is deliberate. The company says developers can tune a controllable “thinking effort” setting to trade latency and token use against quality. It also argues that a broad model is a better substrate for specialized systems than one optimized around a single benchmark. A smaller preview model, Inkling-Small, has 276 billion total parameters and 12 billion active parameters; on some reported evaluations it even edges the larger release, reinforcing the idea that the family is meant to be tuned and selected by workload rather than treated as a one-size-fits-all flagship.

Independent measurements make the promise more concrete, and more conditional. Artificial Analysis data reported by The Decoder puts Inkling at 41 on its Intelligence Index, ahead of other U.S. open-weights models such as Nemotron 3 Ultra. On agentic knowledge-work tests, it also surpasses Kimi K2.6 and DeepSeek v4 Flash max. The model generates fewer output tokens than several peers on the same benchmark suite, which supports Thinking Machines’ efficiency narrative.

But the same measurements identify the constraint that matters for deployment. Inkling scored just +2 on Artificial Analysis’ factuality-focused Omniscience benchmark, with 40 percent accuracy and a 63 percent hallucination rate. Its listed API pricing is also above comparable Chinese open models for certain context lengths. In other words, token efficiency does not automatically translate into lower total cost or suitable reliability for factual, high-stakes workflows.

Inkling’s significance is therefore less about a clean leaderboard victory than about the shape of the offer. Thinking Machines is releasing weights, a multimodal base and a customization platform together. For teams building agents, internal tools or domain-specific assistants, that package could be compelling—provided they bring robust evaluation, grounding and human review. The launch makes customization easier; it does not remove the need to verify what a customized model says.


Discover more from TekCrispy

Subscribe to get the latest posts sent to your email.

Leave a comment

Leave a Reply