Skip to content

Thinking Machines Ships Inkling: The Top US Open Weights Model Is a Fine Tuning Play

July 20, 2026. Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati, released its first production model on July 15. Inkling is a 975 billion parameter open weights Mixture of Experts model, and unlike most open announcements this month, the weights are downloadable on Hugging Face today. The announcement is unusually honest about positioning: Inkling is pitched not as the strongest model available but as a base for fine tuning on Tinker, the company's training platform.

What shipped

  1. The model. 975B total parameters with 41B active, native text, image, and audio input, a context window up to 1 million tokens (64K and 256K tiers on Tinker), pre trained on 45 trillion tokens, with continuously adjustable thinking effort.
  2. The benchmark position. Artificial Analysis scores Inkling 41 on its Intelligence Index, making it the leading US open weights model, ahead of Nemotron 3 Ultra at 38, Gemma 4 31B at 29, and gpt-oss-120b at 24, though still behind the best Chinese open models. On the agentic GDPval-AA v2 benchmark it reaches an Elo of 1,238, edging Kimi K2.6 at 1,190 and DeepSeek v4 Flash max at 1,189.
  3. The caveat. Per The Decoder's analysis of the Artificial Analysis data, Inkling posts 40 percent accuracy and a 63 percent hallucination rate on factual recall, scoring just +2 on AA Omniscience. It is not a knowledge oracle.
  4. The economics. Hosted pricing starts at $1.87 per million input and $4.68 per million output tokens at 64K context, and Inkling is notably token efficient: about 25,000 output tokens per Intelligence Index task versus 37,000 to 43,000 for comparable open models.
  5. The preview. Inkling-Small, 276B total with 12B active, beats its big sibling on several benchmarks (88.3 versus 87.2 on GPQA Diamond) and gets open weights after testing.

Open now beats open soon

The contrast with Kimi K3, which announced open weights that do not arrive until around July 27, matters. Inkling's weights are live on day one. For US and European businesses that ruled out Chinese hosted models over data residency, a leading US open weights model, covered in our cost shift analysis, removes the main objection to self hosting cheap intelligence.

What it means for operators

Fine tuning as the business model is the interesting part for SMBs and agencies. A tuned mid size open model beats a generic frontier model on plenty of narrow, high volume jobs: document extraction, call summarization, product categorization, structured agent steps. Open weights mean you own the artifact, control the hosting, and stop renting the capability per token. The hallucination number draws the boundary: use Inkling for structured and agentic work where outputs are verified by systems, keep it grounded with retrieval, and do not deploy it as a standalone facts engine. If that maps to a workflow you run at volume, an AI engineer can scope the fine tune, and our automation team builds the verification rails around it.

Want a custom model tuned to your workflow?

We design, build, and run it for you, integrated with the tools you already use. Free audit in 24 hours.

Get Your Free Audit

Frequently Asked Questions

Inkling is the first production model from Thinking Machines Lab, founded by former OpenAI CTO Mira Murati. Released July 15, 2026, it is a 975 billion parameter Mixture of Experts model with 41 billion active parameters, native text, image, and audio input, up to 1 million tokens of context, and open weights available on Hugging Face.

Yes, in the sense that matters today: the full weights are downloadable on Hugging Face now, under the company's release terms. Kimi K3, announced July 16, promised open weights around July 27. Inkling-Small, the 276B preview model, gets its weights published after testing completes.

Artificial Analysis ranks Inkling 41 on its Intelligence Index, the best US open weights score, ahead of Nemotron 3 Ultra at 38, but still behind the top Chinese open models. It wins on agentic knowledge work benchmarks against Kimi K2.6 and DeepSeek v4 Flash max while costing slightly more per token, and it uses fewer output tokens per task than comparable models.

Strong: agentic tasks, tool use, multimodal input, token efficiency, and serving as a fine tuning base on Tinker. Weak: factual recall, with 40 percent accuracy and a 63 percent hallucination rate on Artificial Analysis testing. Use it for structured, verified workflows with retrieval grounding, not as a standalone knowledge engine.

Free Strategy Audit

Ready to put this to work?

Join 200+ businesses already scaling with AI and automation. Get your free audit and a custom roadmap within 48 hours.

Website & marketing performance analysis
AI & automation opportunity mapping
Custom growth roadmap with ROI estimates
Delivered within 48 hours, 100% free
200+
Clients served
48hr
Turnaround
100%
Free, no strings

Get Your Free Audit

Takes 30 seconds. No credit card required.

Prefer to chat?

WhatsApp us