
Inkling and Inkling Small: The Open-Weights Bet of Mira Murati and Thinking Machines
22 months after leaving OpenAI, Mira Murati has delivered the first model from Thinking Machines Lab: Inkling, with 975 billion parameters in open weights. A lighter version, Inkling Small (276 B), will follow. What it is, where it's available, and why it positions Murati against Anthropic and OpenAI.
On July 15, 2026, Thinking Machines Lab — the company founded by Mira Murati, former CTO of OpenAI — released its first model: Inkling. 22 months after her dramatic departure from OpenAI (September 2024), Murati finally delivers the long-awaited product. And she does so in a radically different way from her former employer: with open weights.
Following this, the company announces Inkling Small, a lighter version of the same model, in advanced testing — imminent release once internal evaluations are completed.
Who is Thinking Machines Lab?
Thinking Machines is the new AI startup by Mira Murati, co-founded with several former OpenAI employees. Key milestones:
- Foundation: late 2024, a few weeks after Murati's departure from OpenAI
- Seed funding: $2 billion in June 2025, led by Andreessen Horowitz — one of the largest seed rounds ever recorded
- Nvidia deal: gigawatt-scale compute agreement signed early 2026 for training
- Positioning: open-weights + enterprise specialization, similar to Meta LLaMA rather than a consumer product like ChatGPT
Inkling: The Numbers
The published model is a substantial Mixture-of-Experts:
- 975 billion parameters in total
- About 41 billion active parameters per token
- Trained on 45 trillion multimodal tokens (text, image, audio, video)
- Released with open weights under a permissive license
It's on the scale of a Kimi K3 or a Llama 4.1, directly targeting the closed frontier models of OpenAI and Anthropic.
Inkling Small: What We Know
The lighter version, nearing completion:
- 276 billion parameters in total
- About 12 billion active parameters per token
- The weights will be published once internal tests are completed — no firm date yet
This is the model that will interest enterprise AI teams and independent developers the most: light enough to run on reasonable hardware (a DGX or a cluster of H100), capable enough to cover most professional use cases.
Where to Find It?
Inkling is distributed through the usual open model channels:
- Official release: thinkingmachines.ai/news/introducing-inkling
- Weights download: Hugging Face (Thinking Machines repository)
- A commercial API from Thinking Machines is also available for those who prefer not to host
- The model is compatible with major inference frameworks (vLLM, TensorRT-LLM, llama.cpp for certain quantizations)
What Inkling Offers
1. A Base Model Designed for Fine-Tuning
Thinking Machines makes it clear: Inkling is not a consumer assistant. It's a base model optimized to be specialized by enterprise clients in their domain (legal, medical, engineering, customer service…).
2. Native Multimodality
The 45 trillion training tokens cover text, image, audio, video. Unlike many open models that retrofit multimodality after the fact, Inkling incorporates it into its training architecture — which should yield better results on tasks involving multiple modalities.
3. Deep Customization
Murati has been emphasizing her thesis for months: “one-size-fits-all AI” is a losing bet in the medium term. Every company, every sector, every team has specific needs that require deep customization — not just prompt engineering. Inkling is designed for this.
4. Data Sovereignty and Control
An open-weights model hosted in-house = no data leaves. A massive selling point for banks, hospitals, governments, large industries that refuse to send their data to OpenAI or Google.
Strategic Positioning
Murati deliberately distances herself from:
- OpenAI, where she comes from — closed model, monolithic consumer product
- Anthropic — closed model, positioned on safety and reasoning quality
- xAI (Musk) — semi-open model, politically aligned
And aligns with:
- Meta LLaMA — open-weights, positioned for enterprise base and fine-tuning
- Mistral — European open-weights model with commercial offering
- DeepSeek and Kimi — Chinese players betting on openness
The bet: in 24 months, the market will be segmented between a few closed public champions (OpenAI, Anthropic, Google) and a rich ecosystem of specialized open-weights models where companies build their custom AI stack. Thinking Machines aims to become the Western reference in this second segment, an open counterweight to both closed American and open Chinese models.
The Stakes
Three elements to watch over the next 12 months:
- Measured quality. Independent benchmarks will soon be available. Does Inkling hold up against Llama 4.1, Kimi K3, DeepSeek V4 on real tasks? The community is currently assessing this.
- The release pace of Inkling Small. A model with 12 B active, if it lives up to Thinking Machines' promises, will be a game-changer for developers. The current delay is starting to cause frustration.
- Economic viability. Selling fine-tuning, support, API from open weights is a more challenging business model than selling ChatGPT-like subscriptions. Thinking Machines will need to prove it's viable beyond the $2 billion raised.
One thing is certain: Mira Murati's return to the forefront of AI, with a tangible product and clear positioning, adds depth to the debate between closed and open models. And Inkling — with its upcoming Small — could become by year's end one of the default references for enterprise AI teams.