Former OpenAI CTO Mira Murati unveiled her first AI model

Former OpenAI CTO Mira Murati unveiled her first AI model

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
17. 7. 2026
6 minutes reading
Former OpenAI CTO Mira Murati unveiled her first AI model

Thinking Machines Lab has released its first model. It is called Inkling and stands out from the competition at first glance. Unlike the flagship models from OpenAI, Anthropic, or Google, it has open weights, allowing external developers and companies to download and directly modify it.

For the company founded a year and a half ago by former OpenAI CTO Mira Murati, this is the first public evidence of its work. For most of that time, it built its infrastructure out of the public eye. A little had already leaked out in May, when the team demonstrated so-called interaction models. They can listen and speak in real time, even interrupting the user instead of waiting for them to finish a sentence like conventional chatbots.

What Inkling can do and how it is built

The model is powered by a Mixture-of-Experts architecture. It has a total of 975 billion parameters but uses only a fraction of them for each task, roughly 41 billion. This approach keeps even very large models fast and inexpensive to run. Its context window supports up to one million tokens.

Training took place on an enormous volume of data. The company trained the model on 45 trillion tokens of text, images, audio, and video, and says it reasons directly over all these types of input. It processes audio as spectrograms and divides images into small tiles. Both are then processed together with text in a single stream.

Its approach to reliability is also noteworthy. Inkling is designed to provide calibrated answers. When it is uncertain, it admits it instead of making things up. Users can also adjust a dial for so-called reasoning effort. Those who need speed can turn it down, trading accuracy for pace.

The Thinking Machines team is not aiming for first place on the leaderboards. In the model documentation, it explicitly states that Inkling is not the strongest model available, whether open or closed. Instead, it targets balanced, broad performance across disciplines and serves as a foundation that can be readily customized.

Its efficiency, however, is worth noting. In one test, the company says Inkling achieved the same programming performance as Nvidia's Nemotron 3 Ultra while using one-third as many tokens. Alongside the main model, the team also released a preview of a smaller version. Inkling-Small has 276 billion parameters, 12 billion of them active, and matches or even surpasses its larger sibling on a range of tests.

Focus on customization

One idea underpins the entire release. Thinking Machines believes that models organizations can customize themselves will ultimately beat the general-purpose solutions currently sold by the largest labs.

The company is therefore not offering Inkling as a finished product, but as a starting point. Customers are expected to fine-tune it themselves through Tinker, a model customization platform. This does mean, however, that they are responsible for the safety of their own modifications, and fine-tuning requires a capable machine learning team. OpenAI, Anthropic, and Google took the opposite route. They first built ChatGPT, Claude, and Gemini as general-purpose chatbots and only added standalone agentic features later.

As of today, Inkling is available for fine-tuning on Tinker with context options of 64,000 and 256,000 tokens. For a limited time, the company is offering a 50 percent discount. The full weights are available on Hugging Face. thinkingmachinesthinkingmachines

An argument gaining momentum

Recent days have shown that Thinking Machines is not alone in this view. Microsoft CEO Satya Nadella, whose company has invested billions in both OpenAI and Anthropic, warned in a Sunday post that businesses using closed models are effectively paying twice. Once for the subscription, and a second time by handing over their corporate knowledge embedded in thousands of prompts and corrections, which may then end up in future versions of the model.

Hugging Face CEO Clem Delangue makes a similar argument. In his view, cutting-edge models will increasingly be used for experiments and the most valuable tasks, while routine production work will be taken over by private or open solutions. That is precisely the division Thinking Machines is targeting.

The strongest evidence recently came from finance. The company's researchers, together with Bridgewater Associates, the world's largest hedge fund, took an existing open model and further trained it on Bridgewater's financial knowledge. The result scored 84.7 percent on financial tests, beat leading closed models, and cost roughly fourteen times less to run. Those figures, however, come from the two companies' own evaluation, not from independent measurement.

The question of distillation

Thinking Machines likes to emphasize speed. It took OpenAI approximately five years and Anthropic roughly three years to bring their technology to market and demonstrate revenue. Thinking Machines claims it accomplished the same in about nine months.

But an uncomfortable question has also emerged. Was Inkling trained on outputs from competing models? This practice is known as distillation and has attracted criticism in the industry. According to the company, the answer is: partly. The base training was conducted from scratch, but the company used other open models, including Moonshot AI's Kimi K2.5, to prepare some of the early post-training data. It says its next model will use a fully self-contained process.

The company devoted considerable attention to safety. On FORTRESS, a test that evaluates refusals of dangerous queries involving weapons and violence, Inkling demonstrated the strongest built-in safeguards of all the open models compared. At the same time, it did not unnecessarily refuse harmless queries that looked similar. External testers verified the results.

Money, compute, and people

Thinking Machines is more cautious when discussing costs. In March, it entered into a partnership with Nvidia to deploy one gigawatt of Vera Rubin computing capacity and says Inkling was trained exclusively on Nvidia GB300 NVL72 systems. However, the company has not yet said how it balances this against revenue, which has so far received little attention.

Last November, the company was reportedly preparing a $50 billion funding round, but according to several media outlets, it had stalled by January. Since then, the company has remained silent about its financing, although Nvidia said when announcing the March partnership that it had made a significant investment in Thinking Machines.

The logic behind the focus itself is also interesting. Once the weights are released, no one who downloads them has to pay Thinking Machines to run them. Revenue must therefore come from Tinker—in other words, from training, fine-tuning, and now also a share of the hosting surrounding the model.

The staffing situation looks calmer today. The company employs roughly 200 people, more than after the departures early this year, when two co-founders also left for OpenAI. According to a source inside the company, Thinking Machines is deliberately built around continuity rather than a single name. If you do not put anyone on a pedestal, their departure does not hurt as much.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok