OpenAI Unveils GPT-5.3-Codex: The First AI Model to Help Develop Itself

OpenAI Unveils GPT-5.3-Codex: The First AI Model to Help Develop Itself

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
9. 2. 2026
4 minutes reading
OpenAI Unveils GPT-5.3-Codex: The First AI Model to Help Develop Itself

OpenAI has introduced a new model, GPT-5.3-Codex, which represents a major milestone in the development of artificial intelligence. It is the company's first model to play a key role in its own creation. The Codex team used early versions of the model to debug its own training, manage its own deployment, and diagnose test results. According to OpenAI, the team was amazed by how much Codex was able to accelerate its own development.

GPT-5.3-Codex combines the frontier coding performance of the previous GPT-5.2-Codex model with the reasoning capabilities and expertise of GPT-5.2, all in a single model that is also 25% faster. The model can take on long-running tasks involving research, tool use, and complex execution.

State-of-the-art benchmark performance

GPT-5.3-Codex achieved a new industry high in several key benchmarks. In SWE-Bench Pro, it scored 56.8%, a rigorous evaluation of real-world software development. Unlike SWE-bench Verified, which only tests Python, SWE-Bench Pro covers four languages.

In Terminal-Bench 2.0, which measures the terminal skills required by a coding agent, the model scored 77.3%, a significant improvement over GPT-5.2-Codex at 64.0%. It also achieves this using fewer tokens than any previous model.

In OSWorld-Verified, which measures computer-use capabilities, GPT-5.3-Codex scored 64.7%, compared with just 38.2% for GPT-5.2-Codex. Humans score approximately 72% on this benchmark.

GPT-5.3-Codex benchmark results.
GPT-5.3-Codex benchmark results.

Creating complex games and applications

The model can create highly functional, complex games and applications from scratch in just a few days. OpenAI asked GPT-5.3-Codex to create two games: a racing game with various racers, eight maps, and items that can be used with the space bar, and a diving game in which players explore different reefs, collect fish in a codex, and keep an eye on oxygen, pressure, and hazards. The model gradually improved the games autonomously over millions of tokens using general follow-up prompts such as “fix the bug” or “improve the game.”

GPT-5.3-Codex is evolving from an agent that can write and review code into one that can do almost anything on a computer that developers and professionals can do. The model is designed to support all work throughout the software lifecycle – debugging, deployment, monitoring, writing PRDs, editing copy, user research, testing, metrics, and more. Its agentic capabilities extend beyond software and can help create anything, from presentations to spreadsheet data analysis. In the GDPval evaluation, which measures model performance on precisely specified knowledge tasks across 44 professions, GPT-5.3-Codex matches GPT-5.2 at 70.9%.

An interactive collaborator

The model provides frequent updates, keeping users informed about key decisions and progress throughout the work. Instead of waiting for the final output, users can communicate in real time – asking questions, discussing approaches, and steering the solution. GPT-5.3-Codex explains what it is doing, responds to feedback, and keeps users informed from start to finish.

GPT-5.3-Codex is the first model that OpenAI has classified as highly capable for cybersecurity-related tasks under its Preparedness Framework. It is also the first model trained directly to identify software vulnerabilities.

OpenAI is deploying its most comprehensive cybersecurity safety suite to date. Measures include safety training, automated monitoring, trusted access to advanced capabilities, and enforcement procedures including threat intelligence. The company is launching Trusted Access for Cyber Defense, a pilot program aimed at accelerating cyber defense research. OpenAI is also allocating $10 million (approximately CZK 205 million) in API credits to accelerate cyber defense using its most capable models.

Strong competition

Interestingly, on the same day, Anthropic introduced its Claude Opus 4.6 model. According to an analysis published on Every.to, the two models were released just tens of minutes apart. Anthropic even moved its planned release from 10:00 to 9:45 PST, while GPT-5.3-Codex launched at 10:01.

According to tests conducted by the Every team, the two models are converging. Opus 4.6 has acquired the thorough and precise style that made Codex the preferred choice for demanding coding tasks. GPT-5.3-Codex, meanwhile, has gained speed and a willingness to simply get things done without constantly asking for permission. According to the Every team's tests, Opus 4.6 has a higher ceiling as a model, but also greater variance. It is more parallelized and more creative. One team member used it for a feature in the Monologue app that the team had been working on for two months – the model simply built it. However, Opus sometimes reports success when it has actually failed, or makes changes that were not requested.

GPT-5.3-Codex is an excellent model with more reliable output. It is highly intelligent and can work autonomously for long periods on difficult coding tasks. It is very fast – faster than Opus – and does not make the silly mistakes that Opus does.

Availability

GPT-5.3-Codex is available with paid ChatGPT plans everywhere Codex can be used: in the app, command-line interface, IDE extension, and on the web. OpenAI is working to make the API safely available soon.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok