AI models that can simulate gameplay for up to four players in real time, including sound

AI models that can simulate gameplay for up to four players in real time, including sound

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
20. 5. 2026
5 minutes reading
AI models that can simulate gameplay for up to four players in real time, including sound

    It would be easy to overlook a startup called Odyssey, based in Palo Alto. But behind it are two people with very specific résumés. Oliver Cameron, CEO and co-founder, previously led product development at Cruise, one of the largest autonomous vehicle companies. His partner Jeff Hawke, CTO, spent 15 years building AI for autonomous driving at Wayve and earned his doctorate at Oxford.

    In 2023, they both left the automotive industry to establish a company focused on general world models. In other words, causal, multimodal systems that learn to predict the world and interact with it in real time. Odyssey set out to build AI that not only generates video but directly simulates how the world behaves.

    The team has attracted experts from DeepMind, Tesla, Waymo, Meta, and Wayve. Incidentally, some of them were behind models such as DeepMind Gemini, DeepMind Veo, and Tesla FSD autonomous systems. Since its founding, the company has run more than 610,000 simulations for users in 184 countries around the world.

    From Hollywood filmmaking to simulating reality

    Odyssey did not start where it is today. Its original vision was different. The company wanted to create tools for professional filmmakers, "Hollywood-grade" AI that would enable animators to accomplish with a much smaller team and at a fraction of the cost what currently requires hundreds of people and hundreds of millions of dollars. At the time, Oliver Cameron said that a film like Avatar could be made by five people in six months.

    But then the vision shifted. Odyssey stopped focusing solely on filmmaking tools and began working on a more general problem: teaching a model to truly simulate the world. Not merely generating a short video clip, but continuously predicting what will happen next, responding to user input, and maintaining a consistent physical reality. The result is the Odyssey-1 and Odyssey-2 world models, and now two new additions: Agora-1 and Starchild-1.

    Agora-1

    Why GoldenEye? This 1997 shooter for the Nintendo 64 is an old classic, and many people at Odyssey grew up playing it. But that is not the only reason. Games have long served as testing environments for AI research, from Atari and Minecraft to StarCraft. GoldenEye was simply the next logical step.

    Agora-1 is the first world model that allows up to four participants to share a single simulated reality at once. Until now, world models were limited to a single active participant. Up to four people can meet in a shared deathmatch simulation, with each seeing an AI-generated world in real time. The model acts as a learned game engine. It does not use traditional code for physics or graphics; instead, it learned all the game’s dynamics directly from GoldenEye’s internal game state.

    How did they achieve this breakthrough? Odyssey separated two things: simulation and rendering. One model continuously computes the shared state of the world—what is happening where, how characters are moving, and what is changing. A second, diffusion-based model converts this state into visuals, separately for each player from their own perspective. Because the state of the world is managed explicitly, Agora-1 can generate entirely new levels without losing the mechanics of the original game.

    Earlier attempts at multi-agent approaches, such as Multiverse or Solaris, ran into problems mainly when players lost sight of one another. Agora-1 aims to provide a consistent representation of the same world from multiple independent perspectives simultaneously. A demo is available to try directly on Odyssey’s website as an early research preview. But the goal is not to build games. Cameron and Hawke see the future in other areas: training AI agents in fully simulated environments and collaborative robotics, where multiple robots need to reason jointly about space and actions.

    Starchild-1: A world model that can hear and speak

    Alongside Agora-1, Odyssey also introduced a second model. Starchild-1 is the first real-time multimodal world model, representing a different kind of breakthrough.

    Traditional world models learned only from visual data. Starchild-1 autoregressively generates synchronized video and audio while continuously responding to text input from the user. Simply put, it not only sees the world but hears it as well, and the user can alter it during the simulation through written text or voice commands.

    Traditional audio-video models such as DeepMind Veo generate clips in advance, offline, with a fixed duration. Once generation begins, the future of the output is fixed. Starchild-1 works differently: it predicts every video frame and every segment of audio based on what came before and what the user has just entered. The trajectory changes continuously.

    Technically, this required entirely new approaches to synchronized audio-video output and stability over long time horizons. Starchild-1 runs at up to 24 frames per second on modern hardware. Unlike Agora-1, it focuses on a single user, but adds a layer of sound and spoken language that no real-time world model has previously had.

    A public demo is not yet available; Odyssey has released only video compilations and a technical report. But the description itself is clear enough: a system that not only simulates the world but also gives it sound and responds to what you tell it.

    Category:AI
    Did you enjoy this article?
    Discover more interesting posts on our blog
    Back to blog

    Related posts

    Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
    OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
    6 min read
    28. 7. 2026
    AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
    Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
    6 min read
    28. 7. 2026
    Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
    A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
    4 min read
    28. 7. 2026
    Přihlaste se k odběru našeho newsletteru
    Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
    CodedTrip

    Operated by CodedTrip LLC, USA.

    YouTube
    TikTok