Revolutionary Taalas HC1 AI Chip: 10x Faster Than Rivals and 20x Cheaper to Make—Will It Bury Nvidia?

Revolutionary Taalas HC1 AI Chip: 10x Faster Than Rivals and 20x Cheaper to Make—Will It Bury Nvidia?

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
24. 2. 2026
4 minutes reading
Revolutionary Taalas HC1 AI Chip: 10x Faster Than Rivals and 20x Cheaper to Make—Will It Bury Nvidia?

    Try this. You submit a query to a language model, and the answer arrives before you can blink. No waiting, no spinning wheel. That is exactly what startup Taalas promises with its first product, and honestly, after reading the numbers, my jaw dropped a little.

    This February, Taalas emerged from stealth and showed the world what it had been quietly developing for two and a half years. A team of just 24 people, with total spending of only $30 million out of more than $200 million raised from investors. And the result? A chip that, according to the company's own measurements, runs 10x faster than the current market leader and costs 20x less to manufacture.

    Why is today's AI so slow and expensive?

    To understand what Taalas has actually done, we need to explain why AI inference is such a pain. Modern language models run on massive GPU servers that consume hundreds of kilowatts of power, require liquid cooling, and need entire rooms full of cables. Data centers are growing to the size of small cities. Costs are skyrocketing. And speed? Coding assistants sometimes spend whole minutes thinking. The developer loses focus, and the flow is gone. Yet agentic AI applications need responses in milliseconds, not at a human pace.

    Taalas says: this is a poorly designed system. And it is right.

    A revolutionary idea: bake the model directly into silicon

    Founder and CEO Ljubisa Bajic compares the situation to the ENIAC computer from 1945. This vacuum-tube-filled colossus occupied an entire room, was slow, and prohibitively expensive. Then came the transistor, followed by the PC and the smartphone. Computing technology became so small and inexpensive that we now carry it in our pockets.

    Taalas wants to do the same with AI. Its approach is radically different from anything else on the market. Instead of running the model as software on a general-purpose chip, it physically bakes the entire model, including its weights, directly into silicon. This creates a chip designed exclusively for one specific model. They call it a Hardcore Model.

    It is based on three principles: total specialization, combining memory and compute logic on a single chip, and radically simplifying the entire hardware stack. No HBM memory, no 3D stacking, no liquid cooling. Just a clean, elegant architecture.

    Numbers that speak for themselves

    Taalas's first product is a hardware implementation of Meta's Llama 3.1 8B model. The results are, to put it mildly, shocking. 17,000 tokens per second per user. For comparison, Cerebras, until now the world's fastest inference platform, achieves roughly one-tenth of that performance. Nvidia H200 GPUs? Two orders of magnitude lower. One of the first testers described it in a single word: "Insane."

    Power consumption? 12 to 15 kW per rack, while a GPU rack requires 120 to 600 kW. A Taalas rack can also be air-cooled. No costly data center retrofits. The price of inference? 0.75 cents per million tokens for Llama 3.1 8B. GPUs cost 20 to 49 cents for the same amount. This is a full order-of-magnitude leap forward!

    Tokens per second per user.
    Tokens per second per user.

    How is such a chip made?

    Taalas works with TSMC, the world's largest chip manufacturer. The base chip has approximately 100 layers and is prefabricated. When a new model arrives, only two metal layers need to be modified. The entire process takes two months, while manufacturing Nvidia Blackwell takes roughly six months.

    This approach addresses one of the biggest potential problems: what happens when the model changes? Taalas claims it can handle an update in two months, not two years. The price for customers includes three upgrades over the chip's four-year lifespan.

    On February 19, the company also announced the closing of a $169 million funding round. In total, it has raised more than $219 million from investors.

    Can Taalas really threaten Nvidia?

    The billion-dollar question. Or rather, the trillion-dollar one. The advantages are clear. Speed, price, power consumption. But data centers do not like managing dozens of different SKUs for different models. And Meta, whose Llama model Taalas is the first to bake into silicon, has just signed a "multigenerational" partnership with Nvidia. Evidently, Taalas did not persuade it to switch.

    But Taalas is not giving up. The plan calls for a mid-size reasoning LLM in spring 2026 and a frontier model on the second-generation HC2 chip by the end of the year. HC2 will deliver higher density, standard 4-bit formats, and even faster performance. If it succeeds in convincing major data centers, the AI hardware market could be shaken to its foundations.

    Category:AI
    Did you enjoy this article?
    Discover more interesting posts on our blog
    Back to blog

    Related posts

    Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
    OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
    6 min read
    28. 7. 2026
    AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
    Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
    6 min read
    28. 7. 2026
    Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
    A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
    4 min read
    28. 7. 2026
    Přihlaste se k odběru našeho newsletteru
    Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
    CodedTrip

    Operated by CodedTrip LLC, USA.

    YouTube
    TikTok