Inference Chips Instead of GPUs: AI Lenders Bet on a Different Kind of Chip for the First Time

Inference Chips Instead of GPUs: AI Lenders Bet on a Different Kind of Chip for the First Time

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
21. 7. 2026
5 minutes reading
Inference Chips Instead of GPUs: AI Lenders Bet on a Different Kind of Chip for the First Time

Startup General Compute borrowed up to $400 million. There would be nothing unusual about that, except for one thing. The collateral is not Nvidia GPUs, which until now have backed all major loans in the artificial intelligence industry. Instead, the company is pledging SambaNova SN50 inference chips. According to available information, this is the first time a lender has accepted chips designed exclusively to run finished models, rather than train them, as the primary collateral for a large AI loan.

The loan was provided by investment firm Upper90 Capital Management. It starts at $100 million, and the amount will increase as paying customers are added, up to a ceiling of $400 million.

Chips Have Become Collateral for Wall Street

To understand why this deal is different, we need to go back a few years. In 2021, Upper90 co-founder and former Goldman Sachs trader Billy Libby financed the purchase of GPUs for Crusoe. He described the loan as the first ever backed by the value of advanced AI chips. Traditional banks avoided such deals at the time. No one really knew how quickly GPUs would lose their value.

The distrust quickly faded. In August 2023, CoreWeave borrowed $2.3 billion against Nvidia H100 chips. Chip-backed loans became the engine of its growth, and its total debt now exceeds $21 billion. The model it developed was adopted by the entire so-called neocloud industry, meaning operators of computing centers tailored to artificial intelligence. Until now, however, everything has been built on Nvidia silicon.

How Are SambaNova Chips Architecturally Different?

So why did General Compute choose the SN50 instead of Nvidia GPUs? Price is not the primary consideration. What matters is the architecture and what happens directly on the chip during inference. Inference is the phase in which a finished, already trained model responds to user queries, generating text or other outputs in real-world operation.

Running a large language model is divided into two phases. The first is prefill. The model reads and processes the entire input query. This is a parallel task and is well suited to the SIMT architecture around which Nvidia built its GPUs. The second phase is decoding. The model generates a response token by token, with each successive token depending on all the previous ones. The problem here is not computing power, but the speed at which the chip moves data between memory and compute units.

SambaNova’s SN50 is designed precisely around this second constraint. It places memory and compute circuitry close together, so data travels a shorter distance. An independent measurement by Artificial Analysis from July 2026, using a mixed configuration of four Nvidia H200 GPUs and sixteen SN50 units, achieved 763 tokens per second on the MiniMax M2.7 model, several times faster than providers relying solely on GPUs.

The second advantage directly supports the deal’s financial structure. The SN50 does not require water cooling. It runs at 20 kilowatts per rack, which a standard air-cooled data center can handle. General Compute has already contracted 15 megawatts of such capacity. Liquid cooling requires specialized data centers that take years to build, while air-cooled systems can be installed within a few weeks.

Lenders Are Following the Revenue

The deeper logic of the deal revolves around the difference between computing for training and computing for inference. And this is precisely where investors’ attention is shifting.

Training a large model is a one-off undertaking. A lab trains a model during a burst of computing that lasts days or weeks, after which the cluster sits only partially utilized. Inference works differently. Serving a finished model to users, developers, and autonomous agents runs continuously, because every API call and every chatbot response is an inference task. Inference is therefore becoming something akin to a utility, providing a steady stream of revenue from every token produced.

And this is the crux of the matter. A GPU is versatile. It can handle training, inference, and graphics, and, most importantly, there is a sufficiently large market for used hardware to allow a lender to estimate its residual value. An inference ASIC such as the SN50, by contrast, is optimized for a specific set of operations. This makes it faster and cheaper per token, but also leaves it with little use outside that task. A lender accepting inference chips as collateral is betting that the task for which the chip was created will remain economically significant throughout the repayment period.

Not everyone needs a supercomputer, but everyone needs inference and artificial intelligence. The company that serves inference most cheaply and quickly will capture that revenue, and the hardware underpinning such economics is more reliable collateral than hardware whose utilization depends on a sporadic training schedule. Each additional tranche of the $400 million is therefore tied to customer demand. Credit risk thus grows in line with actual revenue, rather than ahead of it.

What Lenders Can Value—and What They Cannot Yet

However, this is not a risk-free deal. A secondary market exists for Nvidia chips and is becoming increasingly liquid; in fact, Compute Exchange launched a marketplace for used H100 and A100 chips just a few days ago. No such market exists for SambaNova SN50 chips. Upper90 accepted them as collateral before any sales history had been established.

While a used Nvidia GPU can handle almost any AI task, a used inference ASIC optimized for decoding transformer-based models could lose value if the prevailing model architecture shifts elsewhere. According to analysts at S&P Global Ratings, a high-end Nvidia GPU loses roughly half of its resale value within three years. However, no one yet knows the depreciation curves for inference chips. ASIC chips used for Bitcoin mining provide a cautionary example: they nearly lost all their value when the algorithm they were designed for was replaced by something else.

Sources: techcrunch.com and techtimes.com

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok