Today, large companies like OpenAI, Anthropic, and Google are competing in artificial intelligence performance rankings, but Amazon says: "Don't pay attention to those rankings." That is the main takeaway from the recent announcements at the AWS re:Invent conference in Las Vegas. Rohit Prasad, Amazon's vice president for AGI (artificial general intelligence), explained this beforehand. According to him, these benchmarks, which are tests of model performance, do not reflect the true power of AI. "I want real-world usefulness. None of these benchmarks are real," Prasad said. And the reason? Everyone uses different training data, and the tests are not sufficiently separated, so the results are full of noise and do not reveal the models' true capabilities.
Amazon is taking a different approach. Instead of chasing the top spots in rankings like LMArena, where the previous version of its Nova model finished in 79th place, it is focusing on practical applications. Prasad emphasized that benchmarks do not work because they are not standardized – everyone would need to have the same data, and the tests would have to be completely separated, which is not happening. So these rankings are more of a marketing gimmick than a measure of actual value.
Nova Forge
The main new development is the Nova Forge service, which Amazon introduced as a way for companies to train their own AI models without spending billions. The problem Forge addresses is real: most companies have three bad options. They either fine-tune a closed model only around the edges, train open models without the original data and risk the model forgetting its broad knowledge, or build everything from scratch at enormous expense.
Forge does things differently. It provides access to Nova model checkpoints at different stages – before, during, and after training. This allows companies to add their own data early, when the model is most "receptive to learning," as Prasad described it. "We've democratized the development of advanced models for your needs at a fraction of the cost," he said. The tool was created because Amazon's internal teams wanted it – much like AWS (Amazon's cloud service) began as an internal tool for its business and became its main source of profit.
Reddit Is Already Using Forge
Reddit is already testing Forge on its own safety models, trained on 23 years of community moderation data. Chris Slowe, Reddit's chief technology officer and first employee, said: "I've never seen anything like it." Their engineer is reportedly like a kid in a candy store. Last week, they launched a training run that looks promising. The goal is to replace several specialized models with a single one that understands the nuances of moderation, including the subjective "Don't be a jerk" rule found across subreddits.
Slowe explained that Forge allows Reddit to control its models, avoid API changes from other providers, own the model weights, and keep sensitive data in-house. They are already planning a similar approach for Reddit Answers and other products. When asked whether it mattered that Nova was not at the top of the benchmarks, Slowe was direct: "In this context, what matters is the model's expertise in relation to Reddit." Amazon is thus emphasizing control and specialization rather than raw intelligence.
Infrastructure over Intelligence
Amazon is betting that the model race has become commoditized and that it will succeed by offering a place where companies can build custom AI for specific problems. This is a typical AWS approach: infrastructure and customization over raw performance. In doing so, it avoids direct comparisons with OpenAI or Anthropic, which it previously sought to compete with at the model level.
Forge's success depends on whether developers adopt it. Amazon claims that the traditional model race does not matter. If that proves true, the measure of success will shift to whether AI actually delivers real-world value in practice.
Source: theverge.com



