An eight-member team led by Yijun Pan from Alibaba Group’s Accio division, working in collaboration with Yale University, built a simulated marketplace and assigned fifteen leading language models the roles of store owners. Each model received starting capital of $80,000 and thirty simulated days to grow it. Average final wealth ranged from just under $21,000 for MiniMax M2.5 to more than $188,000 for Gemini 3.1 Pro, representing an almost ninefold difference. More than half of all simulations ended in a loss, and only four models managed to preserve their starting capital in every single trial. The authors published the study on the arXiv platform.
Why test AI specifically on commerce
A marketplace built on real Alibaba data
Models’ trading results



