GEN-1.5 Model: A Robot Learns a New Task from a Single Seconds-Long Demonstration

GEN-1.5 Model: A Robot Learns a New Task from a Single Seconds-Long Demonstration

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
21. 8. 2026
6 minutes reading
GEN-1.5 Model: A Robot Learns a New Task from a Single Seconds-Long Demonstration

American company Generalist AI has introduced the GEN-1.5 model, which needs only a three- to twelve-second demonstration for a robot to begin a task it has never seen before. No additional training, code rewriting, or hours of data collection are required. You load the demonstration into the model’s memory, and the robot starts working.

GEN-1.5 model

It is a large multimodal model that processes video together with sensor data, speech, and information about the positions of its own joints. It retains the last thirty seconds of activity around it in memory and generates motion trajectories at a rate of one hundred commands per second. Instead of text, images, touch, and motion go in, while commands for robotic arms come out. The company calls the supplied demonstration a physical prompt. It can be recorded by a person using controllers, or it can come from a robot that has successfully completed the task once. As soon as it enters the model’s memory, the robot starts working immediately.

Generalist previously introduced the GEN-0 model, for which the company described predictable scaling laws under which larger amounts of data and computing power reliably improved results. Five months later came GEN-1, which could be fine-tuned to achieve a success rate of more than ninety-nine percent on a specific task, and the team observed the first signs of improvisation. 

Meanwhile, the GEN-1.5 model was trained continuously for eight months. The developers kept it running for a simple reason: every metric they monitored continued to improve. The model absorbed more data, made better use of computing power, and achieved major leaps after architectural changes. It learned new tasks faster and with less input data. This was when something unexpected happened. The team tested how little fine-tuning was needed to master a new task. At first, it took several hundred training steps, then dozens, and finally a single step on one minute of recorded data. When even that worked, they asked whether the model could learn a new task without any training at all. The result was successful.

Fifty-nine percent without a single training step

Tests were conducted on ten different tasks. The robot unscrewed a jar lid, took money out of a wallet, stacked cups, swept up a mess, opened a book, unzipped a pencil case, and removed a suction cup. With a single demonstration in memory and no training whatsoever, it succeeded in an average of fifty-nine percent of attempts. When the model received an additional five minutes of data for the given task and ten fine-tuning steps, the success rate rose to eighty-three percent. When the team used a single training step on one minute of recorded data, the success rate on a new task remained around sixty-six percent. Ten fine-tuning steps alter the model’s weights by less than fifteen hundredths of a percent, suggesting that they merely rearrange what the model already knows.

The authors point out that the tasks are simple and short and that the results are not particularly impressive. Nevertheless, they say this is the first model in which the ability to learn a wide range of dexterous physical tasks from one or several demonstrations has emerged broadly.

The most interesting aspect is that nobody deliberately built this capability into the model. The team did not modify the architecture to support in-context learning, add any layer for rapid adaptation, or introduce auxiliary objectives for improvisation. It all emerged on its own from the enormous amount of physical activity data on which the model was trained.

For now, the developers can only speculate about the reasons for this phenomenon. One hypothesis draws on an analogy with language, where in-context learning is associated with the uneven distribution of phenomena in data. The second assumes that physical work is naturally cyclical and that the model learned to recognize these patterns. Curiously, the model was trained on continuous recordings from homes, warehouses, and factories, and had never before encountered the temporal jumps caused by a physical prompt.

Demonstrations can be combined. When the team loaded two separate recordings into the model’s memory—unzipping a pencil case and taking out money—the robot combined both actions into one smooth sequence. It devised the transition between them on its own, including regrasping, repositioning, and correcting its own mistakes—movements that appeared in neither demonstration. The company calls this physical prompt composition and sees it as a practical approach. Instead of recording one long demonstration of a complex procedure, the entire sequence can be assembled from a library of short segments.

Transfer from simulation also works. A real robot can use a demonstration recorded in a simulator as a template, even though the model saw no simulated images during training. It can also handle different object positions and sizes and perform the task with both hands. For some activities, this means that no one needs to record demonstrations in the real world. In some cases, the model can even bridge the difference between human and robotic bodies. A person demonstrates the task with their own hands in front of cameras, and the machine immediately repeats it.

A banana instead of a broom and a dustpan that calls for a different approach

The clearest demonstration that the model does not copy movements but understands the goal comes from a sweeping experiment. The developers fine-tuned it on five minutes of human demonstrations in which a broom swept a block into a bowl. They then gave it a banana, and the model used it as a substitute broom. With a dustpan, however, it chose a completely different approach. Instead of imitating the sweeping motion, it scooped the block onto the dustpan and tipped it into the bowl. The dustpan was not used this way in either the training or fine-tuning data, and the closest similar situations among nearly two million scenes bore no resemblance to the experiment.

They collected several similar surprises. The model removed a piece of paper that someone had used to cover the bowl and occasionally put it back after completing the task. When a Lego brick became stuck on its fingers, it removed it with the other hand. It sometimes unscrewed a jar lid with both hands, even though the demonstrations used only one. Models trained to place a single block into a single bowl then occasionally began sorting blocks by color.

One interesting detail is that improvisation becomes stronger the fewer fine-tuning steps the model receives. The developers explain this by saying that a lightly modified model remains closer to its original capabilities and has more to draw on in an unfamiliar situation.

Applications in manufacturing

The trade magazine Assembly focuses primarily on the practical implications of the results. Introducing a new robotic operation today usually means programming it or directly training a model for it. Generalist’s approach replaces this with a simple demonstration. For facilities with varied and frequently changing production, this could represent significant time savings. 

Robots have been sold as universal machines for years, but this universality has always depended on an expert spending several months programming them. If simply showing a robot a task is enough, two things change at once: how quickly the machine can be put into operation and the range of people capable of working with it.  For now, there are limitations. Skills learned from context are more fragile than fine-tuned ones, and a success rate of around sixty percent is far from operational reliability. But the team says it has no idea where the improvement curve will level off. 

Source: interestingengineering.com

Category:Robotics
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

China Unveils First Robotic Centaur. It Will Work Where Humans DieChina Unveils First Robotic Centaur. It Will Work Where Humans Die
At a Shanghai exhibition center this July, a machine appeared that at first glance defied classification. Shanghai-based Run Robotics unveiled a robot at the WAIC 2026 conference that it calls
3 min read
27. 7. 2026
Not for homes, but robots for industry. English firm Humanoid reaches billion-dollar valuationNot for homes, but robots for industry. English firm Humanoid reaches billion-dollar valuation
London-based robotics company Humanoid, officially known as SKL Robotics, has joined the ranks of so-called unicorns—young companies valued at more than one billion dollars. According to...
6 min read
20. 7. 2026
A Breakthrough in Robotics: Robots Get Artificial Hands That Work Like Ours!A Breakthrough in Robotics: Robots Get Artificial Hands That Work Like Ours!
California-based 1X from Palo Alto has unveiled new hands for its NEO robot. They have 25 degrees of freedom, are tendon-driven like human hands and, according to the manufacturer, approach human dexterity, strength and reliability.
5 min read
14. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok