
In case you have missed it, companies are using human videos to train models for AI and robots in general. Take Dyna-2 for instance: it is a world action model that is trained on one million hours of human video. This model has better language following capabilities. The model has zero-shot environment generalization and continues to work under various disturbances. According to Dyna Robotics, this model achieved “87% quality and throughput rating during zero-shot deployment at novel customer sites.”
Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws:
• world-action models exhibit scaling law on human data across four orders of magnitude, from 1000… pic.twitter.com/wZamR0axzS
— Dyna Robotics (@DynaRobotics) August 10, 2026
As the researchers explain:
Dyna-2 is a world-action model: a single generative model that can denoise future video and future actions jointly or separately, built on a video-diffusion backbone. Architecturally it is a mixture of transformers
[HT]












































