Induction Labs Photon-1: A Single Pretraining Run Powers Desktop Simulation, Checkers, and Billiard Physics

2026 aicomputer use benchmarkfoundation modelimagination modelsinduction labsmixture-of-expertsmoe transformerphoton-1video pretraining

Most AI agents that learn from video footage require knowledge of which action produced each frame. Induction Labs argues this assumption is the fundamental bottleneck. In July 2026, the company released imagination models—a novel foundation model architecture that pretrains on raw video without any action labels whatsoever.


Their flagship system, Photon-1, is a sparse 106-billion-parameter, 5-billion-active-parameter mixture-of-experts (MoE) transformer, trained on 18 years of computer demonstration video. On internal benchmarks for computer use, Induction Labs reports that Photon-1 outperforms Gemini 3.1 Flash-Lite while using significantly less pretraining compute.


Photon-1 demonstrates remarkable emergent abilities from a single pretraining run. It can simulate interactive desktop environments, play checkers at a competent level, and model realistic billiard ball physics—all without explicit training on those specific tasks. This suggests that learning purely from observation, without action labels, allows the model to internalize rich causal and physical models of the world.


The implications for 2026's AI landscape are significant. By removing the need for action-labeled data, imagination models could dramatically expand the scale of usable training video, reduce annotation costs, and unlock more general-purpose physical reasoning in foundation models. Induction Labs has open-sourced the architecture and training pipeline to accelerate research in this direction.


As the field moves toward more autonomous systems, the ability to learn by watching—rather than by being told what actions to take—may prove a crucial step on the path to general intelligence.

via MarkTechPost

Related