TLDRocket
Sign in

Induction Labs Photon-1 Simulates Desktops, Plays Checkers, and Models Billiard Physics From One Pretraining Run

MarkTechPost Michal Sutter

Induction Labs released Photon-1, a 106B-parameter vision model trained on 18 years of unlabeled computer screen recordings using next-latent-token prediction instead of action labels. The model used approximately 30,000 H200 GPU-hours for pretraining and achieves lower inference costs than Gemini 3.1 Flash-Lite on an internal computer use benchmark. After finetuning on tasks outside its pretraining domain, Photon-1 demonstrated improved performance on checkers and billiard physics simulation compared to LLM and vision encoder baselines.

Why it matters

Most agents that learn from video need to know what action produced each frame. Induction Labs is arguing that this requirement is the bottleneck. Last week, they released imagination models, a foundation model architecture that pretrains on raw video with no action labels at all. Their test system is Photon-1, a sparse 106B-A5B mixture-of-experts (MoE) […] The post Induction Labs Photon-1 Simulates Desktops, Plays Checkers, and Models Billiard Physics From One Pretraining Run appeared first on MarkTechPost.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.