Standard Intelligence: Training General Intelligence in Pixel Space
Sequoia by Sonya Huang ● Covered by 2 sources
Standard Intelligence is pursuing a general computer agent by pre-training a model directly on raw video of computer use rather than predicting text tokens. It reports an 11-million-hour computer action dataset and a video encoder about 50× more token-efficient, so nearly 2 hours of 30 FPS video fits into a 1-million-token context window. The approach changes agent training by learning next mouse movements, clicks, and keystrokes from pixels and aims to produce models like FDM-1 that can perform tasks after fine-tuning.
Why it matters
The future of useful agents may begin not with text, but with pixels.
Related stories
Generalist AI Releases GEN-1.5: A Robot Foundation Model That Learns New Tasks From One 3–12 Second Demo
MarkTechPost · 1 week ago ·
7
Genie 3: A new frontier for world models
Google DeepMind · 10 months ago ·
12
Intelligence is Free, Now What? Data Systems for, of, and by Agents
BAIR · 1 month ago ·
36