TLDRocket
Sign in

Standard Intelligence: Training General Intelligence in Pixel Space

Sequoia by Sonya Huang Covered by 2 sources

Standard Intelligence is pursuing a general computer agent by pre-training a model directly on raw video of computer use rather than predicting text tokens. It reports an 11-million-hour computer action dataset and a video encoder about 50× more token-efficient, so nearly 2 hours of 30 FPS video fits into a 1-million-token context window. The approach changes agent training by learning next mouse movements, clicks, and keystrokes from pixels and aims to produce models like FDM-1 that can perform tasks after fine-tuning.

Why it matters

The future of useful agents may begin not with text, but with pixels.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.