Standard Intelligence announced its approach to train a general computer agent from raw video (pixel-based training)
Research publication Provisional 55% confidence first seen
Standard Intelligence reported that it is training a general computer agent, FDM-1, by pre-training on raw video of computer use instead of predicting text tokens. The company says it built an 11-million-hour computer action dataset and developed a video encoder designed to be substantially more token-efficient, enabling long video sequences to fit within a large context window.