AGI Is Not Multimodal
The Gradient Benjamin A. Spiegel
An AI researcher argues that multimodal AI models will fail to achieve artificial general intelligence because they lack embodied physical understanding of the world, relying instead on learned syntactic patterns rather than genuine world models. Large language models achieve language proficiency through statistical rules and heuristic memorization of training data rather than understanding physical reality, as evidenced by their inability to solve sensorimotor tasks like sweeping a floor or repairing a car. Achieving true AGI requires systems grounded in physical interaction with environments rather than scaling multimodal networks that treat modalities as separate components to be combined.
Why it matters
"In projecting language back as the model for thought, we lose sight of the tacit embodied understanding that undergirds our intelligence." –Terry WinogradThe recent successes of generative AI models have convinced some that AGI is imminent. While these models appear to capture the essence of human