TLDRocket
Sign in

SmolVLA: Efficient Vision-Language-Action Model trained on Lerobot Community Data

Hugging Face Blog

SmolVLA is a 450-million-parameter open-source vision-language-action model for robotics that runs on consumer hardware and uses only publicly available datasets. The model was trained on fewer than 30,000 episodes—roughly one-tenth the data of comparable systems—yet matches or exceeds the performance of much larger models on simulation and real-world robotics tasks. The asynchronous inference system enables 30% faster response times and 2× task throughput by decoupling action execution from perception processing, allowing robots to respond more quickly to changing environments.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.