TLDRocket
Sign in

SmolVLM2: Bringing Video Understanding to Every Device

Hugging Face Blog

Hugging Face released SmolVLM2, a suite of video understanding models in three sizes—2.2B, 500M, and 256M parameters—designed to run efficiently on devices from phones to servers. The 2.2B model matches the performance of comparable frontier models on the Video-MME benchmark, while the 500M model achieves similar video understanding capabilities with less than a quarter of the parameters. The release includes demo applications such as an iPhone app, VLC media player integration, and a video highlight generator, alongside libraries for inference and fine-tuning through Transformers and MLX.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.