Hugging Face releases smaller SmolVLM models and adds vision language model support to smolagents framework
Model release Provisional 92% confidence first seen
Hugging Face released two new smaller versions of SmolVLM with 256M and 500M parameters, with the 256M model being the smallest vision language model ever released. Additionally, the Hugging Face smolagents framework now natively supports vision language models in agent pipelines, enabling capabilities like autonomous web browsing agents that can understand visual elements.