SmolVLM Grows Smaller – Introducing the 256M & 500M Models!
Hugging Face Blog ● Covered by 2 sources
Hugging Face released SmolVLM-256M and SmolVLM-500M, vision language models with 256 million and 500 million parameters respectively. The 256M model is the smallest vision language model ever released and uses a 93-million-parameter vision encoder processing images at 4096 pixels per token, compared to 1820 pixels per token in the previous 2B version. These models enable vision language capabilities on consumer devices and low-cost data processing while maintaining performance on tasks like image captioning, document question-answering, and visual reasoning.