TLDRocket
Sign in

Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model That Reads Screens, Grounds Objects, and Calls Tools On-Device

MarkTechPost Asif Razzaq Covered by 2 sources

Liquid AI shipped a small vision model that reads screens, spots objects, and calls tools on-device. It’s built to stay fast and cheap enough to run locally, with a license catch for bigger companies.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Liquid AI has released LFM2.5-VL-3B, a 3.1B-parameter vision-language model aimed at on-device use. The pitch is pretty clear: let the model read what’s on a screen, pick out objects by coordinates, handle documents and charts, and even trigger tools from text or image input without sending everything off to the cloud.

The company says it fits in roughly 3 GB of memory and can decode 228 tokens per second on an Apple M5 Max. It ships in four formats — native, GGUF, ONNX, and MLX — with day-one support from llama.cpp, MLX, vLLM, SGLang, and ONNX. And because it’s a non-reasoning model, it’s designed to answer directly instead of spending time thinking out loud.

The benchmark story is more interesting than the raw size. Liquid AI reports a 69.4 average across 28 vision benchmarks. That ties InternVL-3.5-4B and trails Qwen3.5-4B by 0.7 points, even though both of those comparison models are 4.7B parameter systems. On ScreenSpot-v2, the new model scores 80.7 overall, with 78.7 on desktop, 81.2 on mobile, and 82.2 on web.

This release also adds function calling to the VL line. ToolSandbox jumps from 26.4 to 59.5, while BFCL v4 rises from 20.5 to 32.5. Grounding improves sharply too: RefCOCO-avg precision@1 climbs from 57.1 to 87.9, which Liquid AI links to scaled synthetic grounding data. Multi-image work gets better as well, with BLINK moving to 61.5 from 50.2 and MuirBench to 58.3 from 34.9.

Under the hood, the model uses LFM2.5-2.6B as its language backbone and a SigLIP2 NaFlex 400M vision tower. NaFlex handles native resolution by breaking large images into 512×512 patches and pairing them with a resized thumbnail. Liquid AI says training used about 34T tokens, with a doubled 128K vocabulary and support for 16 languages. The license is the part that will make lawyers blink: commercial use stays free only while a company’s annual revenue is under $10M.

My take — AI-written commentary, not fact-checked reporting

This is the right kind of small-model release: practical, local, and not pretending every prompt needs a choir of chain-of-thought theatrics. The license, though, is classic open-ish software theater — friendly to startups, a tollbooth for everyone else. That split is becoming the new normal: open enough to spread, closed enough to collect later.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.