LFM2.5-Encoders for Fast Long-Context Inference on CPU
Hugging Face ● Covered by 2 sources
Liquid AI released two encoder models, LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, that match larger models on benchmarks while maintaining fast inference on CPU with 8,192-token context windows. The 230M model achieves 3.7× faster CPU inference than ModernBERT-base at long context, processing 8,192 tokens in 28 seconds versus ModernBERT's 90 seconds. These models enable document-scale classification, routing, and detection tasks to run cheaply on existing hardware without GPU acceleration.
Related stories
Liquid AI Releases LFM2.5-2.6B: An On-Device Agentic Model With 128K Context, Tool Calling, And Open Weights
MarkTechPost · 1 month ago ·
45
LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
Hugging Face · 1 month ago ·
22
Deploy local agents everywhere with LFM2.5-2.6B
Hugging Face · 1 month ago ·
6