A tutorial provides an end-to-end evaluation workflow for PerceptionBench, a multimodal vision model benchmark covering OCR, counting, localization, and other visual tasks. The implementation uses a three-stage fallback data loader, processes base64-encoded images, and creates a balanced subset of 120 examples (12 per capability category) from the full 3,000-example benchmark. The workflow enables reproducible evaluation across local and API-based models with automated judging and comparative analysis of visual perception capabilities.
Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model accepting text, image, and video input, with open weights coming next week. The hosted API costs $2 per million input tokens and $6 per million output tokens, with a 1-million-token context window and support for cached inputs at $0.25 per million tokens. The smaller 27B checkpoint will be the practical option for on-premise deployment, while performance gains over the previous version are largest in multimodal and agentic tasks rather than reasoning benchmarks.
Researchers conducted a comprehensive study examining how preference alignment techniques affect multimodal large language models that process both text and images. The study focuses on reducing hallucination—when models generate responses inconsistent with image content—through alignment methods that encourage outputs to match visual information more closely. The findings suggest alignment techniques improve MLLM performance on image understanding tasks, establishing a foundation for better multimodal model development.
Every AI story that matters,
in your inbox by 8am.
TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the
day in two minutes. Follow companies and topics for alerts, or get the
briefing in Slack. Free, no spam, unsubscribe anytime.
Reading TLDRocket needs no cookies, and the readership counts we rely on come from
our own cookieless analytics. Google Analytics is the exception: it sets cookies and
reports to Google, so it stays switched off until you allow it. You can change your
mind any time from “Cookie settings” in the footer.
Strictly necessary
Session security and form protection (tldrocket-session,
XSRF-TOKEN, 2 hours). The site cannot work without them,
so they need no consent.
Always on
Google Analytics 4 (_ga,
_ga_<id>, up to 2 years). Measures which
stories and sections readers use. Google acts as a third-party processor and may
store the data outside the EU. No advertising, no profiling, no data sold.