A tutorial provides an end-to-end evaluation workflow for PerceptionBench, a multimodal vision model benchmark covering OCR, counting, localization, and other visual tasks. The implementation uses a three-stage fallback data loader, processes base64-encoded images, and creates a balanced subset of 120 examples (12 per capability category) from the full 3,000-example benchmark. The workflow enables reproducible evaluation across local and API-based models with automated judging and comparative analysis of visual perception capabilities.
Interconnects launched two free data tools for tracking open-source AI models: the Artifacts Hub covers 792 models released in the last two years with metrics from Hugging Face and Open Router, while an Adoption Dashboard tracks downloads and derivatives by geography and organization, particularly highlighting US-China adoption patterns. The hub provides adoption scores, intelligence indices, and similarity metrics for popular models like GLM-5.2 and DeepSeek R1. These tools aim to increase transparency in the open model ecosystem and help developers understand which models are gaining traction.
Onton released Ontology 1, a neurosymbolic search model for e-commerce product discovery that achieved a precision@10 score of 0.630 compared to Google Shopping's 0.543 and Amazon's 0.469 on a 90-query benchmark. The model uses an inspectable knowledge graph to reason about product attributes rather than relying on seller labels or vector embeddings, and won 52 of 90 queries outright while indexing only 1% of competitor catalogs. The system is available as a live product on Onton.com with case-by-case partner access, but no public API or open weights, meaning adoption requires partnership rather than standard deployment methods.
Researchers conducted a comprehensive study examining how preference alignment techniques affect multimodal large language models that process both text and images. The study focuses on reducing hallucination—when models generate responses inconsistent with image content—through alignment methods that encourage outputs to match visual information more closely. The findings suggest alignment techniques improve MLLM performance on image understanding tasks, establishing a foundation for better multimodal model development.
Every AI story that matters,
in your inbox by 8am.
TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the
day in two minutes. Follow companies and topics for alerts, or get the
briefing in Slack. Free, no spam, unsubscribe anytime.
Reading TLDRocket needs no cookies, and the readership counts we rely on come from
our own cookieless analytics. Google Analytics is the exception: it sets cookies and
reports to Google, so it stays switched off until you allow it. You can change your
mind any time from “Cookie settings” in the footer.
Strictly necessary
Session security and form protection (tldrocket-session,
XSRF-TOKEN, 2 hours). The site cannot work without them,
so they need no consent.
Always on
Google Analytics 4 (_ga,
_ga_<id>, up to 2 years). Measures which
stories and sections readers use. Google acts as a third-party processor and may
store the data outside the EU. No advertising, no profiling, no data sold.