Every AI story that matters — and the intelligence behind it.
TLDRocket reads all relevant sources, removes duplicate coverage, and publishes a short neutral
summary of every story, linking back to the original. Free, no spam, unsubscribe anytime.
Mistral released Voxtral Transcribe 2, a pair of speech-to-text models including Voxtral Mini for batch processing and Voxtral Realtime for live applications, with the latter available as open-weights software. Voxtral Mini achieves 4% word error rate at $0.003 per minute, outperforming competitors like GPT-4o mini and Gemini 2.5 Flash while processing audio 3x faster than ElevenLabs' Scribe v2 at one-fifth the cost. The models support 13 languages with speaker diarization, context biasing, and configurable latency down to sub-200 milliseconds, enabling applications from live subtitling to voice agents and contact center automation.
Researchers developed Sequential Attention, a technique for efficiently selecting the most important features, blocks, or components in neural networks during training rather than through expensive post-hoc analysis. The method uses attention scores to sequentially identify redundant components and achieved state-of-the-art results on standard benchmarks while maintaining accuracy. This approach enables building smaller, faster models and can be applied to feature selection, network pruning, large language models, and drug discovery without significantly increasing training costs.
Codex released an App Server that uses a bidirectional JSON-RPC API to embed the Codex agent into applications. The API supports streaming progress updates, tool execution, approval workflows, and diff visualization. Developers can now integrate Codex agents directly into their own applications with standardized communication protocols.
Moonshot AI released Kimi K2.5, an open-source multimodal model trained on 15 trillion mixed visual and text tokens with capabilities for understanding text, images, and video. The model includes agent orchestration features allowing multiple agents to work together in coordinated swarms. The release expands available open-source options for multimodal AI applications with agentic capabilities.
Rime Arcana V3 Turbo and V3 text-to-speech models are now available on Together AI for multilingual voice applications with code-switching capabilities. V3 Turbo achieves approximately 120 milliseconds time-to-first-audio for English-Spanish switching, while V3 extends support to 11 languages at roughly 160 milliseconds latency. Voice agents can now handle mid-sentence language switching with consistent prosody and cadence, reducing the need to route between separate language-specific models.
Hugging Face launched a decentralized evaluation system where benchmark datasets can host leaderboards and any community member can submit model evaluation results via pull requests. The initial rollout includes four benchmarks—MMLU-Pro, GPQA, and HLE among them—with results stored as YAML files in model repositories and aggregated automatically across the Hub. This creates a transparent record of evaluation sources and methodology, allowing the community to track and build upon scores rather than relying on conflicting reports from papers and closed platforms.
Every AI story that matters,
in your inbox by 8am.
TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the
day in two minutes. Follow companies and topics for alerts, or get the
briefing in Slack. Free, no spam, unsubscribe anytime.