TLDRocket
Sign in
Latest Cohere’s chief AI officer: US and China will hold us ‘by the throat’ i... — Sifted Swedish A.I. Startup Klang Raises 1.5 Million Euros and Sets Its Sight... — Trending Topics Fighter Jets Without Pilots: U.S. Air Force Takes Delivery of First A.... — Trending Topics European tech weekly recap: Over €2B invested across 65+ deals — Tech.eu Complaion raises €13.5M to simplify compliance for European SMEs — Tech.eu Walmart CEO says that the retail giant and its AI shopping assistant a... — Fortune Power Shortages From A.I. Data Centers: The Gas Power Comeback Has Alr... — Trending Topics Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About... — MarkTechPost

The AI intelligence platform

Every AI story that matters — and the intelligence behind it.

TLDRocket reads all relevant sources, removes duplicate coverage, and publishes a short neutral summary of every story, linking back to the original. Free, no spam, unsubscribe anytime.

Add to Slack

Every story also updates live profiles event timelines weekly rankings the AI Market Index

Friday, 24 April 2026

Accelerate RL rollouts by up to 50% with distribution-aware speculative decoding

Together AI 5 months ago 51

Distribution-aware speculative decoding (DAS) reduces rollout time in RL post-training by up to 50% through an adaptive suffix tree drafter and length-aware scheduling that mitigates the long-tail problem where slow generations block entire training batches. Testing on math reasoning with DeepSeek-R1-Distill-Qwen-7B achieved over 50% rollout speedup while maintaining identical reward curves, and on code generation with Qwen3-8B achieved approximately 25% speedup. DAS eliminates GPU idle time during training without changing model outputs or requiring separate neural drafters, allowing practitioners to reduce compute costs while preserving model quality across varying sequence lengths and batch sizes.

DeepSeek-V4: a million-token context that agents can actually use

Hugging Face 5 months ago 50

DeepSeek released V4 with two model variants—V4-Pro at 1.6 trillion total parameters (49 billion active) and V4-Flash at 284 billion total (13 billion active)—both supporting a 1 million token context window. At 1M tokens, V4-Pro requires 27% of the single-token inference compute of V3.2 and uses 10% of its KV cache memory, while V4-Flash drops to 10% of the FLOPs and 7% of the cache through hybrid attention mechanisms that compress key-value entries by 4x to 128x. The efficiency gains enable practical long-context agentic workflows where tool results accumulate across multiple rounds without exhausting GPU memory or degrading performance mid-task.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.