Blaize has developed a Graph Streaming Processor (GSP) designed to reduce unnecessary data movement during AI inference by scheduling computations to keep intermediate data flowing through the processor rather than shuttling it to external memory. The company positions its GSP for the first stage of multi-stage inference pipelines, handling routine tasks efficiently while routing complex cases to GPUs, with the same silicon packaged across embedded modules, PCIe cards, and rack-mount servers. This approach reduces power consumption and hardware costs by matching processor type to workload complexity rather than using expensive GPUs for every computation.
AMD acquired Taalas, a startup that embeds AI model weights directly into silicon to accelerate inference. Taalas's approach delivers performance improvements of 10x or greater compared to standard inference. The acquisition gives AMD proprietary technology to optimize its AI chip offerings against competitors like NVIDIA.
Researchers from UC Berkeley and affiliated labs developed ARBITRAGE, a technique that accelerates large language model reasoning by using a lightweight router to dynamically choose between a fast draft model and a slower but more capable target model based on their relative performance. The method achieves up to 2× speedup in inference latency on mathematical reasoning benchmarks compared to prior step-level speculative decoding approaches. This allows LLMs to generate longer reasoning chains more efficiently without sacrificing accuracy.
Researchers compared the performance characteristics of diffusion language models (DLMs) and autoregressive language models (ARMs) across inference scenarios. DLMs achieve higher arithmetic intensity through parallel token generation but fail to scale effectively with longer contexts, while ARMs maintain superior throughput in batched inference. The key finding is that reducing sampling steps in DLMs is essential for them to achieve lower latency than ARMs in practical deployments.
Every AI story that matters,
in your inbox by 8am.
TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the
day in two minutes. Follow companies and topics for alerts, or get the
briefing in Slack. Free, no spam, unsubscribe anytime.
Reading TLDRocket needs no cookies, and the readership counts we rely on come from
our own cookieless analytics. Google Analytics is the exception: it sets cookies and
reports to Google, so it stays switched off until you allow it. You can change your
mind any time from “Cookie settings” in the footer.
Strictly necessary
Session security and form protection (tldrocket-session,
XSRF-TOKEN, 2 hours). The site cannot work without them,
so they need no consent.
Always on
Google Analytics 4 (_ga,
_ga_<id>, up to 2 years). Measures which
stories and sections readers use. Google acts as a third-party processor and may
store the data outside the EU. No advertising, no profiling, no data sold.