TLDRocket
Sign in
Latest Nebius looks to raise $4.5BN through bond issue — Tech.eu Also’s $3,500 e-bike is a $1 billion Trojan horse for autonomous trans... — Fortune Unitree, famous for its dancing robots, surges by 460% on its trading... — Fortune Exclusive: Replit taps OpenAI's low-cost Luna model for new 'Free Mode... — Fortune Adronite launches Codistry AI coding platform, claims half the token c... — SiliconANGLE Rundoo raises $30M to expand its AI-native operating system for small... — SiliconANGLE Temporal is in talks to raise $500M at a $12B pre-money valuation, mor... — Tech Funding News Etched raises $700M led by Jane Street, doubling to $21B and it still... — Tech Funding News

The AI intelligence platform

Every AI story that matters and the intelligence behind it.

TLDRocket reads all relevant sources, removes duplicate coverage, and publishes a short neutral summary of every story, linking back to the original. Free, no spam, unsubscribe anytime.

Add to Slack

Every story also updates live profiles event timelines weekly rankings the AI Market Index

Thursday, 6 June 2024

Hello Qwen2

GitHub Pages 2 years ago 50

Alibaba released Qwen2, the successor to Qwen1.5, with five model sizes ranging from 0.5B to 72B parameters. The models support 29 languages total, achieve state-of-the-art results on multiple benchmarks, and the largest versions handle context windows up to 128K tokens. The release expands Alibaba's language model offerings with improved capabilities in coding, mathematics, and multilingual support.

Generalizing an LLM from 8k to 1M Context using Qwen-Agent

GitHub Pages 2 years ago 35

Alibaba's Qwen team built a multi-level agent system that extends an 8k-token context model to handle 1-million-token documents by combining retrieval-augmented generation, chunk-by-chunk reading, and step-by-step reasoning rather than relying on native long-context models. The system was evaluated on NeedleBench and LV-Eval benchmarks designed for 256k-context tasks, where the 4k-Agent consistently outperformed both a 32k-context model extended via RoPE extrapolation and basic RAG approaches. The agent framework is being released as open-source infrastructure to generate synthetic fine-tuning data for training long-context models.

Extracting Concepts from GPT-4

OpenAI 2 years ago 17

Researchers used scaled sparse autoencoders to extract 16 million distinct computational patterns from GPT-4's operations. The technique identified 16 million individual concepts that the model uses during processing. This capability enables better understanding of how large language models compute internally and may improve interpretability of AI systems.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.