TLDRocket
Sign in
Latest Nebius looks to raise $4.5BN through bond issue — Tech.eu Also’s $3,500 e-bike is a $1 billion Trojan horse for autonomous trans... — Fortune Unitree, famous for its dancing robots, surges by 460% on its trading... — Fortune Exclusive: Replit taps OpenAI's low-cost Luna model for new 'Free Mode... — Fortune Adronite launches Codistry AI coding platform, claims half the token c... — SiliconANGLE Rundoo raises $30M to expand its AI-native operating system for small... — SiliconANGLE Temporal is in talks to raise $500M at a $12B pre-money valuation, mor... — Tech Funding News Etched raises $700M led by Jane Street, doubling to $21B and it still... — Tech Funding News

The AI intelligence platform

Every AI story that matters and the intelligence behind it.

TLDRocket reads all relevant sources, removes duplicate coverage, and publishes a short neutral summary of every story, linking back to the original. Free, no spam, unsubscribe anytime.

Add to Slack

Every story also updates live profiles event timelines weekly rankings the AI Market Index

Wednesday, 20 March 2024

GaLore: Advancing Large Model Training on Consumer-grade Hardware

Hugging Face 2 years ago 35

GaLore reduces memory requirements for training large language models by projecting gradients into lower-dimensional subspaces before optimizer processing. The technique achieves an 82.5% reduction in memory for optimizer states and enables training of 7-billion-parameter models on consumer GPUs like the NVIDIA RTX 4090. When combined with 8-bit quantization, GaLore allows researchers with limited computational resources to train larger models or use larger batch sizes on standard hardware.

Cosmopedia: how to create large-scale synthetic data for pre-training Large Language Models

Hugging Face 2 years ago 8

Researchers at Hugging Face created Cosmopedia, an open synthetic dataset containing 25 billion tokens generated using Mixtral-8x7B to replicate the training data behind Microsoft's Phi-1.5 language model. The dataset comprises over 30 million files across textbooks, blog posts, stories, and WikiHow articles, with less than 1% duplicate content achieved through extensive prompt engineering across 145 web-clustered topics and curated educational sources. The release includes the generation code, the full dataset, and a 1-billion-parameter model trained on it, enabling the community to reproduce high-performance language model training without proprietary data or models.

A Chatbot on your Laptop: Phi-2 on Intel Meteor Lake

Hugging Face 2 years ago 18

Microsoft's Phi-2, a 2.7-billion parameter language model, can now run on standard laptops using Intel's Meteor Lake processor with 4-bit quantization applied through the Optimum Intel library. The quantized model achieved adequate generation speed on a mid-range Core Ultra 7 155H laptop while maintaining high output quality for tasks like physics explanations and code generation. Local inference eliminates the need for cloud API calls, reducing latency and costs while enabling offline work and data privacy.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.