TLDRocket
Sign in
Latest Nebius looks to raise $4.5BN through bond issue — Tech.eu Also’s $3,500 e-bike is a $1 billion Trojan horse for autonomous trans... — Fortune Unitree, famous for its dancing robots, surges by 460% on its trading... — Fortune Exclusive: Replit taps OpenAI's low-cost Luna model for new 'Free Mode... — Fortune Adronite launches Codistry AI coding platform, claims half the token c... — SiliconANGLE Rundoo raises $30M to expand its AI-native operating system for small... — SiliconANGLE Temporal is in talks to raise $500M at a $12B pre-money valuation, mor... — Tech Funding News Etched raises $700M led by Jane Street, doubling to $21B and it still... — Tech Funding News

The AI intelligence platform

Every AI story that matters and the intelligence behind it.

TLDRocket reads all relevant sources, removes duplicate coverage, and publishes a short neutral summary of every story, linking back to the original. Free, no spam, unsubscribe anytime.

Add to Slack

Every story also updates live profiles event timelines weekly rankings the AI Market Index

Tuesday, 5 December 2023

Goodbye cold boot - how we made LoRA Inference 300% faster

Hugging Face 2 years ago 2

Hugging Face implemented dynamic LoRA loading on its Inference API, where a single base model remains active while different LoRA adapters are swapped in and out for each user request instead of maintaining separate deployments for each adapter. The warm-up time for serving a LoRA decreased from 25 seconds to 3 seconds, and overall response time dropped from 35 seconds to 13 seconds, while the system now serves hundreds of distinct LoRAs on fewer than 5 A10G GPUs. This approach reduced compute resource requirements since the vast majority of the 2,500 public LoRAs share only a few base models, allowing efficient reuse of GPU capacity across different adapter requests.

Optimum-NVIDIA Unlocking blazingly fast LLM inference in just 1 line of code

Hugging Face 2 years ago 36

Hugging Face released Optimum-NVIDIA, an inference library that accelerates large language model performance on NVIDIA GPUs through a modified single line of code. The library achieves up to 28 times faster throughput and generates 1,200 tokens per second using FP8 quantization on Ada Lovelace and Hopper architectures. Currently supporting LLaMA-based models, the library will expand to other architectures while adding techniques like In-Flight Batching and INT4 quantization.

AMD + 🤗: Large Language Models Out-of-the-Box Acceleration with AMD GPU

Hugging Face 2 years ago 45

AMD and Hugging Face expanded their partnership to enable large language models to run on AMD Instinct server GPUs without requiring code changes compared to NVIDIA equivalents. An AMD Instinct MI250 GPU with 128GB of memory delivers 2.33x more decode throughput and half the prefill latency of an NVIDIA A100, while also fitting larger workloads that exceed the A100's 80GB capacity. Text Generation Inference is now available in production on AMD Instinct GPUs, with future work planned for AMD Radeon consumer GPUs and Ryzen AI laptop processors.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.