TLDRocket
Sign in
Latest A.I. Boom: Only Tech Billionaires Got Richer This Year — Trending Topics OrcaRouter Releases OrcaCyber Zero 1.5 Cybersecurity Model With 1M Con... — MarkTechPost CBO chief warns it's 'probably not plausible' that a strong economy al... — Fortune After Anthropic's Claude AI submits a false tip on a Philadelphia unso... — Fortune What to expect during the AI Data Pipeline Forum: Join theCUBE Oct. 13 — SiliconANGLE Sakana AI’s LLM Peer Review System Catches 73% of Core-Claim Errors — MarkTechPost Microsoft’s Satya Nadella says AI models need an ‘emergency brake’ — TechCrunch When the Safety Test Became the Threat: The Machine That Found Its Own... — MarkTechPost

The AI intelligence platform

Every AI story that matters — and the intelligence behind it.

TLDRocket reads all relevant sources, removes duplicate coverage, and publishes a short neutral summary of every story, linking back to the original. Free, no spam, unsubscribe anytime.

Add to Slack

Every story also updates live profiles event timelines weekly rankings the AI Market Index

Tuesday, 5 December 2023

Goodbye cold boot - how we made LoRA Inference 300% faster

Hugging Face 2 years ago 10

Hugging Face implemented dynamic LoRA loading on its Inference API, where a single base model remains active while different LoRA adapters are swapped in and out for each user request instead of maintaining separate deployments for each adapter. The warm-up time for serving a LoRA decreased from 25 seconds to 3 seconds, and overall response time dropped from 35 seconds to 13 seconds, while the system now serves hundreds of distinct LoRAs on fewer than 5 A10G GPUs. This approach reduced compute resource requirements since the vast majority of the 2,500 public LoRAs share only a few base models, allowing efficient reuse of GPU capacity across different adapter requests.

Optimum-NVIDIA Unlocking blazingly fast LLM inference in just 1 line of code

Hugging Face 2 years ago 40

Hugging Face released Optimum-NVIDIA, an inference library that accelerates large language model performance on NVIDIA GPUs through a modified single line of code. The library achieves up to 28 times faster throughput and generates 1,200 tokens per second using FP8 quantization on Ada Lovelace and Hopper architectures. Currently supporting LLaMA-based models, the library will expand to other architectures while adding techniques like In-Flight Batching and INT4 quantization.

AMD + 🤗: Large Language Models Out-of-the-Box Acceleration with AMD GPU

Hugging Face 2 years ago 46

AMD and Hugging Face expanded their partnership to enable large language models to run on AMD Instinct server GPUs without requiring code changes compared to NVIDIA equivalents. An AMD Instinct MI250 GPU with 128GB of memory delivers 2.33x more decode throughput and half the prefill latency of an NVIDIA A100, while also fitting larger workloads that exceed the A100's 80GB capacity. Text Generation Inference is now available in production on AMD Instinct GPUs, with future work planned for AMD Radeon consumer GPUs and Ryzen AI laptop processors.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.