TLDRocket
Sign in

Hardware & Infrastructure

191 summarised stories in Hardware & Infrastructure, each linking back to the original source. Browse all topics →

Thursday, 22 January 2026

Scaling PostgreSQL to power 800 million ChatGPT users

OpenAI Blog 6 months ago

OpenAI scaled PostgreSQL to handle the database demands of supporting 800 million ChatGPT users through replicas, caching, rate limiting, and workload isolation techniques. The system processes millions of queries per second across their infrastructure. This approach allows OpenAI to maintain database performance without replacing PostgreSQL entirely, demonstrating how traditional relational databases can support massive-scale applications.

Optimizing inference speed and costs: Lessons learned from large-scale deployments

Together AI 6 months ago

Together AI describes practical methods for reducing inference latency and cost through optimization techniques including quantization achieving 20-40% throughput improvement, distillation delivering 2-5× lower cost, speculative decoding providing 20-50% faster decoding, and dynamic GPU capacity shifting across endpoints. Teams can reduce TTFT by 50-100ms using regional inference proxies, eliminate GPU compute stalls through kernel fusion and better scheduling, and improve utilization on newer hardware like NVIDIA Blackwell through appropriate parallelism strategies. Organizations implementing these optimizations can achieve faster responses with lower cost per token and better predictability without requiring proportionally larger hardware clusters.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.