TLDRocket
Sign in

Hardware & Infrastructure

261 summarised stories in Hardware & Infrastructure, each linking back to the original source. Browse all topics →

Friday, 19 November 2021

Accelerating PyTorch distributed fine-tuning with Intel technologies

Hugging Face 4 years ago 34

Intel has published guidance on accelerating PyTorch distributed fine-tuning using Intel Xeon Scalable CPU servers with Ice Lake architecture and optimized software libraries. The baseline single-node BERT training job on the MRPC dataset takes 7 minutes and 46 seconds before distributed optimization. The approach involves using Intel's extension for PyTorch and the oneAPI Collective Communications Library to reduce network bottlenecks when distributing training across multiple CPU-based server clusters.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.