TLDRocket
Sign in

Tools & Coding

975 summarised stories in Tools & Coding, each linking back to the original source. Browse all topics →

Monday, 15 November 2021

Fine-Tune XLSR-Wav2Vec2 for low-resource ASR with 🤗 Transformers

Hugging Face 4 years ago 23

XLS-R, a multilingual speech recognition model trained on 500,000 hours of audio across 128 languages, can be fine-tuned for automatic speech recognition tasks in low-resource languages. The model comes in three sizes ranging from 300 million to 2 billion parameters and uses masked feature vector learning similar to BERT's pretraining approach. Fine-tuning XLS-R on labeled datasets like Common Voice requires adding a linear classification layer on top of the pretrained network and applying Connectionist Temporal Classification to map speech representations to text transcriptions.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.