TLDRocket
Sign in

Thinking about High-Quality Human Data

Lilian Weng

High-quality human-annotated data is essential for training deep learning models, including classification tasks and RLHF labeling for LLM alignment. The community recognizes data quality's importance but faces a cultural preference for model development over the unglamorous work of careful data collection and annotation. This imbalance risks compromising model performance when the foundational data work receives insufficient attention and resources.

Why it matters

[Special thank you to Ian Kivlichan for many useful pointers (E.g. the 100+ year old Nature paper “Vox populi”) and nice feedback. 🙏 ] High-quality data is the fuel for modern data deep learning model training. Most of the task-specific labeled data comes from human annotation, such as classification task or RLHF labeling (which can be constructed as classification format) for LLM alignment training. Lots of ML techniques in the post can help with data quality, but fundamentally human data collection involves attention to details and careful execution. The community knows the value of high quality data, but somehow we have this subtle impression that “Everyone wants to do the model work, not the data work” (Sambasivan et al. 2021).

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.