TLDRocket
Sign in

How to Write Data Labeling/Annotation Guidelines

Eugene Yan

Eugene Yan breaks down how to actually write data labeling guidelines, using Google and Bing's real search-rating docs as examples. Turns out good annotation instructions aren't just task lists — they need to explain why, what, and how, or your labels get messy.

Based on reporting by Eugene Yan — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Eugene Yan has written annotation guidelines a handful of times, and each time he started from zero, unsure how to structure the thing. So he sat down and reverse-engineered what makes a guideline actually work, pulling real examples from Google's and Bing's search-rating documentation to make the point concrete rather than theoretical.

His framework boils down to five questions a solid guideline should answer. Why does the task matter — annotators work harder when they understand how their labels feed into search results or model training, not just that a spreadsheet needs filling. What exactly is the task — this means putting annotators in the user's shoes, spelling out intents like "know," "do," or "visit in person," and showing the actual labeling options upfront. What do the terms mean — Google's guidelines define words as basic as "query" and "results," because ambiguous vocabulary quietly wrecks consistency, especially in technical or newly-invented UX contexts.

The biggest chunk of any guideline, Yan argues, is the how-should-annotators-decide section. Google walks reviewers through page purpose, site reputation, and ad placement before they rate page quality. Bing goes further with an actual decision tree for match quality and a step-by-step flow for location quality — Yan notes that shared decision processes like these tend to boost inter-rater agreement, not just annotator confidence. Worked examples with explanations help too, since they calibrate everyone toward the same judgment calls.

Then there's the operational layer: how the task should be performed. Simple things — don't overthink each rating, here's how to break a tie, here's what to do when you're genuinely stuck. Google's own guidelines tell annotators to "use your judgment" nineteen times, which Yan treats as proof that some subjectivity is fine, even expected — that's the whole reason you use human reviewers instead of a rulebook. He also suggests giving annotators an explicit "Unsure" option rather than forcing a guess.

Beyond the guideline text itself, Yan points to task design as a lever: swap numeric scales for binary choices to boost speed and precision, and reword subjective prompts like "is this adult content?" into narrower, more objective ones like "is this nudity?" To know if any of this is working, he recommends tracking Cohen's kappa across iterations — if agreement between raters holds steady or climbs as you refine the guidelines, you're moving in the right direction.

My take — AI-written commentary, not fact-checked reporting

This is the kind of unglamorous infrastructure work that never gets a keynote slide but quietly decides whether your model is any good — garbage labels in, garbage model out, no amount of parameter-count bragging fixes that. I'd bet most AI teams shipping mediocre fine-tunes never wrote a guideline half this rigorous, and it shows in the outputs. If you're serious about data quality, steal Google's decision-tree approach before you steal their transformer architecture.

Read more about this at: Eugene Yan

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.