TLDRocket
Sign in

Mailbag: How to Bootstrap Labels for Relevant Docs in Search

Eugene Yan

An engineer asks how companies determine the total number of relevant documents when evaluating semantic search systems using Recall@K metrics, since manually annotating large datasets is expensive. The response suggests starting with lexical matching in production, then bootstrapping labels from user click data rather than hiring annotators, with human annotation reserved for edge cases and defects. This practical approach avoids the cost of large-scale annotation while building evaluation data from real user behavior.

Why it matters

Building semantic search; how to calculate recall when relevant documents are unknown.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.