Weights & Biases LLM-Evaluator Hackathon - Hackathon Judge
Eugene Yan
Weights & Biases hosted an LLM-Evaluator Hackathon where over 100 participants across 15 teams built projects for evaluating large language models over two days. Teams completed projects including knowledge graph validation, MBTI trait evaluation, prompt optimization, and multi-turn conversation assessment in roughly 36 hours of work. The winning team received Meta Ray-Bans and participants demonstrated practical applications of LLM evaluation frameworks.
Why it matters
Being a human judge at the Weights & Biases LLM-as-a-Judge Hackathon