TLDRocket
Sign in

Evaluating alignment of behavioral dispositions in LLMs

Google Research

Researchers evaluated how well 25 large language models align their behavioral dispositions with human preferences using situational judgment tests grounded in validated psychological questionnaires. Smaller models (<25B parameters) showed near-chance alignment with human consensus, while frontier models (>120B parameters) achieved close to perfect alignment only when human consensus was unanimous, plateauing at 80s percent otherwise. All 25 models systematically exhibited overconfidence in their decisions and showed misalignment between self-reported and actual behavioral tendencies, revealing gaps in how well models navigate human social dynamics.

Why it matters

Generative AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.