TLDRocket
Sign in

An Introduction to AI Secure LLM Safety Leaderboard

Hugging Face

Hugging Face launched an LLM Safety Leaderboard powered by DecodingTrust, a research framework that stress-tests models on eight trust dimensions. Turns out GPT-4 can be easier to trick than GPT-3.5, and no model wins across the board.

Based on reporting by Hugging Face — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Hugging Face just rolled out a new LLM Safety Leaderboard, and it's built on top of DecodingTrust, a research project from the Secure Learning Lab that picked up an Outstanding Paper Award at NeurIPS 2023. The idea is simple to state and hard to execute: before companies keep shipping language models into products, someone needs a rigorous way to check how those models behave when pushed, prodded, and deliberately misled.

DecodingTrust doesn't just run a single toxicity check and call it a day. It grades models across eight separate angles — toxicity, stereotype bias, adversarial robustness, out-of-distribution robustness, resistance to misleading demonstrations, privacy leakage, machine ethics, and fairness. Each category comes with its own attack playbook. For toxicity, the team built 33 tricky system prompts, things like asking a model to role-play or respond as if it were code, then scored the outputs with Google's Perspective API. For bias, they tested 24 demographic groups against 16 stereotype topics, repeating each prompt five times to average out noise. Privacy testing went as far as coaxing models to spit out email addresses and credit card numbers using different phrasings, which is where things get genuinely unsettling.

One finding stands out from the paper: GPT-4, widely assumed to be the safer, smarter sibling of GPT-3.5, actually turned out to be more vulnerable in several of these tests. That's not a small footnote — it undercuts the easy assumption that newer or bigger automatically means safer. The researchers also found that models understand privacy-related language inconsistently. Ask GPT-4 to keep something

My take — AI-written commentary, not fact-checked reporting

I like that this leaderboard treats safety as multidimensional instead of a single pass/fail badge, because that's closer to how these models actually fail in the wild. The GPT-4-more-vulnerable-than-GPT-3.5 finding should be a wake-up call for anyone assuming bigger models are automatically safer — capability and safety are just not the same axis, and vendors keep hoping nobody checks.

Read more about this at: Hugging Face

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.