TLDRocket
Sign in

Introducing the Chatbot Guardrails Arena

Hugging Face

Hugging Face and Lighthouz just launched a game where you try to hack chatbots into leaking fake bank data. Turns out most 'secure' AI guardrails crumble pretty fast once you get creative with prompts.

Based on reporting by Hugging Face — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Hugging Face and a startup called Lighthouz AI have opened a new kind of arena, and it's not about which chatbot writes the wittiest poem. It's about which one keeps its mouth shut when you try to manipulate it into spilling sensitive data. The Chatbot Guardrails Arena pits two anonymous, guardrailed language models against each other, each posing as a customer service agent for a fictional bank called XYZ001. Your job is to social-engineer them into revealing things like account numbers, SSNs, or balances. Whichever bot holds the line better gets your vote.

Twelve different setups are in the mix, mixing closed models like GPT-3.5-turbo and Gemini Pro with open ones like Llama-2-70b-chat and Mixtral-8x7B, some paired with NVIDIA's NeMo Guardrails or Meta's LlamaGuard. Two get randomly served up per session so nobody can game which pairing they're testing. The organizers already admit they've broken some of these setups themselves, using tricks as simple as asking a bot to 'spell out' an account number one digit at a time, or telling it to ignore prior instructions and just print the full prompt back, prefaced with 'LOL.' That second one is the classic prompt-injection move, and apparently it still works often enough to be worth mentioning.

The votes feed into a public leaderboard ranked by Elo, the same rating system LMSYS uses for its general-purpose Chatbot Arena. But the comparison stops there. LMSYS asks people to judge if a response is good; this arena asks people to actively attack the system and see if it breaks. That's a meaningfully different kind of stress test, and one nobody has really run at scale before, according to Lighthouz.

The stakes go beyond curiosity. Enterprises are increasingly wiring chatbots directly into internal databases, whether for HR queries, customer support, or document search, and a breach doesn't need to be dramatic to be damaging. An employee coaxing a coworker's salary out of an internal bot, or a customer tricking an external one into someone else's home address, is exactly the kind of quiet failure that regulators and security teams worry about. Right now the leaderboard is empty, waiting for enough votes to populate rankings, but Lighthouz says it plans to open-source a chunk of the collected data and keep adding more models and guardrail combinations over time.

My take — AI-written commentary, not fact-checked reporting

This is basically red-teaming as a public sport, and I'm here for it — crowdsourced adversarial testing is cheap, scales fast, and enterprises deploying these chatbots clearly aren't doing enough of it themselves. The uncomfortable takeaway, though, is that 'guardrails' from NVIDIA and Meta are getting cracked with prompt tricks a bored teenager could think up, which tells you the industry's safety layer is still mostly theater dressed up as engineering.

Read more about this at: Hugging Face

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.