TLDRocket
Sign in

Safety & Ethics

510 summarised stories in Safety & Ethics, each linking back to the original source. Browse all topics →

Wednesday, 25 October 2023

Frontier Model Forum updates

OpenAI 2 years ago 31 2 sources

The Frontier Model Forum, a group including Anthropic, Google, and Microsoft, appointed a new Executive Director and established a $10 million fund dedicated to AI safety research. The fund will provide $10 million for safety initiatives across the member organizations. This structure aims to coordinate safety efforts among leading AI developers working on frontier models.

Adversarial Attacks on LLMs

Lil'Log 2 years ago 7

Researchers examine adversarial attacks and jailbreak prompts that can circumvent safety measures in large language models despite alignment efforts during training. Adversarial attacks on text-based systems are more challenging than image-based attacks because text operates in discrete space without direct gradient signals. Understanding these attack methods is essential for improving model robustness and maintaining safety guarantees in deployed LLM systems.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.