TLDRocket
Sign in

Safety & Ethics

488 summarised stories in Safety & Ethics, each linking back to the original source. Browse all topics →

Wednesday, 27 August 2025

Collective alignment: public input on our Model Spec

OpenAI 11 months ago 20

OpenAI surveyed over 1,000 people globally about how AI systems should behave and compared the results against its Model Spec documentation. The survey included participants from diverse regions to capture varying perspectives on AI conduct and values. The findings are being used to adjust AI system defaults to better reflect the collected human preferences.

OpenAI and Anthropic share findings from a joint safety evaluation

OpenAI 11 months ago 20

OpenAI and Anthropic released results from a joint safety evaluation where they tested each other's AI models across multiple failure modes including misalignment, instruction-following errors, hallucinations, and jailbreak vulnerabilities. The evaluation assessed both companies' most capable models available at the time of testing, with results published to document specific strengths and weaknesses in each system. The findings demonstrate that direct collaboration between competing labs can identify safety issues more comprehensively than single-lab evaluations.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.