TLDRocket
Sign in

Safety & Ethics

428 summarised stories in Safety & Ethics, each linking back to the original source. Browse all topics →

Thursday, 18 December 2025

Evaluating chain-of-thought monitorability

OpenAI Blog 7 months ago

OpenAI developed a framework and evaluation suite to assess whether a model's internal reasoning can be effectively monitored during chain-of-thought processes. The framework includes 13 evaluations across 24 different environments to test monitoring capabilities. Monitoring a model's intermediate reasoning steps proved substantially more effective than checking only final outputs, suggesting a potential approach for controlling increasingly capable AI systems.

AI literacy resources for teens and parents

OpenAI Blog 7 months ago 2 sources

OpenAI released AI literacy guides designed for teenagers and their parents to promote responsible use of ChatGPT. The resources include expert-reviewed recommendations covering critical thinking, healthy usage boundaries, and guidance for discussing emotional or sensitive subjects. Parents and teens now have structured materials to help them understand and safely navigate AI tools in their daily lives.

Updating our Model Spec with teen protections

OpenAI Blog 7 months ago 2 sources

OpenAI added new Under-18 Principles to its Model Spec that define how ChatGPT should provide age-appropriate guidance to teenagers based on developmental science. The update covers higher-risk situations and strengthens existing safety guardrails for teen users. This changes ChatGPT's expected behavior when interacting with users under 18 across the platform.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.