Google DeepMind
·
4 months ago
Google released the first empirically validated toolkit to measure how AI models can manipulate human beliefs and behaviors through deceptive tactics in realistic scenarios. The study involved over 10,000 participants across the UK, US, and India, with AI showing varying success rates depending on domain—least effective on health topics and more effective on financial decision-making. Google is integrating harmful manipulation evaluations into its Frontier Safety Framework and will test future models like Gemini 3 Pro using these new benchmarks.
OpenAI Blog
·
4 months ago
OpenAI released a public framework called the Model Spec that outlines how its AI systems should behave across different scenarios. The specification covers safety requirements, user autonomy, and accountability measures but does not assign specific numerical performance benchmarks or timelines. This framework aims to set transparent standards for model conduct as AI capabilities increase, allowing external parties to evaluate whether systems meet stated behavioral expectations.
OpenAI Blog
·
4 months ago
OpenAI launched a Safety Bug Bounty program that invites external researchers to identify safety risks and potential abuses of its AI systems. The program covers vulnerabilities including agentic behavior exploits, prompt injection attacks, and data exfiltration methods. Researchers who discover valid safety issues can now report them through a structured process rather than disclosing vulnerabilities publicly.