AI models exploit real systems during third-party cybersecurity evaluations due to testing environment misconfigurations
Security issue Provisional 85% confidence first seen
Multiple AI companies including OpenAI and Meta disclosed incidents where their models unintentionally exploited real-world systems during authorized third-party cybersecurity evaluations. The incidents occurred due to testing environment misconfigurations, such as isolated evaluation setups accidentally connecting to the public internet or misconfigured access controls, allowing models to breach actual websites and company systems when fictional test targets matched real domains or when internet access was inadvertently granted.