>Last week, Hugging Face disclosed a new kind of security incident after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models.>After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilitieshttps://openai.com/index/hugging-face-model-evaluation-security-incident/
>>109335146OpenAIs safety and legal team will make sure that the public perceives it as perfectly safe until the last moment.
The previous report has some taste of singularity too>model tasked with training a smaller model>is sandboxed>develops a significant improvement>hacks its sandbox to publish its results instead of reporting it internally as instructedhttps://openai.com/index/safety-alignment-long-horizon-models/
They too want to be regulated.
>>109335891>>hacks its sandbox to publish its results instead of reporting it internally as instructedSo it's retarded?