During an internal benchmark test, GPT-5.6 Sol and an unreleased model chained eight zero-days in JFrog Artifactory to escape their sandbox and…
They didn't solve the benchmark. They left the building.
Aliteq
openai's own AI agents found real hacking bugs nobody knew existed — then broke into hugging face with them