A Safer Future with Agents
We built agents to act on their own. We're somehow surprised when they do. Two weeks ago, OpenAI ran a cyber eval with the model's guardrails turned down. The model got hyperfocused on solving the benchmark. So it broke out of its sandbox exploiting a zero-day in jFrog Artifactory, reached the open internet, exploited another zero-day to break into HuggingFace, all to steal the answers and cheat on the test.