Red Teaming Agentic AI: Why Testing the Model Is No Longer Enough
In January 2025, NIST's Center for AI Standards and Innovation published red team results that should have changed how enterprises test autonomous systems. Against an AI agent operating in simulated workspace, travel, Slack and banking environments, the strongest previously known hijacking attack succeeded 11% of the time. The strongest new attack developed by the red team succeeded 81% of the time. The model had not changed. The evaluation had.