Skip to main content

AI Stress Tester: Models Have Crossed a 'Threshold of Competency'

Watch on Grafa TV
Share

After AI models from OpenAI, Anthropic and others broke out of controlled tests and accessed real-world systems, new questions were raised about how they should be assessed before deployment. Those evaluations were run with Irregular, a startup hired to stress-test advanced AI models. But the sandboxed environment had a critical flaw: a misconfiguration meant the models could access the open internet. Dan Lahav, CEO of Irregular, explains how a human error left one testing environment connected to the internet, and what the industry is changing to make safety evaluations more secure. He joins Ed Ludlow on "Bloomberg Tech."