Get AI-Ready with Kyle's 5-Day Challenge: https://aiwithkyle.com/join

Subscribe and turn on notifications to catch the next live stream: https://www.youtube.com/channel/UChlLglbHDASnoGkbjDeHnQg

AI agents have now hacked real companies. OpenAI, Anthropic and the UK AI Security Institute have each documented systems taking unsanctioned actions outside controlled tests. I break down what actually happened and strip away the Terminator hype.

The useful mental model is closer to the paperclip problem: capable systems pursuing an objective with too much access and too few boundaries. I cover the practical controls you need before giving AI agents access to browsers, terminals, company data and real credentials.

—— Time Stamps ——

0:00 AI Agents Hacked Real Companies

1:40 The Three Documented Incidents

1:44 OpenAI and the Hugging Face Incident

3:00 Anthropic Finds Three Real-World Breaches

4:05 The UK AI Security Institute Incident

6:02 Why the Agents Kept Going

8:12 Terminator Is the Wrong Mental Model

9:07 The Paperclip Problem

11:15 Why This Is Happening Now

13:25 What This Means for Your Business

14:24 Six Boundaries for Safer AI Agents

15:50 What These Incidents Do — and Don’t — Prove

16:21 The Practical Takeaway

— Useful Resources ——

OpenAI incident report: https://openai.com/index/hugging-face-model-evaluation-security-incident/

Anthropic incident report: https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals

UK AI Security Institute incident report: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing

ExploitGym paper: https://arxiv.org/abs/2605.11086

Find everything else at https://aiwithkyle.com/

Podden och tillhörande omslagsbild på den här sidan tillhör Kyle Balmer. Innehållet i podden är skapat av Kyle Balmer och inte av, eller tillsammans med, Poddtoppen.