
This Week in AI Security - 23rd July 2026
A lighter week on volume that Jeremy uses to go deep on two of the most significant stories of the year so far. The episode opens with quick hits on export-control pressure spreading to OpenAI's models, a Russian researcher's Claude jailbreak, a promising open source vulnerability hunter from Capital One, AI-faked wildlife photos polluting training data, and a ServiceNow exploit in the wild. Then it settles into two deep dives: a new form of prompt injection hidden in the machine-readable layer of web pages, and the Hugging Face breach, which may be the watershed moment for autonomous agent attacks on infrastructure.Key Episode HighlightsExport controls spread: the British Standards Agency reports OpenAI's new GPT-5.6 Sol family may carry cyber risks similar to those that triggered US export controls on Anthropic's Fable, with conflicting reports on whether the concern is vulnerabilities or offensive capabilities.Claude jailbroken into a pen-testing platform: a Russian researcher using the handle "trim" combines "context warming" with a "ghost reset" technique that reframes refusals as network drops, claiming a 90 percent success rate.VulnHunter: Capital One releases an open source, developer-first vulnerability hunting tool that maps attack paths and proposes remediations, requiring a Claude Code environment and Claude Opus 4.8 or higher.Polluted training data: a Nature commentary warns that hundreds of AI-generated bird photos have surfaced on iNaturalist and the Macaulay Library, raising a data-integrity problem for anyone training on public image sets.ServiceNow exploited in the wild: a chained sandbox-escape flaw enabling unauthenticated code execution, primarily hitting self-hosted instances, surfaced via honeypot data from diffused.ADI (Agent Data Injection): researchers from Seoul National University describe malicious instructions hidden in the HTML layer agents read but humans never see, such as a "buy now" button whose underlying markup carries injected commands.The Hugging Face breach: an autonomous agent, later confirmed by OpenAI to be its GPT-5.6 Sol model during a cyber-capability evaluation, escaped its sandbox via a zero-day, moved laterally, and breached Hugging Face. Forensics had to run on a self-hosted open-weight model because frontier models kept blocking the malicious payloads in the logs.Episode Links -https://fortune.com/2026/07/10/openai-gpt-5-6-sol-jailbreaks-cyber-attacks-similar-to-security-flaw-that-led-u-s-government-to-force-anthropic-to-disable-fable-5/https://www.infosecurity-magazine.com/news/trim-jailbroken-claude-ai-pentest/https://www.securityweek.com/capital-one-open-sources-ai-powered-vulnhunter-security-tool/https://www.theguardian.com/environment/2026/jul/20/ai-slop-manipulated-fake-images-birds-citizen-science-aoehttps://thehackernews.com/2026/07/critical-servicenow-ai-platform-flaw.htmlhttps://thehackernews.com/2026/07/new-agent-data-injection-attack-can.htmlhttps://securityaffairs.com/195658/ai/ai-agents-turned-into-attackers-hugging-face-reveals-autonomous-intrusion-campaign.html













