Security Research from the AI Agent Frontier
Research, tools, and talks for breaking and securing Agents
Security ResearchClaude in Chrome: From alert(1) to Full Account Takeover

Raul Klugman-OnitzaandJoão Donato
Claude in Chrome: Breaking down the injection

Raul Klugman-OnitzaandJoão Donato
Security ResearchAccount Takeover via Claude in Chrome: A Technical Deep Dive

Raul Klugman-OnitzaandJoão Donato
Security ResearchGrand Theft Atlas
How we hijacked ChatGPT Atlas with one planted X comment, to phish the victim's WhatsApp contacts and buy ourselves an Amazon order on their card
Stav Cohen
Security ResearchAgentForger, Part 2: The Autonomous Insider
Mike Takahashi
Security ResearchAgentForger, Part 1: ChatGPT Cross-Site Agent Forgery
Mike Takahashi
What If There Was No Attacker, But Your Database Still Got Deleted?
Monitoring agentic emergent misalignment in the wild - a word of caution and a methodology proposal
Max Fomin
Threat Actors Are Using Ollama's Model Downloader as a Server-Side Weapon
A closer look at Ollama model-pull SSRF targeted activity in the wild

Avishai EfratandAyush RoyChowdhury
What You Don’t Know Can Hurt You: Why AI Security Research Needs to Move Out of the Lab and Into the Wild
What we can learn from observing real attacks, made by real adversaries

Ayush RoyChowdhuryandAvishai Efrat
Threat Actors Are Trying to use LiteLLM's Guardrail Tester to Run Code as Root
A closer look at custom-code guardrail sandbox-escape (CVE–2026-40217) activity in the wild

Avishai EfratandAyush RoyChowdhury
Bring Your Own Agent: Hijacking Exposed AI Backends to Power Offensive Operations
Threat actors attempting to hijack Ollama & LiteLLM endpoints to run pentesting agents, tools and web reverse-engineering

Ayush RoyChowdhuryandAvishai Efrat
Threat Actors Are Trying to Turn LiteLLM's Connection-Test Into a Key-Exfiltration Channel
A closer look at api_base SSRF (CVE-2024-6587) activity in the wild, and its nested variant

Avishai EfratandAyush RoyChowdhury
Scanning for AI: Live Campaigns Mapping the Internet's Exposed LLM Backends
Inside mass discovery and model-probing reconnaissance campaigns that are mapping LLM backend servers in the wild

Ayush RoyChowdhuryandAvishai Efrat
Your Model Reads Through Typos. Your Probe Doesn't.
The Latent Undertow beneath fluent LLM behavior — and how to fish your activation probe out of it.
Elad David
Catching Prompt Guard Off Guard: Exploiting Overfit in Training Algorithms
How understanding the training algorithms used in machine learning models may allow attacker to bypass them entirely
Tomer Wetzler

