Security Research from the AI Agent Frontier
Research, tools, and talks for breaking and securing Agents

Inside the Agent Stack: Securing Agents in Amazon Bedrock AgentCore
An in-depth examination of emerging risks and effective mitigation techniques for protecting AI agents operating within the Bedrock AgentCore ecosystem.
Lana Salameh
Inside the Agent Stack: Securing Microsoft Foundry-Built Agents
A deep dive into realistic threat scenarios and practical strategies for securing enterprise AI agents built in Microsoft Foundry.
Lana Salameh
Enabling Safety in AI Agents via Choice Architecture
How adding a single safety labeled tool to an LLM's toolset can sharply increase its defense
Tomer Wetzler
Tools of the Trade
0-click indirect prompt injection with tool use - a look through attribution graphs
Max Fomin
Modeling LLMs via Structured Self-Modeling (SSM)
How using structured prompts present findings of self-modeling in LLMs, which may benefit both attackers and defenders
Tomer Wetzler
Data-Structure Injection (DSI) in AI Agents
How controlling the structure of the prompt, not just the semantics, can exploit your AI agents and their tools
Tomer Wetzler
AgentFlayer: Versión en español.
Inbar Raz
Security ResearchExploring the Risks of ChatGPT’s Atlas Browser

Tamir Ishay SharbatandRaul Klugman-Onitza
Security ResearchInterpreting Jailbreaks and Prompt Injections with Attribution Graphs
Max Fomin
Security ResearchAppendix: Interpreting Jailbreaks and Prompt Injections with Attribution Graphs
Max Fomin
Security ResearchBreaking down AgentKit's Guardrails
A deep dive into OpenAI's AgentKit guardrails, how they are implemented, and where they fail
Stav Cohen
Security ResearchAnalyzing The Security Risks of OpenAI's AgentKit

Stav CohenandRaul Klugman-Onitza
Exhibit & Exploit: Two DEF CON 33 Highlights from the Past & Future of Hacking
Humans, hacker culture and AI: Notes from Hacker Summer Camp
Avishai Efrat
Security ResearchPrompt Mines: 0-Click Data Corruption In Salesforce Einstein
Tamir Ishay Sharbat
Security ResearchAgentFlayer: Minimum Clicks, Maximum Leaks: Tilling ChatGPT’s Attack Surface
Exploiting ChatGPT with Language Alone: A Deep Dive into 0Click and 1Click Attacks.
Dmitry Lozovoy

