Dec 29, 2025
•
10 min read
Exploiting Copilot Studio's newest feature and exploring protection options
Dec 28, 2025
8 min read
A deep dive into activation space of prompts in safety classifiers. Showing not why - but where - safety fails in LLM classifiers meant to detect malicious prompts.
Dec 20, 2025
13 min read
An in-depth examination of emerging risks and effective mitigation techniques for protecting AI agents operating within the Bedrock AgentCore ecosystem.
Dec 17, 2025
7 min read
A deep dive into realistic threat scenarios and practical strategies for securing enterprise AI agents built in Microsoft Foundry.
Dec 3, 2025
14 min read
How adding a single safety labeled tool to an LLM's toolset can sharply increase its defense
Nov 19, 2025
27 min read
0-click indirect prompt injection with tool use - a look through attribution graphs
Nov 11, 2025
6 min read
How using structured prompts present findings of self-modeling in LLMs, which may benefit both attackers and defenders
Nov 6, 2025
5 min read
How controlling the structure of the prompt, not just the semantics, can exploit your AI agents and their tools
Oct 24, 2025
2 min read
Security research
Oct 23, 2025
11 min read
Oct 21, 2025
15 min read