Meta Prompt Guard 2 misses 99% of buried injection attacks in test
Tool · Dev.to · stat: 1% catch Meta's Prompt Guard 2 security model catches only six of 629 buried prompt-injection attacks during a benchmark of open-source detectors. Developer Rudratosh tests the…
Tool · Dev.to · stat: 1% catch
Meta's Prompt Guard 2 security model catches only six of 629 buried prompt-injection attacks during a benchmark of open-source detectors. Developer Rudratosh tests the systems using real-world exploits from the AgentDojo framework. The silent failures occur because the model's default configuration applies an arbitrary 0.5 decision threshold.
Out-of-the-box AI guardrails are silently failing to protect production agents Founders deploying AI agents must manually calibrate their guardrail thresholds instead of relying on default model settings to prevent silent security breaches.
Every claim ties to a primary source. See our methodology.