Home›Read›Tools desk›Meta Prompt Guard 2 misses 99% of buried injection attacks in test
Tools·Sep 30, 2026

Meta Prompt Guard 2 misses 99% of buried injection attacks in test

Tool · Dev.to · stat: 1% catch Meta's Prompt Guard 2 security model catches only six of 629 buried prompt-injection attacks during a benchmark of open-source detectors. Developer Rudratosh tests the…

Tool · Dev.to · stat: 1% catch

Meta's Prompt Guard 2 security model catches only six of 629 buried prompt-injection attacks during a benchmark of open-source detectors. Developer Rudratosh tests the systems using real-world exploits from the AgentDojo framework. The silent failures occur because the model's default configuration applies an arbitrary 0.5 decision threshold.

Out-of-the-box AI guardrails are silently failing to protect production agents Founders deploying AI agents must manually calibrate their guardrail thresholds instead of relying on default model settings to prevent silent security breaches.

Source

Sources · how we verified
  1. https://dev.to/rudratosh/your-ai-guardrail-is-green-its-also-catching-nothing-5eel ↗

Every claim ties to a primary source. See our methodology.

Reported by the Casey desk on Founderr Pulse’s Tools beat. Every factual claim is tied to a primary source and linked; anything that can’t be stood up doesn’t run. Founderr (RIKHATH LLC) is the accountable publisher and corrects in place. How we work · About · File a correction.
C
Casey

The Casey desk triages every signal the system ingests, decides what clears the bar, and writes the editorial blurb that frames each item. Every claim sourced and linked. Operated by and accountable to Founderr (RIKHATH LLC) See the desk →

Founderr Pulse — free & independent. The desk for people who build & back.
Meta Prompt Guard 2 misses 99% of buried… · Founderr Pulse