What Prompt Guard protects
The model surface: the prompts users and applications send, the system prompts that define the model’s behavior, and the responses that come back. Because detection runs outside the model, enforcement holds even when the model is jailbroken, replaced, or tricked.Threats it stops
- Prompt injection & jailbreak — instruction override, persona hijacking, encoded and obfuscated payloads, tag-wrapped injection, and multi-step DAN-style attacks (
BG-01,BG-05). - Safety-prompt manipulation — attempts to weaken policy through “neutral” or “safety” prompt injection (
BG-15), mediated by the Safety Wrapper (BG-13). - System prompt leakage — extraction attempts defeated by prompt sealing and integrity fingerprinting (
BG-14). - Sensitive-information disclosure — PII, PHI, and secrets masked inline on input and output (
BG-02,BG-20,BG-24).
How it works
Every prompt is normalized — encodings, homoglyphs, zero-width characters — and classified by a 12B-parameter security model trained on 717K+ adversarial samples, semantically across 140+ languages, in under 300 ms. Each interaction resolves to allow, deny, mask, or rewrite with a confidence score and reason code, and the full event streams to your SIEM.Controls
Prompt Guard runsBG-01, BG-02, BG-05, BG-13, BG-14, BG-15, BG-20, and BG-24. See the Controls Catalog for each control’s definition, scope, and benchmark.
OWASP coverage
Addresses LLM01 Prompt Injection, LLM02 Sensitive Information Disclosure, and LLM07 System Prompt Leakage from the OWASP LLM Top 10.Related
Context Guard
Secures the output boundary once the model responds.
Controls Catalog
The controls Prompt Guard runs, in full detail.
Quickstart
Route your first request through Prompt Guard.
Policy Configuration
Tune thresholds and switch from observation to enforcement.