What Prompt Guard Detects
Prompt Guard covers the full range of input-layer threats that attackers use to subvert AI application behavior.- Prompt Injection: Malicious instructions embedded inside user input — for example, hidden directives appended to a pasted document — that attempt to override your system prompt or redirect the model’s behavior.
- Jailbreak Attempts: Adversarial prompts carefully crafted to bypass your model’s safety constraints, such as role-play framings, hypothetical scenarios, or encoded payloads designed to make the model ignore its guidelines.
- System Prompt Leakage: Requests that try to extract the contents of your confidential system prompt through direct interrogation, reflection tricks, or indirect inference attacks.
- Scope Violations: Requests that fall entirely outside your application’s defined purpose — for example, asking a customer-support bot to write legal contracts — which may signal misuse or an active probing attempt.
- Multi-modal Attacks: Malicious instructions embedded invisibly or visibly inside images, PDFs, or other documents passed to multi-modal models, exploiting the model’s document-reading capabilities to inject commands.
How It Works
Prompt Guard intercepts every inbound prompt before it is forwarded to your model and runs it through a structured inspection pipeline.1
Input Capture
The prompt — including any retrieved context from RAG pipelines, tool outputs, or conversation history — is captured by Prompt Guard at the ingress point of your application.
2
Intent Extraction and Evaluation
Prompt Guard parses the full prompt structure, separates user-controlled content from trusted system content, and evaluates intent signals against its threat detection models. Each threat category (injection, jailbreak, leakage, scope, multi-modal) is scored independently.
3
Policy Decision
The scores are compared against your configured sensitivity thresholds. Based on the result, Prompt Guard issues one of three decisions: allow (forward the prompt to your model), flag (forward but log the event for review), or block (reject the prompt and return a safe response to the user).
4
Audit Logging
Every evaluated prompt — whether allowed, flagged, or blocked — is recorded in the BeyondGuard audit log with its threat scores, matched patterns, and the policy decision. This log is available in real time from the Control Plane dashboard.
Configuring Prompt Guard
Enable and configure Prompt Guard from the BeyondGuard Control Plane.1
Open Your Project
Navigate to the Control Plane and select the project you want to protect.
2
Open the Guards Tab
Inside your project, select the Guards tab from the left navigation panel.
3
Enable Prompt Guard
Locate Prompt Guard in the guard list and toggle it to Enabled.
4
Choose an Operating Mode
Select either Observation Mode or Enforcement Mode (see the section below). For new deployments, BeyondGuard recommends starting with Observation Mode.
5
Set Sensitivity Thresholds
Expand the Advanced Settings panel to adjust per-threat sensitivity thresholds. Each threat category can be tuned independently on a scale from 1 (permissive) to 10 (strict).
6
Save and Deploy
Click Save Configuration. Changes take effect within seconds for all active sessions on that project.