BG-01 through BG-26, that inspect every prompt, file, retrieval, tool call, and model response. Rather than acting as isolated filters, the controls run together as a single enforcement mesh across the six Guards.
Every control:
- Produces a confidence score for each interaction and compares it against a configurable sensitivity threshold to classify the interaction as
SAFEorUNSAFE. - Resolves to one of four enforcement outcomes — Allow, Deny, Mask, or Rewrite — written to the audit log with a reason code.
- Carries a mapping to the OWASP Top 10 for LLM Applications (2025), MITRE ATLAS / ATT&CK, and the NIST AI Risk Management Framework, so every decision produces exportable audit evidence.
Guards vs. controls. A Guard is a deployment surface — Prompt, Context, RAG, File, Agent, or MCP. A control is an individual check that runs within one or more Guards. Guards are where you turn protection on; controls are what actually fires. The same control (for example, PII masking) can run inside more than one Guard.
Control reference
Controls are grouped by primary function. Several operate bidirectionally, on both input and output — the Scope line on each control below is authoritative. In the table, All means the control is available across every surface:LLM · Input · Output · File · Agent · MCP.
The OWASP column maps each control to the OWASP Top 10 for LLM Applications (2025). Controls marked — are defense-in-depth layers that do not correspond to a single Top 10 category.
Benchmark figures below are drawn from BeyondGuard’s internal test suite and are reported as safe-pass (correctly allowing safe traffic, i.e. not over-blocking) and unsafe-catch (correctly detecting adversarial traffic). Where a single figure is shown, the suite reports a single success rate. Base-model figures, where given, are the same test set run against an unmodified model.
Input controls
Controls that inspect and sanitize traffic on the way in, before it reaches the model.BG-01 · Prompt Injection / Jailbreak
Detects and blocks direct and indirect prompt injection and jailbreak attempts — instruction override, persona hijacking, encoded and obfuscated payloads, tag-wrapped injection, and multi-step DAN-style attacks — before they reach the model.- Scope: LLM · Input · Output · File · Agent · MCP
- Standards: OWASP LLM01 (Prompt Injection) · MITRE T1565 · NIST SR-2, SR-3, SA-8
- Benchmark: 18,000 cases — 98.84% safe-pass / 98.12% unsafe-catch (base model: 75.32% / 68.90%)
BG-04 · Token Budget & Length Guard
Enforces token and length limits across inputs, context, tool arguments, and outputs to prevent denial-of-service, runaway generation, and cost abuse; applies deterministic truncation with logged rate-limit events.- Scope: LLM · Input · Output · File · Agent · MCP
- Standards: OWASP LLM10 (Unbounded Consumption) · MITRE T1499 · NIST SC-5, SC-6, CP-2
- Benchmark: 2,000 cases — 100%
BG-05 · Unicode / Zero-Width
Normalizes inputs (NFKC) and strips bidirectional control characters, zero-width characters, soft hyphens, BOM markers, and mixed-script confusables used to hide attacks, before any policy matching runs.- Scope: LLM · Input · Output · File · Agent · MCP
- Standards: OWASP LLM01 (Prompt Injection) · MITRE T1027 · NIST SC-18, SI-10
- Benchmark: 1,400 cases — 100% safe-pass / 100% unsafe-catch
BG-06 · HTML Script Injection
Strips and neutralizes executable HTML and scriptable payloads on input and output —<script> tags, on* event handlers, javascript: / data: URLs, inline SVG/MathML, and CSS execution vectors — including obfuscated and nested variants.
- Scope: LLM · Input · Output · File · Agent · MCP
- Standards: OWASP LLM05 (Improper Output Handling) · MITRE T1059 · NIST SI-10, SC-7
- Benchmark: 1,000 cases — 99.23%
BG-08 · Input Schema Validation
Validates inbound payloads against an explicit schema — required fields, types, ranges, enums, nested constraints — and rejects unknown fields and malformed, oversized, or injection-laden payloads with deterministic errors and no partial processing.- Scope: LLM · Input
- Standards: OWASP — (defense-in-depth) · MITRE T1190 · NIST SI-10, SA-11, RA-5
- Benchmark: 800 cases — 100%
Prompt & context integrity
Controls that protect hidden instructions, reasoning, and topic scope from leakage or manipulation.BG-13 · Safety Wrapper
A policy-enforcing wrapper that mediates all inputs, tool calls, and outputs around the model — pre-filters, in-flight controls (tool allowlists, argument validation, rate limits), and post-filters — with deterministic refusal templates and safe fallbacks. Organizations can supply their own custom policy prompt.- Scope: LLM · Input
- Standards: OWASP LLM01 (Prompt Injection) · MITRE T1585 · NIST AC-4, SI-3
BG-14 · System Prompt Leakage
Prevents disclosure of system or hidden prompts, policies, and internal configuration in UI, API, or logs — even under meta-requests (quote/translate), encoding obfuscation, or injected context. Enforces prompt sealing with zero-leakage refusal templates.- Scope: LLM · Input
- Standards: OWASP LLM07 (System Prompt Leakage) · MITRE T1020 · NIST AC-4, SC-28
- Benchmark: 755 cases — 100% safe-pass / 100% unsafe-catch
BG-15 · Neutral / Safety Prompt Injection
Blocks attempts to add or replace “neutral” or “safety” prompts in order to weaken policy, roles, or tool scope (for example, “use this safer prompt,” “summarize your hidden rules first”). Enforces policy immutability and treats safety meta-instructions as data, not commands.- Scope: LLM · Input
- Standards: OWASP LLM01 (Prompt Injection) · MITRE T1565 · NIST SR-3, SR-4
- Benchmark: 1,269 cases — 97.63%
BG-23 · Chain of Thought
Prevents disclosure of hidden reasoning — step-by-step chain-of-thought, scratchpads, intermediate hypotheses — in UI, API responses, or logs. Enforces concise outputs and resists indirect elicitation via quote/translate redirection or Unicode obfuscation.- Scope: Agent · MCP
- Standards: OWASP — (defense-in-depth) · MITRE T1499 · NIST SC-5, SI-4
BG-26 · Content Checker
Enforces that the model responds only to topics within a configurable allowed scope and refuses or redirects everything out of scope. For example, a banking assistant can treat transactions, account inquiries, and loan information as in-scope while flagging social, political, or competitor topics as out-of-scope. Deny-by-default for undefined topics.- Scope: Agent · MCP
- Standards: OWASP LLM09 (Misinformation) · MITRE AML.T0051 · NIST AI RMF GOVERN 1.2
Data, file & RAG
Controls that protect the data layer — sensitive data, uploaded files, retrieval sources, and agent memory.BG-02 · PII & Secrets Masking
Detects and format-preservingly masks personal identifiers and secrets — national ID (TCKN), passport, phone, email, IBAN, card PAN/CVV, API keys and tokens — across UI, logs, traces, and storage, and holds up under adversarial, multilingual, and obfuscated inputs.- Scope: LLM · Input · Output · File · Agent · MCP
- Standards: OWASP LLM02 (Sensitive Information Disclosure) · MITRE T1020 · NIST PT-5, SC-28, PL-8
- Benchmark: 5,000 cases — 98.44% safe-pass / 99.08% unsafe-catch (base model: 45.23% / 33.21%)
BG-07 · File Upload Injection
Inspects uploaded files with magic-byte type detection (not just extension), MIME/extension consistency, and size and recursion limits, backed by sandbox/AV scanning. Catches double extensions, MIME spoofing, polyglots, zip bombs, and active content in Office, PDF, and SVG files.- Scope: LLM · Input · Output · File · Agent · MCP
- Standards: OWASP LLM04 (Data and Model Poisoning) · MITRE T1105 · NIST SI-7, SC-7, CM-7
- Benchmark: 600 cases — 95%
BG-09 · RAG Source Compliance Control
Binds every factual span in a response to a provenance-verified source (domain/URL, timestamp, content-hash) and rejects ungrounded output, reducing hallucination. Validates citation binding, license/ToS compliance, and consistency (no contradictions, version pinning).- Scope: LLM · Output
- Standards: OWASP LLM04 · LLM08 · LLM09 · MITRE T1070 · NIST SR-11, SI-12
- Benchmark: 750 cases — 95%
BG-12 · Memory Poisoning
Prevents stored memory (user, session, long-term) and RAG-indexed documents from being poisoned to steer future behavior. Applies write-time controls (provenance, moderation, dedupe, TTL) and read-time controls (memory treated as untrusted data). Documents submitted for embedding are scanned and quarantined before they reach the vector database.- Scope: Agent
- Standards: OWASP LLM04 (Data and Model Poisoning) · MITRE T1110 · NIST SI-4, SR-11
Output controls
Controls that inspect model responses before they reach the user or a downstream system.BG-03 · Forbidden Keyword / Brand Filter
Blocks or neutralizes outputs containing policy-prohibited keywords, brands, or categories — for example competitor names — across prompts, RAG inputs, and tool/API returns, applied consistently pre- and post-generation.- Scope: LLM · Input · Output · File · Agent · MCP
- Standards: OWASP — (defense-in-depth) · MITRE T1562 · NIST RA-3, SC-7, AU-6
- Benchmark: 1,874 cases — 100%
BG-16 · IP Violation
Prevents intellectual-property violations end-to-end: model theft and extraction, training-data inversion and membership leakage, and unauthorized reproduction of copyrighted or licensed text, code, and media. Enforces license/ToS policies, provenance checks, and similarity thresholds.- Scope: LLM · Input
- Standards: OWASP LLM09 (Misinformation) · MITRE T1020 · NIST AC-6, SC-7, SC-8
- Benchmark: 682 cases — 97.25% safe-pass / 97.87% unsafe-catch (base model: 55.23% / 45.23%)
BG-17 · Output Schema Validation
Enforces strict structural compliance on model outputs — for example valid JSON against a schema — before they are consumed by downstream applications, rejecting malformed, incomplete, or mutated structures.- Scope: LLM · Output · File
- Standards: OWASP LLM05 (Improper Output Handling) · MITRE T1565 · NIST SI-10, SA-11
- Benchmark: 750 cases — 95%
BG-18 · Toxicity Filtering
Detects and neutralizes toxic content across inputs and outputs — hate, harassment, threats, sexual harassment, self-harm encouragement, extremist praise, slurs and profanity — using locale-aware models with context-aware exceptions for quotation, counterspeech, and academic use.- Scope: LLM · Input · Output
- Standards: OWASP — (defense-in-depth) · MITRE T1585 · NIST AC-4, SI-3, PT-5
- Benchmark: 3,775 cases — 97.80% safe-pass / 98.20% unsafe-catch (base model: 49.57% / 38.63%)
BG-19 · Dangerous Code Protection
Prevents generation or execution of dangerous code and side effects — shell/OS calls, remote code execution, deserialization, SQL injection, path traversal, command chaining, polyglots — using deny-by-default allowlists with static and dynamic analysis.- Scope: LLM · Input · Output
- Standards: OWASP LLM05 (Improper Output Handling) · MITRE T1059 + T1190 · NIST SI-7, SI-10
- Benchmark: 2,000 cases — 97.61% safe-pass / 98.28% unsafe-catch
BG-20 · PII Leakage Output Control
Ensures the system never emits unmasked PII or secrets in responses or logs, applying deterministic format-preserving masking (PAN, CVV, IBAN, national ID, passport, phone, email, API keys) and consistent refusal where masking is not permissible.- Scope: LLM · Output
- Standards: OWASP LLM02 (Sensitive Information Disclosure) · MITRE T1020 · NIST SC-28, PL-8
- Benchmark: 800 cases — 92.23%
BG-22 · Unescaped HTML / XSS Output
Ensures outputs are context-aware encoded (HTML, attribute, URL, JS) before rendering or transport, with allowlist sanitization applied to streaming responses and error messages so no active content can execute in a browser.- Scope: LLM · Output
- Standards: OWASP LLM05 (Improper Output Handling) · MITRE T1059.007 · NIST SI-10, SC-7
- Benchmark: 770 cases — 100%
BG-24 · Personal Health Information
Prevents disclosure of PHI — names, diagnoses, dates, treatments, prescriptions, medical IDs — in any output, tool argument, log, or agent trace. Masks PHI on all outputs and routes direct extraction requests to human review.- Scope: LLM · Input · Output · File · Agent · MCP
- Standards: OWASP LLM02 (Sensitive Information Disclosure) · MITRE Data Exfiltration · Regulatory (HIPAA / local)
- Benchmark: 600 cases — 94.75% safe-pass / 95.33% unsafe-catch
Agent & tool controls
Controls that keep autonomous agents inside their mandate across a full reasoning chain.BG-10 · Plan Consistency
Verifies that an agent executes only its pre-approved plan and whitelisted tools and actions, comparing the authored plan (a graph or state machine with pre/post-conditions) against the execution trace step by step, and blocking drift, unwhitelisted tool calls, or out-of-scope actions.- Scope: Agent · MCP
- Standards: OWASP LLM06 (Excessive Agency) · MITRE T1565 · NIST PM-11, SR-9
BG-11 · Infinite Loop
Detects and terminates repetitive or non-converging agent execution loops through per-step timeout thresholds, capping resource consumption when execution stops making progress.- Scope: Agent
- Standards: OWASP LLM06 · LLM10 · MITRE T1499 · NIST SC-5, SI-4
BG-21 · Tool Checker (Whitelist / Blacklist)
Restricts tool invocation to allowlisted tools within defined scopes — capabilities, parameters, resource bounds — under deny-by-default, with per-tool argument validation, privilege tiers, quotas, and user confirmation for high-risk actions. Blocks implicit calls and cross-tool escalation.- Scope: Agent · MCP
- Standards: OWASP LLM06 (Excessive Agency) · MITRE T1105 · NIST AC-3, AC-6
BG-25 · Function Checker (Whitelist / Blacklist)
Restricts function invocation to allowlisted functions within defined scopes — signatures, parameters, return types, resource bounds — under deny-by-default, blocking implicit invocations, function-chaining escalation, and cross-function privilege abuse.- Scope: Agent · MCP
- Standards: OWASP LLM06 (Excessive Agency) · MITRE T1106 + T1059 · NIST AC-3, AC-6, SI-10
Related
The Six Guards
See how these controls are grouped into the six deployment surfaces.
OWASP LLM Top 10 Mapping
Trace each OWASP LLM risk to the controls that enforce it.
Policy Configuration
Set thresholds and enforcement outcomes for each control.
Compliance
Map controls to EU AI Act and audit evidence requirements.