How Guards work
Every Guard follows the same core pipeline: inspect → evaluate → act. When an interaction reaches a protected layer, the Guard intercepts it before it proceeds. It evaluates the interaction against your configured policies — checking for known threat signatures, contextual anomalies, scope boundaries, and compliance rules — using a purpose-built security model that produces a confidence score. Based on that score and your thresholds, the Guard resolves the interaction to one of four enforcement outcomes: Allow, Deny, Mask, or Rewrite, each written to the audit log with a reason code. This pipeline runs at every node independently. A Guard at the MCP/Tool layer does not rely on the Prompt Guard having already caught a threat upstream — it makes its own determination based on what it can observe at its own layer.The six Guards
Prompt Guard
Inspects every prompt and retrieved context before inference. Stops prompt injection, jailbreaks, system prompt leakage, and sensitive-data disclosure.
Context Guard
Treats model responses as untrusted input to the rest of your stack. Validates output schemas, neutralizes script and markup, and enforces topic scope.
RAG Guard
Protects the retrieval pipeline and vector database. Binds answers to authorized sources and blocks RAG poisoning and embedding-layer attacks.
File Guard
Screens every uploaded and ingested file. Catches embedded instructions, malicious active content, and poisoned documents before they reach the model.
Agent Guard
Monitors autonomous agent behavior across the whole reasoning chain. Catches plan deviation, memory poisoning, and infinite loops.
MCP Guard
Validates every tool call and MCP server interaction. Blocks tool poisoning, parameter injection, and schema violations under deny-by-default.