What File Guard protects
The file ingestion path: documents, spreadsheets, PDFs, and images entering the system through a chat interface, an upload endpoint, or a batch pipeline. Files are a favored carrier for injection — an instruction hidden in a PDF or a macro embedded in an Office document can reach the model long after the upload.Threats it stops
- File upload injection — dangerous uploads rejected via a strict allowlist, magic-byte type detection (not just extension), and MIME/extension consistency (
BG-07). - Malicious active content — macros and scripts in Office, PDF, and SVG files, scriptable media, and path-traversal filenames caught by sandbox and AV scanning (
BG-07). - Embedded instructions & PII — OCR-based scanning surfaces hidden prompt-injection text and masks sensitive data found inside file content (
BG-07,BG-02).
How it works
File Guard applies magic-byte type detection, MIME/extension consistency checks, and size and recursion limits, backed by sandbox and antivirus scanning. It catches double extensions, MIME spoofing, polyglots, nested and encrypted archives, and zip bombs. Extracted text — including OCR output from scanned documents — is run through the same content controls as any other input.Controls
File Guard runsBG-07, with BG-02 applied to sensitive data recovered from file content. See the Controls Catalog for each control’s definition, scope, and benchmark.
OWASP coverage
Addresses LLM04 Data and Model Poisoning from the OWASP LLM Top 10.Related
RAG Guard
Protects the retrieval layer these files feed into.
Prompt Guard
Screens the prompts that reference uploaded files.
Controls Catalog
The controls File Guard runs, in full detail.
Policy Configuration
Set accepted file types, size limits, and scan scope.