Skip to main content
BeyondGuard is deployed as a containerized microservice stack that sits inline in your AI traffic and inspects every interaction before it reaches a model, a tool, or a user. The platform runs fully on-premise with no cloud dependency — no data has to leave your environment — and is organized into six cooperating layers, each with a distinct role.

The six layers

The Core Layer runs a pool of dedicated guard workers that execute the platform’s inspection controls in parallel — one worker per control in the Controls Catalog — backed by an OCR service (Tesseract plus vision-language models) that extracts text from images and scanned documents before inspection. The Artificial Layer hosts the embedded BG Model: an on-device security model, served locally via vLLM, that performs threat scoring, classification, and response shaping without any external API call. This is what keeps detection independent of the model being protected.

End-to-end request flow

1

Client request

A chat, LLM, RAG, agent, proxy, or mail request arrives at the system boundary.
2

Adapter normalization

The matching adapter normalizes the request into the standard internal format.
3

Gateway

Rate limiting is applied and traffic is dispatched to the Core Layer via the Gateway, ICAP Server, or Mail Server.
4

Core processing

The Service Engine distributes work across the guard workers; the cache, OCR, policy engine, and rule checker are evaluated in parallel.
5

Routing decision

The AI Router and RAG Router select the target — cloud LLM, on-prem LLM, cloud vector DB, or on-prem vector DB — based on routing policy, compliance rules, and data-residency requirements.
6

Response and logging

The result is returned to the caller; events are logged to PostgreSQL, streamed via Kafka, cached in Redis, and forwarded to your SIEM.

Deployment targets

The stack is fully containerized and runs on:
  • Red Hat OpenShift — Helm chart deployment, Route & Ingress support, OpenShift OAuth, SCC/SCM compliance, built-in image registry.
  • Kubernetes (vanilla K8s, EKS, GKE, AKS) — Helm 3 chart, namespace isolation, HPA auto-scaling per service, ConfigMap/Secret management, NetworkPolicy, Persistent Volumes.
  • Docker Compose — single-node deployment with a provided docker-compose.yml, volume mounts, and health-check/restart policies.
  • Google Cloud Anthos — Config Management, multi-cluster fleet support, Workload Identity, OPA Policy Controller, Anthos Service Mesh.

Built-in Red Team Worker

A Celery-powered adversarial-testing subsystem is built into the pipeline for continuous security validation. It runs a five-stage loop — scenario Builder, Test Builder, Runner, Evaluator, and Reporter — executing attack suites against the live pipeline, scoring policy effectiveness, and pushing findings to the Log DB and SIEM integrations. See Red Teaming.

On-Prem Requirements

Backing services, network, storage, and GPU prerequisites.

Infrastructure Sizing

Reference CPU, memory, and GPU sizing.

ICAP & Proxy

Inline inspection as a proxy via ICAP.

The Six Guards

How the layers map to the guard model.