Skip to main content
BeyondGuard’s inline inspection is powered by a security model served on GPU via vLLM. GPU capacity is the primary lever for end-to-end throughput, and it scales with your expected request volume.
The RPS figures below are end-to-end application throughput — the number of requests the full BeyondGuard pipeline completes — not raw GPU inference throughput. They are internal benchmark results measured on a Kubernetes microservice stack with the GPU on a separate node and input lengths varying from 50 to 8,000 tokens.

Single-GPU results

Memory bandwidth and the CPU–GPU interconnect matter most under high concurrency, long context, and large KV caches — which is why the higher-bandwidth Hopper and Blackwell parts pull ahead.

Horizontal scaling (1–6 GPUs)

Model services scale horizontally behind a load balancer, with the application layer expanded in parallel. Throughput is close to linear across the range; the bottleneck stays in the GPU / model-serving tier rather than the application layer. Multi-GPU values are rounded measurements from internal load tests. In the recorded 6× GH200 test, the system sustained 550 RPS at 600 ms average latency with no latency increase as load rose.

Application resource profile

At full horizontal scale, the microservice stack (separate from the GPU node) requested roughly 95 vCPU / 224 GB RAM, running across 3 worker nodes of 32 vCPU / 128 GB RAM each, with the GPU kept on its own node.

Sizing guidance

  • Start on the minimum tier (48 GB vRAM, ~20–30 concurrent requests) for pilots and smaller workloads.
  • Scale up to H100- or H200-class hardware, or add GPUs behind the load balancer, as concurrency grows.
  • Because scaling is near-linear, you can size to a target RPS by multiplying the single-GPU figure for your chosen accelerator.
Hardware specifications are compiled from NVIDIA’s official product documentation. RPS and latency figures are internal end-to-end test results and will vary with input length, policy configuration, and your infrastructure.

Infrastructure Sizing

CPU, memory, and GPU-tier reference profiles.

High Availability

Active-active, multi-site GPU serving.

On-Prem Requirements

GPU and infrastructure prerequisites.

Architecture Overview

Where the GPU model service fits.