> ## Documentation Index
> Fetch the complete documentation index at: https://docs.beyondguard.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# GPU Sizing and End-to-End Performance Benchmarks

> How to size the BeyondGuard GPU tier, with internal end-to-end throughput benchmarks across NVIDIA accelerators and 1–6 GPU horizontal scaling.

BeyondGuard's inline inspection is powered by a security model served on GPU via vLLM. GPU capacity is the primary lever for end-to-end throughput, and it **scales with your expected request volume**.

<Info>
  The RPS figures below are **end-to-end application throughput** — the number of requests the full BeyondGuard pipeline completes — not raw GPU inference throughput. They are internal benchmark results measured on a Kubernetes microservice stack with the GPU on a separate node and input lengths varying from 50 to 8,000 tokens.
</Info>

## Single-GPU results

| Accelerator | Architecture | GPU memory | App RPS | Avg latency |
| - | - | - | - | - |
| **NVIDIA L40** | Ada Lovelace | 48 GB GDDR6 | 37 | 600 ms |
| **NVIDIA H100 PCIe** | Hopper | 80 GB HBM2e | 50 | 600 ms |
| **NVIDIA H200** | Hopper | 141 GB HBM3e | 112 | 550 ms |
| **NVIDIA GH200** | Grace + Hopper | 96–144 GB HBM3/HBM3e | 87 | 600 ms |
| **NVIDIA B300** | Blackwell Ultra | 288 GB HBM3e | 240 | 500 ms |

Memory bandwidth and the CPU–GPU interconnect matter most under high concurrency, long context, and large KV caches — which is why the higher-bandwidth Hopper and Blackwell parts pull ahead.

## Horizontal scaling (1–6 GPUs)

Model services scale horizontally behind a load balancer, with the application layer expanded in parallel. Throughput is close to linear across the range; the bottleneck stays in the GPU / model-serving tier rather than the application layer.

| Accelerator | 1 GPU | 2 GPU | 3 GPU | 4 GPU | 5 GPU | 6 GPU |
| - | - | - | - | - | - | - |
| **L40 48 GB** | 37 | 77 | 116 | 155 | 195 | 234 |
| **H100 80 GB PCIe** | 50 | 103 | 156 | 210 | 263 | 316 |
| **H200** | 112 | 231 | 350 | 470 | 589 | 708 |
| **GH200** | 87 | 180 | 272 | 365 | 457 | 550 |
| **B300** | 240 | 495 | 751 | 1,006 | 1,262 | 1,517 |

Multi-GPU values are rounded measurements from internal load tests. In the recorded 6× GH200 test, the system sustained **550 RPS at 600 ms average latency with no latency increase** as load rose.

## Application resource profile

At full horizontal scale, the microservice stack (separate from the GPU node) requested roughly **95 vCPU / 224 GB RAM**, running across **3 worker nodes** of 32 vCPU / 128 GB RAM each, with the GPU kept on its own node.

## Sizing guidance

* Start on the **minimum tier** (48 GB vRAM, \~20–30 concurrent requests) for pilots and smaller workloads.
* Scale up to **H100- or H200-class** hardware, or add GPUs behind the load balancer, as concurrency grows.
* Because scaling is near-linear, you can size to a target RPS by multiplying the single-GPU figure for your chosen accelerator.

<Note>
  Hardware specifications are compiled from NVIDIA's official product documentation. RPS and latency figures are internal end-to-end test results and will vary with input length, policy configuration, and your infrastructure.
</Note>

## Related

<CardGroup cols={2}>
  <Card title="Infrastructure Sizing" icon="gauge" href="/deployment/infrastructure-sizing">
    CPU, memory, and GPU-tier reference profiles.
  </Card>

  <Card title="High Availability" icon="shield-halved" href="/deployment/high-availability">
    Active-active, multi-site GPU serving.
  </Card>

  <Card title="On-Prem Requirements" icon="server" href="/deployment/on-prem-requirements">
    GPU and infrastructure prerequisites.
  </Card>

  <Card title="Architecture Overview" icon="diagram-project" href="/deployment/architecture">
    Where the GPU model service fits.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.