The RPS figures below are end-to-end application throughput — the number of requests the full BeyondGuard pipeline completes — not raw GPU inference throughput. They are internal benchmark results measured on a Kubernetes microservice stack with the GPU on a separate node and input lengths varying from 50 to 8,000 tokens.
Single-GPU results
Memory bandwidth and the CPU–GPU interconnect matter most under high concurrency, long context, and large KV caches — which is why the higher-bandwidth Hopper and Blackwell parts pull ahead.
Horizontal scaling (1–6 GPUs)
Model services scale horizontally behind a load balancer, with the application layer expanded in parallel. Throughput is close to linear across the range; the bottleneck stays in the GPU / model-serving tier rather than the application layer.
Multi-GPU values are rounded measurements from internal load tests. In the recorded 6× GH200 test, the system sustained 550 RPS at 600 ms average latency with no latency increase as load rose.
Application resource profile
At full horizontal scale, the microservice stack (separate from the GPU node) requested roughly 95 vCPU / 224 GB RAM, running across 3 worker nodes of 32 vCPU / 128 GB RAM each, with the GPU kept on its own node.Sizing guidance
- Start on the minimum tier (48 GB vRAM, ~20–30 concurrent requests) for pilots and smaller workloads.
- Scale up to H100- or H200-class hardware, or add GPUs behind the load balancer, as concurrency grows.
- Because scaling is near-linear, you can size to a target RPS by multiplying the single-GPU figure for your chosen accelerator.
Hardware specifications are compiled from NVIDIA’s official product documentation. RPS and latency figures are internal end-to-end test results and will vary with input length, policy configuration, and your infrastructure.
Related
Infrastructure Sizing
CPU, memory, and GPU-tier reference profiles.
High Availability
Active-active, multi-site GPU serving.
On-Prem Requirements
GPU and infrastructure prerequisites.
Architecture Overview
Where the GPU model service fits.