Service high availability
- Backing services — PostgreSQL, Redis, and Kafka run in HA (see Infrastructure Sizing).
- Proxy server —
bg-proxy-serveris the only traffic-facing service and runs in HA so that interception is never a single point of failure. - Guard microservices — stateless and horizontally scaled (HPA on Kubernetes), so capacity and resilience grow with replica count.
Active-active model serving
The vLLM model-serving tier uses a two-active-site design with a passive DR site, fronted by a Global Load Balancer (GSLB) that performs health checks and failover routing.
How it works:
- DC1 and DC2 run active-active and each carry 50% of production traffic.
- Every site runs model serving only — stateless request handling with a local model cache and no additional application components.
- Model artifacts are preloaded to all sites, including the DR site, so failover is fast.
- The DRC site takes no traffic under normal conditions and is brought online only on disaster or the loss of an active site.
Failover behavior
The GSLB continuously health-checks the active sites and routes around a failure. Because request handling is stateless and models are already resident at every site, traffic can shift between DC1, DC2, and — in a disaster — DRC without a cold start.Related
GPU & Performance
Per-site GPU throughput and scaling.
Infrastructure Sizing
HA database and application sizing.
Architecture Overview
The full layered architecture.
On-Prem Requirements
Prerequisites for a resilient deployment.