> ## Documentation Index
> Fetch the complete documentation index at: https://docs.beyondguard.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# High Availability and Disaster Recovery

> BeyondGuard's high-availability model — HA backing services and an active-active, multi-site vLLM model-serving architecture with a passive disaster-recovery site.

BeyondGuard is designed to run without single points of failure. Stateful and traffic-facing services run in high availability, guard microservices scale horizontally, and the GPU model-serving tier can be deployed **active-active across multiple sites** with a passive disaster-recovery site.

## Service high availability

* **Backing services** — PostgreSQL, Redis, and Kafka run in HA (see [Infrastructure Sizing](/deployment/infrastructure-sizing)).
* **Proxy server** — `bg-proxy-server` is the only traffic-facing service and runs in HA so that interception is never a single point of failure.
* **Guard microservices** — stateless and horizontally scaled (HPA on Kubernetes), so capacity and resilience grow with replica count.

## Active-active model serving

The vLLM model-serving tier uses a two-active-site design with a passive DR site, fronted by a **Global Load Balancer (GSLB)** that performs health checks and failover routing.

| Site | Role | Normal traffic | GPU servers |
| - | - | - | - |
| **DC1** | Active site | 50% | 2 |
| **DC2** | Active site | 50% | 2 |
| **DRC** | Passive DR site | 0% | 1 |

How it works:

* **DC1 and DC2 run active-active** and each carry 50% of production traffic.
* Every site runs **model serving only** — stateless request handling with a local model cache and no additional application components.
* **Model artifacts are preloaded to all sites, including the DR site**, so failover is fast.
* The **DRC site takes no traffic under normal conditions** and is brought online only on disaster or the loss of an active site.

## Failover behavior

The GSLB continuously health-checks the active sites and routes around a failure. Because request handling is stateless and models are already resident at every site, traffic can shift between DC1, DC2, and — in a disaster — DRC without a cold start.

## Related

<CardGroup cols={2}>
  <Card title="GPU & Performance" icon="microchip" href="/deployment/gpu-performance">
    Per-site GPU throughput and scaling.
  </Card>

  <Card title="Infrastructure Sizing" icon="gauge" href="/deployment/infrastructure-sizing">
    HA database and application sizing.
  </Card>

  <Card title="Architecture Overview" icon="diagram-project" href="/deployment/architecture">
    The full layered architecture.
  </Card>

  <Card title="On-Prem Requirements" icon="server" href="/deployment/on-prem-requirements">
    Prerequisites for a resilient deployment.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.