Application servers
Databases
Each backing datastore runs in HA and can be a managed service or integrated into the deployment.GPU worker
The security model runs on a dedicated GPU node, separate from the application workers. Choose a tier based on expected concurrency:
GPU allocation scales with your intended application and expected throughput. The minimum tier deploys easily on commodity GPUs; scale up to H100-class hardware as capacity demands grow. See GPU & Performance for measured throughput across accelerator types.
Full-scale resource profile
At full horizontal scale, the microservice stack (excluding the GPU node) has been measured at:- Requests: 95 vCPU / 224 GB RAM
- Limits: 429 vCPU / 594 GB RAM
- Recommended worker nodes: 3, each 32 vCPU / 128 GB RAM
- Control plane: a separate control plane or managed Kubernetes
- GPU node: kept separate from the microservice worker nodes
Customer prerequisites
- Network connectivity, including public IPs and internet access as required by your use cases
- SSL certificates
- A load balancer
The operating system and all required third-party software licenses are provided by BeyondGuard, on either physical or virtual infrastructure.
Related
On-Prem Requirements
Backing services, network, storage, and GPU prerequisites.
GPU & Performance
Measured throughput per GPU configuration.
High Availability
Multi-site active-active model serving.
Installation
Deploy with Helm or Docker Compose.