Compute Delivery
Three ways to deliver compute, matched to the workload
Choose how completed compute infrastructure is delivered: a dedicated group of bare-metal servers, shared and schedulable virtual machine and container resources, or token-metered model calls. Each model has its own specifications and pricing.
Three delivery models for three workload patterns
Use bare metal for sustained full-load workloads, cloud resources for elastic demand, or token-metered access when only model capabilities are needed. Choose one model or combine them.
Bare-metal delivery
Exclusive use of a dedicated group of GPU servers. Physical machines run without a virtualization layer, are supplied under a long-term contract, and provide full performance.
Virtual machine and container delivery
Virtualized clusters with container orchestration, billed by the minute and ready to use with frameworks and drivers preconfigured.
Token-metered delivery
Model calls billed by token. One key provides access to multiple models, every request is attributed to an account, and quotas flow from the organization down.
Still deciding which delivery model fits?
We will recommend a delivery model and pricing method based on your workload and data requirements.
