Skip to main content
KONST

GPU Bare-Metal Pricing

No middleware and no virtualization overhead. These are the current published hourly prices per GPU, together with the complete hardware configurations.

Models available to rent

Configurations range from one to eight GPUs. Eight-GPU systems use 3.2 Tbit/s InfiniBand and support distributed training across nodes.

$3.00 / GPU / hour

NVIDIA B200

A next-generation training platform. Its 180GB of VRAM accommodates long contexts and large models without splitting model weights.

  • Intel Emerald Rapids
  • 1× or 8× B200 · 180GB SXM VRAM
  • 16× or 128× vCPU
  • 224 or 1,792 GB DDR5
  • 3.2 Tbit/s InfiniBand
$2.30 / GPU / hour

NVIDIA H200

A balanced option for training and inference. Its 141GB of high-bandwidth memory is especially useful for high-throughput large-model inference.

  • Intel Sapphire Rapids
  • 1× or 8× H200 · 141GB SXM VRAM
  • 16× or 128× vCPU
  • 200 or 1,600 GB DDR5
  • 3.2 Tbit/s InfiniBand
$2.00 / GPU / hour

NVIDIA H100

The industry-standard model, with mature ecosystem and framework support. It is well suited to teams with established training workflows.

  • Intel Sapphire Rapids
  • 1× or 8× H100 · 80GB SXM VRAM
  • 16× or 128× vCPU
  • 200 or 1,600 GB DDR5
  • 3.2 Tbit/s InfiniBand

Hourly GPU bare-metal list prices

These are the current published starting prices for one GPU, in USD per GPU per hour. Actual pricing depends on availability, contract term, and location.

ModelHourly price fromGPU memoryCPU platformAvailable configurationsSystem memoryCluster network
NVIDIA GB200 NVL72Pre-orderTBDTBDRack-scale NVL72TBDNVLink + InfiniBand
NVIDIA B200$3.00180GB SXM VRAMIntel Emerald Rapids1× or 8× B200224 or 1,792 GB DDR53.2 Tbit/s InfiniBand
NVIDIA H200$2.30141GB SXM VRAMIntel Sapphire Rapids1× or 8× H200200 or 1,600 GB DDR53.2 Tbit/s InfiniBand
NVIDIA H100$2.0080GB SXM VRAMIntel Sapphire Rapids1× or 8× H100200 or 1,600 GB DDR53.2 Tbit/s InfiniBand

Next-generation rack-scale options

Pre-order

NVIDIA GB200 NVL72

Reserve the GB200 NVL72 for access to next-generation AI performance. This rack-scale system uses NVLink to operate the rack's GPUs as one compute unit, supporting trillion-parameter model training and large-scale inference.

  • Rack-scale NVL72 integration
  • Direct-to-chip liquid cooling (DTC)
  • Pre-orders and delivery scheduling available
Coming soon

Storage pricing

Storage is billed only for capacity used, with no extra charges for uploads, downloads, expansion, or migration. Capacity uses binary units: 1 GB = 2^30 bytes (GiB).

  • No data transfer fees
  • No capacity expansion fees
  • No data migration fees

From quote to server access in four steps

1

Confirm requirements

1 to 2 business days

Confirm the model, GPU count, expected training duration, multi-node interconnect needs, data residency, and compliance requirements.

2

Confirm pricing and availability

We provide the actual unit price and delivery date based on the contract term. For terms of 3 months or longer, we also compare reserved options.

3

Prepare the contract and environment

After signing, we initialize accounts, segment the network, configure firewall rules, mount storage, and grant access.

4

Provision and accept

Before handover, we validate node and interconnect performance and provide the test results. The service then moves into 24×7 operations and incident reporting.

Need monthly pricing and delivery dates for bare-metal GPUs?

Tell us the model, quantity, and contract term, and we will provide available configurations and pricing.