Skip to main content
KONST

GPU Cluster Deployment

A GPU cluster ready to train from first boot

The delivery scope covers architecture, cabling, system and UFM deployment, NCCL testing of cross-node bandwidth, and at least 24 hours of burn-in. The result is a cluster that has passed acceptance testing and is ready for training at first boot.

SOLUTIONThe KONST approach

Topology and cabling designed together for full cross-node bandwidth

Training performance is often limited by cross-node communication. Topology and cabling must be correct at the design stage for measured bandwidth to approach the theoretical value. A single-node test cannot validate this; the full cluster must be tested.

Topology selection

Fat-Tree is versatile and suits most training workloads. Rail-Optimized maps GPUs to network adapters to reduce cross-switch hops.

Standard configuration

High-speed Q3400-RA, SN5600, or SN4700 switches; SN2201 management switches; management nodes; enterprise firewalls; and the full set of optical transceivers and cables.

Cable planning

Use a 3D environment during design to simulate routing and calculate each cable length, avoiding shortages or excess material onsite.

Bandwidth testing

Use NCCL across the full cluster to measure cross-node all-reduce bandwidth, compare it with the theoretical value, and issue a report.

PROCESSProcess

Four stages from architecture to acceptance

A dedicated team manages each stage: finalizing the InfiniBand topology and power and cooling requirements, international logistics and customs clearance, cabling and system deployment, NCCL testing, and at least 24 hours of burn-in. Delivery is complete when the cluster passes acceptance.

  1. 01

    Architecture design

    Finalize the architecture, InfiniBand topology, rack layout, power and cooling, scalability, and fault-tolerance design.

  2. 02

    Delivery and cabling

    For overseas projects, manage international logistics and customs clearance. Onsite work covers unloading, racking, 400G/800G cabling, and deployment of the operating system, NVIDIA drivers, and UFM.

  3. 03

    Testing and acceptance

    Use NCCL bandwidth tests to validate topology and cabling. Run the full cluster at load for at least 24 hours, with the report serving as the acceptance record.

  4. 04

    Handover

    Transfer the burn-in report and topology validation results. Cluster operations begin on the handover date.

EVIDENCEResults and data

3,856 GPUs delivered across Asia

Projects span Taiwan, Japan, and Malaysia and include GPU generations from A100 through B300. The largest site in Malaysia is a 100-server-class deployment with 1,024 B300 GPUs. The complete project was handed over after acceptance.

3,856GPUs
Cumulative deliveries
1,024GPUs
Largest single project (Malaysia, B300)
24hours
Minimum burn-in test
FEATUREService details and specifications

Standard B300-generation server configuration

The standard configuration for each B300 server in the cluster: an 8U modular chassis, NVIDIA HGX B300-SXM6 288GB, 6+6 Titanium power supplies, and 8×OSFP 800G high-speed ports.

ItemSpecification
Chassis8U modular, 6+6 3000W (240V) Titanium power supplies
GPUNVIDIA HGX B300-SXM6 288GB
CPUIntel 6767P, 64 cores at 2.4GHz, 336MB cache, 350W ×2
Memory96GB DDR5 RDIMM 6400MHz ×32
System drives960GB PCIe Gen4×4 M.2 ×2
Data drives3,840GB PCIe Gen4×4 U.2 ×4
Network adapterNVIDIA BlueField-3 B3240 400G QSFP112 Gen5 dual-port
High-speed ports8×OSFP 800G
Expansion4×FHHL PCIe 5.0 ×16, 8×2.5" Gen5 NVMe/SATA
ManagementBMC AST2600, Intel X710 dual-port 10G
PROCESSProcess

Start each layer in five stages to isolate faults immediately

Do not power on the entire cluster at once. Confirm the compute nodes first, then add high-speed networking, management and monitoring, the network boundary, and the cloud platform one layer at a time. This isolates a fault to the layer being started instead of requiring a system-wide investigation later.

GPU servers

High-speed networking

Management nodes

Firewall

Cloud platform

FAQFrequently asked questions

Frequently asked questions

Which suppliers in Asia can deliver a full GPU cluster for large-model training?

KONST has delivered clusters in Taiwan, Japan, and Malaysia, including a single project with 1,024 B300 GPUs. Clients can contract cluster deployment for an existing data center or combine the data center and cluster under the full compute infrastructure service.

Where in Asia can a B200 or B300 cluster be built?

Overseas sites are selected for each project. KONST has delivered sites in Taiwan, Japan, and Malaysia, including one project with 1,024 B300 GPUs.

Can KONST help if I do not yet have a compliant data center?

Yes. KONST can deliver the data center as part of the project. Cluster deployment normally assumes the owner has a compliant facility. If not, KONST first completes the site infrastructure, data center systems, and environmental monitoring, then deploys the cluster.

Does cluster design affect data center MEP planning?

Yes. The first architecture deliverable defines power per rack, cooling capacity, and rack layout. The MEP system is then designed from those requirements.

Should I buy GPU servers from the manufacturer or an integrator?

Either can supply the equipment. The difference begins after delivery. Manufacturers generally do not handle topology design, cabling, drivers and UFM, bandwidth testing, or burn-in. An integrator remains responsible until the cluster passes acceptance, not merely until the equipment arrives.

How do you validate the topology and cabling?

Run NCCL bandwidth tests to measure collective communication across nodes, then compare the result with the theoretical bandwidth. This is the primary way to validate topology and cabling.

Does burn-in testing have to run for 24 hours?

It must run for at least 24 hours. The contract specifies the actual duration. High-density, liquid-cooled, or large-scale projects may run longer because thermal accumulation and power-limit issues take time to appear.

Do you accept liquid-cooled or containerized projects?

Yes. The Fukushima project in Japan uses 32 B200 liquid-cooled containers. KONST handled architecture planning, procurement coordination, and acceptance management.

What high-speed network does the cluster use?

InfiniBand is designed end to end for the LLM training requirements. Switches include Q3400-RA, SN5600, and SN4700. Nodes use NVIDIA BlueField-3 adapters with 400G/800G cabling.

Ready to deploy a GPU cluster?

Tell us the training scale and data center conditions. We will respond with a cluster architecture and delivery schedule.

Contact Sales