Request-for-quote configurator

Build your system.
Spec it your way.

Start from a proven template, tune it to your workload, and send it straight to our team as a quote request. Every template locks the components that define compatibility — GPU, interconnect, CPU platform — and leaves everything else, including cooling, open to configure.

This page

Configure a build

Start from a template, tune the spec, get a quote. Best when the exact configuration matters more than lead time.

Ready now

Ready-to-ship inventory

New and used systems in stock, pre-configured and ready for pickup or fast shipment — no build queue.

Browse in-stock systems

10 templates, 4 GPU generations

// CHASSIS: SUPERMICRO ONLY  |  COOLING: AIR OR LIQUID, YOUR CHOICE

Each template locks the identity-defining components and leaves storage, memory, networking, power redundancy, and cooling open to configure. Pick a template to see the full breakdown and submit a quote request.

TEMPLATE: A100 · INFERENCE
Cost-optimized inference
High-throughput serving for established models where budget per token matters more than peak latency.
4-8x A100 80GBAir or Liquid
TEMPLATE: A100 · DEV/TEST
Dev & test rig
Small-footprint build for prototyping, fine-tuning experiments, and pre-production validation.
1-2x A100 80GBAir or Liquid
TEMPLATE: H100 · TRAINING
Training cluster
The mainstream choice for pretraining and large fine-tuning jobs, built for multi-node scale-out.
8x H100 80GB SXMAir or Liquid
TEMPLATE: H100 · INFERENCE
High-throughput inference
Serves demanding production traffic with headroom for batched requests and longer contexts.
4-8x H100 80GBAir or Liquid
TEMPLATE: H100 · EDGE
Edge / on-prem deployment
Compact, single-node build for data-residency-sensitive or connectivity-limited environments.
2-4x H100 80GBAir or Liquid
TEMPLATE: H200 · TRAINING
Large-context training cluster
Extra HBM headroom for larger models and longer context windows during pretraining.
8x H200 141GBAir or Liquid
TEMPLATE: H200 · INFERENCE
High-memory LLM inference
Built for serving the largest open-weight models without splitting them across too many nodes.
4-8x H200 141GBAir or Liquid
TEMPLATE: B300 · TRAINING
Frontier-scale training cluster
Our highest-density training build, engineered for the largest pretraining and RL workloads.
8x B300 SXMAir or Liquid
TEMPLATE: B300 · INFERENCE
Next-gen inference server
Lowest-latency serving tier, sized for flagship-model production traffic at scale.
4-8x B300 SXMAir or Liquid
TEMPLATE: B300 · SCALE-OUT
Multi-node scale-out cluster
Pre-wired for multi-node fabric — spec once, we quote the full rack-scale deployment.
8x B300/nodeInfiniBand-ready
// ALSO AVAILABLE: RTX Pro 6000 servers, and non-Supermicro builds on Asus, Dell, Lenovo, and HPE chassis. Not offered as a preset template — contact us for a custom quote.
Configure

A100 cost-optimized inference

Locked specs define this template's identity and can't be changed. Everything else — storage, memory, networking, power redundancy, and cooling — is yours to configure.

TEMPLATE: A100 · INFERENCE
A100 cost-optimized inference
Locked — defines this template
GPU4-8x A100 80GB SXM
GPU interconnectNVLink (per-node)
CPU platformDual-socket, PCIe Gen4
Swappable — adjust to fit your build
Storage (NVMe)
System memory
Networking
Power redundancy
Chassis
Cooling
Tell us about your deployment
▶ Request quote