Back to Projects you can startLocal LLMMulti-GPURTX PRO 6000

Four RTX PRO 6000 GPUs on one PCIe switch

jamesob put four RTX PRO 6000 Blackwell cards behind a PCIe Gen4 switch on an ASRock Rack ROMED8-2T and serves a 594B model at about 80 tokens/s. Parts list with prices, BIOS settings, power and measured results.

7 min read ai StoreLabs

Project card: 4x RTX PRO 6000 on one PCIe switch, about 80 tokens/s, 50.4 GB/s GPU to GPU, 350 W per GPU

Four RTX PRO 6000 Blackwell cards give 384 GB of GPU memory, enough to serve a 594-billion-parameter model on one machine. The developer jamesob built exactly that on a used AMD EPYC server board and put a PCIe Gen4 switch between the CPU and the GPUs, so the cards talk to each other directly. The result is about 80 tokens per second on GLM-5.2 at a 460k-token context.[1][2] The write-up on GitHub has the full parts list with prices, the BIOS settings, the Linux tweaks and the measured numbers. Here is what it contains, checked against the manufacturers' specs.

The project at a glance

Item Detail
Built by jamesob, repository jamesob/local-llm (created 3 July 2026)[1]
Goal Run near state-of-the-art language models locally[1]
GPUs 4 × NVIDIA RTX PRO 6000 Blackwell Workstation, 96 GB each, 384 GB in total[1]
Host ASRock Rack ROMED8-2T with an AMD EPYC 7313P[1]
Key trick A Microchip Switchtec PCIe Gen4 switch from c-payne, with ACS disabled so GPU traffic stays in the switch[1]
Result About 80 tokens/s on GLM-5.2 594B at 460k context; 50.4 GB/s GPU to GPU[1]

Why a PCIe switch

When several GPUs split a model, they exchange data constantly. On this build that traffic goes card to card through the switch rather than up to the CPU and back down. By default, PCIe Access Control Services (ACS) sends peer-to-peer traffic through the CPU's root port, so the author turns ACS off at every boot.[1] The author also chose an older PCIe Gen4 host instead of a Gen5 platform and estimates that saved about $10,000 on the host.[1]

Diagram: ROMED8-2T host connects over two SlimSAS 8i cables to a Switchtec PM40100 PCIe switch feeding four RTX PRO 6000 GPUs
Figure 1. How the parts connect, drawn by us from the author's parts list and notes.[1]

Parts list

Prices are the author's, in US dollars, mostly for parts bought second hand on eBay.[1]

Part Model Author's price
Motherboard ASRock Rack ROMED8-2T $715
CPU AMD EPYC 7313P (Milan, 16 cores, 3.0 GHz) $504
Memory 8 × 16 GB Crucial CT16G4RFD4213 DDR4 ECC RDIMM (128 GB) $642
CPU cooler Noctua NH-U14S $140
Case AAAWave Sluice V2 open frame $100
Power supplies 2 × Super Flower 1700 W $750
PCIe switch kit c-payne Switchtec PM40100 Gen4 (details below) ~$1,330
Boot SSD 4 TB M.2 $291
Model storage 2 × 8 TB M.2 $1,200
Fans 3 × 120 mm PWM $15
Base system total $5,687
GPUs 4 × NVIDIA RTX PRO 6000 Blackwell Workstation, 96 GB ~$46,000

The switch kit, bought from c-payne.com:[1][5]

Part Qty What it does
PCIe Gen4 switch, Microchip Switchtec PM40100 1 5 × x16 slots for the GPUs, 2 × SlimSAS 8i upstream
SlimSAS host adapter x16 with redriver (DS160PR810) 1 Sits in a ROMED8-2T x16 slot and feeds the switch
SlimSAS SFF-8654 8i cable, PCIe Gen4 2 Each carries x8; together they make the x16 uplink

The board: ASRock Rack ROMED8-2T

From ASRock Rack's user manual:[3]

  • ATX, 305 × 244 mm, one SP3 (LGA 4094) socket. The latest BIOS supports AMD EPYC 7003 and 7002 series only.
  • Seven PCIe 4.0 x16 slots. PCIE2 shares lanes with the M2_1 slot, the OCuLink ports and four SATA ports, set by jumpers.
  • Eight DDR4 slots for RDIMM, LRDIMM or NVDIMM at up to 3200 MHz; RDIMMs up to 64 GB each.
  • Two M.2 slots (PCIe 4.0 x4 or SATA), two OCuLink ports, eight SATA ports through two mini-SAS HD connectors.
  • Two 10GBASE-T ports on an Intel X550-AT2, plus a dedicated IPMI port. ASPEED AST2500 management controller.

The GPU: RTX PRO 6000 Blackwell Workstation Edition

NVIDIA lists 96 GB of GDDR7 with ECC, 1,792 GB/s of memory bandwidth, a maximum power of 600 W and a dual-slot, double flow-through cooler.[4] We sell this card: PNY NVIDIA RTX PRO 6000 Blackwell Workstation Edition.

BIOS settings

The author's ROMED8-2T settings for the slot that holds the switch adapter:[1]

Setting Value Why, per the author
AMD PCIE Link Width (switch slot) x16 (not x8/x8) Bifurcation split the slot and the uplink trained at x8. Needs both SlimSAS cables connected.
PCIe Link Speed (switch slot) Gen4 (not Auto) The Gen5 Blackwell cards could fail to train through the Gen4 switch and fall back to Gen1
ASPM Disabled Removes idle link downclocking that made lspci report "downgraded" links
Re-Size BAR Enabled Exposes the full 96 GB per card; needed for GPU peer-to-peer
SR-IOV Disabled Bare-metal inference, no virtualisation overhead
Preferred IO Auto Pinning the switch bus is a small optimisation, not a fix

Linux settings

  1. Boot parameters: iommu=off amd_iommu=off nomodeset. Without iommu=off, NCCL hangs on multi-GPU peer-to-peer.[1]
  2. NVIDIA UVM: the repository sets uvm_disable_hmm=1 for the nvidia_uvm module as a peer-to-peer fix.[1]
  3. Disable ACS at boot: a small script, started by a systemd service, clears ACS with setpci on every PCIe bridge. A kernel override would need a patched kernel.[1]
  4. Check it worked: lspci -vvv | grep ACSCtl shows only minus signs, and nvidia-smi topo -m shows PIX between all four GPUs instead of PHB or NODE.[1]
  5. Cap the power: persistence mode on and nvidia-smi -pl 350 for each card, applied at boot.[1]

Power

Chart: 2,400 W at NVIDIA rating, 1,400 W at the 350 W cap, 1,320 W whole machine on one PSU
Figure 2. NVIDIA's 600 W rating against the author's cap.[1][4]

At NVIDIA's 600 W per card, four GPUs could draw 2,400 W on their own.[4] The author runs the rig on a single 110 V circuit rather than install a 220 V one, calls that probably unwise, and caps each card at 350 W: 1,400 W for the four GPUs, sized for the power supply budget. In an earlier phase on one 1700 W supply, the cards ran at about 260 W each: about 1,040 W for the GPUs plus about 280 W for the rest, roughly 1,320 W.[1] Check what your own circuit can deliver before you copy this.

Measured results

Measurement Result[1]
Uplink to the CPU Gen4 x16, about 30 GB/s
GPU to GPU through the switch, one direction 27.5 GB/s
GPU to GPU through the switch, both directions 50.4 GB/s
GPU to GPU latency 0.37 to 0.45 µs
GLM-5.2-Int8Mix-NVFP4-REAP-594B with vLLM About 80 tokens/s at 460k context

The author describes the peer-to-peer numbers as Gen4 line rate.[1] The model runs in vLLM from a docker-compose file in the repository's runners folder, which you can reuse as a starting point.[2]

Lessons from the build

  • Buy the exact SlimSAS cables. The author ordered too few from c-payne, bought what looked like the same cable elsewhere, and a small difference caused problems until the cables were reordered.[1]
  • Redriver gain matters. On c-payne's advice the gain was lowered to level 3 with c-payne's own tool, which the author found the fiddliest step. The right level depends on the cable length.[1]
  • "Downgraded" links at idle can be harmless. With ASPM on, idle links show 2.5 GT/s and retrain to Gen4 under load.[1]
  • Plan the enclosure. The author built a wooden enclosure for the switch and GPUs.[1]

Want to build it?

We sell the RTX PRO 6000 Blackwell Workstation Edition used in this project. We don't stock the ROMED8-2T, the c-payne switch kit or the other host parts yet. Ask us for a quote on the parts in the author's list.

Credit: build, settings and measurements by jamesob. Diagrams are our own drawings of the published build.

Sources

  1. jamesob, local-llm: guide to running SOTA LLMs locally (GitHub README, independent)
  2. jamesob, runners: ready-to-run serving configs (GitHub)
  3. ASRock Rack, ROMED8-2T user manual (PDF)
  4. NVIDIA, RTX PRO 6000 Blackwell Workstation Edition
  5. c-payne, PCIe switches and adapters

Keep reading

More projects.