Four RTX PRO 6000 Blackwell cards give 384 GB of GPU memory, enough to serve a 594-billion-parameter model on one machine. The developer jamesob built exactly that on a used AMD EPYC server board and put a PCIe Gen4 switch between the CPU and the GPUs, so the cards talk to each other directly. The result is about 80 tokens per second on GLM-5.2 at a 460k-token context.[1][2] The write-up on GitHub has the full parts list with prices, the BIOS settings, the Linux tweaks and the measured numbers. Here is what it contains, checked against the manufacturers' specs.
The project at a glance
| Item | Detail |
|---|---|
| Built by | jamesob, repository jamesob/local-llm (created 3 July 2026)[1] |
| Goal | Run near state-of-the-art language models locally[1] |
| GPUs | 4 × NVIDIA RTX PRO 6000 Blackwell Workstation, 96 GB each, 384 GB in total[1] |
| Host | ASRock Rack ROMED8-2T with an AMD EPYC 7313P[1] |
| Key trick | A Microchip Switchtec PCIe Gen4 switch from c-payne, with ACS disabled so GPU traffic stays in the switch[1] |
| Result | About 80 tokens/s on GLM-5.2 594B at 460k context; 50.4 GB/s GPU to GPU[1] |
Why a PCIe switch
When several GPUs split a model, they exchange data constantly. On this build that traffic goes card to card through the switch rather than up to the CPU and back down. By default, PCIe Access Control Services (ACS) sends peer-to-peer traffic through the CPU's root port, so the author turns ACS off at every boot.[1] The author also chose an older PCIe Gen4 host instead of a Gen5 platform and estimates that saved about $10,000 on the host.[1]
Parts list
Prices are the author's, in US dollars, mostly for parts bought second hand on eBay.[1]
| Part | Model | Author's price |
|---|---|---|
| Motherboard | ASRock Rack ROMED8-2T | $715 |
| CPU | AMD EPYC 7313P (Milan, 16 cores, 3.0 GHz) | $504 |
| Memory | 8 × 16 GB Crucial CT16G4RFD4213 DDR4 ECC RDIMM (128 GB) | $642 |
| CPU cooler | Noctua NH-U14S | $140 |
| Case | AAAWave Sluice V2 open frame | $100 |
| Power supplies | 2 × Super Flower 1700 W | $750 |
| PCIe switch kit | c-payne Switchtec PM40100 Gen4 (details below) | ~$1,330 |
| Boot SSD | 4 TB M.2 | $291 |
| Model storage | 2 × 8 TB M.2 | $1,200 |
| Fans | 3 × 120 mm PWM | $15 |
| Base system total | $5,687 | |
| GPUs | 4 × NVIDIA RTX PRO 6000 Blackwell Workstation, 96 GB | ~$46,000 |
The switch kit, bought from c-payne.com:[1][5]
| Part | Qty | What it does |
|---|---|---|
| PCIe Gen4 switch, Microchip Switchtec PM40100 | 1 | 5 × x16 slots for the GPUs, 2 × SlimSAS 8i upstream |
| SlimSAS host adapter x16 with redriver (DS160PR810) | 1 | Sits in a ROMED8-2T x16 slot and feeds the switch |
| SlimSAS SFF-8654 8i cable, PCIe Gen4 | 2 | Each carries x8; together they make the x16 uplink |
The board: ASRock Rack ROMED8-2T
From ASRock Rack's user manual:[3]
- ATX, 305 × 244 mm, one SP3 (LGA 4094) socket. The latest BIOS supports AMD EPYC 7003 and 7002 series only.
- Seven PCIe 4.0 x16 slots. PCIE2 shares lanes with the M2_1 slot, the OCuLink ports and four SATA ports, set by jumpers.
- Eight DDR4 slots for RDIMM, LRDIMM or NVDIMM at up to 3200 MHz; RDIMMs up to 64 GB each.
- Two M.2 slots (PCIe 4.0 x4 or SATA), two OCuLink ports, eight SATA ports through two mini-SAS HD connectors.
- Two 10GBASE-T ports on an Intel X550-AT2, plus a dedicated IPMI port. ASPEED AST2500 management controller.
The GPU: RTX PRO 6000 Blackwell Workstation Edition
NVIDIA lists 96 GB of GDDR7 with ECC, 1,792 GB/s of memory bandwidth, a maximum power of 600 W and a dual-slot, double flow-through cooler.[4] We sell this card: PNY NVIDIA RTX PRO 6000 Blackwell Workstation Edition.
BIOS settings
The author's ROMED8-2T settings for the slot that holds the switch adapter:[1]
| Setting | Value | Why, per the author |
|---|---|---|
| AMD PCIE Link Width (switch slot) | x16 (not x8/x8) | Bifurcation split the slot and the uplink trained at x8. Needs both SlimSAS cables connected. |
| PCIe Link Speed (switch slot) | Gen4 (not Auto) | The Gen5 Blackwell cards could fail to train through the Gen4 switch and fall back to Gen1 |
| ASPM | Disabled | Removes idle link downclocking that made lspci report "downgraded" links |
| Re-Size BAR | Enabled | Exposes the full 96 GB per card; needed for GPU peer-to-peer |
| SR-IOV | Disabled | Bare-metal inference, no virtualisation overhead |
| Preferred IO | Auto | Pinning the switch bus is a small optimisation, not a fix |
Linux settings
-
Boot parameters:
iommu=off amd_iommu=off nomodeset. Withoutiommu=off, NCCL hangs on multi-GPU peer-to-peer.[1] -
NVIDIA UVM: the repository sets
uvm_disable_hmm=1for thenvidia_uvmmodule as a peer-to-peer fix.[1] -
Disable ACS at boot: a small script, started by a systemd service, clears ACS with
setpcion every PCIe bridge. A kernel override would need a patched kernel.[1] -
Check it worked:
lspci -vvv | grep ACSCtlshows only minus signs, andnvidia-smi topo -mshows PIX between all four GPUs instead of PHB or NODE.[1] -
Cap the power: persistence mode on and
nvidia-smi -pl 350for each card, applied at boot.[1]
Power
At NVIDIA's 600 W per card, four GPUs could draw 2,400 W on their own.[4] The author runs the rig on a single 110 V circuit rather than install a 220 V one, calls that probably unwise, and caps each card at 350 W: 1,400 W for the four GPUs, sized for the power supply budget. In an earlier phase on one 1700 W supply, the cards ran at about 260 W each: about 1,040 W for the GPUs plus about 280 W for the rest, roughly 1,320 W.[1] Check what your own circuit can deliver before you copy this.
Measured results
| Measurement | Result[1] |
|---|---|
| Uplink to the CPU | Gen4 x16, about 30 GB/s |
| GPU to GPU through the switch, one direction | 27.5 GB/s |
| GPU to GPU through the switch, both directions | 50.4 GB/s |
| GPU to GPU latency | 0.37 to 0.45 µs |
| GLM-5.2-Int8Mix-NVFP4-REAP-594B with vLLM | About 80 tokens/s at 460k context |
The author describes the peer-to-peer numbers as Gen4 line rate.[1] The model runs in vLLM from a docker-compose file in the repository's runners folder, which you can reuse as a starting point.[2]
Lessons from the build
- Buy the exact SlimSAS cables. The author ordered too few from c-payne, bought what looked like the same cable elsewhere, and a small difference caused problems until the cables were reordered.[1]
- Redriver gain matters. On c-payne's advice the gain was lowered to level 3 with c-payne's own tool, which the author found the fiddliest step. The right level depends on the cable length.[1]
- "Downgraded" links at idle can be harmless. With ASPM on, idle links show 2.5 GT/s and retrain to Gen4 under load.[1]
- Plan the enclosure. The author built a wooden enclosure for the switch and GPUs.[1]
Want to build it?
We sell the RTX PRO 6000 Blackwell Workstation Edition used in this project. We don't stock the ROMED8-2T, the c-payne switch kit or the other host parts yet. Ask us for a quote on the parts in the author's list.
Credit: build, settings and measurements by jamesob. Diagrams are our own drawings of the published build.
Sources
- jamesob, local-llm: guide to running SOTA LLMs locally (GitHub README, independent)
- jamesob, runners: ready-to-run serving configs (GitHub)
- ASRock Rack, ROMED8-2T user manual (PDF)
- NVIDIA, RTX PRO 6000 Blackwell Workstation Edition
- c-payne, PCIe switches and adapters