Back to Projects you can startAMD Strix HaloClusteringNetworking

A two-node Strix Halo cluster over 100 GbE RDMA

kyuz0 linked two Framework Desktop boards with Intel E810 100 GbE cards and one DAC, and cut node-to-node latency from 74 µs to 5.23 µs with RDMA. The parts, settings and measured results from the public guide.

6 min read ai StoreLabs

Project card: a two-node Strix Halo cluster over 100 GbE RDMA, 5.23 µs latency, 50.64 Gb/s

Two AMD Strix Halo boards can split one large model between them with vLLM, but only if the link between them is fast. The developer kyuz0 put an Intel E810 100 GbE card in each of two Framework Desktop boards, joined them with one cable and switched the link to RDMA. Node-to-node latency dropped from 74 µs to 5.23 µs.[1] The guide is public, step by step, and lists every part and setting. This is a summary of it, with the manufacturer specs next to it.

The project at a glance

Item Detail
Built by kyuz0, who maintains the amd-strix-halo-vllm-toolboxes repository on GitHub[1][7]
Posted Framework Community, 8 February 2026; setup guide in the same repository[1][2]
Goal Run one model across two nodes with vLLM tensor parallelism (TP = 2)[1]
Nodes 2 × Framework Desktop Mainboard, AMD Ryzen AI Max+ "Strix Halo", 128 GB unified memory[1]
Link Intel E810-CQDA1 in each node, one 100G QSFP28 DAC, no switch[1]
Result 5.23 µs latency and 50.64 Gb/s over RDMA[1]

Why RDMA matters here

With TP = 2, each node holds half of every layer. After every layer the two nodes swap partial results, which the guide says happens thousands of times a second.[1] Over a normal TCP/IP connection that costs about 70 to 100 µs each time; with RDMA it is about 5 µs.[1] RDMA over Converged Ethernet (RoCE v2) lets one node write straight into the other node's memory without going through the CPU and the kernel network stack.[1]

Diagram: two Framework Desktop boards, each with an Intel E810-CQDA1 on a x4 to x16 riser, linked by one 100G QSFP28 DAC
Figure 1. Our drawing of kyuz0's setup. Addresses and settings are the ones in kyuz0's guide.[1]

Parts list

Part What kyuz0 used Notes
Mainboard × 2 Framework Desktop Mainboard, Ryzen AI Max+, 128 GB[1] Framework sells the board with a PCIe x4 slot[5]
Network card × 2 Intel E810-CQDA1, or a similar 100 GbE QSFP28 card[1] Firmware 4.91 or newer recommended[1]
Cable × 1 100G QSFP28 direct attach copper (DAC); the guide gives a QSFPTEK cable as an example[1] Point to point; no switch needed for two nodes[1]
Riser × 2 PCIe x4 to x16 extender[1] The Framework slot is physically x4, so a x16 card needs a riser[1]

One of kyuz0's boards had its slot cut open by Framework so a x16 card fits directly. kyuz0 does not recommend that: risers are cheaper, safer and easier, and the measured result was the same either way, about 50 Gb/s and about 5 µs.[1]

E810-CQDA1 or E810-CQDA2?

Intel builds both cards on the same Intel Ethernet Controller E810. Both are PCIe 4.0, run each port at 100, 50, 25 or 10 GbE, and support RoCE v2 and iWARP. The CQDA1 has one QSFP28 port; the CQDA2 has two.[3][4] kyuz0's guide used the CQDA1. In the forum thread, another reader ordered two E810-CQDA2 cards for the same kind of build, but has not posted results.[2] In the same thread, a user reported that E810 cards link at PCIe 4.0 x4 on these boards, with more than 50 Gb/s and under 5 µs, while the Mellanox cards they tried dropped to PCIe 3.0 x4 and 28 Gb/s.[2]

Software

  • Operating system: Fedora Linux 43 on both nodes. The kernels kyuz0 verified are 6.18.5-200.fc43 (node 1) and 6.18.6-200.fc43 (node 2).[1]
  • Drivers: the in-kernel ice and irdma drivers. The guide says you do not need Intel's own drivers.[1]
  • Packages: rdma-core, libibverbs-utils and perftest.[1]
  • Container: kyuz0's toolbox image includes a custom-built librccl.so with Strix Halo (gfx1151) support for RDMA, which upstream ROCm packages did not have when the guide was written.[1]

Setup, step by step

  1. Check the card firmware. Run ethtool -i on the E810 interface. If the firmware is older than 4.91, update it with Intel's NVM Update Utility for the E810 series.[1][6]
  2. Give the link fixed addresses. Node 1 (head) gets 192.168.100.1/30, node 2 (worker) 192.168.100.2/30. Set MTU 9000 on both and put the interface in the firewall's trusted zone.[1]
  3. BIOS: set the iGPU memory allocation to 512 MB.[1]
  4. Kernel parameters: add the five below to GRUB_CMDLINE_LINUX.[1]
  5. Passwordless SSH between the two nodes.[1]
  6. Install the toolbox with the repository's refresh_toolbox.sh. It detects the RDMA devices and passes them into the container.[1]
  7. Prove RDMA works with the comparison script before running a model (results below).[1]
  8. Start the cluster with start-vllm-cluster: Ray head on node 1, worker on node 2, then vLLM with tensor parallelism 2 and "Force Eager Mode" on.[1]
Kernel parameter What it does, per the guide[1]
iommu=pt IOMMU pass-through mode, for performance
pci=realloc Reallocates PCI BARs so large devices such as the E810 map properly
pcie_aspm=off Turns off PCIe power saving, which can cause latency spikes
amdgpu.gttsize=126976 Caps GPU GTT memory at about 124 GiB
ttm.pages_limit=32505856 Limits the TTM memory manager to about 124 GiB

Measured results

Chart: latency 74 µs, 68 µs and 5.23 µs; bandwidth 0.94, 55.70 and 50.64 Gb/s for 1G Ethernet, TCP over the E810 and RDMA
Figure 2. kyuz0's measurements, drawn by us.[1]
Path Latency Bandwidth
Ethernet, 1G LAN 0.074 ms (74 µs) 0.94 Gb/s
Ethernet (TCP) over the E810 0.068 ms (68 µs) 55.70 Gb/s
RDMA (RoCE v2) over the E810 5.23 µs 50.64 Gb/s

The fast card alone barely moves latency: TCP over the E810 is 68 µs, close to the 1G LAN. RDMA is what brings it down to 5.23 µs.[1] The guide gives example models for the cluster (Meta-Llama-3.1-8B-Instruct, and gemma-2-27b-it as a gated model) but does not publish tokens per second.[1]

Known issues

  • Deadlocks with CUDA graphs: on distributed APU clusters they can hang, so the guide runs vLLM in eager mode. kyuz0 estimates you might gain 1 to 3% by turning it off, at your own risk.[1]
  • Gated models such as Gemma need a Hugging Face token exported before launch.[1]
  • First run: each node downloads the model weights on its own.[1]
  • Link problems: update the E810 firmware first.[1]
  • No E810? The guide also covers a Thunderbolt/USB4 link. It uses the normal TCP/IP stack, so it does not reach RDMA latency.[1]

Want to build it?

We don't stock the Intel E810 cards, the DAC or the Framework boards for this build in our shop yet. Ask us and we will quote the parts the guide lists. We only quote parts named in the guide or in the manufacturers' own documents.

Credit: project, guide and measurements by kyuz0. Diagrams are our own drawings of the published setup.

Sources

  1. kyuz0, AMD Strix Halo RDMA cluster setup guide (GitHub)
  2. kyuz0 and replies, Low-latency Strix Halo cluster with RDMA (RoCE/Intel E810) and vLLM (Framework Community, independent)
  3. Intel, Ethernet Network Adapter E810-CQDA1 specifications
  4. Intel, Ethernet Network Adapter E810-CQDA2 specifications
  5. Framework, Framework Desktop
  6. Intel, NVM Update Utility for the E810 series (Linux)
  7. kyuz0, amd-strix-halo-vllm-toolboxes (GitHub repository)

Keep reading

More projects.