Two AMD Strix Halo boards can split one large model between them with vLLM, but only if the link between them is fast. The developer kyuz0 put an Intel E810 100 GbE card in each of two Framework Desktop boards, joined them with one cable and switched the link to RDMA. Node-to-node latency dropped from 74 µs to 5.23 µs.[1] The guide is public, step by step, and lists every part and setting. This is a summary of it, with the manufacturer specs next to it.
The project at a glance
| Item | Detail |
|---|---|
| Built by | kyuz0, who maintains the amd-strix-halo-vllm-toolboxes repository on GitHub[1][7] |
| Posted | Framework Community, 8 February 2026; setup guide in the same repository[1][2] |
| Goal | Run one model across two nodes with vLLM tensor parallelism (TP = 2)[1] |
| Nodes | 2 × Framework Desktop Mainboard, AMD Ryzen AI Max+ "Strix Halo", 128 GB unified memory[1] |
| Link | Intel E810-CQDA1 in each node, one 100G QSFP28 DAC, no switch[1] |
| Result | 5.23 µs latency and 50.64 Gb/s over RDMA[1] |
Why RDMA matters here
With TP = 2, each node holds half of every layer. After every layer the two nodes swap partial results, which the guide says happens thousands of times a second.[1] Over a normal TCP/IP connection that costs about 70 to 100 µs each time; with RDMA it is about 5 µs.[1] RDMA over Converged Ethernet (RoCE v2) lets one node write straight into the other node's memory without going through the CPU and the kernel network stack.[1]
Parts list
| Part | What kyuz0 used | Notes |
|---|---|---|
| Mainboard × 2 | Framework Desktop Mainboard, Ryzen AI Max+, 128 GB[1] | Framework sells the board with a PCIe x4 slot[5] |
| Network card × 2 | Intel E810-CQDA1, or a similar 100 GbE QSFP28 card[1] | Firmware 4.91 or newer recommended[1] |
| Cable × 1 | 100G QSFP28 direct attach copper (DAC); the guide gives a QSFPTEK cable as an example[1] | Point to point; no switch needed for two nodes[1] |
| Riser × 2 | PCIe x4 to x16 extender[1] | The Framework slot is physically x4, so a x16 card needs a riser[1] |
One of kyuz0's boards had its slot cut open by Framework so a x16 card fits directly. kyuz0 does not recommend that: risers are cheaper, safer and easier, and the measured result was the same either way, about 50 Gb/s and about 5 µs.[1]
E810-CQDA1 or E810-CQDA2?
Intel builds both cards on the same Intel Ethernet Controller E810. Both are PCIe 4.0, run each port at 100, 50, 25 or 10 GbE, and support RoCE v2 and iWARP. The CQDA1 has one QSFP28 port; the CQDA2 has two.[3][4] kyuz0's guide used the CQDA1. In the forum thread, another reader ordered two E810-CQDA2 cards for the same kind of build, but has not posted results.[2] In the same thread, a user reported that E810 cards link at PCIe 4.0 x4 on these boards, with more than 50 Gb/s and under 5 µs, while the Mellanox cards they tried dropped to PCIe 3.0 x4 and 28 Gb/s.[2]
Software
- Operating system: Fedora Linux 43 on both nodes. The kernels kyuz0 verified are 6.18.5-200.fc43 (node 1) and 6.18.6-200.fc43 (node 2).[1]
-
Drivers: the in-kernel
iceandirdmadrivers. The guide says you do not need Intel's own drivers.[1] -
Packages:
rdma-core,libibverbs-utilsandperftest.[1] -
Container: kyuz0's toolbox image includes a custom-built
librccl.sowith Strix Halo (gfx1151) support for RDMA, which upstream ROCm packages did not have when the guide was written.[1]
Setup, step by step
-
Check the card firmware. Run
ethtool -ion the E810 interface. If the firmware is older than 4.91, update it with Intel's NVM Update Utility for the E810 series.[1][6] -
Give the link fixed addresses. Node 1 (head) gets
192.168.100.1/30, node 2 (worker)192.168.100.2/30. Set MTU 9000 on both and put the interface in the firewall's trusted zone.[1] - BIOS: set the iGPU memory allocation to 512 MB.[1]
-
Kernel parameters: add the five below to
GRUB_CMDLINE_LINUX.[1] - Passwordless SSH between the two nodes.[1]
-
Install the toolbox with the repository's
refresh_toolbox.sh. It detects the RDMA devices and passes them into the container.[1] - Prove RDMA works with the comparison script before running a model (results below).[1]
-
Start the cluster with
start-vllm-cluster: Ray head on node 1, worker on node 2, then vLLM with tensor parallelism 2 and "Force Eager Mode" on.[1]
| Kernel parameter | What it does, per the guide[1] |
|---|---|
iommu=pt |
IOMMU pass-through mode, for performance |
pci=realloc |
Reallocates PCI BARs so large devices such as the E810 map properly |
pcie_aspm=off |
Turns off PCIe power saving, which can cause latency spikes |
amdgpu.gttsize=126976 |
Caps GPU GTT memory at about 124 GiB |
ttm.pages_limit=32505856 |
Limits the TTM memory manager to about 124 GiB |
Measured results
| Path | Latency | Bandwidth |
|---|---|---|
| Ethernet, 1G LAN | 0.074 ms (74 µs) | 0.94 Gb/s |
| Ethernet (TCP) over the E810 | 0.068 ms (68 µs) | 55.70 Gb/s |
| RDMA (RoCE v2) over the E810 | 5.23 µs | 50.64 Gb/s |
The fast card alone barely moves latency: TCP over the E810 is 68 µs, close to the 1G LAN. RDMA is what brings it down to 5.23 µs.[1] The guide gives example models for the cluster (Meta-Llama-3.1-8B-Instruct, and gemma-2-27b-it as a gated model) but does not publish tokens per second.[1]
Known issues
- Deadlocks with CUDA graphs: on distributed APU clusters they can hang, so the guide runs vLLM in eager mode. kyuz0 estimates you might gain 1 to 3% by turning it off, at your own risk.[1]
- Gated models such as Gemma need a Hugging Face token exported before launch.[1]
- First run: each node downloads the model weights on its own.[1]
- Link problems: update the E810 firmware first.[1]
- No E810? The guide also covers a Thunderbolt/USB4 link. It uses the normal TCP/IP stack, so it does not reach RDMA latency.[1]
Want to build it?
We don't stock the Intel E810 cards, the DAC or the Framework boards for this build in our shop yet. Ask us and we will quote the parts the guide lists. We only quote parts named in the guide or in the manufacturers' own documents.
Credit: project, guide and measurements by kyuz0. Diagrams are our own drawings of the published setup.
Sources
- kyuz0, AMD Strix Halo RDMA cluster setup guide (GitHub)
- kyuz0 and replies, Low-latency Strix Halo cluster with RDMA (RoCE/Intel E810) and vLLM (Framework Community, independent)
- Intel, Ethernet Network Adapter E810-CQDA1 specifications
- Intel, Ethernet Network Adapter E810-CQDA2 specifications
- Framework, Framework Desktop
- Intel, NVM Update Utility for the E810 series (Linux)
- kyuz0, amd-strix-halo-vllm-toolboxes (GitHub repository)