Every GB10 system has an NVIDIA ConnectX-7 network card with two QSFP ports running at up to 200 Gb/s. Connect two systems directly with one cable and they can share a model too big for either one alone: roughly 405 billion parameters at 4-bit, against about 200 billion on a single unit. No switch is needed for two nodes.
What you need
- Two GB10 systems. You can mix brands; they all use the same ConnectX-7 hardware.
- One short QSFP112 direct-attach cable, such as the 0.5 m QSFP112 cable we stock, which works with every GB10 system. One cable gives you the full bandwidth.
- Both systems already set up and updated (see getting started), ideally with the same username on both.
Buying from scratch? Our DGX Spark, HP ZGX Nano and Lenovo ThinkStation PGX 2-node bundles include both systems and the cable.
1. Plug in the cable
- Power both systems off before inserting or removing the cable.
- Use the same QSFP port on both systems (left to left, or right to right). Mismatched ports are a common cause of failing tests later.
- The connector only fits one way up. If it doesn't slide in, don't force it.
- To remove it, pull the pull-tab straight back. Don't twist the connector or pull on the cable itself.
2. Check the link
Power both systems on and run ibdev2netdev on each. The ports with the cable attached should show as (Up). Note the interface names; you'll need them next.
3. Give the ports addresses
The two systems need fixed IP addresses on the direct link. NVIDIA's playbook uses a small netplan file (for example /etc/netplan/40-cx7.yaml) with addresses such as:
- Node 1:
192.168.100.10/24 - Node 2:
192.168.100.11/24
Apply it with sudo netplan apply and confirm with ip addr show <interface>. You can also assign addresses with ip addr add for a quick test, but those are lost on reboot, so use netplan for anything permanent.
4. Set up passwordless SSH
Multi-node tools start processes on both machines over SSH. Either run NVIDIA's discovery script from the playbook, or copy your key across manually with ssh-copy-id in both directions. Test it: ssh 192.168.100.11 hostname should answer without asking for a password, and the same from the other side.
5. Prove the bandwidth, then run a model
- Run the NCCL test from NVIDIA's playbooks to confirm the GPUs can talk at full speed over the link.
- Then follow the multi-node vLLM or TensorRT-LLM playbook to split a large model across both systems.
All the official steps are at build.nvidia.com/spark, under Connect Two Sparks and NCCL.
Going beyond two
For three or more systems, connect them through a switch instead of point to point. Have a look at our switches and small-lab networking guide, or ask us to size a cluster with you.
Troubleshooting
- Port shows Down: reseat the cable with both systems off, and check you used the same port on both.
- SSH asks for a password: the key wasn't copied in both directions, or the usernames differ.
-
Addresses gone after reboot: they were set with
ip addr; move them into netplan.