Every GB10 desktop we sell (the NVIDIA DGX Spark, Asus Ascent GX10, HP ZGX Nano and Lenovo ThinkStation PGX) is built on the same NVIDIA GB10 Grace Blackwell superchip with 128 GB of unified memory. That means the first hour looks almost identical on all of them. This guide walks you through it.
1. Before you power on
- Plug in everything first. GB10 systems start as soon as they get power, so connect your monitor (HDMI), keyboard, mouse and network cable before you plug in the power adapter.
- Use a wired connection if you can. The first-boot setup downloads updates, and a 10GbE or gigabit cable is faster and more reliable than Wi-Fi.
- Give it air. The box is small but draws up to around 240 W. Keep the vents clear and don't stack anything on top of it.
Planning to run it headless (without a monitor)? You can, but doing the very first setup with a screen attached is the least fiddly route.
2. First-boot setup
On first start a setup assistant walks you through language, time zone, network, and creating your user account. It then downloads and installs the latest updates, which can take a while and may restart the system. Let it finish before you install anything else.
Tip: pick the same username on every GB10 system you own. If you ever connect two of them into a cluster, matching usernames make the setup much simpler.
3. What is already installed
GB10 systems ship with NVIDIA DGX OS, an Ubuntu-based Linux with the NVIDIA AI software stack preinstalled: GPU drivers and CUDA, Docker with GPU support, access to NVIDIA NGC containers, and the DGX Dashboard, a browser-based page for monitoring the system that also opens JupyterLab. You don't need to install drivers yourself.
4. Reach it from your own laptop
Most people end up using their GB10 system as a quiet AI box on the desk or in a cupboard, and working on it from their normal computer:
- NVIDIA Sync (Windows, macOS and Linux) connects to the system by name, sets up SSH for you, and gives one-click access to the DGX Dashboard, a terminal and editors like VS Code or Cursor.
-
Plain SSH works too:
ssh yourname@your-spark.local. - To open the DGX Dashboard without Sync, tunnel its port over SSH and open it in your browser.
5. Run your first model
Three good starting points, from easiest to most production-like:
- Ollama with Open WebUI. The fastest way to chat with a model. Pull a model such as a Llama, Qwen or Gemma variant and you get a ChatGPT-style interface in your browser, served entirely from your desk.
- LM Studio or llama.cpp. Good for trying many quantised (GGUF) models quickly and comparing them.
- vLLM, TensorRT-LLM or NVIDIA NIM. For serving a model to a team or an application with an OpenAI-compatible API, with higher throughput.
NVIDIA publishes step-by-step playbooks for each of these (and for fine-tuning, image generation and multi-node setups) at build.nvidia.com/spark. They apply to every GB10 system, not just NVIDIA's own.
6. How big can you go?
With 128 GB of unified memory, one GB10 system can run models up to roughly 200 billion parameters when they are quantised to 4-bit (FP4). For a feel of what fits, read our memory sizing guide. Need more? Two systems linked with a QSFP112 cable handle models of around 405 billion parameters. See how to connect two GB10 systems.
Quick checklist
- Peripherals and network connected before power
- First-boot updates fully installed
- Same username planned for every unit
- NVIDIA Sync installed on your laptop
- First model running through Ollama or LM Studio