VorticPanel

Installation

Adding a GPU container host

A GPU container host is a node that runs customers’ GPU containers instead of KVM servers, the way GPU rental sites such as vast.ai work. A container starts in seconds because there’s no virtual machine to boot, and it gets whole cards straight from the NVIDIA driver on the node.

Adding one works like adding a node: enroll it in the panel, then run one command on it. The command installs the NVIDIA driver, Docker and the NVIDIA Container Toolkit instead of KVM.

Once the host is ready, set up what customers rent on it (images, offers and instances) in GPU containers.

How it differs from a KVM host

KVM host GPU container host
Runs Servers from your packages and OS templates Customers’ GPU containers on Docker. Servers are never placed on it.
Its cards Go to one server whole, through vfio-pci Stay on the NVIDIA driver; containers are given them by their ID
Needs VT-x or AMD-V, and IOMMU for cards NVIDIA cards. No virtualization or IOMMU.
Network The br-public bridge and IP pools The node’s own network. No bridge and no IP pools.

Which one a node is shows as a GPU containers badge next to its name on the Nodes page and on its own page. A node is one or the other, never both, so the NVIDIA driver and vfio-pci never fight over a card.

What the node needs

  • Ubuntu 22.04 / 24.04, Debian 12, or Rocky / Alma Linux 9, freshly installed.
  • One or more NVIDIA cards. Cards from Turing on (RTX 20 and newer, T4, A-series, L4, L40S, H100) use NVIDIA’s open driver, the default. Older cards (GTX 10, Tesla P100 and V100) need NVIDIA’s closed driver: add --nvidia-proprietary to the command.
  • Secure Boot off, or a console at the next boot to enroll the driver’s signing key. Otherwise the driver installs but doesn’t load.
  • Root access over SSH, and outbound HTTPS to the panel and to the package servers it installs from: developer.download.nvidia.com, nvidia.github.io and download.docker.com.

1. Enroll it in the panel

Go to Nodes → Enroll node. Under What it runs, choose GPU containers, then fill in the name, management IPv4, region and group as for any node. A name such as dfw-gpu-1 keeps it easy to tell apart.

Choose Get install command. The command ends in --gpu-containers; like any install command, it works once and expires after 60 minutes. A node waiting to be installed is marked GPU container host at the top of the Nodes page.

2. Run the command on the node

SSH into the node as root and paste the command:

curl -fsSL https://panel.example.com/agent/install.sh | sudo sh -s -- --token enr_… --gpu-containers

For cards older than Turing, add --nvidia-proprietary at the end. --storage-disk /dev/nvme1n1 works as on a KVM host, and makes an empty disk into storage.

What it does

  1. Checks that the panel is reachable. It doesn’t check for hardware virtualization, which containers don’t need.
  2. Installs nftables, LVM tools, an SSH server and Node.js 24. It doesn’t install KVM or libvirt.
  3. Installs the NVIDIA driver from NVIDIA’s own repository (nvidia-open, or cuda-drivers with --nvidia-proprietary), with the kernel headers it’s built against. If nvidia-smi already works, it leaves the driver you have.
  4. Installs Docker Engine from Docker’s repository, unless Docker is already there. It writes /etc/docker/daemon.json if there isn’t one: containers keep running while Docker restarts or updates, and their logs are capped at three files of 20 MB.
  5. Installs the NVIDIA Container Toolkit, makes it Docker’s NVIDIA runtime, and adds panel-nvidia-cdi.service, which describes the cards for containers (/etc/cdi/nvidia.yaml) at every boot, so a new or swapped card is picked up.
  6. Enrolls the node and starts the panel-agent service. There’s no network bridge, so nothing about the network changes.

Reboot if it asks

If the script installed the driver and the open-source nouveau driver still has the cards, it ends with Reboot this machine to load the NVIDIA driver. Reboot; the node comes back online by itself and describes its cards for containers as it starts.

3. Check it’s ready

The node shows as online in the panel within a few seconds. On its Overview, under Host, it shows the NVIDIA driver (with the newest CUDA it supports), Docker and the toolkit, each with its version.

Until everything is there, the node’s page says This GPU container host isn’t ready for containers yet, followed by what’s missing:

It says What to do
The NVIDIA driver isn’t loaded Reboot if the driver was just installed. Otherwise run the setup command. With Secure Boot on, enroll the driver’s key at the console, or switch Secure Boot off.
Docker isn’t running systemctl enable --now docker, or run the setup command
The NVIDIA Container Toolkit isn’t installed Run the setup command
The toolkit hasn’t described the cards for containers yet nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml, or reboot
No NVIDIA card was found on it Fit one, then choose Look again on the node’s Assets tab
Its agent hasn’t said what’s installed yet The agent is older than this feature. It updates itself; the message clears within a few minutes.

The agent reports what’s installed when it connects and when you choose Look again on the Assets tab.

4. Switch on its cards

The node’s Assets tab lists its cards, each with its memory, the driver it’s on and its ID from the NVIDIA driver (GPU-…). New cards start switched off: Let containers use it makes one available.

A card can’t go to containers when:

  • It isn’t on the NVIDIA driver. A card on vfio-pci was set aside for passthrough (often with vfio-pci.ids= on the kernel command line or in /etc/modprobe.d). Take it out of those settings and reboot.
  • It isn’t an NVIDIA card. Only NVIDIA cards work in containers for now.

Cards on GPU container hosts aren’t offered to packages, since packages are for servers. Next, create an offer for them: see GPU containers.

Turning an existing node into one

A node can change what it runs only while no servers are on it: move them off first (see Moving servers).

  1. On the node’s Settings tab, under What it runs, choose Make it a GPU container host. Its cards are switched off, so you decide again which ones containers can use.
  2. Install what it needs. On a node that’s already enrolled, run the install script with --gpu-containers and no token. It only installs the driver, Docker and the toolkit, then restarts the agent:
curl -fsSL https://panel.example.com/agent/install.sh | sudo sh -s -- --gpu-containers

The node’s page shows this command, with your panel’s address, for as long as something is missing. KVM and libvirt stay installed but unused, and the node’s IP pools and bridge settings no longer apply.

Make it a KVM host on the same tab goes the other way. The node then needs KVM, libvirt and the br-public bridge; if it was enrolled as a GPU container host, the simplest way is to remove it and enroll it again as a KVM host.

The files it adds

Besides the agent’s own files:

Path What
/etc/docker/daemon.json Docker’s settings: live restore, capped logs, and the NVIDIA runtime the toolkit adds
/etc/cdi/nvidia.yaml The cards, described for containers (CDI)
/etc/systemd/system/panel-nvidia-cdi.service Writes /etc/cdi/nvidia.yaml again at every boot
/etc/apt/sources.list.d/docker.list, nvidia-container-toolkit.list, and NVIDIA’s cuda-keyring The package repositories on Debian and Ubuntu, so apt upgrade keeps them current. On Rocky and Alma they’re in /etc/yum.repos.d/.

Every word has to appear. ↑ ↓ to move, Enter to open.