VorticPanel

Administration

Nodes

Nodes lists every hypervisor with its status, load, servers and agent version. To add one, see Adding a node.

What a node runs

A node is either a KVM host, which runs servers, or a GPU container host, which runs customers’ GPU containers on Docker with its cards on the NVIDIA driver. You choose when you enroll it (What it runs); nodes from before there was a choice are KVM hosts. GPU container hosts have a GPU containers badge next to their name. See Adding a GPU container host.

  • Servers are never placed on a GPU container host, not even when staff pick a node by hand or move a server. Node groups and the placement preview list it as Runs GPU containers, not servers.
  • To change what a node runs, use Settings → What it runs. It works only while no servers are on the node, and switches its cards off so you decide again who can use them.
  • On a GPU container host, the settings for servers alone are hidden: CPU mode, nested virtualization, bridges, IPv6 and IP pools. Its Overview shows the NVIDIA driver, Docker and the NVIDIA Container Toolkit in place of libvirt and QEMU, and its page lists anything still missing before it can run containers.

Status

Status Meaning
Online The agent is connected
Offline No connection for a minute. Its servers keep running; telemetry catches up when it’s back.
Draining No new servers are placed. Existing servers keep running and can be moved off.
Maintenance No new servers. Customers see a notice and can’t run power, reinstall or console actions on its servers.

Putting a node into maintenance notifies every customer with a server on it. Maintenance windows can do this for exactly the window.

Settings

Each node’s Settings tab:

Setting Default Allowed
Group From enrollment A group in the node’s region
Server limit No limit 1–2000
CPU overcommit 4× 1–16, in steps of 0.5
Memory overcommit 1×, or 1.1× on ZFS 1–2, in steps of 0.05
Memory reserved for the host About an eighth of memory: 16 GB from 128 GB, 32 GB from 768 GB Up to half the memory
CPU mode host-model host-passthrough, host-model, qemu64
Nested virtualization Off
Public bridge br-public, or the first bridge the node has A bridge on the node
Uplink 10 Gbit/s (25 Gbit/s from 96 cores) 1, 10, 25, 40 or 100 Gbit/s
IPv6 On
Notes Up to 2000 characters, staff only

Credential

Rotate credential issues the agent a new secret. The agent must be online. The old one keeps working until the agent confirms it saved the new one, so a rotation never cuts a node off. Credentials are valid for a year and renew automatically within 30 days of expiry.

Removing a node

Move its servers off and delete its GPU instances first; a node with either can’t be removed. Settings → Remove node takes it out of the panel: its credential is revoked, so its agent can never connect again, its storage pools are forgotten, and IP pools stop placing addresses through it. You can enroll the machine again later with a new enrollment token.

Nothing on the machine itself is touched. The agent stays installed and keeps trying to connect (it’s refused each time), and the bridge, firewall tables and storage it set up stay as they are. To use the machine for something else, such as another panel, uninstall the agent too.

Uninstalling the agent

To keep the OS, remove the node in the panel first, then run these on the machine as root.

1. Keep the network backup, if you want it. When the installer moved the public network onto br-public, it saved the old network files in /var/lib/panel-agent/network-backup-…, which the next step deletes. Copy them somewhere else first if you might put the network back as it was:

cp -a /var/lib/panel-agent/network-backup-* /root/ 2>/dev/null; ls -d /root/network-backup-*

2. Stop the agent and remove its service and files:

systemctl disable --now panel-agent
systemctl disable panel-firewall 2>/dev/null
rm -f /etc/systemd/system/panel-agent.service /etc/systemd/system/panel-firewall.service
systemctl daemon-reload
rm -rf /opt/panel-agent /etc/panel-agent /var/lib/panel-agent

3. Remove its firewall tables (servers’ firewall, private network tunnels, and on a GPU container host the containers’ isolation):

for t in "bridge panel" "inet panel_vxlan" "inet panel_containers"; do nft delete table $t 2>/dev/null; done; true

4. Remove its SSH lines. Moves between nodes add lines to root’s SSH files, marked panel-migration (keys of other nodes let in) and panel-node (other nodes’ host keys). A finished move takes its key out again, but an interrupted one can leave it:

sed -i '/ panel-migration$/d' /root/.ssh/authorized_keys 2>/dev/null
sed -i '/ panel-node$/d' /root/.ssh/known_hosts 2>/dev/null; true

Do step 4 on the panel’s other nodes too, for lines letting this machine’s key in, if a move from it was ever interrupted: grep panel-migration /root/.ssh/authorized_keys lists them, each with the address it’s allowed from.

What’s left on purpose:

  • The br-public bridge. The machine’s public address is on it, so taking it down cuts the network. Other hypervisor panels need a bridge for their servers too: point the new one at br-public (or rename it in the network configuration, from the console in case the network drops). To go back to the network as it was, copy the files from the backup in step 1 back to where their names say (etc__netplan__50-cloud-init.yaml is /etc/netplan/50-cloud-init.yaml), delete /etc/cloud/cloud.cfg.d/99-panel-network.cfg if it’s there (it stops cloud-init rewriting the network), and reboot from the console.
  • Storage. The LVM volume group panel (if the node was installed with --storage-disk), ZFS pools, and /var/lib/libvirt/images/. With no servers left they hold nothing of use; remove them with vgremove panel, zpool destroy or rm once you’re sure.
  • Packages: libvirt, QEMU, nftables, Node.js and the rest, which other hypervisor panels use too. On a GPU container host: Docker, the NVIDIA driver and Container Toolkit, /etc/docker/daemon.json, and panel-nvidia-cdi.service (systemctl disable panel-nvidia-cdi and delete its unit file if you don’t want it). See the files it adds. Leftover GPU containers and the panel-gpu Docker network go with docker rm -f $(docker ps -aq --filter label=panel.instance) and docker network rm panel-gpu.
  • A few lines in /etc/lvm/lvmlocal.conf marked panel:, which stop the host from picking up LVM that servers made on their own disks. They’re harmless, and worth keeping if the machine keeps hosting virtual machines.

Usage by server

A node’s Usage tab shows what fills it and who’s using the most, averaged over the last 5 minutes, hour or 24 hours:

  • Four pie charts: CPU (cores kept busy, against the node’s cores, with the free cores as their own slice), memory (what each server is given, against the node’s memory), network (in and out) and disk I/O (IOPS). Every server is a slice of its own. Hover any slice to see which server it is, its customer, its figure and its share; its row lights up in the list below. The four heaviest servers across the charts get a colour, the same one in every chart and in the list; the rest alternate between two greys.
  • The list: every server on the node with its CPU (as a share of its own vCPUs and in cores), memory, network in and out, disk IOPS and MB/s, and traffic this month against its allowance. Sort by any column, search by server or customer, or show only the ones that stand out.
  • Stands out: servers averaging 90% CPU or more, 500 Mbit/s or more outbound, or 5,000 IOPS or more over the whole window are flagged, and listed at the top of the tab. That’s where to look for mining, floods or a runaway database. Open a server to see its graphs, limit its speed or suspend it.

Memory is what each server is given: how much of it the guest actually uses isn’t visible from the host. Figures come from the node’s agent every few seconds, and the 24-hour view from the per-minute history.

Assets and GPUs

A node’s Assets tab lists hardware servers can be given besides CPU, memory and storage. For now that’s graphics cards, of any make or model (consumer or data-centre), passed through to one server whole.

The board’s own video chip isn’t listed (the Matrox in HP iLO and Dell iDRAC, ASPEED on most server boards, which often share a slot with the management controller), and neither is graphics built into the CPU (Intel UHD, AMD APUs).

The agent finds every card when it connects; Look again asks it again after you fit one. Each card shows its PCI address and its other functions (such as its HDMI audio, which goes to the server with it), its IOMMU group, the driver the node has it on, and whether it can be split. New cards start switched off: Let servers use it makes one available.

What the node needs

  • IOMMU on. In the BIOS, turn on VT-d (Intel) or AMD-Vi (AMD). Add intel_iommu=on iommu=pt or amd_iommu=on iommu=pt to GRUB_CMDLINE_LINUX in /etc/default/grub, run update-grub and reboot. Until then the tab says IOMMU is off and no card can be given out.
  • A card alone in its IOMMU group, apart from its own functions and bridges. If something else shares it (a USB controller, say), it would go to the server too: the card lists it. Use another slot, or turn on ACS in the BIOS.
  • The node not using the card. libvirt moves the card from the node’s driver to vfio-pci when a server starts and back when it stops. If the node’s own driver (nvidia, amdgpu, nouveau) won’t let go, blacklist it on the node. A card the node’s screen runs on is marked: giving it away takes the local console with it.
  • lspci (the pciutils package) for card names. Nodes installed from now on have it; without it a card shows by its vendor and device IDs.

Giving a card to a server

  • On the server’s Hardware tab, under GPUs, pick a free card on its node and choose Give it this card. The server picks it up the next time it’s stopped and started, like other hardware changes; Take off works the same way. A package can also give its servers cards; see Packages.
  • Customers see the card’s name on their server’s Overview, not where it is.
  • For NVIDIA cards the server is hidden from the card’s driver as a VM, which older consumer drivers refused to run in (error 43).
  • A server with a card can’t move to another node yet: take the card off, move it, then give it one on the new node. Emptying a node lists such servers as blocked.
  • Some AMD cards don’t reset when a server stops (the “reset bug”); if one stops working after its server restarts, install vendor-reset on the node.
  • Cards that can be split between servers (NVIDIA vGPU or MIG, AMD MxGPU, Intel SR-IOV) are marked, but for now go to one server whole.

Cards on a GPU container host

On a GPU container host the Assets tab is for containers: Let containers use it switches a card on. Nothing about passthrough applies, so IOMMU groups aren’t shown. Each card shows its memory and its ID from the NVIDIA driver (GPU-…), which containers are given it by. A card can’t go to containers unless it’s an NVIDIA card on the nvidia driver; one on vfio-pci says so. These cards aren’t offered to packages, and can’t be given to servers.

Agent versions

The Nodes page, the fleet overview and the node’s page mark agents that are behind the one the controller hands out. Nodes update their agent on their own after you upgrade the controller, by default one node first and the rest after it has run the new agent for 5 minutes, 3 at a time (change that under Settings → Agent → How updates go out: all at once, one at a time with a wait between, or a longer first look; see Upgrading), or when you press Update agent under the node’s Settings → Agent; servers keep running while it restarts. The node’s Overview also shows the Panel address its agent connects to, marked old address after the panel moved until the agent has switched (see Moving to a domain). See Upgrading. An agent too old to talk to the controller is refused with “This controller speaks agent protocol 1; update the agent.”

Emptying a node

On a draining or in-maintenance node, Move servers off plans where every server goes, using the group’s placement rule and taking earlier moves into account. It then moves them two at a time, with a progress banner and a way to stop after the current moves. Servers that don’t fit anywhere stay put, and the dialog says why. See Moving servers.

Every word has to appear. ↑ ↓ to move, Enter to open.