VorticPanel

Administration

GPU containers

GPU containers rent your NVIDIA cards out the way vast.ai and RunPod do: a customer gets a container with whole cards, ready over SSH or Jupyter in the time it takes to pull an image, with no virtual machine to boot. They run on GPU container hosts, never on your KVM nodes, and they sit beside servers in the panel rather than replacing anything.

Three things make them up, all under GPU containers in the admin sidebar:

What it is Like, for servers
Images A Docker image an instance runs, and how it’s reached OS templates
Offers What an instance gets: its cards, CPU, memory, /workspace and network, and where it can be ordered Packages
Instances A customer’s running container Servers

There’s no billing in the panel. Your billing portal sells the offer and creates, suspends and deletes instances through the API, as it does servers: see Connecting a billing panel.

Before you start

  1. Add a GPU container host, and on its Assets tab choose Let containers use it for each card you rent out.
  2. Put the host in a node group of its own, such as Frankfurt GPU. Offers are sold per node group, and servers are never placed on GPU container hosts anyway.
  3. Give the host storage for /workspace. The installer’s --storage-disk makes an LVM thin pool, where each instance gets a volume of exactly its size. Without LVM or ZFS, /workspace lives in Docker’s own volumes, with no size limit per instance.
  4. For offers with their own IPv4 address, attach an IPv4 pool to the host (IP pools).

Images

GPU containers → Images lists the images instances can run. A new panel starts with five: CUDA 12.4 on Ubuntu, PyTorch, TensorFlow with Jupyter, vLLM and Ollama. Change, hide or remove them, and add any Docker image from Docker Hub or another registry (ghcr.io/owner/name:tag). Pin a version tag rather than latest, so every host runs the same thing. Changing images needs the Packages and templates permission.

Field What it does
Docker image What the host pulls, the first time an instance needs it
How it’s reached SSH: an SSH server starts with the instance, with the customer’s SSH keys for root, installed in the container if the image has none. Jupyter and SSH: JupyterLab on port 8888 with a login token as well, installed with pip if missing. Its own command: the image runs as it’s made to, such as a server on a port, with no SSH.
Other ports Ports inside the container customers reach, such as 8000 for vLLM. Each takes one of the offer’s forwarded ports.
Oldest CUDA it needs Hosts whose NVIDIA driver supports an older CUDA aren’t used for it
Environment variables Set in every instance of the image; customers can read them. Names starting PANEL_ are the panel’s own.
Start-up script Runs as root in /workspace each time the instance starts, before SSH or Jupyter, such as pip install transformers

A hidden image can’t be picked for new instances or a change of image; instances running it keep it. An image instances run, or an offer lists, can only be hidden, not removed.

Offers

GPU containers → Offers → New offer, on a page of its own:

Setting Allowed
GPUs per instance and Model 1–8 cards, all on the same host; a model as hosts report it (such as NVIDIA GeForce RTX 4090), or any card
vCPU, Memory 1–256 vCPU; 1 GB to 2 TB, in steps of 0.25 GB. Hard limits for the container.
/workspace 10–32,000 GB, kept when the instance stops or changes image, removed when it’s deleted
Network Forwarded ports (1–100 per instance) on the host’s own address, or Own IPv4 from the host’s IP pools. See Reaching an instance.
Locations The node groups it can be ordered in
Images The images its instances can run; none ticked allows every available image, including ones added later
Billing product ID Your billing portal’s product, such as whmcs:41
Availability Orderable, or hidden (staff and the API can still use it). Archive an offer to stop new instances; existing ones keep running.

The offers list shows, per location, how many more instances fit now, or why none do: no GPU container host in the group, no free card of the model, not enough CPU, memory or storage, no forwarded ports or addresses left, or a host that isn’t ready (Docker or the driver missing). Changing an offer applies to new instances; existing ones keep what they were given.

Placement

A new instance goes to a GPU container host in its location that is online, not draining or in maintenance, has Docker, the NVIDIA driver and the toolkit, and has a new enough CUDA for the image. Of those, it takes the one with the most free memory that has enough free cards of the offer’s model, vCPUs, memory, storage and ports or addresses. vCPUs and memory count against the host’s overcommit settings, as for servers; cards are never shared. Staff can pick the host by hand when creating an instance.

An instance keeps its cards, ports and address while it’s stopped or suspended: the customer is paying for them. They’re freed when it’s deleted.

Instances

GPU containers → Instances lists every customer’s instance, its status, cards, image and host. Each has a number (#), as servers do, counted separately from servers and never reused; search for #12 or 12 to find instance 12. Tick several (or all) to start, stop or restart them at once from the bar at the bottom: instances an action doesn’t apply to (already in that state, busy or suspended) are skipped and counted. Staff need See all servers to see them; the server permissions decide the rest:

To Permission
Create one (New instance) Create servers
Start, stop, restart Power
Change its image Reinstall
Open its terminal, read its output Console
Suspend and unsuspend Suspend
Delete it, with its /workspace Terminate

An instance’s page is laid out like a server’s: its number and ID, name (with a pencil to rename it), state, cards, CPU, memory and uptime at the top with Start, Stop and Restart (and JupyterLab while it runs), then tabs: Overview (live usage, how to connect and its details, with the latest activity), Terminal (the terminal and its output), Image, Activity (every job), and for staff Manage (suspend, unsuspend or delete). A tab can be linked to with ?tab=terminal and the like. It shows how to connect, its cards and resources, its host and customer, a Terminal (a root shell inside the container, in the browser), Live usage while it runs (each card’s load, memory, temperature and power from nvidia-smi, CPU and memory from Docker, how full /workspace is, and network since it started; read from the host every few seconds while the page is open), its Output (the last 500 lines: the SSH server and Jupyter starting, the start-up script, its own program), and Image, to change it.

  • Changing the image (or Start afresh with the same one) makes the container again: everything outside /workspace is lost. /workspace, the address, the ports and the SSH keys stay.
  • SSH keys (Manage keys, under the SSH command): which keys on the customer’s account root’s login takes. A new instance takes all of them unless the order names some (sshKeyIds). The keys are a file on the host, /var/lib/panel-agent/state/containers/<id>.ssh/authorized_keys, mounted read-only at /etc/panel-ssh inside, so a change applies at once, running or stopped, and keys the customer adds to /root/.ssh/authorized_keys themselves are left alone. Instances made before this release pick the file up after one Start afresh.
  • Suspending stops it with a reason the customer sees; they can’t start it or open its terminal. Unsuspending starts it again.
  • Deleting frees its cards, ports and address straight away; the host removes the container and /workspace, at once or when it’s next connected.
  • A node with instances can’t be removed or made a KVM host, and a card an instance has can’t be switched off.

Activity on the instance’s page (and the latest few on its Overview) lists what its host has done to it: creating it, starting, stopping, restarting, changing the image, suspending and unsuspending. Each is a job under Operations → Jobs (kind GPU instance), with a live log of the image pull and the agent’s steps; a failed one says why and notifies staff. The customer sees their own instance’s jobs, without the host’s name. Every change to instances, offers and images (and every refused attempt) is in the Audit log under the GPU containers category; the instance page’s Audit log link shows just that instance’s.

Notifications (the bell, and email by each person’s notification settings): the customer hears when their instance is ready (with how to connect), when its image has changed, when it’s suspended, unsuspended or deleted, when its host goes into maintenance, and when it couldn’t be set up or change image (without the host’s own error, which staff get). Staff hear when work on an instance fails, when a GPU container host stops answering (counting the instances on it), and when a host that could run containers can’t any more, for example a driver update waiting for a reboot or Docker stopped, and again when it can.

Reaching an instance

Forwarded ports share the host’s own address. Each instance gets the offer’s number of ports from 40000–49999 on the host, TCP and UDP: SSH first (to port 22 inside), then Jupyter (8888), then the image’s own ports; the rest are spare, reaching the same port number inside. The customer sees, for example, ssh -p 40000 root@203.0.113.5. The address is the host’s public IPv4, or its management address when it has none.

Own IPv4 gives the instance an address from an IPv4 pool attached to the host, on a macvlan network over the host’s default-route interface, so every port on it is reachable at that address (ssh root@198.51.100.42). The address counts as used in the pool, like a server’s. Your provider has to accept another MAC address on the port, as it does for bridged servers.

How they run

  • Each instance is a Docker container named after its ID (gi_…), with only its own cards (by their NVIDIA UUID), CPU and memory limits, swap off, a /dev/shm of half its memory (up to 32 GB, for PyTorch’s data loaders), and a limit of 8,192 processes.
  • It isn’t privileged, can’t gain privileges, and can’t use raw sockets (so it can’t send from addresses it doesn’t have). Containers share the host’s kernel, so rent hosts only to customers you’d sell a server to.
  • /workspace is an ext4 volume of the instance’s own (gi_…-ws in the host’s LVM or ZFS), mounted by Docker, so its size is a hard limit. On a host whose only storage is a folder, /workspace is a plain Docker volume with no size limit: give GPU hosts an LVM thin pool or ZFS (the installer’s --storage-disk) before selling them. Docker restarts the container if the host restarts, unless it was stopped.
  • Instances with forwarded ports run on the panel’s own Docker network, panel-gpu, with traffic between containers switched off, so one customer can’t reach another’s services. A firewall table on the host (panel_containers) drops new connections from them to the host itself, and their traffic to cloud metadata (169.254.0.0/16) and to the internal ranges under Settings → Network, as for servers. Instances with an IPv4 address of their own sit directly on the public network (macvlan), so the host can’t filter them: keep that network’s other machines firewalled.
  • On the host: docker ps --filter label=panel.instance lists them, docker logs gi_… shows an instance’s output, and /var/lib/panel-agent/state/containers/ holds what the agent last set up for each.

Every word has to appear. ↑ ↓ to move, Enter to open.