One GPU for AI, shared by many users: virtual machines and containers on Proxmox

A 48 GB NVIDIA Quadro RTX 8000 GPU shared between virtual machines and LXC containers on Proxmox: why the standard options fall short, what the "merged" driver solves and what it costs, plus a remote power button built on an ESP32 microcontroller.

ProxmoxNVIDIALXCESPHomeMQTT
Contents
  1. The problem: one card, different needs
  2. Two ways to share a GPU
  3. vGPU and "merged" drivers
  4. Letting a container see the GPU
  5. Switching it on remotely: an ESP32 on the button
  6. What changes with the NanoKVM

In short

  • A 48 GB graphics card (GPU) to be shared by several people working with AI models, without giving everyone access to the whole machine.
  • On Proxmox a GPU can be split between virtual machines (each gets a fixed "slice") or shared between containers, which all use it together.
  • With a particular driver, known as "merged", you can do both on the same card.
  • In containers the critical point is the driver version: it must match the host's exactly, down to the last digit.
  • Until a remote console (NanoKVM) is in place, I wired a small microcontroller (ESP32) to the power button to switch the server on and off remotely. It knows whether the server is on by reading the power LED line, without asking the server anything.

At muHack, the hackerspace in Brescia, we have a server with an NVIDIA Quadro RTX 8000: a professional card with 48 GB of memory, enough to run mid-sized language models or train small ones. There is only one card, though, and it has to serve several people, each with their own projects and libraries.

The server runs Proxmox VE, the Debian-based virtualisation platform I also use for work. This post covers how I made the GPU available to both virtual machines and containers, and how to switch the server on without being there.

The problem: one card, different needs

People using the server don't all need the same thing. Some want a machine of their own, perhaps with Windows or a different kernel, where they can do whatever they like without disturbing anyone. Others just want to run a Jupyter notebook or a model for a few hours, and for them a lightweight container is more convenient and wastes fewer resources. The GPU has to serve both, and on Proxmox the standard options force a choice.

  • PCI passthrough. The whole card is assigned to a single virtual machine. It's the simplest option with the best performance, but while that virtual machine (VM) is running nobody else, the host included, can use the GPU. With 48 GB and one user at a time, most of the memory sits idle.
  • Normal driver on the host, GPU to the containers. Every container sees the card and they share it. This works well, but virtual machines get no GPU: anyone who needs a different operating system or real isolation is left out.
  • vGPU driver on the host. The card is split into virtual GPUs, one per VM, each with its own reserved memory. But, as explained below, the host can no longer use it for compute, and so neither can the containers.

None of the three covers every case. The "merged" driver is how you get the second and third together.

Two ways to share a GPU

Proxmox offers two kinds of guest: virtual machines and LXC containers. For a GPU the difference is large.

Virtual machine with vGPULXC container
How it sees the GPUa virtual slice, with reserved memorythe whole card, shared with other containers
Isolationstrong: its own kernel and driverweak: same kernel as the host
Driver in the guestvGPU guest driveronly the user-space part of the driver, same version as the host
Performanceclose to native, within its slicenative
When it fitsseparate environments, different operating systemsmany trusted users, intermittent GPU use

The difference, in plain words. A virtual machine is a complete fake computer with its own operating system: with vGPU it gets a fixed piece of the card, like a flat in a block of flats. A container is lighter: it shares the host's operating system and sees the whole GPU, like flatmates sharing the same kitchen.

vGPU and "merged" drivers

vGPU is NVIDIA's technology for splitting a card into several virtual GPUs, each assigned to a VM. It's normally reserved for datacentre cards and needs specific drivers and licences. The RTX 8000 is one of the professional cards that supports it natively, so no unlock is needed.

The host vGPU driver has a limitation, though: it turns the card into a "factory" of virtual GPUs, and the host itself can no longer use it for compute. That's a problem, because containers don't have a driver of their own: they use the host's.

The solution I used is a "merged" driver, put together by the community working on vGPU for Proxmox. It combines the host vGPU driver and the regular driver in one package. With it, the same card can create virtual GPUs for VMs and, at the same time, work for the host and therefore for the containers.

The steps on the host are:

  1. enable IOMMU, the feature that lets devices be assigned to VMs, in the BIOS and in the kernel parameters;
  2. install the merged driver (version 550.90.07) with DKMS, so the module is rebuilt automatically when the Proxmox kernel is updated;
  3. reboot and check with nvidia-smi that the card is visible.
Virtual machine vGPU: a fixed slice guest driver LXC container whole GPU, shared user-space driver only mdev /dev/nvidia* Proxmox host merged driver 550.90.07 vGPU for VMs + regular driver for host and containers Quadro RTX 8000 48 GB, vGPU supported
The same card serves virtual machines, through virtual GPUs, and containers, through the host's driver.

The cost of the merged driver

It works well, but it isn't what NVIDIA intends, and it has to be managed with that in mind:

  • It isn't an official package. The community builds it by combining two NVIDIA drivers. If something breaks there's no support to turn to, and each new version arrives when someone prepares it.
  • It's tied to exact versions. The driver must be compatible with the Proxmox kernel. A kernel update can make the module rebuild fail, and after a reboot the GPU is gone. So kernel updates aren't applied blindly: first you check that the driver builds, and until then the running kernel stays pinned to the one that works.
  • It drags the containers along. As explained below, every container needs exactly the same driver version. Changing the driver on the host means updating every container.
  • Containers aren't isolated from each other. VMs have their own slice of memory; containers don't. A process in one container can take all 48 GB and leave the others with nothing. It takes a few rules agreed among the people using the server, and an eye on nvidia-smi to see who is using what.

For a shared server in a hackerspace these are acceptable trade-offs: flexibility is worth more than official support we wouldn't have anyway. In a company the same choice would have to be weighed differently.

Letting a container see the GPU

An LXC container has no kernel of its own: it uses the host's. That's why you don't install kernel modules in the container. Instead you need two things: make the GPU's device files visible to the container, and install only the user-space part of the driver inside it.

Device files

On Linux the GPU is used through special files in /dev: nvidia0 for the card, nvidiactl for control, nvidia-uvm for the unified memory used by CUDA (NVIDIA's compute platform), and /dev/dri/renderD128 for rendering. The container configuration needs two kinds of lines:

  • permissions: rules that allow the container to access devices with a given major number, the number the kernel assigns to each type of device;
  • mounts: each host device file is made visible at the same path inside the container.

The detail not to get wrong is that some major numbers are fixed (195 for the main NVIDIA devices, 226 for /dev/dri) while others, such as nvidia-uvm's, are assigned at boot and differ from one machine to another. Read them on the host with ls -l /dev/nvidia* before writing the configuration. Copying lines from a guide only works if the guide was written on the same machine.

Driver version

Inside the container you install NVIDIA's regular driver, the exact same version as the host's (550.90.07), with the option that skips the kernel modules. The reason is that the driver libraries in the container talk to the module in the host, and each side only accepts its own version. Even a difference in the last digit produces a version mismatch error and nvidia-smi won't run.

This has a practical consequence: every driver update on the host has to be repeated in every container. So it's better to have a few well-kept containers, or a template to create them from.

CUDA and PyTorch

For most uses you don't need the full CUDA Toolkit. PyTorch, for example, installs with its own CUDA runtime (here the CUDA 12.4 build) and only needs the driver. The toolkit is for people who need to compile CUDA code. So I didn't install it in the template container: I left a script in root's home directory (/root) that installs it, and only those who need it run it. The container stays smaller, and people who don't compile don't carry several gigabytes they don't use.

Switching it on remotely: an ESP32 on the button

The final solution for switching the server on and off remotely will be a NanoKVM: a small remote console (a "KVM over IP", from keyboard-video-mouse) that gives access to the screen, the keyboard and the power button through a browser, even when the operating system isn't responding. Until we install it, I built a temporary solution.

It's an ESP32-S2 (a Lolin S2 mini), a Wi-Fi microcontroller that costs a few euros, programmed with ESPHome: you describe the behaviour in a configuration file instead of writing the firmware from scratch. The ESP32 does two things: it presses the power button for us and, above all, it knows whether the server is on.

Phone or script MQTT over TLS MQTT broker Wi-Fi ESP32-S2 (ESPHome) presses reads Motherboard front panel power switch pins · power LED
The ESP32 is wired in parallel with the power button and reads the LED that shows whether the server is on.

State from the power LED

The most interesting part is how the ESP32 knows whether the server is on. It doesn't ask the server anything and doesn't use the network: it reads the power LED line on the front panel header, the same one that lights the little LED on the case.

That line is driven directly by the motherboard, so it reflects the real state of the hardware, not of the operating system. The difference matters:

MethodWhat it actually tells you
ping or a monitoring servicethat the operating system is up and the network works
software agent on the serverthat the operating system is up and the agent is running
power LEDthat the machine is powered, whatever the software is doing

If the server has hung, or stopped during boot, a ping would say "off" while the LED says "on". And that's exactly the information you need before pressing the button: pressing it on a machine you think is off, but is actually on, switches it off.

Electrically, the reading is simple. The line the ESP32 reads sits at about 3.3 V when the server is off and 0 V when it's on. The input is configured with no pull-up or pull-down resistors, so the ESP32 only listens and doesn't alter the LED circuit. Two filters keep the reading stable: the "on" state must last 300 milliseconds before it counts, the "off" state 2 seconds, so brief changes don't produce false transitions.

In plain words. Instead of asking the server "are you on?", a question a hung server can't answer, the ESP32 looks at its power light. It's what a person standing next to the machine would do.

Pressing the button without breaking anything

A PC's power button is just a contact that, when pressed, connects a motherboard pin to ground. The ESP32 does the same with an output in open-drain mode: when it "presses" it pulls the pin to ground, and when it doesn't it disconnects completely rather than driving a voltage. That way it never pushes current into the motherboard, and the physical button on the case keeps working as before.

Commands

The ESP32 connects to a private MQTT broker, over an encrypted TLS connection with authentication. It subscribes to a few topics, named channels, and publishes its state on another:

MessageWhat it does
short pressholds for 1.5 seconds: switches the server on, or asks it to shut down cleanly
long pressholds for 12 seconds, like holding the button by hand when the system doesn't respond
togglepicks the action by itself, based on the state read from the LED
statepublished by the ESP32 on every change: ON or OFF

The state is published as a retained message: the broker keeps the last value and delivers it straight away to anyone who connects, without waiting for the next change. So whoever opens an MQTT app on their phone sees at once whether the server is on.

What MQTT is. A very lightweight messaging system, common in home automation. There's a central server, the broker. Devices don't talk to each other directly: they publish messages on named channels and receive messages from the channels they subscribe to. The ESP32 doesn't need to be reachable from the internet: it's the one that connects to the broker.

What changes with the NanoKVM

The NanoKVM will do everything the ESP32 does and also give access to the screen and keyboard, and so to the BIOS and the boot process. It also connects to the front panel, for the power button and the power LED. So the idea of reading the state from the LED stays the same: only the device that does it changes.

← All posts