The Question Every Proxmox Evaluation Starts With
“Is Proxmox actually enterprise grade?”
It comes up in nearly every migration conversation, and the worry underneath it is almost never the web interface. Nobody seriously fears that a browser dashboard will corrupt their data. What people are asking is whether the thing standing between a virtual machine and the hardware — the component that has to keep one tenant out of another tenant’s memory, forever, without a single mistake — is a serious piece of engineering or a community project that got popular.
That is exactly the right thing to be nervous about. It is just aimed at the wrong layer, because Proxmox VE does not contain a hypervisor.
The hypervisor is KVM. It is part of Linux, it has been there since 2007, and if your organisation uses EC2, Google Cloud, Oracle Cloud, Alibaba Cloud, DigitalOcean or Nutanix, you are already running it in production today — you have just never had to think about it, because somebody else owned the layers above.
This post is about what that shared foundation actually means. Both halves of it: the part of the argument that genuinely holds, and the part that gets overclaimed in vendor slides.
Proxmox VE Is a Management Layer
Virtualisation on Linux is four distinct layers, built and maintained by four different sets of people.
1. Hardware virtualisation extensions. Intel VT-x with EPT, or AMD-V with NPT. Silicon. This is what makes a guest able to run its own kernel at native speed with its own page tables, without anything emulating instructions.
2. KVM — the hypervisor.
Kernel modules: kvm.ko for the architecture-independent core, plus kvm-intel.ko or kvm-amd.ko for the vendor extensions. This is the component that owns the isolation boundary. It sets up the guest’s virtual machine control structures, handles VM exits, manages the second-level page tables, and delivers interrupts.
3. The VMM — the virtual machine monitor, in user space. On Proxmox VE this is QEMU. It builds the virtual motherboard: chipset, PCIe topology, disks, NICs, serial ports, firmware. KVM runs the CPU; QEMU decides what hardware the guest thinks it has.
4. The management layer.
This is Proxmox VE: pve-manager and pveproxy for the API and interface, qemu-server to turn a VM config file into a QEMU command line, pve-container for LXC, pmxcfs on top of Corosync for the replicated cluster configuration, pve-ha-manager for fencing and restart, plus the firewall and SDN stack.
Those four layers exist on the platform you are migrating from too, doing the same four jobs — vCenter is layer 4, VMkernel is layer 2 — and I will come back to that comparison once the pieces are on the table.
What Proxmox Actually Maintains
Proxmox VE owns layer 4 outright. It would be wrong to say it just packages layers 2 and 3, though, and that is the most common misreading of what the company does.
pve-qemu carries 78 patches against upstream QEMU in its series file at the time of writing:
debian/patches/series. Most of the queue is Proxmox’s own engineering, and the whole of it sits in layer 3 — none of these patches touch kvm.ko.They build their own kernel as well, and maintain packaging or patches for most of the surrounding stack — pve-edk2-firmware for OVMF, plus lxc, zfsonlinux, openvswitch, libiscsi, corosync-pve, lvm and ceph.
So the real boundary is: all of layer 4, substantial engineering inside layer 3, and a kernel they compile. What they have not done is write a hypervisor. Layer 2’s KVM code is upstream Linux.
And none of this is a criticism of Proxmox. As such, it is the reason a company of Proxmox’s size can be trusted with the job at all. A small company in Vienna did not sit down and write a hypervisor from scratch; they built on one that Intel, AMD, Red Hat, Google, Amazon and IBM were already paying engineers to maintain, and spent their own effort on the layer above it and the integration work that layer needs. That is where a team that size can make a difference, and the patch queue shows them making it.
What KVM Actually Is
KVM stands for Kernel-based Virtual Machine, and the name is exact: it is a kernel feature, not a program.
Load the modules and Linux gains a character device, /dev/kvm, plus a small set of ioctl calls on it — KVM_CREATE_VM, KVM_CREATE_VCPU, KVM_SET_USER_MEMORY_REGION, KVM_RUN.
That interface is the entire hypervisor API. Anything that can open a file descriptor and call ioctl can create virtual machines.
It was written by Avi Kivity at Qumranet and merged into Linux 2.6.20, released in February 2007 — nineteen years ago. Red Hat acquired Qumranet in 2008, and KVM has shipped in every kernel release since, on the kernel’s usual nine-to-ten-week cadence. It has been ported well beyond x86: arm64, POWER, s390 on IBM Z, and RISC-V.
The important architectural point is why it is small.
KVM does not have a scheduler, because Linux has one. A vCPU is an ordinary host thread, and the completely fair scheduler puts it on a core like any other thread. It does not have a memory manager, because Linux has one. Guest RAM is a normal user-space mapping, so it can be paged, backed by hugepages, or placed on a NUMA node using the same machinery as any other process. It does not have a driver stack, a block layer, a network stack, or a filesystem, because Linux already had all of them and they are the ones your hardware vendor is testing against.
That is the actual maturity argument, and it is much stronger than a version number. Every NUMA balancing improvement, every io_uring change, every network driver, every new CPU errata workaround that lands in Linux lands underneath your virtual machines, because there is no separate hypervisor kernel for anyone to port it to.
The Type 1 Argument Draws the Line in the Wrong Place
The objection that follows this, reliably, is that KVM is “only a type 2 hypervisor”. It runs on a host OS, unlike ESXi, which runs on bare metal.
That taxonomy is older than hardware virtualisation by three decades, and the thing it was drawing a line around is no longer where anyone thinks it is.
What Actually Happens When a vCPU Runs
QEMU calls ioctl(vcpu_fd, KVM_RUN). Control passes into kvm.ko, which loads the guest CPU state and executes VMLAUNCH. From that instruction until the next VM exit, the guest is executing directly on the physical core, in guest mode, with its own page tables active through EPT, at full hardware speed. There is nowt beneath it interpreting anything. The “host OS” is not in the path. It is not even running on that core.
When the guest does something that needs handling, the CPU exits to the host — and lands in kvm.ko, in the kernel, at exactly the privilege level ESXi’s VMkernel occupies. Most exits are resolved right there and re-entered without user space ever being involved.
ESXi Has the Same Split
Now look at the platform that supposedly proves the distinction.
VMware’s own architecture documentation describes VMkernel as “a POSIX-like operating system” which provides “process creation and control, signals, file system, and process threads”. That is an operating system, by its author’s description. And a running VM on ESXi is not a thing inside the kernel — it is a group of userworld processes: a VMM per virtual CPU, which virtualises the guest’s instructions and manages its memory, and a VMX process per VM, which handles I/O to the devices that are not performance-critical and talks to the snapshot manager and the remote console.
Read that with QEMU in mind. Per-vCPU execution context in the kernel, per-VM user-space process doing device emulation and management. VMware split it for the same reason everyone else did.
Hyper-V is no different. Microsoft’s documentation is explicit that the Virtual Machine Worker Process, vmwp.exe, is “a user mode component of the virtualization stack”, spawned per VM, and that all emulated devices are implemented in it — running in the root partition, which is Windows. Even Xen, the architecture the taxonomy fits best, needs a general-purpose Linux in dom0 to function, and gets its device model for fully virtualised guests from QEMU.
So Where Is the Line?
The criterion was never “does it use user space”. Goldberg’s taxonomy, from the early 1970s, asks whether the hypervisor is an application running on a pre-existing operating system that already owns the hardware and does the scheduling. That is a type 2, hosted hypervisor: VMware Workstation, VirtualBox, Parallels Desktop, plain QEMU with no acceleration. You install a general-purpose OS, then you install a program onto it, and that program asks the OS for memory and CPU time like any other program.
That is not what KVM is. kvm.ko is not a program on top of Linux. It is part of Linux, executing at the same privilege level as the code that owns the hardware, and when a guest exits it lands there directly. There is no host OS underneath the hypervisor. The kernel is the hypervisor. Proxmox VE ships that kernel as the system, exactly the way ESXi ships VMkernel as the system.
And the “it needs user space, so it is type 2” version cannot be rescued, because applied consistently it catches everything. No VM runs on ESXi without its VMX process, none on Hyper-V without vmwp.exe, none on Xen as an HVM guest without QEMU. A test that puts every shipping hypervisor in one bucket is not distinguishing anything.
| Kernel that owns CPU and memory virtualisation | Per-VM user-space device model | Needs a pre-existing host OS? | |
|---|---|---|---|
| VMware ESXi | VMkernel | VMX per VM | No |
| Microsoft Hyper-V | hypervisor plus the Windows root partition | vmwp.exe | No |
| Xen | Xen hypervisor plus dom0 Linux | QEMU, in dom0 or a stub domain | No |
| KVM | kvm.ko | QEMU | No |
| VirtualBox, VMware Workstation | the host’s kernel, via an installed driver | the application itself | Yes |
Four products, one shape — and then a fifth row that is genuinely different. That last row is what “type 2” was coined to describe, and it is the only one where something else was already in charge of the hardware.
So the taxonomy does still draw a line. It just does not draw it anywhere near where the argument assumes: KVM and ESXi are on the same side of it. Calling KVM type 2 borrows a word from the VirtualBox category and applies it to something architecturally in the ESXi category.
Which leaves the label doing no useful work in an evaluation, because both of the products you are choosing between sit in the same box. What differs is not the type number. It is that one vendor also wrote the kernel and will not let you read it — a licensing and transparency distinction wearing an architecture diagram’s clothes. Nutanix shows the point commercially: it ships the same KVM code Proxmox does and describes AHV as a bare-metal type 1 hypervisor. Same code, opposite label, different marketing department.
There is a real concern hiding inside the accusation, and it deserves a better name: a general-purpose kernel is doing a thousand jobs a purpose-built one is not, which is more code and more attack surface beside the isolation boundary. That is legitimate and measurable, and I come back to it near the end.
Where the Work Actually Happens
The one place the “it’s on a host OS” instinct has a real point is exit handling, so it is worth being clear about which exits go where.
| The guest does this | Handled by | Cost |
|---|---|---|
| Touches a page not yet mapped in EPT/NPT | kvm.ko, in kernel | One exit, microseconds |
| Writes to its local APIC | The CPU itself, via APICv/AVIC | Often no exit at all |
| Sends an inter-processor interrupt | Posted interrupts in hardware | Often no exit |
Transmits on a virtio-net queue with vhost-net | Kernel thread, no user-space hop | One doorbell |
| Reads a register on an emulated e1000 or IDE controller | All the way out to QEMU | Exit plus a user-space round trip |
Only the last row looks anything like the type 2 caricature — and it is also the row you engineer away, by using virtio devices and not presenting emulated legacy hardware you do not need. That is the same reasoning behind choosing Q35 over i440fx: fewer trapped register accesses, fewer legacy devices to walk.
The Host OS Is a Feature, Not Baggage
The other half of that ledger never makes it into the argument, so here it is: a Proxmox node is a machine you can actually work on. It is Debian, so the whole Debian archive is one apt install away.
- Monitoring you already run — a Prometheus node exporter,
smartmontools, your existing agent — rather than whatever the appliance chooses to expose. - Diagnostics when something is slow:
fio,iperf3,nvme-cli,perf,bpftrace. - Backup agents from any vendor shipping a Linux binary.
- Configuration management, so the hypervisor sits in the same Ansible inventory as everything else instead of being a special case.
fwupdfor firmware, on hardware whose vendor supports LVFS.
None of that needs a plugin format, a signed bundle, or vendor blessing. Compare it with the ESXi model, where the shell is deliberately restricted, third-party code arrives as a VIB, and there is no package manager to reach for at all.
It Reaches Into the Storage Stack
Monitoring agents are the boring version of this. Proxmox’s own storage documentation lists the native plugins — dir, NFS, CIFS, CephFS, ZFS, BTRFS, LVM, LVM-thin, iSCSI, FC/SAS, RBD, ZFS-over-iSCSI, PBS — and then adds a sentence worth taking literally:
you may use all storage technologies available for Debian Linux
That is a statement about where the boundary is, and the mechanism behind it is generic: get a block device onto every node, put LVM on it, add it as an LVM storage with shared enabled. That is exactly how the supported Fibre Channel and iSCSI paths work, so anything that can produce a shared block device can use the same route.
NVMe over TCP or RDMA is the case worth knowing about, because it is fast, current, and absent from that plugin list. nvme-tcp and nvme-rdma are in-tree Linux host drivers — NVMe/TCP has been in mainline since 5.0 — so there is nothing to compile. nvme-cli is a Debian package (nvme discover, then nvme connect), and nvmetcli configures the in-kernel nvmet target at the other end. A connected namespace appears as /dev/nvmeXnY, and from there it is an ordinary block device.
ATA over Ethernet makes the same point from the opposite end of the spectrum — ancient, obscure, equally unsupported as a plugin. The aoe driver is in mainline; aoetools gives you aoe-discover and aoe-stat; vblade turns any file or block device on another machine into a target. Same route, same outcome.
Neither is a plugin API, an SDK, or a certification programme. It is what happens when the hypervisor’s storage layer is the Linux block layer.
The honest caveats, because this reads as a party trick until it is 3 a.m. Neither transport is a tested Proxmox storage type, so the integration and its failure modes — reconnect behaviour, multipath, timeouts under load — are yours to own and yours to test before anything important lives on it. And AoE is a bare Layer 2 protocol, non-routable and with no authentication, so it belongs on an isolated storage VLAN and nowhere else. NVMe/TCP at least has a discovery model and can be routed, which is much of why it is the one to reach for now.
Who Else Runs KVM
Here is where the “already running it” claim comes from. Every platform below runs the same kernel module.
| Platform | Where you meet it | The KVM part | The user-space VMM |
|---|---|---|---|
| Amazon EC2 (Nitro) | Public cloud | KVM core module | Custom — QEMU removed, device model in the Nitro cards |
| AWS Lambda, Fargate | Serverless | /dev/kvm | Firecracker — a minimal microVM monitor in Rust |
| Google Compute Engine | Public cloud | KVM since launch | Google’s own VMM, deliberately not QEMU |
| Alibaba Cloud ECS | Public cloud | Simplified KVM (X-Dragon) | Custom, with net and storage offloaded to a MoC card |
| Oracle Cloud (OCI) | Public cloud | Oracle Linux KVM — the same stack Oracle ships on-premises | QEMU lineage |
| DigitalOcean, Linode/Akamai, Vultr, Hetzner, OVHcloud, Scaleway, UpCloud | Public cloud | Stock KVM | QEMU |
| Nutanix AHV | On-premises HCI | Stock KVM | QEMU with libvirt and Open vSwitch |
| OpenStack (Nova) | Private cloud | Stock KVM | QEMU via libvirt — the default and best-tested driver |
| Apache CloudStack, OpenNebula, oVirt | On-premises | Stock KVM | QEMU via libvirt |
| OpenShift Virtualisation, SUSE Harvester | Kubernetes | Stock KVM | QEMU inside a pod, via KubeVirt |
| Proxmox VE | On-premises | Stock KVM | QEMU, with LXC alongside for containers |
Everyone Keeps the Kernel Half and Rewrites the User-Space Half
The hyperscalers did not fork KVM. They forked QEMU’s job.
AWS moved EC2 off Xen onto the Nitro hypervisor, which is built on the KVM core kernel module with QEMU thrown out entirely. The device model lives in dedicated Nitro cards instead, which is how they get performance indistinguishable from bare metal. Every current-generation EC2 instance type runs this. Separately, Lambda and Fargate run Firecracker, a purpose-built VMM in Rust that talks to the same /dev/kvm.
Google has run every Compute Engine VM on KVM since Compute Engine launched, and wrote its own user-space VMM rather than using QEMU, explicitly to avoid QEMU’s huge matrix of guests, devices and modes. They also went the other way and hardened the kernel module upstream, removing emulated devices nobody needed and narrowing the set of emulated instructions. That work is in the KVM you are running.
Alibaba did the same shape of thing with X-Dragon: a stripped-down KVM hypervisor with the network and storage virtualisation offloaded onto an FPGA-based MoC card.
Nutanix AHV is KVM plus libvirt plus QEMU plus Open vSwitch plus Nutanix’s orchestration — of every commercial product on that list, the closest relative Proxmox VE has. A different layer 4, and a very different invoice.
Proxmox VE Keeps QEMU On Purpose
It is tempting to read the table as a ranking, with AWS and Google at the top for having replaced QEMU. That is the wrong reading, because their constraint is not yours.
AWS and Google run one hardware profile, at a scale where a single device-emulation bug is a fleet-wide event, and they control every guest image boundary they care about. In that world QEMU’s breadth is nearly all liability, so deleting it is obviously correct.
You are not in that world.
You have an appliance image from 2013 that wants an e1000. You have a Windows VM whose machine type must stay pinned for the rest of its life. You have a GPU to pass through, an emulated SAS controller to satisfy an installer, a UEFI variable store to preserve. QEMU is exactly what lets a general-purpose platform say yes to all of that.
What AWS calls attack surface is what you call a compatibility matrix. Both descriptions are accurate; the difference is whether you get to choose your workloads.
It also explains the shape of that patch queue. AWS and Google solved backup and snapshots outside the VMM, in their own storage services. Proxmox had no storage service to solve it in, so they put it in QEMU — which is why savevm-async and the PBS block driver exist as patches rather than as products.
And that is worth knowing for a practical reason: the parts of Proxmox VE you would miss most are the parts that are not upstream. Your VM configs are plain text and your disk images are standard formats, so a machine will move. But a snapshot including RAM state, and a Proxmox Backup Server incremental chain, depend on Proxmox’s QEMU fork. That is a much lighter dependency than a proprietary hypervisor — the fork is public, AGPL, and you can read every patch in it — but it is not zero, and “no lock-in at the software level” should carry that footnote.
So — Is It Enterprise Grade?
The lazy form of this argument does not work, and it is worth saying so plainly. “AWS uses KVM, therefore Proxmox VE is enterprise grade” is a non-sequitur: it takes a claim about one layer and quietly applies it to an entire product.
Here is the version that does hold.
What is shared is the layer that is hardest to get right and most dangerous to get wrong. CPU and memory virtualisation, and the isolation boundary between tenants, is the part where a bug is a breach rather than an outage. That code is reviewed by engineers paid by Amazon, Google, Red Hat, Intel, AMD, IBM and Alibaba, and its bugs are found by the outfits running the largest fleets in existence — usually before the kernel reaches you. When a guest-escape vulnerability does land, the fix arrives through the normal kernel update you were going to apply anyway. You are not waiting on one vendor’s release cycle for a hypervisor only that vendor can see.
What is not shared is everything above the boundary. Which means the question collapses into two much more answerable ones:
- Is the management layer good enough for the way you operate?
- Can you buy support for it with terms you can live with?
Both can be tested in a proof of concept and written into a contract. Neither requires faith in a hypervisor.
That is a much better position than the one the original question assumes — that you are being asked to trust a novel hypervisor from a small vendor. You are not. The Proxmox-specific code is a management layer mostly in Perl and increasingly in Rust, plus that patch queue against QEMU, and note where the patches land: the device model and the backup path, not the isolation boundary.
It is also worth knowing what the Proxmox-specific layer failing costs you. The QEMU processes are ordinary independent processes on the host, so pveproxy falling over does not stop a single virtual machine. That is a very different blast radius from losing the component that owns the isolation boundary.
What “The Same Hypervisor” Does Not Buy You
This is where vendor confidence documents tend to stop, which tells you something about who they are written for. It is the more useful half.
It does not buy you AWS’s reliability. Nitro’s availability has very little to do with KVM. It comes from the control plane, the network fabric, the storage service, the capacity management and the operational practice around it. Your cluster’s reliability will come from your Corosync quorum design, your fencing configuration, your storage choice and your network redundancy. KVM has no opinion on any of those. Sharing a hypervisor with a hyperscaler does not inherit their operations.
It does not buy you Nitro’s attack surface, and this is where the legitimate half of the type 2 accusation lands.
You are running QEMU on a general-purpose kernel, and historically QEMU’s device model is where the memorable VM escapes lived — VENOM, in an emulated floppy controller nobody was using, being the canonical example. Google could note at the time that Compute Engine was unaffected exactly because it does not run QEMU. You do run it, so the compensating controls are yours: prefer virtio over emulated hardware, do not present devices you do not need, leave the AppArmor profiles alone, and patch QEMU on the same discipline as the kernel. Proxmox’s extra/ patches are them doing exactly that on your behalf, which is a reasonable thing to check they are still doing.
It does not make VM behaviour portable between platforms.
CPU model selection, machine type versioning, live migration compatibility and clock behaviour are all decided in layers 3 and 4, and they differ everywhere. A cluster of mixed CPU generations will still punish you for setting the CPU type to host, and guest clocks still drift whatever the logo on the platform is. Same hypervisor is not same behaviour.
It does not answer the support question — which is the one procurement actually cares about, and rightly so. Proxmox VE is AGPLv3 and free to run in production. The subscription buys the enterprise-tested repository and vendor support, starting at €120 per socket per year. Proxmox’s own support is delivered during Austrian business hours, so round-the-clock coverage comes from partners rather than from Vienna. croit, where I work, is one of those partners, and covers 24/7, 365 days a year. That is a commercial negotiation, not a technical risk — and having it be a commercial negotiation is the point of everything above.
What To Actually Evaluate Instead
If the hypervisor is settled, a proof of concept should spend its time on the layer that is genuinely specific to Proxmox VE:
- Fencing and HA. Pull the power on a node with running HA guests and time the restart. Then do it to two nodes and confirm the remaining ones behave the way you expect when quorum is lost.
- Live migration across CPU generations. With the CPU model you actually intend to standardise on, not
host. - Backup and, more importantly, restore. Restore times under load, not backup times. Nobody has ever been thanked for a fast backup.
- The permissions model. Whether you can hand an application team console and power control over their own VMs and nothing else, granularly enough to satisfy an auditor.
- The API. Everything the web interface does is an API call; if your automation cannot drive it, the platform will not fit how you work.
- The upgrade path. A major-version upgrade on the PoC cluster, before you have 400 VMs on it.
None of those are questions about KVM. That is rather the point.
The hypervisor is the settled part. Spend the proof of concept on the parts that are not.
References
KVM itself
- KVM API documentation — the
/dev/kvmioctl interface, which is the whole hypervisor contract - Linux 2.6.20 release notes — the release KVM was merged into, February 2007
- Some KVM developments — LWN, January 2007, on KVM in the days after it landed in mainline
How the other hypervisors are built
- Interpreting virtual machine monitor and executable failures — Broadcom’s own description of a running ESXi VM as “several processes or userworlds”, with one VMM per vCPU and one VMX per VM
- The Architecture of VMware ESXi (PDF) — the whitepaper calling VMkernel “a POSIX-like operating system” with processes, signals, a file system and threads. Linked via a mirror because the original VMware URL did not survive the Broadcom reorganisation, which is its own small comment on vendor continuity
- Hyper-V architecture — Microsoft on the Virtual Machine Worker Process as a user-mode component, one per VM, where all emulated devices live
- The Nutanix Bible — AHV architecture — AHV described as KVM with libvirt, QEMU and Open vSwitch
Who runs KVM, and how they changed it
- 7 ways we harden our KVM hypervisor at Google Cloud — Google on running KVM without QEMU, and the upstream hardening they did
- The AWS Nitro System — which EC2 instance types run the Nitro hypervisor
- AWS EC2 Virtualization 2017: Introducing Nitro — Brendan Gregg’s write-up of the Xen-to-KVM transition, still the clearest account of it
- Firecracker — AWS’s minimal KVM-based VMM behind Lambda and Fargate
- Alibaba Cloud’s sixth-generation ECS instances — the X-Dragon hypervisor and MoC offload
- KubeVirt — the QEMU/KVM-in-a-pod model behind OpenShift Virtualisation and Harvester
- OpenStack Nova hypervisor support matrix — KVM as the reference driver
What Proxmox maintains
- The Proxmox GitHub organisation — 89 repositories, including
pve-qemu,pve-kernel,pve-edk2-firmwareand packaging forlxc,zfsonlinux,openvswitch,libiscsi,corosync-pveandceph pve-qemupatch series — the 78 patches counted above, and the fastest way to see exactly what Proxmox adds to QEMU- Proxmox VE storage documentation — the native plugin list, which storage types are shared, and the line about using all storage technologies available for Debian Linux
nvme-cliandnvmetcli— NVMe-oF host and target tooling;aoetoolsandvbladefor the ATA over Ethernet equivalent. All in Debian trixie, which is what Proxmox VE 9 is built on- Proxmox VE pricing and subscription tiers — what the subscription covers
Disclosure: I work for croit, a Proxmox Gold Partner. The technical claims above are sourced and checkable; the commercial paragraph is the part where I have an interest, so treat it accordingly.