The Problem With Copied Boot Lines
Search for Proxmox tuning and you will find a single long GRUB_CMDLINE_LINUX line, presented as a unit, with no indication of which flags apply where.
That matters more than it sounds. The hypervisor and the guest are solving opposite problems.
The host wants deterministic access to real hardware: IOMMU behaviour, PCIe link states, physical idle states. The guest wants to stop pretending it has hardware at all — its timers are approximations, its idle states are fiction, and its stalls are usually somebody else’s scheduler. As such, the same flag can be correct on one side, pointless on the other, and occasionally harmful.
Below is where each one actually belongs.
First: Are You Even Editing The Right File?
A Proxmox install on ZFS root boots with systemd-boot, where /etc/default/grub is read by nobody. Editing it and rebooting produces no change and no error. That is a frustrating hour.
proxmox-boot-tool status # tells you which bootloader is in use
# systemd-boot: edit /etc/kernel/cmdline, then
proxmox-boot-tool refresh
# GRUB: edit /etc/default/grub, then
update-grub
Either way, verify rather than assume:
cat /proc/cmdline
And on the GRUB side, use GRUB_CMDLINE_LINUX_DEFAULT, not GRUB_CMDLINE_LINUX. The latter applies to every boot entry including recovery — and recovery is exactly when you want stock behaviour, not mitigations disabled and C-states pinned.
The Host Line
IOMMU and passthrough
iommu=pt amd_iommu=pgtbl_v2 pcie_acs_override=downstream,multifunction
iommu=pt puts the IOMMU in passthrough mode: devices assigned to VMs get translated, host-native devices bypass translation. It is real and it is handled in arch/x86/kernel/pci-dma.c, which calls iommu_set_default_passthrough(true). The kernel documents it as equivalent to iommu.passthrough=1.
amd_iommu=on is not a thing. This is the most widely copied non-existent parameter in Proxmox guides. The kernel’s parse_amd_iommu_options() accepts fullflush, force_enable, off, force_isolation, pgtbl_v1, pgtbl_v2, irtcachedis, nohugepages and v2_pgsizes_only. Anything else lands here:
pr_notice("Unknown option - '%s'\n", str);
AMD-Vi is enabled by default when the firmware advertises it. Check your own log and you will find the parameter was never doing the job it was credited with:
dmesg | grep -i "AMD-Vi\|Unknown option"
amd_iommu=pgtbl_v2 is valid — it selects the v2 DMA page-table format, which shares the CPU page-table structure rather than using AMD’s own. Two things to know: the documentation scopes it to the DMA-API, meaning the host’s own device domains rather than the VFIO domains used for passthrough; and it fails safe with a log line you should check for:
if (amd_iommu_pgtable == PD_MODE_V2) {
if (!amd_iommu_v2_pgtbl_supported()) {
pr_warn("Cannot enable v2 page table for DMA-API. Fallback to v1.\n");
amd_iommu_pgtable = PD_MODE_V1;
}
}
So it is worth measuring on a node with heavy host-side IO, and worth verifying you actually got it.
pcie_acs_override=downstream,multifunction is Proxmox’s out-of-tree patch. It splits IOMMU groups by asserting isolation the hardware does not advertise, which is what makes passthrough possible on consumer boards. It is also, exactly, telling the kernel something untrue about the topology. Fine on a box whose guests you trust as much as the host. Not fine otherwise. There is more on why in the IOMMU tax article.
Latency and jitter
pcie_aspm=off processor.max_cstate=1 amd_pstate=disable
pcie_aspm=off keeps PCIe links out of low-power states so an arriving IO never waits for one to wake. It costs a few watts per link and removes a latency tail that is hard to diagnose. See PCIe ASPM and passthrough.
processor.max_cstate=1 caps ACPI idle at C1. Note the driver: this is the processor/acpi_idle knob, so on Intel you need intel_idle.max_cstate=1 as well, because intel_idle takes precedence. On AMD this is the right one.
There is a real counter-argument. Deep sleep on idle cores is what gives the package thermal and power headroom to boost the busy ones, so pinning everything at C1 can lower your peak single-thread frequency while raising idle draw. On a latency-sensitive host that trade is usually worth it. On a host chasing throughput it might not be. Measure it rather than inheriting it.
amd_pstate=disable falls back to acpi-cpufreq. Worth knowing the documented alternatives before reaching for it: passive (driver requests a performance level), active (the EPP driver, biasing towards performance or efficiency), and guided. active with a performance bias, or passive plus the performance governor, often gets the same latency while keeping CPPC’s finer control. And if you do disable it, set a governor deliberately — landing on acpi-cpufreq with schedutil may be a step back.
Memory
default_hugepagesz=1G hugepages=64
default_hugepagesz=1G on its own reserves nothing. The kernel documents it as setting “the size of the default HugeTLB page… the default hugetlb size used for shmget(), mmap() and mounting hugetlbfs” — a unit, not an allocation. The allocation comes from hugepages=, documented as “Number of HugeTLB pages to allocate at boot”.
That matters far more for 1 GiB than 2 MiB pages, because contiguous 1 GiB regions are effectively unobtainable once the host has been up and fragmented memory. Boot is your only reliable chance.
Then the guest has to opt in (hugepages: 1024 in the VM config). Reserved pages nothing uses are just memory you cannot have back, and you lose ballooning and KSM on the VMs that use them.
The security trade
mitigations=off
This is not one switch. The kernel expands it to a list, and on a hypervisor these are the entries that matter:
l1tf=off mds=off mmio_stale_data=off kvm.nx_huge_pages=off
gather_data_sampling=off retbleed=off spec_rstack_overflow=off
nospectre_v2 nopti indirect_target_selection=off
The kernel’s own summary is “improves system performance, but it may also expose users to several CPU vulnerabilities”. L1TF, MDS and MMIO stale data are specifically guest-to-host and guest-to-guest leak paths, and kvm.nx_huge_pages is the iTLB-multihit mitigation inside KVM itself.
Defensible on a single-tenant box where every guest is as trusted as the host. Not defensible where guests are untrusted or belong to different tenants. And note it stacks with pcie_acs_override: two independent isolation guarantees removed on the same line. Worth doing on purpose rather than by inheritance.
PCIe Over USB4 Changes Two Of These
If your PCIe devices arrive over USB4 or Thunderbolt — an external GPU or NVMe enclosure — two of the answers above change.
MPS tuning stops being free
On fixed slots, pci=pcie_bus_perf is a small free win. The kernel describes it as:
Set device MPS to the largest allowable MPS based on its parent bus. Also set MRRS (Max Read Request Size) to the largest supported value… for best performance.
The catch is that it configures bridges at boot, from the topology present at boot. Over USB4 hot-plug is the normal case, and a device added later may support a smaller MPS than the bridge was already set to.
The kernel says the quiet part out loud while advertising a different policy:
pcie_bus_peer2peer— Set every device’s MPS to 128B, which every device is guaranteed to support… This also guarantees that hot-added devices will work.
Only one policy carries that guarantee, and it is the one that pins everything to 128 bytes — exactly what MaxPayloadSize tuning sets out to escape. For a hot-plug topology, pcie_bus_safe (the largest value supported by all devices below the root complex) or simply leaving tune_off is the safer starting point. The tunnel’s own switch caps the achievable MPS anyway, so the ceiling was never yours to raise.
The security stack-up gets serious
External PCIe means someone can plug a DMA-capable device into your hypervisor. The kernel’s Thunderbolt documentation is direct about it:
…the connected devices can be DMA masters and thus read contents of the host memory without CPU and OS knowing about it. There are ways to prevent this by setting up an IOMMU but it is not always available for various reasons.
The IOMMU is the defence. Now count what the host line does to it. iommu=pt gives host-owned devices untranslated identity domains, pcie_acs_override asserts isolation that is not there, and mitigations=off disables the guest-isolation mitigations. Each is defensible alone. Together, on a machine with a physically reachable USB4 port, they stack up.
Check where you stand:
cat /sys/bus/thunderbolt/devices/domain*/security # none | user | secure | dponly | usbonly
none alongside that boot line is an open door. If the ports are reachable by people you would not give root to, iommu=pt is the first thing I would reconsider.
The Guest Line
These are the ones that belong inside the VM, and three of them mean something different here than they would on the host.
nmi_watchdog=0 softlockup_panic=0 cpuidle.off=1
cpuidle.off=1 disables the cpuidle subsystem. In a guest that is close to free: the guest’s idle states are emulation, and there is no physical core to put to sleep, so all the framework buys you is wake-up latency. On the host the same flag is a real power and boost-headroom trade, and it overlaps processor.max_cstate=1. Guest side only.
softlockup_panic=0 stops a soft lockup panicking the guest. This is genuinely protective in a VM, because a soft lockup there is frequently not the guest’s fault. A descheduled vCPU looks exactly like a task that refused to yield. That is the same mechanism behind why a VM’s clock cannot be trusted. Check whether you need it, though. It is 0 by default on most builds.
sysctl kernel.softlockup_panic
nmi_watchdog=0 is the interesting one, and it deserves more than a rule.
The Watchdog Question
There are two detectors sharing one threshold:
watchdog_thresh=— Set the hard lockup detector stall duration threshold in seconds. The soft lockup detector threshold is set to twice the value. A value of 0 disables both. Default is 10 seconds.
Hard lockup (nmi_watchdog) fires when a CPU stops taking timer interrupts altogether. Soft lockup fires when a task hogs a CPU for twice as long without scheduling.
In a guest, disabling the hard lockup detector is right. A descheduled vCPU can trip it through no fault of its own, and the detector’s perf-counter work causes VM exits for a signal that was nowt but noise.
On the host it is a judgement call, and it depends on what noise you are actually chasing. Over-provisioning starves guests, not the host kernel — the host’s physical CPUs keep taking interrupts however packed the VMs are. So the log noise a busy hypervisor throws is overwhelmingly soft lockup and RCU stall messages, not NMI hard-lockup reports. If that is the noise, nmi_watchdog=0 will not silence it, and softlockup_panic=0 will not either — that stops the panic, not the messages.
The targeted knobs are:
watchdog_thresh=30 # hard 30s, soft 60s — scale to taste
nowatchdog # honest single flag: disables both detectors
sysctl -w kernel.soft_watchdog=0 # runtime, keeps hard-lockup detection
There is a separate and better reason to disable the hard lockup detector on a busy host, which has nothing to do with noise: it consumes a hardware performance counter per CPU. That is why the parameter accepts rNNN to configure a raw perf event. If you are doing PMU-based profiling, or running a deliberately over-provisioned box where you have already accepted latency variance as the price of density, giving that counter back is a reasonable trade — and small stalls you have consciously signed up for are not incidents.
Just make the choice for that reason rather than for the noise reason, because only one of them is true.
consoleblank=0 does nothing. The kernel documents the console blank timeout as “A value of 0 disables the blank timer. Defaults to 0.” It is already off. Harmless, but it is the second parameter in common circulation that has no effect, and carrying it makes a line look considered when it is copied.
Quick Reference: Where Each Flag Belongs
| Flag | Host | Guest | Notes |
|---|---|---|---|
iommu=pt | yes | no | Host owns the IOMMU. Only applies in a guest if you are running nested passthrough with a vIOMMU |
amd_iommu=pgtbl_v2 | yes | no | Host DMA-API domains. Verify you got v2 rather than the v1 fallback |
amd_iommu=on | — | — | Not a valid option. Kernel logs “Unknown option - ‘on’” |
pcie_acs_override=… | yes | no | Proxmox patch, host topology only. Weakens isolation by design |
pcie_aspm=off | yes | no | There are no real PCIe links in a guest; the host owns the physical link |
pci=pcie_bus_perf | yes | not reliably | Host sets MPS on the wire. Use pcie_bus_safe instead if devices arrive over USB4 |
processor.max_cstate=1 | yes | no | Real idle states are the host’s. Add intel_idle.max_cstate=1 on Intel |
amd_pstate=disable | yes | no | Guests do not control CPU frequency |
cpuidle.off=1 | with care | yes | Free in a guest. On the host it costs boost headroom and overlaps max_cstate |
nmi_watchdog=0 | judgement | yes | Right in a guest. On the host, do it for the PMU counter, not for the noise |
softlockup_panic=0 | no | yes | Guest stalls are often the host’s scheduler. Usually already the default |
consoleblank=0 | — | — | No-op. Kernel default is already 0 |
mitigations=off | both | both | Valid on either, different risk calculus: guest-to-host leaks on the host, process isolation in the guest |
default_hugepagesz + hugepages= | both | both | Host: backing VM memory. Guest: a workload inside the VM that wants them. Different purposes, same flags |
watchdog_thresh= / nowatchdog | both | both | Host: quieten stalls you accepted. Guest: the detector was never trustworthy |
Three genuinely belong on both sides, and it is worth being exact that “both” does not mean “for the same reason”:
mitigations=offon the host is about guest-to-host and guest-to-guest leakage. Inside a guest it is about process isolation within that VM. You can reasonably disable it in one place and not the other.- Hugepages on the host back guest RAM; in a guest they back an application. Reserving them twice for the same memory is waste, so decide which layer wants them.
- Watchdog tuning is a noise decision on the host and a correctness decision in the guest.
Everything else is one side or the other, and two of them are not choices at all.
The Two Lines
Host, on an AMD box doing passthrough, single-tenant:
GRUB_CMDLINE_LINUX_DEFAULT="quiet amd_iommu=pgtbl_v2 iommu=pt pcie_acs_override=downstream,multifunction pcie_aspm=off pci=pcie_bus_perf default_hugepagesz=1G hugepages=64 processor.max_cstate=1 amd_pstate=disable mitigations=off"
Swap pci=pcie_bus_perf for pcie_bus_safe if anything arrives over USB4. Drop mitigations=off if the guests are not all yours. On Intel, intel_iommu=on replaces the AMD flags and intel_idle.max_cstate=1 joins the C-state cap.
Linux guest:
GRUB_CMDLINE_LINUX_DEFAULT="quiet nmi_watchdog=0 softlockup_panic=0 cpuidle.off=1"
And leave kvm-clock alone in the guest — do not force tsc or hpet. The paravirtual clock exists precisely because the counters are not yours. That is the whole argument in the VM timekeeping article.
What To Delete
If you inherited a line from a forum post, these two are the first things to remove, because they cost you nowt and prove the line was never tested:
amd_iommu=on— not a valid option; the kernel logs “Unknown option - ‘on’” and carries onconsoleblank=0— already the default
And check the remainder against /proc/cmdline after a reboot. Every flag on that line should be one you can name a reason for.
If you cannot say what a flag does, it is not tuning. It is superstition. And it will be copied into the next build by someone who trusts you.
References
- Linux kernel — the kernel’s command-line parameters —
mitigations=,default_hugepagesz=,hugepages=,watchdog_thresh=,nowatchdog,consoleblank=,processor.max_cstate=,amd_pstate=,amd_iommu=and thepci=pcie_bus_*policies - Linux kernel source —
drivers/iommu/amd/init.c—parse_amd_iommu_options(), and the v2 page-table capability check that falls back to v1 - Linux kernel source —
arch/x86/kernel/pci-dma.c—iommu=ptcallingiommu_set_default_passthrough() - Linux kernel — Thunderbolt — security levels, and connected devices as DMA masters
- Proxmox VE — System Administration —
proxmox-boot-tool status, editing the kernel command line for systemd-boot versus GRUB