The Problem With Copied Boot Lines

Search for Proxmox tuning and you will find a single long GRUB_CMDLINE_LINUX line, presented as a unit, with no indication of which flags apply where.

That matters more than it sounds. The hypervisor and the guest are solving opposite problems.

The host wants deterministic access to real hardware: IOMMU behaviour, PCIe link states, physical idle states. The guest wants to stop pretending it has hardware at all — its timers are approximations, its idle states are fiction, and its stalls are usually somebody else’s scheduler. As such, the same flag can be correct on one side, pointless on the other, and occasionally harmful.

Below is where each one actually belongs.

Where each kernel boot flag belongsHost onlyreal hardware lives hereiommu=ptamd_iommu=pgtbl_v2pcie_acs_override=…pcie_aspm=offpci=pcie_bus_perf→ pcie_bus_safe if USB4processor.max_cstate=1+ intel_idle.max_cstate on Intelamd_pstate=disableset a governor tooThe guest has no PCIe links,no C-states and no cpufreq.Guest onlystop pretending it is hardwarecpuidle.off=1guest idle states are fictionnmi_watchdog=0a descheduled vCPU trips itsoftlockup_panic=0the stall was the host's faultLeave kvm-clock alone.Do not force tsc or hpet —the counters are not yours.On the host these three allmean something different,and two of them cost yousomething real.Both — different reasonssame flag, separate decisionmitigations=offhost: guest-to-host leaksguest: process isolationdefault_hugepagesz + hugepageshost: backs guest RAMguest: backs an applicationpick one layer, not bothwatchdog_thresh / nowatchdoghost: noise you acceptedguest: never trustworthy"Both" does not mean youshould set it in both places.It means the decision has tobe made twice.Neither — these do nothingamd_iommu=onnot a valid option; the kernel logs "Unknown option - 'on'" and carries onconsoleblank=0already the kernel defaultIf a copied boot line contains either of these, it was never verified against /proc/cmdline —which is the cheapest check available, and the one that catches the wrong-bootloader mistakeas well.
The flags in circulation, sorted. Two of the popular ones do nothing on either side, and the watchdog row is a genuine judgement call rather than a rule.

First: Are You Even Editing The Right File?

A Proxmox install on ZFS root boots with systemd-boot, where /etc/default/grub is read by nobody. Editing it and rebooting produces no change and no error. That is a frustrating hour.

proxmox-boot-tool status          # tells you which bootloader is in use
# systemd-boot: edit /etc/kernel/cmdline, then
proxmox-boot-tool refresh

# GRUB: edit /etc/default/grub, then
update-grub

Either way, verify rather than assume:

cat /proc/cmdline

And on the GRUB side, use GRUB_CMDLINE_LINUX_DEFAULT, not GRUB_CMDLINE_LINUX. The latter applies to every boot entry including recovery — and recovery is exactly when you want stock behaviour, not mitigations disabled and C-states pinned.

The Host Line

IOMMU and passthrough

iommu=pt amd_iommu=pgtbl_v2 pcie_acs_override=downstream,multifunction

iommu=pt puts the IOMMU in passthrough mode: devices assigned to VMs get translated, host-native devices bypass translation. It is real and it is handled in arch/x86/kernel/pci-dma.c, which calls iommu_set_default_passthrough(true). The kernel documents it as equivalent to iommu.passthrough=1.

amd_iommu=on is not a thing. This is the most widely copied non-existent parameter in Proxmox guides. The kernel’s parse_amd_iommu_options() accepts fullflush, force_enable, off, force_isolation, pgtbl_v1, pgtbl_v2, irtcachedis, nohugepages and v2_pgsizes_only. Anything else lands here:

pr_notice("Unknown option - '%s'\n", str);

AMD-Vi is enabled by default when the firmware advertises it. Check your own log and you will find the parameter was never doing the job it was credited with:

dmesg | grep -i "AMD-Vi\|Unknown option"

amd_iommu=pgtbl_v2 is valid — it selects the v2 DMA page-table format, which shares the CPU page-table structure rather than using AMD’s own. Two things to know: the documentation scopes it to the DMA-API, meaning the host’s own device domains rather than the VFIO domains used for passthrough; and it fails safe with a log line you should check for:

if (amd_iommu_pgtable == PD_MODE_V2) {
    if (!amd_iommu_v2_pgtbl_supported()) {
        pr_warn("Cannot enable v2 page table for DMA-API. Fallback to v1.\n");
        amd_iommu_pgtable = PD_MODE_V1;
    }
}

So it is worth measuring on a node with heavy host-side IO, and worth verifying you actually got it.

pcie_acs_override=downstream,multifunction is Proxmox’s out-of-tree patch. It splits IOMMU groups by asserting isolation the hardware does not advertise, which is what makes passthrough possible on consumer boards. It is also, exactly, telling the kernel something untrue about the topology. Fine on a box whose guests you trust as much as the host. Not fine otherwise. There is more on why in the IOMMU tax article.

Latency and jitter

pcie_aspm=off processor.max_cstate=1 amd_pstate=disable

pcie_aspm=off keeps PCIe links out of low-power states so an arriving IO never waits for one to wake. It costs a few watts per link and removes a latency tail that is hard to diagnose. See PCIe ASPM and passthrough.

processor.max_cstate=1 caps ACPI idle at C1. Note the driver: this is the processor/acpi_idle knob, so on Intel you need intel_idle.max_cstate=1 as well, because intel_idle takes precedence. On AMD this is the right one.

There is a real counter-argument. Deep sleep on idle cores is what gives the package thermal and power headroom to boost the busy ones, so pinning everything at C1 can lower your peak single-thread frequency while raising idle draw. On a latency-sensitive host that trade is usually worth it. On a host chasing throughput it might not be. Measure it rather than inheriting it.

amd_pstate=disable falls back to acpi-cpufreq. Worth knowing the documented alternatives before reaching for it: passive (driver requests a performance level), active (the EPP driver, biasing towards performance or efficiency), and guided. active with a performance bias, or passive plus the performance governor, often gets the same latency while keeping CPPC’s finer control. And if you do disable it, set a governor deliberately — landing on acpi-cpufreq with schedutil may be a step back.

Memory

default_hugepagesz=1G hugepages=64

default_hugepagesz=1G on its own reserves nothing. The kernel documents it as setting “the size of the default HugeTLB page… the default hugetlb size used for shmget(), mmap() and mounting hugetlbfs” — a unit, not an allocation. The allocation comes from hugepages=, documented as “Number of HugeTLB pages to allocate at boot”.

That matters far more for 1 GiB than 2 MiB pages, because contiguous 1 GiB regions are effectively unobtainable once the host has been up and fragmented memory. Boot is your only reliable chance.

Then the guest has to opt in (hugepages: 1024 in the VM config). Reserved pages nothing uses are just memory you cannot have back, and you lose ballooning and KSM on the VMs that use them.

The security trade

mitigations=off

This is not one switch. The kernel expands it to a list, and on a hypervisor these are the entries that matter:

l1tf=off   mds=off   mmio_stale_data=off   kvm.nx_huge_pages=off
gather_data_sampling=off   retbleed=off   spec_rstack_overflow=off
nospectre_v2   nopti   indirect_target_selection=off

The kernel’s own summary is “improves system performance, but it may also expose users to several CPU vulnerabilities”. L1TF, MDS and MMIO stale data are specifically guest-to-host and guest-to-guest leak paths, and kvm.nx_huge_pages is the iTLB-multihit mitigation inside KVM itself.

Defensible on a single-tenant box where every guest is as trusted as the host. Not defensible where guests are untrusted or belong to different tenants. And note it stacks with pcie_acs_override: two independent isolation guarantees removed on the same line. Worth doing on purpose rather than by inheritance.

PCIe Over USB4 Changes Two Of These

If your PCIe devices arrive over USB4 or Thunderbolt — an external GPU or NVMe enclosure — two of the answers above change.

MPS tuning stops being free

On fixed slots, pci=pcie_bus_perf is a small free win. The kernel describes it as:

Set device MPS to the largest allowable MPS based on its parent bus. Also set MRRS (Max Read Request Size) to the largest supported value… for best performance.

The catch is that it configures bridges at boot, from the topology present at boot. Over USB4 hot-plug is the normal case, and a device added later may support a smaller MPS than the bridge was already set to.

The kernel says the quiet part out loud while advertising a different policy:

pcie_bus_peer2peer — Set every device’s MPS to 128B, which every device is guaranteed to support… This also guarantees that hot-added devices will work.

Only one policy carries that guarantee, and it is the one that pins everything to 128 bytes — exactly what MaxPayloadSize tuning sets out to escape. For a hot-plug topology, pcie_bus_safe (the largest value supported by all devices below the root complex) or simply leaving tune_off is the safer starting point. The tunnel’s own switch caps the achievable MPS anyway, so the ceiling was never yours to raise.

Why hot-plug changes the MPS policy answerAt boot, pcie_bus_perf configures the bridge from what it can seebridge — MPS 512device A · 512device B · 512present at boothot-added · 256 onlymismatchnothing renegotiatedIn a chassis that never happens — the topology at boot is the topology forever. On a USB4 port itis the normal case.The four policies, and which one the kernel says is hot-plug safepcie_bus_tune_offleave the BIOS values alonepcie_bus_safelargest value all devices below the root complex supportpcie_bus_perflargest the parent bus allows, per device — plus MRRSpcie_bus_peer2peer128 B everywhere —"guarantees that hot-added devices will work"Only one policy carries that guarantee, and it is the one that throws away the payload size youwere tuning for.
The bridge is configured once, at boot, from the devices present then. Everything after that has to live with the decision — which is fine in a chassis and not fine on a port.

The security stack-up gets serious

External PCIe means someone can plug a DMA-capable device into your hypervisor. The kernel’s Thunderbolt documentation is direct about it:

…the connected devices can be DMA masters and thus read contents of the host memory without CPU and OS knowing about it. There are ways to prevent this by setting up an IOMMU but it is not always available for various reasons.

The IOMMU is the defence. Now count what the host line does to it. iommu=pt gives host-owned devices untranslated identity domains, pcie_acs_override asserts isolation that is not there, and mitigations=off disables the guest-isolation mitigations. Each is defensible alone. Together, on a machine with a physically reachable USB4 port, they stack up.

Check where you stand:

cat /sys/bus/thunderbolt/devices/domain*/security   # none | user | secure | dponly | usbonly

none alongside that boot line is an open door. If the ports are reachable by people you would not give root to, iommu=pt is the first thing I would reconsider.

The Guest Line

These are the ones that belong inside the VM, and three of them mean something different here than they would on the host.

nmi_watchdog=0 softlockup_panic=0 cpuidle.off=1

cpuidle.off=1 disables the cpuidle subsystem. In a guest that is close to free: the guest’s idle states are emulation, and there is no physical core to put to sleep, so all the framework buys you is wake-up latency. On the host the same flag is a real power and boost-headroom trade, and it overlaps processor.max_cstate=1. Guest side only.

softlockup_panic=0 stops a soft lockup panicking the guest. This is genuinely protective in a VM, because a soft lockup there is frequently not the guest’s fault. A descheduled vCPU looks exactly like a task that refused to yield. That is the same mechanism behind why a VM’s clock cannot be trusted. Check whether you need it, though. It is 0 by default on most builds.

sysctl kernel.softlockup_panic

nmi_watchdog=0 is the interesting one, and it deserves more than a rule.

The Watchdog Question

There are two detectors sharing one threshold:

watchdog_thresh= — Set the hard lockup detector stall duration threshold in seconds. The soft lockup detector threshold is set to twice the value. A value of 0 disables both. Default is 10 seconds.

Hard lockup (nmi_watchdog) fires when a CPU stops taking timer interrupts altogether. Soft lockup fires when a task hogs a CPU for twice as long without scheduling.

In a guest, disabling the hard lockup detector is right. A descheduled vCPU can trip it through no fault of its own, and the detector’s perf-counter work causes VM exits for a signal that was nowt but noise.

On the host it is a judgement call, and it depends on what noise you are actually chasing. Over-provisioning starves guests, not the host kernel — the host’s physical CPUs keep taking interrupts however packed the VMs are. So the log noise a busy hypervisor throws is overwhelmingly soft lockup and RCU stall messages, not NMI hard-lockup reports. If that is the noise, nmi_watchdog=0 will not silence it, and softlockup_panic=0 will not either — that stops the panic, not the messages.

The targeted knobs are:

watchdog_thresh=30        # hard 30s, soft 60s — scale to taste
nowatchdog                # honest single flag: disables both detectors
sysctl -w kernel.soft_watchdog=0   # runtime, keeps hard-lockup detection

There is a separate and better reason to disable the hard lockup detector on a busy host, which has nothing to do with noise: it consumes a hardware performance counter per CPU. That is why the parameter accepts rNNN to configure a raw perf event. If you are doing PMU-based profiling, or running a deliberately over-provisioned box where you have already accepted latency variance as the price of density, giving that counter back is a reasonable trade — and small stalls you have consciously signed up for are not incidents.

Just make the choice for that reason rather than for the noise reason, because only one of them is true.

consoleblank=0 does nothing. The kernel documents the console blank timeout as “A value of 0 disables the blank timer. Defaults to 0.” It is already off. Harmless, but it is the second parameter in common circulation that has no effect, and carrying it makes a line look considered when it is copied.

Quick Reference: Where Each Flag Belongs

FlagHostGuestNotes
iommu=ptyesnoHost owns the IOMMU. Only applies in a guest if you are running nested passthrough with a vIOMMU
amd_iommu=pgtbl_v2yesnoHost DMA-API domains. Verify you got v2 rather than the v1 fallback
amd_iommu=on——Not a valid option. Kernel logs “Unknown option - ‘on’”
pcie_acs_override=…yesnoProxmox patch, host topology only. Weakens isolation by design
pcie_aspm=offyesnoThere are no real PCIe links in a guest; the host owns the physical link
pci=pcie_bus_perfyesnot reliablyHost sets MPS on the wire. Use pcie_bus_safe instead if devices arrive over USB4
processor.max_cstate=1yesnoReal idle states are the host’s. Add intel_idle.max_cstate=1 on Intel
amd_pstate=disableyesnoGuests do not control CPU frequency
cpuidle.off=1with careyesFree in a guest. On the host it costs boost headroom and overlaps max_cstate
nmi_watchdog=0judgementyesRight in a guest. On the host, do it for the PMU counter, not for the noise
softlockup_panic=0noyesGuest stalls are often the host’s scheduler. Usually already the default
consoleblank=0——No-op. Kernel default is already 0
mitigations=offbothbothValid on either, different risk calculus: guest-to-host leaks on the host, process isolation in the guest
default_hugepagesz + hugepages=bothbothHost: backing VM memory. Guest: a workload inside the VM that wants them. Different purposes, same flags
watchdog_thresh= / nowatchdogbothbothHost: quieten stalls you accepted. Guest: the detector was never trustworthy

Three genuinely belong on both sides, and it is worth being exact that “both” does not mean “for the same reason”:

  • mitigations=off on the host is about guest-to-host and guest-to-guest leakage. Inside a guest it is about process isolation within that VM. You can reasonably disable it in one place and not the other.
  • Hugepages on the host back guest RAM; in a guest they back an application. Reserving them twice for the same memory is waste, so decide which layer wants them.
  • Watchdog tuning is a noise decision on the host and a correctness decision in the guest.

Everything else is one side or the other, and two of them are not choices at all.

The Two Lines

Host, on an AMD box doing passthrough, single-tenant:

GRUB_CMDLINE_LINUX_DEFAULT="quiet amd_iommu=pgtbl_v2 iommu=pt pcie_acs_override=downstream,multifunction pcie_aspm=off pci=pcie_bus_perf default_hugepagesz=1G hugepages=64 processor.max_cstate=1 amd_pstate=disable mitigations=off"

Swap pci=pcie_bus_perf for pcie_bus_safe if anything arrives over USB4. Drop mitigations=off if the guests are not all yours. On Intel, intel_iommu=on replaces the AMD flags and intel_idle.max_cstate=1 joins the C-state cap.

Linux guest:

GRUB_CMDLINE_LINUX_DEFAULT="quiet nmi_watchdog=0 softlockup_panic=0 cpuidle.off=1"

And leave kvm-clock alone in the guest — do not force tsc or hpet. The paravirtual clock exists precisely because the counters are not yours. That is the whole argument in the VM timekeeping article.

What To Delete

If you inherited a line from a forum post, these two are the first things to remove, because they cost you nowt and prove the line was never tested:

  • amd_iommu=on — not a valid option; the kernel logs “Unknown option - ‘on’” and carries on
  • consoleblank=0 — already the default

And check the remainder against /proc/cmdline after a reboot. Every flag on that line should be one you can name a reason for.

If you cannot say what a flag does, it is not tuning. It is superstition. And it will be copied into the next build by someone who trusts you.

References