What MPS Is

PCIe devices transfer data in packets called Transaction Layer Packets (TLPs). Each TLP has a header and a payload. The maximum size of that payload is the MaxPayloadSize (MPS).

MPS is negotiated between a device and its upstream bridge during link training. The negotiated value is the smaller of what the device supports and what the bridge allows. Every bridge and switch in the path between the device and the root complex has its own MPS capability. The final MPS for any device is set by the narrowest point in the chain.

Common MPS values are 128, 256, and 512 bytes. Some devices support 1024 or even 4096 bytes, but in practice 256 or 512 is typical for NVMe controllers and network cards. GPUs often support 256 bytes.

Why It Matters

Larger MPS means fewer packets for the same amount of data. A 4KB IO transferred at MPS 128 requires 32 TLPs. The same transfer at MPS 512 requires 8 TLPs.

Each TLP carries protocol overhead — the header, CRC, framing. Fewer TLPs means less protocol overhead per byte transferred. At high throughput, that difference is measurable. Not dramatic. But real. And it is the same overhead, paid on every packet.

MPS also affects how well the PCIe link is used. Smaller payloads mean the link spends more time on headers relative to data. Larger payloads shift the ratio towards useful data.

The same 4 KB of payload costs 32 packet headers at MPS 128 and 8 at MPS 512MPS 128 — a 4 KB transfer becomes 32 TLPs32 headers × ≈24 B — ≈16% of the bytes on the wire is protocol overheadMPS 512 — the same 4 KB becomes 8 TLPs8 headers × ≈24 B — ≈4% of the bytes on the wire is protocol overheadheader, sequence number, CRC and framingpayloadBoth strips carry the same 4 KB. Larger payloads do not move more data — theyspend less of the link describing it. The measured throughput gain is small.
Both strips carry the same 4 KB. The solid bars are the per-packet headers — at MPS 128 there are 32 of them, at MPS 512 only 8.

The impact on latency is smaller than on throughput. A single 4KB read at MPS 128 vs 512 will not show a latency difference worth measuring, because the TLPs are pipelined. But push high IOPS with many transfers in flight and the smaller overhead from larger MPS adds up.

The Problem Under Passthrough

On bare metal, the BIOS sets MPS during POST based on the PCIe topology. A modern server BIOS typically sets MPS to the maximum the topology supports, usually 256 or 512 bytes.

Under QEMU, the virtual Q35 chipset’s root complex has its own MPS capability. By default, it presents a low MPS. The Linux kernel’s default MPS policy (pcie_bus_default) sets each device’s MPS to match its parent bridge, which in a virtual topology means the QEMU root complex’s default. Often 128 bytes.

As such, a device that can do 512-byte payloads runs at 128 because the virtual root complex set the ceiling.

MPS is set by the narrowest point in the path, which under QEMU is the virtual root complexBare metal — the BIOS sets MPS from the real topologyNVMesupports 512PCIe switchallows 512Root complexallows 512MPS = 512 Bthe smallest in the pathPassed through to a VM — the virtual root complex is now the narrowest pointNVMesupports 512QEMU Q35 virtual root complexpresents 128 by defaultMPS = 128 Bdevice capability unusedHost kernel with pci=pcie_bus_perfNVMesupports 512each device set to its parent bus maximumMRRS raised to matchMPS = 512 Bset on the host, not the guest
MPS is the smallest value in the path. Passing a device through inserts the virtual root complex into that path, and its conservative default becomes everyone’s ceiling.

How to Fix It — Host Side

Tell the kernel to set MPS to the maximum each device’s parent bus supports:

# Add to GRUB_CMDLINE_LINUX_DEFAULT in /etc/default/grub
pci=pcie_bus_perf

This sets each device’s MPS to the largest value its parent bus allows. It also sets MRRS (Max Read Request Size) accordingly. The kernel source says it makes sure a device’s MPS is no larger than its parent’s. That keeps the chain consistent while getting the best transfer sizes.

After reboot, verify the new MPS:

# Check MPS on a specific device
lspci -vv -s XX:00.0 | grep -i "MaxPayload"

You should see MaxPayload 256 bytes or MaxPayload 512 bytes instead of the default 128.

The Other Kernel Options

The kernel offers four MPS policies, each set via the pci= boot parameter:

pcie_bus_tune_off — don’t touch MPS at all. Use whatever the BIOS set. On bare metal with a good BIOS, this is often fine. Under QEMU, the BIOS is OVMF or SeaBIOS, which may not optimise MPS.

pcie_bus_default — the kernel default. Sets each device’s MPS to match its upstream bridge. Conservative and safe, but doesn’t maximise performance.

pcie_bus_safe — sets MPS to the largest value supported by all devices in the system. Useful for closed systems where you know all devices and nowt will be hotplugged. Slightly more aggressive than default.

pcie_bus_perf — sets MPS per-device to the largest value the parent bus allows. Each device gets the best MPS its local topology supports. This is the right choice for passthrough because it optimises each path independently.

pcie_bus_peer2peer — sets MPS to 128 bytes on everything. Every device talks at the lowest common size. Used when devices need to DMA directly to each other (GPU-to-GPU, GPU-to-NIC via RDMA). Not useful for standard passthrough.

pcie_bus_perf is the right one for passthrough.

How to Fix It — Guest Side

You can set pci=pcie_bus_perf in the guest kernel’s boot configuration as well. Whether it has any practical effect depends on how QEMU presents the virtual PCIe topology. The virtual root complex caps what the guest can negotiate.

In testing, the host-side fix is the one that sticks. The host owns the physical device, and its MPS setting sets the actual TLP size on the wire. The guest’s setting only touches the virtual topology inside the VM, and whether that changes real behaviour depends on how QEMU presents the PCIe path for that device.

Set it on the host. Setting it in the guest as well won’t hurt, but don’t rely on it alone.

What MRRS Is

Max Read Request Size (MRRS) is related but separate. MPS limits how much data a device can send in one TLP. MRRS limits how much data a device can request in one read request.

A device with MRRS 4096 can issue a single 4KB read request. The response comes back in multiple TLPs, each up to the MPS in size. Higher MRRS means the device can request more data per transaction, reducing the number of read request TLPs on the bus.

MRRS sizes the request, MPS sizes each packet of the replyNVMe controllerMRRS 4096 Bhost memoryvia the root complex1 × read request — “send me 4 KB”MRRS caps how much one request may ask for8 × completion TLP — 512 B eachMPS caps how large each packet of the reply may beOne request, many packets. pci=pcie_bus_perf raises both, so they do not need tuning separately.
The two are easy to confuse: MRRS limits how much a device may ask for in one request, MPS limits how large each packet of the reply may be.

pci=pcie_bus_perf sets both MPS and MRRS to their optimal values. You don’t need to tune them separately.

Impact vs Other Tuning

The MPS difference between 128 and 512 bytes has a smaller performance impact than NUMA alignment or ASPM. It’s typically a low single-digit percentage improvement on throughput. You won’t see it in latency benchmarks at low queue depths.

But it’s a free optimisation. One kernel parameter, no downside, no compatibility risk. There’s no reason not to set it on any system doing passthrough.

It costs nowt, and you have already paid for the hardware. You may as well have what you bought.

References