Skip to content
Three storage tiers: Optane absorbing writes, NVMe holding Ceph metadata, and HDDs providing bulk capacity

Making HDD-Backed Proxmox Ceph Clusters Fast — NVMe Metadata, One Optane Per Spindle, and What It Costs You

Flash pricing has made all-NVMe hard to justify, and enterprise HDDs are worth another look. Getting acceptable latency out of them on hyper-converged Proxmox Ceph means moving RocksDB off the spindle, pairing each disk with its own Optane through bcache — Optane specifically, because NAND endurance is wrong for that job — and being honest about the failure modes it buys you.

7th August 2026 Â· 35 min Â· 7359 words Â· Damien Dye
Eight 512-byte logical blocks mapped onto one 4K physical sector, beside a one-to-one 4K mapping

4Kn, 512e and 512n — Why Native 4K Wins, and What the Emulation Costs

512e drives present 512-byte sectors they do not have, and the firmware makes up the difference on every unaligned write. What that costs on the medium, in the host and in write amplification — why direct synchronous writes are the worst case — and how to convert a fleet to 4Kn.

7th August 2026 Â· 26 min Â· 5489 words Â· Damien Dye
A curve rising from 36 percent of bare metal at queue depth 1 to within a few percent by queue depth 32

PCIe Passthrough Performance on Proxmox VE — The IOMMU Tax and How to Minimise It

Why PCIe devices lose throughput when passed through to a VM via VFIO, and the practical tuning steps that claw most of it back.

7th August 2026 Â· 19 min Â· 3994 words Â· Damien Dye
Two CPU sockets with a DMA path crossing the inter-socket link between them

NUMA Alignment on Proxmox VE — Why It Matters and How to Get It Right

On multi-socket systems, a VM with its vCPUs on one NUMA node and its passed-through device on another loses 20–30% throughput before you’ve even looked at anything else.

7th August 2026 Â· 9 min Â· 1758 words Â· Damien Dye
A latency-percentile curve that stays flat to p90 then climbs steeply, beside one that stays flat

PCIe ASPM and Why You Should Disable It for Passthrough

Active State Power Management saves a few watts on idle PCIe links. Under VFIO passthrough, it adds latency jitter that’s hard to diagnose and easy to fix.

7th August 2026 Â· 8 min Â· 1661 words Â· Damien Dye
The same 4 KB payload drawn as 32 packets at MPS 128 and 8 packets at MPS 512

PCIe MaxPayloadSize — A Free Performance Win for Passthrough

QEMU’s virtual root complex defaults to 128-byte TLP payloads. Most devices support 256 or 512. One kernel parameter fixes it.

7th August 2026 Â· 7 min Â· 1400 words Â· Damien Dye