[{"content":"The Licence Is Already in the Firmware Plenty of people own a licence for the Fisher-Price OS (Windows) without ever having seen its product key. It came with the machine, and the key lives in that machine\u0026rsquo;s firmware, in a small ACPI table called MSDM.\nThen the machine stops being where the work happens. Either its owner wants that licence in a VM they now rent from you, or the machine itself gets wiped, turned into a Proxmox host, and the copy of the OS it shipped with is wanted back as a VM on top.\nMechanically, it is one QEMU option. That is the problem. QEMU hands any table to a guest without checking it, the guest activates with whatever key it finds, and nothing anywhere in the stack asks whether the licence allows any of it. This post covers what the table is, how to get it into a VM without passing on a corrupt one, what Microsoft\u0026rsquo;s terms say about moving it, and why a volume licence never goes near it.\nWhat the Table Is Since OEM Activation 3.0, a PC maker no longer prints the key on a sticker on the bottom of the case where anybody with a phone camera can copy it, and the key goes into the machine\u0026rsquo;s firmware instead. Microsoft\u0026rsquo;s factory tooling \u0026ldquo;injects the product keys into the firmware\u0026rdquo;, and its validation step checks \u0026ldquo;that the MSDM table exists\u0026rdquo; and that its header and entries \u0026ldquo;comply with the correct formats\u0026rdquo;1. A machine set up that way is \u0026ldquo;activated by using the OA3 DPK in the firmware\u0026rdquo;2. MSDM is an ordinary ACPI table, and on Linux it reads straight off the firmware:\nsudo cat /sys/firmware/acpi/tables/MSDM \u0026gt; msdm.bin The laptop this was written on carries one. 85 bytes.\nMicrosoft\u0026rsquo;s published specification defines the standard 36-byte ACPI header with the signature MSDM, and then stops. Everything after the header is a \u0026ldquo;Proprietary data structure that contains all the licensing data necessary to enable Windows activation\u0026rdquo;3.\nBytes Holds Known from 0 to 35 the standard ACPI header: signature MSDM, length, checksum, OEM ID Microsoft\u0026rsquo;s specification 36 to 55 20 bytes of fields: version, data type, data length real tables, not a published specification 56 to 84 the 29-character product key, five groups of five real tables, not a published specification QEMU Will Pass Anything You Give It QEMU\u0026rsquo;s -acpitable file= takes the \u0026ldquo;whole ACPI table from the specified files, including all ACPI headers (possible overridden by other options)\u0026rdquo;4, and the guest then sees an MSDM table exactly as a laptop\u0026rsquo;s firmware would present it. That was checked with a deliberately fake key. The bytes came out the other side identical, signature to last character.\nQEMU refuses nothing. This is what hw/acpi/core.c does with a broken upload5:\nWhat is wrong with the file What QEMU does What the guest sees Header length disagrees with the file warns, then overwrites the length with the real size a valid header Checksum does not sum to zero recalculates it, on every table, every time a valid checksum Key truncated or mangled nothing, the data past the header is not its business a valid-looking table round a broken key The guest has no way of telling the difference, and neither do you until somebody opens a support ticket. The first anybody hears of it is a customer whose activation failed.\nSo if customers upload these, check them before QEMU ever sees them:\n#!/usr/bin/env python3 \u0026#34;\u0026#34;\u0026#34;Refuse anything that is not a well-formed MSDM table before QEMU sees it.\u0026#34;\u0026#34;\u0026#34; import re, struct, sys t = open(sys.argv[1], \u0026#39;rb\u0026#39;).read() fail = lambda why: sys.exit(f\u0026#34;{sys.argv[1]}: {why}\u0026#34;) if len(t) \u0026lt; 56: fail(f\u0026#34;{len(t)} bytes, too short for an MSDM table\u0026#34;) sig, length = t[:4], struct.unpack_from(\u0026#39;\u0026lt;I\u0026#39;, t, 4)[0] if sig != b\u0026#39;MSDM\u0026#39;: fail(f\u0026#34;signature is {sig!r}, not MSDM\u0026#34;) if length != len(t): fail(f\u0026#34;header says {length} bytes, file is {len(t)}\u0026#34;) if sum(t) \u0026amp; 0xff: fail(\u0026#34;checksum does not sum to zero\u0026#34;) ver, _, dtype, _, dlen = struct.unpack_from(\u0026#39;\u0026lt;5I\u0026#39;, t, 36) if (ver, dtype) != (1, 1): fail(f\u0026#34;version {ver}, data type {dtype}: expected 1 and 1\u0026#34;) if dlen != 29 or 56 + dlen != length: fail(f\u0026#34;data length {dlen}: expected 29\u0026#34;) key = t[56:].decode(\u0026#39;ascii\u0026#39;, \u0026#39;replace\u0026#39;) if not re.fullmatch(r\u0026#39;([0-9A-Z]{5}-){4}[0-9A-Z]{5}\u0026#39;, key): fail(\u0026#34;data is not a 5x5 product key\u0026#34;) print(f\u0026#34;OK OEM {t[10:16].decode(errors=\u0026#39;replace\u0026#39;).strip()!r} key *****-*****-*****-*****-{key[-5:]}\u0026#34;) Checks Hold the file to A failure means Signature, length, checksum Microsoft\u0026rsquo;s documented header the file is broken Version, data type, data length, key shape the shape real tables carry look at this one by hand, not \u0026ldquo;this is forged\u0026rdquo; It never prints the key, only the last group. A product key in a log file is a product key somebody else can use.\nRun against a good table and three broken copies of it:\nOK OEM \u0026#39;EXMPLE\u0026#39; key *****-*****-*****-*****-EEEEE bad-sum.bin: checksum does not sum to zero bad-trunc.bin: header says 85 bytes, file is 70 bad-sig.bin: signature is b\u0026#39;SLIC\u0026#39;, not MSDM Where It Goes, and Who Can Put It There From a customer's upload to the key the guest activates with Customer's table msdm.bin, 85 bytes from their own machine Validator signature, length, checksum, key shape Stored /etc/pve/priv/ root only, every node VM args -acpitable file= set by root only QEMU rewrites the length, recomputes the checksum, refuses nothing Guest MSDM in ACPI, activation reads the key The validator sits in front of QEMU because QEMU will fix up a broken header rather than refuse it. The table lives in the private half of the cluster filesystem, so it follows the VM to any node and only root can read it. A product key is a secret. It does not belong anywhere www-data can read, and pmxcfs gives you exactly one place in /etc/pve where it cannot6:\nWhere the table could live Who can read it On every node the VM can migrate to /etc/pve/priv root only yes anywhere else in /etc/pve group-readable, the web interface\u0026rsquo;s www-data included yes a directory on one node whatever you set no At 85 bytes it is nowhere near the 1 MiB pmxcfs file limit7:\nmkdir -p /etc/pve/priv/msdm check-msdm.py upload.bin \u0026amp;\u0026amp; cp upload.bin /etc/pve/priv/msdm/9000.bin qm set 9000 --args \u0026#34;-acpitable file=/etc/pve/priv/msdm/9000.bin\u0026#34; qm set --args replaces the whole line. If the VM already carries SMBIOS or firmware options in args, write them all out together. And because args is root only, this is a job for your provisioning, not something a customer can do from the web interface. That is the right way round. The customer supplies the file, your tooling checks it and attaches it.\nThe Table Moves a Key, Not a Right This is where the feature needs a clear head. Moving the MSDM table into a VM moves the key. Not the licence.\nMicrosoft\u0026rsquo;s own terms say so, and the lines that matter fit in a table:\nCase What Microsoft says Source OEM licence, moving it transferable to another user \u0026ldquo;only with the licensed device\u0026rdquo; OEM licence terms8 OEM licence, how many installs \u0026ldquo;only one instance of the software for use on one device, whether that device is physical or virtual\u0026rdquo; OEM licence terms8 OEM licence, in plain words \u0026ldquo;locked\u0026rdquo; to the original PC \u0026ldquo;and cannot be transferred to any other PC\u0026rdquo; Microsoft small business blog9 The exception the transfer provisions \u0026ldquo;do not apply\u0026rdquo; where the software was acquired in Germany or a list of other countries OEM licence terms8 Hosting it for customers the Services Provider License Agreement is for \u0026ldquo;hosted applications to end customers\u0026rdquo; SPLA10 Desktop editions as hosted VMs \u0026ldquo;VMs must be hosted by a Qualified Multitenant Hoster (QMTH)\u0026rdquo; Microsoft Learn11 The table always comes from the customer\u0026rsquo;s own licensed machine, never from your hosts. The case that sits squarely inside the wording is the simplest one: the machine the licence was sold with, now running Proxmox, with that same copy of the Fisher-Price OS (Windows) moved into a single VM on it. That is still one instance on one device, and the OEM terms allow that \u0026ldquo;whether that device is physical or virtual\u0026rdquo;8. A second VM, or another machine, and it is not.\nIn that one case the customer\u0026rsquo;s machine and the hypervisor are the same box, so there is nothing to upload and nothing to copy. The Proxmox kernel already exposes the firmware\u0026rsquo;s MSDM table under /sys/firmware/acpi/tables/, readable by root only, and QEMU on Proxmox runs as root, so it can take the table straight from there:\nqm set 100 --args \u0026#34;-acpitable file=/sys/firmware/acpi/tables/MSDM\u0026#34; Pointing at the live table rather than a copy has two consequences. The path is read on whichever node the VM starts on, so this belongs on a standalone host and not in a cluster: migrate the VM and it either gets that node\u0026rsquo;s table, a different licence on a different machine, or does not start at all, because QEMU refuses a missing file with can't open file … No such file or directory. And it is one line, which makes it one line to paste onto a second VM. That second VM is the case the terms do not cover, so one VM gets it and no others.\nSo build the feature. The mechanism is sound, and there are customers entitled to use it; one who bought their licence in Germany may well be among them. But put the entitlement where it belongs: in your terms of service, the customer warrants they hold a licence that covers running it in your VM, and they do it before the upload button does anything. The validator proves the table is well-formed. Nothing proves it is theirs to use.\nVolume Licences Never Go Near the Table Can a customer with a volume licence use the same route? No. There is nothing for the table to carry.\nMSDM belongs to the OEM channel and nothing else. Microsoft\u0026rsquo;s own planning guide says OEM activation \u0026ldquo;is available only for computers that are purchased through OEM channels and have the Windows operating system preinstalled\u0026rdquo;12. Volume licences activate through one of three other models, \u0026ldquo;Multiple Activation Keys (MAK)\u0026rdquo;, \u0026ldquo;KMS\u0026rdquo; and \u0026ldquo;Active Directory-based activation\u0026rdquo;12, and in every one of them the key is installed in the operating system. Nothing reads it from firmware.\nMethod Where the key lives What the guest talks to In a Proxmox VM OEM, OA 3.0 the MSDM table in firmware Microsoft\u0026rsquo;s activation servers the -acpitable route above MAK installed in the guest Microsoft, once, counted against the key\u0026rsquo;s activations works with internet access or a phone KMS a generic key (GVLK) in the guest the customer\u0026rsquo;s KMS host on TCP 1688 needs a route to it, and 25 clients or 5 servers before it activates anything Active Directory a GVLK in the guest the customer\u0026rsquo;s domain controllers, at least every 180 days needs the VM joined to their domain AVMA a generic key in the guest the Hyper-V host\u0026rsquo;s own Datacenter licence not available at all So for a volume customer there is nothing to build on the hypervisor. Their key goes in their image, and your part is the network path: port 1688 to their KMS host, or a VPN to their domain controllers. The MSDM slot stays empty, and it should.\nTwo things catch people out. The generic key on volume media activates nothing by itself, because \u0026ldquo;the GVLK doesn\u0026rsquo;t work unless a valid KMS host key can be found\u0026rdquo;12, so a VM built from the customer\u0026rsquo;s Enterprise ISO with no route home sits unactivated. And a desktop volume licence is not a licence on its own. Volume programmes \u0026ldquo;cover upgrades to Windows client operating systems only\u0026rdquo;, and \u0026ldquo;an existing retail or OEM operating system license is needed for each computer\u0026rdquo;12. The volume licence sits on top of a base licence, very often the same OEM one the last section was about, and whether it may run in your VM is still the hosting question in that table, not an activation one.\nAVMA earns a line of its own, because it is the one that looks like the answer. It is Microsoft\u0026rsquo;s way for a host to activate its own server guests, binding \u0026ldquo;the VM activation to the licensed virtualization host\u0026rdquo;, and it requires \u0026ldquo;a Windows Server Datacenter edition with the Hyper-V server host role installed\u0026rdquo;13. Microsoft are plain about the rest: \u0026ldquo;AVMA doesn\u0026rsquo;t work with other server virtualization technologies\u0026rdquo;13. On Proxmox, server guests activate through your SPLA keys or through the customer\u0026rsquo;s own KMS.\nNothing in the Stack Checks the Paperwork Every layer here does its job and nothing more. The firmware stores a key. QEMU copies the bytes and tidies up the header. The guest reads the key and activates with it. Not one of them can tell whether the machine underneath is the one the licence was sold with, and none of them was ever meant to.\nSo the check lands on whoever runs the hypervisor. A licence is an agreement between the person who bought it and Microsoft, and your platform is not a party to it, but it is the thing that moves the key, and that makes the question yours whether you asked for it or not.\nWrite the answer down before the upload button does anything. One machine, one VM, the customer\u0026rsquo;s own licence and their word that it covers this. It costs nowt, and it protects you as much as them. The software will move whatever key it is handed. Whether it should have is a question only a person can answer, so make sure one has.\nMicrosoft Learn, OA 3.0 on the factory floor — \u0026ldquo;injects the product keys into the firmware\u0026rdquo;; /Validate checks \u0026ldquo;that the MSDM table exists\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft Learn, Validate an OEM Activation key — \u0026ldquo;is activated by using the OA3 DPK in the firmware\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft, Microsoft Software Licensing Tables (SLIC and MSDM) — table 2, offset 36: \u0026ldquo;Proprietary data structure that contains all the licensing data necessary to enable Windows activation.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nQEMU, System Emulation, Invocation — -smbios fields per type; -acpitable \u0026ldquo;For file=, take whole ACPI table from the specified files, including all ACPI headers\u0026rdquo;; -boot \u0026ldquo;Currently Seabios for X86 system support it.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nQEMU v11.0.3, hw/acpi/core.c — \u0026ldquo;ACPI table has wrong length\u0026rdquo; is a warn_report, then the length is overwritten and the checksum recalculated.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\npve-cluster, src/pmxcfs/pmxcfs.c — paths under priv are masked to 0777700, everything else to 0777750.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\npve-cluster, src/pmxcfs/memdb.h — #define MEMDB_MAX_FILE_SIZE (1024 * 1024) // 1 MiB.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft, Windows 11 OEM licence terms — \u0026ldquo;you may transfer the license to use the software directly to another user, only with the licensed device\u0026rdquo;; the transfer provisions \u0026ldquo;do not apply if you acquired the software in Germany\u0026rdquo; or the listed countries. Archived copy; the live page refuses automated requests.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft, small business blog archive — \u0026ldquo;the OEM Windows license is \u0026rsquo;locked\u0026rsquo; to the original PC it comes with and cannot be transferred to any other PC.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft, Services Provider License Agreement — \u0026ldquo;for service providers and software development companies licensing eligible Microsoft products to provide software services and hosted applications to end customers.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft Learn, Windows subscription activation for VDA — \u0026ldquo;VMs must be hosted by a Qualified Multitenant Hoster (QMTH).\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft Learn, Plan for volume activation — the three volume activation models; the KMS thresholds of \u0026ldquo;at least five computers\u0026rdquo; for servers and \u0026ldquo;at least 25 computers\u0026rdquo; for clients, on TCP port 1688; Active Directory-based activation needing the domain \u0026ldquo;at least once every 180 days\u0026rdquo;; \u0026ldquo;the GVLK doesn\u0026rsquo;t work unless a valid KMS host key can be found\u0026rdquo;; volume licences \u0026ldquo;cover upgrades to Windows client operating systems only\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft Learn, Automatic Virtual Machine Activation in Windows Server — \u0026ldquo;AVMA requires a Windows Server Datacenter edition with the Hyper-V server host role installed\u0026rdquo;; \u0026ldquo;AVMA doesn\u0026rsquo;t work with other server virtualization technologies.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/proxmox/moving-an-oem-licence-into-a-proxmox-vm/","summary":"An OEM licence for the Fisher-Price OS (Windows) lives in the PC\u0026rsquo;s firmware as an ACPI table called MSDM, and QEMU will pass any such table to a guest without checking it. This covers what the table holds, a validator that refuses malformed tables before QEMU silently repairs their headers, storing them in the private half of the Proxmox cluster filesystem, Microsoft\u0026rsquo;s terms on moving an OEM licence, the one case they plainly allow, taking the table straight from the host when the licensed machine is the hypervisor, and why MAK, KMS, Active Directory and AVMA activation never touch the table.","title":"Moving an OEM Licence Into a Proxmox VM, and What Moves With It"},{"content":"What Your Customer Sees When Your VM Boots You sell virtual machines under your own name. A customer powers one on, and the first thing on the screen is somebody else\u0026rsquo;s logo.\nThat is a stock Proxmox VM. Nothing is broken, and Proxmox are doing nothing wrong: it is their product and their name, they wrote the firmware and the interface it appears in, and they are entitled to put it there, the same way a server vendor puts its badge on the front of the box. It just is not yours.\nThe logo is just the start. This is what a stock UEFI guest reports on Proxmox\u0026rsquo;s current firmware, pve-edk2-firmware 4.2026.08-1, read from inside the guest with dmidecode:\nWhere the guest looks What it reads Whose name Boot screen the Proxmox logo Proxmox Firmware vendor (SMBIOS type 0) Proxmox distribution of EDK II Proxmox Firmware version 4.2026.08-1 Proxmox\u0026rsquo;s package version System manufacturer (type 1) QEMU QEMU Product name (type 1) Standard PC (Q35 + ICH9, 2009) QEMU Chassis manufacturer (type 3) QEMU QEMU Baseboard (type 2) not present at all nobody Management web interface the Proxmox logo, top left Proxmox The boot logo does not stop at the firmware, either. UEFI firmware hands the image it drew to the operating system through an ACPI table called BGRT, the Boot Graphics Resource Table, which exists to say \u0026ldquo;an image was drawn on the screen during boot\u0026rdquo;1, and the operating system then draws it again on its own boot screen. The Fisher-Price OS (Windows) puts it above its spinner, and Microsoft calls BGRT \u0026ldquo;the standard interface that Windows uses to access the logo\u0026rdquo;2. Fedora\u0026rsquo;s default Plymouth theme does the same on Linux3. So a guest on Proxmox\u0026rsquo;s firmware shows Proxmox\u0026rsquo;s logo twice before anybody has logged in.\nNone of this is hard to change. Keeping it changed is. The next apt full-upgrade will quietly put half of it back, and the half it does not put back is the half that ought to worry you, which is what most of this post is about.\nWhere each piece of branding enters the boot Power on QEMU builds SMBIOS and ACPI from the config Firmware OVMF draws its built-in logo, SeaBIOS a splash Hand-over SMBIOS tables, ACPI with BGRT and MSDM OS boot screen redraws the firmware logo out of BGRT Running guest reads the strings, activation reads the MSDM key VM config package file VM config package file VM config lives in /etc/pve, survives every upgrade lives under /usr/share, apt replaces it Branding enters the boot in five places. Three of them come out of the VM\u0026rsquo;s own config and survive anything apt does. Two come from files owned by a Proxmox package, and those are the ones an upgrade takes back. The MSDM table in that hand-over is not branding at all. It carries a licence key, and moving one into a VM has its own post: Moving an OEM Licence Into a Proxmox VM.\nEverything Under /usr/share Belongs to apt One rule decides every choice below. A file a package installed is the package\u0026rsquo;s file. Edit it in place and the next upgrade of that package overwrites your edit without a word, because as far as dpkg is concerned it is only putting its own file back where it left it.\nSo each piece of branding has to live somewhere apt does not own. A Proxmox node has three such places, and each costs something different:\nWhere it lives How it gets there Survives an upgrade What it costs you The VM config in /etc/pve smbios1, and args for everything else Yes, config is never touched by a package args is root only, and is not shown in the GUI A file dpkg has been told to leave alone dpkg-divert Yes, the package\u0026rsquo;s copy goes to a .distrib name instead Upgrades no longer reach the file, which matters when the file is firmware Your own directory, outside the package tree /usr/local, referenced from args Yes, no package owns it It has to be put on every node yourself Debian\u0026rsquo;s own description of a diversion is the clearest: \u0026ldquo;a way of forcing dpkg not to install a file into its location, but to a diverted location\u0026rdquo;4. That covers the web interface. For the firmware it is one of two options, and the more dangerous one.\nThe Chassis Is a Handful of Strings and a Permission Check SMBIOS is the table a machine uses to describe itself: who made it, what model it is, its serial number, what the board and the case are5. dmidecode reads it, every asset inventory tool you have ever pointed at a network reads it, the Fisher-Price OS (Windows) system information panel reads it, and so do most of the licence checks that decide whether a piece of software will run on a given machine at all. On a VM, QEMU writes it. Hence QEMU and Standard PC in the table above.\nType 1 through smbios1 Proxmox exposes one SMBIOS structure in the VM config: type 1, System Information, as smbios16. It takes manufacturer, product, version, serial, sku, family and uuid.\nThe catch is in qemu-server, not the docs. Every field except uuid has to match a base64 pattern, [A-Za-z0-9+\\/]+={0,2}, so a plain value with a space in it is refused outright. Real strings get base64-encoded with base64=1 set, and qemu-server decodes them again before it builds the QEMU command line7. The GUI\u0026rsquo;s SMBIOS editor does this on every save, with the comment \u0026ldquo;smbios values can be arbitrary, so encode and mark config as such\u0026rdquo;8. On the command line it is your job:\nb() { printf %s \u0026#34;$1\u0026#34; | base64 -w0; } qm set 9000 --smbios1 \u0026#34;uuid=$(qm config 9000 | sed -n \u0026#39;s/.*uuid=\\([0-9a-f-]*\\).*/\\1/p\u0026#39;),base64=1,\\ manufacturer=$(b \u0026#39;Example Cloud Ltd\u0026#39;),product=$(b \u0026#39;EC Compute Instance\u0026#39;),\\ version=$(b \u0026#39;2026.10\u0026#39;),family=$(b \u0026#39;General Purpose\u0026#39;),sku=$(b \u0026#39;ec-gp-4c16g\u0026#39;)\u0026#34; Note the uuid going back in. It is the guest\u0026rsquo;s machine identity, and whatever you pass to --smbios1 is stored as the whole option, so read it first and keep it.\nBrand the template, not each VM. A Proxmox clone gets a freshly generated UUID and keeps every other smbios1 field, so every clone comes out with its own identity and your strings9, with one exception, below.\nTypes 0, 2, 3 and 11 through args smbios1 stops at type 1. The firmware vendor, the baseboard, the chassis and the OEM strings all go through args, the line Proxmox passes straight to QEMU, which its own documentation calls \u0026ldquo;for experts only\u0026rdquo;6. QEMU\u0026rsquo;s -smbios option takes the fields for each type10:\nargs: -smbios \u0026#39;type=0,vendor=Example Cloud Ltd,version=EC-FW 1.0,date=10/04/2026,uefi=on\u0026#39; -smbios \u0026#39;type=2,manufacturer=Example Cloud Ltd,product=EC Virtual Board,version=1.0\u0026#39; -smbios \u0026#39;type=3,manufacturer=Example Cloud Ltd,version=1.0,asset=EC-ASSET-123,sku=ec-gp\u0026#39; -smbios \u0026#39;type=11,value=example-cloud:instance=123\u0026#39; Booted on QEMU with those lines, the guest\u0026rsquo;s dmidecode reads:\nStructure Field Stock Branded Type 0 Vendor Proxmox distribution of EDK II Example Cloud Ltd Type 0 Version 4.2026.08-1 EC-FW 1.0 Type 1 Manufacturer QEMU Example Cloud Ltd Type 1 Product Name Standard PC (Q35 + ICH9, 2009) EC Compute Instance Type 1 Family not specified General Purpose Type 2 Manufacturer structure absent Example Cloud Ltd Type 2 Product Name structure absent EC Virtual Board Type 3 Manufacturer QEMU Example Cloud Ltd Type 3 Asset Tag not specified EC-ASSET-123 Type 11 String 1 structure absent example-cloud:instance=123 Four traps turned up on the way. All four are silent.\nSupplying type 0 drops \u0026ldquo;UEFI is supported\u0026rdquo;. OVMF writes its own type 0 only when QEMU has not supplied one; the loop that walks QEMU\u0026rsquo;s tables sets NeedSmbiosType0 = FALSE the moment it meets a type 011. Your vendor string replaces Proxmox\u0026rsquo;s, which is the point. But QEMU\u0026rsquo;s type 0 replaces OVMF\u0026rsquo;s firmware characteristics as well, and without uefi=on the line \u0026ldquo;UEFI is supported\u0026rdquo; disappears from dmidecode. Put it back with uefi=on. For UEFI guests there is a better home for the vendor string anyway, inside the firmware itself, further down.\nThe chassis type cannot be set. QEMU\u0026rsquo;s type 3 takes manufacturer, version, serial, asset and sku, and nothing else10. It stays Other.\nTwo definitions of one type merge, and the later one wins. qemu-server puts its -smbios type=1 early on the command line and your args at the very end12, so a second -smbios type=1,manufacturer=… in args does not throw the first away: QEMU merged them field by field, the manufacturer came from args, and the UUID stayed where smbios1 put it. That matters for the fourth trap.\nThe two halves have different owners. qemu-server checks permissions option by option. smbios1 sits with the hardware options and needs VM.Config.HWType, while args falls through to the catch-all at the bottom, which says \u0026ldquo;only root can set\u0026rdquo;13. The board, the chassis, the firmware vendor and the OEM strings are yours alone. Type 1 is not. Anyone you have given hardware rights to can rewrite it, and on a hosted product that may well be the customer. If the manufacturer string matters, set it in args too. The later definition wins.\nThe Serial Number Is Already in the Config Most of what belongs in type 1 is already sitting in the VM config. The name makes a good serial number, because qemu-server only accepts a DNS name there14: short, printable, never a comma in it, and the one thing about the VM a customer will recognise.\nType 1 field Taken from For a 4-core, 16 GiB VM called web-01, tagged production;web Serial Number name web-01 SKU Number cores × sockets, and memory ec-4c16g Family the first entry in tags production UUID the smbios1 UUID it already has unchanged Manufacturer, Product your own fixed strings Example Cloud Ltd, EC Compute Instance Two things stay out. The node changes on every migration, and vmgenid is meant to change: its whole job is telling the guest it has been restored from a snapshot or built from a template15. Neither is an identity.\n#!/bin/bash # /usr/local/sbin/ec-smbios VMID: rebuild smbios1 from the VM\u0026#39;s own config. set -euo pipefail id=$1 cfg=$(qm config \u0026#34;$id\u0026#34; --current) get() { sed -n \u0026#34;s/^$1: //p\u0026#34; \u0026lt;\u0026lt;\u0026lt;\u0026#34;$cfg\u0026#34;; } b() { printf %s \u0026#34;$1\u0026#34; | base64 -w0; } name=$(get name) [ -n \u0026#34;$name\u0026#34; ] || { echo \u0026#34;VM $id has no name to use as a serial\u0026#34; \u0026gt;\u0026amp;2; exit 1; } cores=$(get cores); sockets=$(get sockets); vcpu=$(( ${cores:-1} * ${sockets:-1} )) mem=$(get memory); mem=${mem#current=}; mem=${mem%%,*}; mem=${mem:-512} (( mem % 1024 )) \u0026amp;\u0026amp; size=\u0026#34;${mem}m\u0026#34; || size=\u0026#34;$(( mem / 1024 ))g\u0026#34; tag=$(get tags); tag=${tag%%;*} uuid=$(get smbios1 | grep -o \u0026#39;uuid=[0-9a-fA-F-]*\u0026#39; | cut -d= -f2 || true) uuid=${uuid:-$(cat /proc/sys/kernel/random/uuid)} s=\u0026#34;uuid=$uuid,base64=1,manufacturer=$(b \u0026#39;Example Cloud Ltd\u0026#39;),product=$(b \u0026#39;EC Compute Instance\u0026#39;)\u0026#34; s+=\u0026#34;,serial=$(b \u0026#34;$name\u0026#34;),sku=$(b \u0026#34;ec-${vcpu}c${size}\u0026#34;)\u0026#34; [ -n \u0026#34;$tag\u0026#34; ] \u0026amp;\u0026amp; s+=\u0026#34;,family=$(b \u0026#34;$tag\u0026#34;)\u0026#34; qm set \u0026#34;$id\u0026#34; --smbios1 \u0026#34;$s\u0026#34; The defaults are qemu-server\u0026rsquo;s own, one core, one socket and 512 MiB16, so a config that leaves them out still gets a true SKU. Run against the config in the table, with qm stood in by a script that served it, the line it wrote went to QEMU decoded the way qemu-server decodes it, and the guest\u0026rsquo;s dmidecode read back web-01, ec-4c16g, production and the UUID it already had. A VM with no name is refused rather than given an empty serial.\nThe obvious home for this is a hookscript. It is the wrong one. qemu-server runs the pre-start hook inside the config lock it took to start the VM, and builds the QEMU command line from the config it loaded before the hook ran17. A hook that rewrites the serial changes the next start, not this one.\nSo run it from your provisioning, at the four points its inputs change: create, clone, rename and resize. Clone is the one that bites. A clone keeps every smbios1 field but the UUID9, so skip it and every VM built from web-template reports the template\u0026rsquo;s name as its serial.\nThe Boot Logo Lives in Two Different Places Which logo a guest shows depends on its firmware. The two firmwares Proxmox ships get it from completely different places:\nSeaBIOS (bios: seabios, the default) OVMF (bios: ovmf, UEFI) Where the logo comes from a JPEG QEMU hands over at boot a bitmap compiled into the firmware Proxmox\u0026rsquo;s file /usr/share/qemu-server/bootsplash.jpg, 640×480 Logo.bmp, 400×120, 8 bit, built into OVMF_CODE_4M*.fd Owning package qemu-server pve-edk2-firmware-ovmf Change it per VM args: -boot splash=… args pointing the VM at other firmware Change it without rebuilding anything yes no Reaches the guest OS through BGRT no yes SeaBIOS: a JPEG on the command line qemu-server puts the same -boot option on every VM it starts, menu=on,strict=on,reboot-timeout=1000,splash=/usr/share/qemu-server/bootsplash.jpg18. QEMU\u0026rsquo;s documentation says the picture is shown \u0026ldquo;when option splash=sp_name is given and menu=on, If firmware/BIOS supports them. Currently Seabios for X86 system support it\u0026rdquo;10. OVMF takes nothing from it but splash-time, which it uses as the boot menu timeout, and there is no reference to the splash file anywhere in OvmfPkg19.\nReplacing it is one line in the VM config:\nargs: -boot splash=/etc/pve/branding/splash.jpg A second -boot does not fight the first. QEMU merges it the same way it merges -smbios, the later value wins, and with two splash files given the screenshot showed the second. /etc/pve is the right home for the file, because it is the cluster filesystem and every node sees the same splash, and its 1 MiB per-file limit is nowhere near a 640×480 JPEG20.\nThe image itself has to meet three conditions, and getting any of them wrong costs you the logo or its colours:\nCondition Why What happens otherwise JPEG or 24-bit BMP QEMU checks the file before the VM starts splash file … format not recognized; must be JPEG or 24 bit BMP Baseline JPEG, 4:2:0 chroma SeaBIOS\u0026rsquo;s decoder handles nothing else21 ERR_NOT_SEQUENTIAL_DCT or ERR_NOT_YCBCR_221111, and a blank screen 640×480, like Proxmox\u0026rsquo;s own SeaBIOS asks the VGA BIOS for a mode of exactly the picture\u0026rsquo;s size no matching mode, no splash Most image tools write baseline 4:2:0 by default, so the usual way to lose the logo is one of the \u0026ldquo;save for web\u0026rdquo; options that quietly switches on progressive encoding, which looks identical in every image viewer you own and leaves SeaBIOS drawing nothing at all.\nThe colour trap This one cost an afternoon. The same JPEG, rendered by two builds of SeaBIOS 1.17.0, comes out in two different colours:\nIt is in SeaBIOS\u0026rsquo;s jpeg.c. The 24 bits-per-pixel writer has a little-endian branch that puts blue in the first byte of each pixel, and the 32 bits-per-pixel writer, PIC_32, has no such branch and writes red there instead21. On a little-endian framebuffer that swaps red and blue. Blue comes out gold, and orange would come out blue.\nWhich writer runs depends on the video mode the VGA BIOS offers:\nFirmware VBE mode for 640×480 Bits per pixel Colours QEMU upstream\u0026rsquo;s prebuilt SeaBIOS 1.17.0, as pve-qemu-kvm 11.0.3-4 ships it 0x111 16 right, quantised to 65,536 Fedora 44\u0026rsquo;s own seabios 1.17.0-10 build 0x142 32 red and blue swapped The mode numbers are SeaBIOS\u0026rsquo;s own table22, and the Proxmox row was checked against the very blobs in the package: byte-identical to QEMU\u0026rsquo;s prebuilt rel-1.17.0-0-gb52ca86e094d. So on Proxmox your colours are right, but there are only 65,536 of them, and a subtle gradient will band. And if your logo ever comes up in the wrong colours on some other hypervisor, it is not your JPEG.\nThe UEFI Logo Is Compiled Into the Firmware UEFI guests are the harder half, and on a modern Proxmox they are most guests.\nThere is no splash option to override. The logo is a bitmap inside the firmware image, and Proxmox\u0026rsquo;s build puts it there with one line in debian/rules:\ndebian/setup-build-stamp: cp -a debian/Logo.bmp MdeModulePkg/Logo/Logo.bmp That copies their 400×120, 8-bit Logo.bmp over the TianoCore one before edk2 is built23. The same file sets the firmware vendor as a build-time constant, PcdFirmwareVendor=L\u0026quot;Proxmox distribution of EDK II\u0026quot;, which is where the type 0 vendor in the first table comes from23.\nNew logo, new firmware. The way to get one that behaves exactly like Proxmox\u0026rsquo;s is to build it exactly the way Proxmox does: their tree at the tag that matches the package they ship, their patches, their flags, and one bitmap swapped. pve-edk2-firmware 4.2026.08-1 pins edk2 at 2970e56, which is the upstream edk2-stable202608 tag23.\ngit clone https://git.proxmox.com/git/pve-edk2-firmware.git cd pve-edk2-firmware git checkout f37039e52d228a7844a90aa2aaf1e161ebd04fcd # 4.2026.08-1 git submodule update --init --recursive cd edk2 QUILT_PATCHES=../debian/patches quilt push -a cp /path/to/your/Logo.bmp MdeModulePkg/Logo/Logo.bmp . ./edksetup.sh \u0026amp;\u0026amp; make -C BaseTools F=\u0026#34;-DNETWORK_HTTP_BOOT_ENABLE=TRUE -DNETWORK_IP6_ENABLE=TRUE -DNETWORK_TLS_ENABLE -DSECURE_BOOT_ENABLE=TRUE -DPVSCSI_ENABLE=TRUE -DTPM2_ENABLE=TRUE --pcd PcdUninstallMemAttrProtocol=TRUE -DFD_SIZE_4MB\u0026#34; P=(--pcd \u0026#34;PcdFirmwareVendor=LExample Cloud Ltd\\\\0\u0026#34; --pcd \u0026#34;PcdFirmwareVersionString=L4.2026.08-1+ec1\\\\0\u0026#34;) build -a IA32 -a X64 -t GCC -p OvmfPkg/OvmfPkgIa32X64.dsc $F \u0026#34;${P[@]}\u0026#34; -b RELEASE cp Build/Ovmf3264/RELEASE_GCC/FV/OVMF_CODE.fd OVMF_CODE_4M.fd rm -rf Build/Ovmf3264 build -a IA32 -a X64 -t GCC -p OvmfPkg/OvmfPkgIa32X64.dsc $F -DSMM_REQUIRE=TRUE \u0026#34;${P[@]}\u0026#34; -b RELEASE cp Build/Ovmf3264/RELEASE_GCC/FV/OVMF_CODE.fd OVMF_CODE_4M.secboot.fd The flags are lifted from debian/rules as it stands at that commit. Mind the vendor string: their Makefile writes it in double quotes, make\u0026rsquo;s shell strips them, and edk2 receives it bare, so that is how it is passed here, and quoting it a second time is a build error. Keep the bitmap at 400×120 and 8 bits like theirs and nothing else on the boot screen moves.\nYou need both images. qemu-server boots any Q35 VM with a 4M EFI disk on OVMF_CODE_4M.secboot.fd, the build with SMM required, and uses the plain OVMF_CODE_4M.fd only for i440fx24. On a current Proxmox the secboot one is what nearly every UEFI guest runs.\nBuilt that way, with nothing changed from Proxmox\u0026rsquo;s tree but the bitmap and two strings, the SMM image boots like this:\nAnd the guest\u0026rsquo;s own view of its firmware changes with it, without a single -smbios option on the command line:\nRead inside the guest Proxmox\u0026rsquo;s 4.2026.08-1 Rebuilt from the same tree Firmware vendor Proxmox distribution of EDK II Example Cloud Ltd Firmware version 4.2026.08-1 4.2026.08-1+ec1 \u0026ldquo;UEFI is supported\u0026rdquo; listed listed ACPI tables BGRT, WSMT and the rest the same set That makes the build-time vendor the better answer for UEFI guests. It is OVMF\u0026rsquo;s own type 0, so the UEFI bit stays put without anybody having to remember uefi=on. The args route for type 0 is for SeaBIOS guests, and for anyone not rebuilding firmware at all.\nTwo Ways to Hand the New Firmware to a VM qemu-server hardcodes where it looks for firmware: OVMF.pm holds a table of paths under /usr/share/pve-edk2-firmware/, keyed by machine type and Secure Boot options, and there is no option anywhere in the VM config, the GUI or the API that chooses a different file24. Two routes are left.\nBuilding branded OVMF and the two ways to hand it to a VM Proxmox's tree pve-edk2-firmware 4.2026.08-1 Your branding Logo.bmp + vendor string Same build, same flags OVMF_CODE_4M.fd + .secboot.fd A: divert on every node every UEFI VM gets it, no VM config changes B: per VM through args opt in VM by VM, Proxmox's files untouched apt full-upgrade Proxmox's new firmware lands, yours stays old Post-Invoke hook versions differ, so rebuild before the next start One build, two routes in. Either way, the next upgrade installs Proxmox\u0026rsquo;s new firmware beside yours and leaves yours as it was, which is why the hook at the bottom exists. Route A: divert it on every node. Tell dpkg that Proxmox\u0026rsquo;s two code images now belong somewhere else, and put yours where they were:\nD=/usr/share/pve-edk2-firmware for f in OVMF_CODE_4M.fd OVMF_CODE_4M.secboot.fd; do dpkg-divert --package example-branding --add --rename --divert \u0026#34;$D/$f.distrib\u0026#34; \u0026#34;$D/$f\u0026#34; install -m 0644 \u0026#34;/root/branding/$f\u0026#34; \u0026#34;$D/$f\u0026#34; done From its next start, every UEFI VM on the node boots your firmware. No VM config changes at all.\nRoute B: point chosen VMs at it. Keep the images in your own directory and override the firmware one VM at a time. Since moving to -blockdev, qemu-server attaches the code image as a block node called pflash0 and names it in -machine24, and QEMU merges a second -machine the same way it merges -boot, so args can add a node of its own and point pflash0 at that instead:\nargs: -blockdev driver=raw,node-name=brandcode,read-only=on,file.driver=file,file.filename=/usr/local/share/example-branding/OVMF_CODE_4M.secboot.fd -machine pflash0=brandcode That was tested the long way round. Proxmox\u0026rsquo;s secboot firmware was attached as pflash0 exactly the way qemu-server does it, with SMM on, and then that args line was added after it, and the guest booted the branded image with both the logo and the vendor string, which is where the screenshot above came from.\nRoute A, divert Route B, per-VM args Which VMs get it every UEFI VM on the node only VMs you set it on Proxmox\u0026rsquo;s files moved to .distrib untouched VM config unchanged an args line, root only Machine type must match handled, both images are diverted your job: secboot image for Q35 Undo dpkg-divert --remove --rename delete the line Migration target node must be diverted too target node must have the file The code image is about 3.5 MB, against a 1 MiB per-file limit in pmxcfs20. As such it cannot go in /etc/pve, and whichever route you take, it has to be put on every node. Migrate a VM to a node without it and you get one of two results, depending on the route: Proxmox\u0026rsquo;s own firmware and logo back again under route A, or under route B a VM that will not start at all because QEMU cannot open a file that is not there.\nThe Upgrade That Leaves You Behind A diversion is the right tool for a logo. Firmware is different.\nA diversion means Proxmox\u0026rsquo;s upgrades no longer reach the file. That is exactly what you asked for. It is also exactly what stops edk2 security fixes reaching your guests: Proxmox\u0026rsquo;s changelog for 4.2026.08-1 opens \u0026ldquo;Besides many bug and security fixes\u0026rdquo; and goes on to list a fix for CVE-2024-1374525, and a branded build from the release before has none of it. Nothing tells you.\nThat was tested, not assumed. In a Debian trixie container with Proxmox\u0026rsquo;s repository, pve-edk2-firmware-ovmf 4.2025.05-3 went in, both code images were diverted, and the package was upgraded to 4.2026.08-1:\nAfter the upgrade Result dpkg-query -W pve-edk2-firmware-ovmf 4.2026.08-1 Proxmox\u0026rsquo;s new image installed as OVMF_CODE_4M.fd.distrib, hash changed The branded image unchanged, still built from 4.2025.05-3 Anything on the console about it nothing, until the hook below So add a hook. apt runs a list of DPkg::Post-Invoke commands after every dpkg run26. Record which Proxmox version you built from, and compare after each one:\ncat \u0026gt; /usr/local/sbin/example-branding-check \u0026lt;\u0026lt;\u0026#39;EOF\u0026#39; #!/bin/sh built=$(cat /usr/local/share/example-branding/ovmf.built-from 2\u0026gt;/dev/null) now=$(dpkg-query -W -f \u0026#39;${Version}\u0026#39; pve-edk2-firmware-ovmf 2\u0026gt;/dev/null) [ \u0026#34;$built\u0026#34; = \u0026#34;$now\u0026#34; ] \u0026amp;\u0026amp; exit 0 echo \u0026#34;W: branded OVMF was built from pve-edk2-firmware-ovmf $built, Proxmox now ships $now.\u0026#34; \u0026gt;\u0026amp;2 echo \u0026#34;W: guests still boot the old firmware. Rebuild before the next VM restart.\u0026#34; \u0026gt;\u0026amp;2 exit 0 EOF chmod +x /usr/local/sbin/example-branding-check echo \u0026#39;DPkg::Post-Invoke { \u0026#34;/usr/local/sbin/example-branding-check\u0026#34;; };\u0026#39; \\ \u0026gt; /etc/apt/apt.conf.d/80example-branding On that same upgrade it printed:\nW: branded OVMF was built from pve-edk2-firmware-ovmf 4.2025.05-3, Proxmox now ships 4.2026.08-1. W: guests still boot the old firmware. Rebuild before the next VM restart. It exits 0 deliberately. apt aborts if a Post-Invoke command fails26, and a branding check is no reason to leave a node half upgraded. A loud warning will do.\nOne more thing. Restarts. A guest only picks up new firmware when its QEMU process restarts, so a rebuild means a stop and start from Proxmox rather than a reboot from inside the guest, which is exactly how Proxmox\u0026rsquo;s own firmware updates reach a running VM as well. They just do not need you to remember a rebuild first.\nOr Leave Proxmox\u0026rsquo;s Firmware Alone Every logo route so far changes the firmware, and the last section is the bill for that. There is one that does not, and it was prototyped for this post.\nA PCI device in QEMU can carry an option ROM, any file you name with romfile=27, and OVMF runs the EFI driver it finds there. Proxmox\u0026rsquo;s build runs it whether it is signed or not: OVMF sets PcdOptionRomImageVerificationPolicy to 0x00, always execute28. OVMF draws its own logo late in boot device selection29 and builds the BGRT only at ReadyToBoot, the moment before it hands over to a boot loader. And edk2\u0026rsquo;s BGRT driver takes a replacement image through EDKII_BOOT_LOGO2_PROTOCOL, rebuilding the table at ReadyToBoot whenever the image has changed30.\nThat is all it needs. The driver waits for ReadyToBoot at TPL_NOTIFY, which runs ahead of the BGRT driver\u0026rsquo;s own TPL_CALLBACK handler, clears the screen, draws its logo and hands the same image over. Built, it is one 8 KB file, attached with one line:\nargs: -device pci-testdev,romfile=/etc/pve/branding/brandrom.rom At 8 KB it sits well inside the 1 MiB pmxcfs limit20, so unlike a 3.5 MB firmware image it can live in /etc/pve and follow the VM to every node.\nIt was tested against Proxmox\u0026rsquo;s own firmware, straight out of pve-edk2-firmware-ovmf 4.2026.08-1: OVMF_CODE_4M.secboot.fd with Microsoft\u0026rsquo;s keys pre-enrolled in OVMF_VARS_4M.ms.fd, on Q35 with SMM, and booted through Fedora\u0026rsquo;s signed shim so that Secure Boot was on and enforcing:\nRead from the guest Without the ROM With the ROM SecureBoot variable 1 1 On screen Proxmox\u0026rsquo;s logo Proxmox\u0026rsquo;s for about 250 ms, then yours BGRT image, 400×120 at 440,340 Proxmox\u0026rsquo;s, SHA-256 be5afe4b… the ROM\u0026rsquo;s, SHA-256 6c494576… Firmware files changed none none The last row is the point. Proxmox\u0026rsquo;s firmware is untouched, so their next security release reaches your guests with the next restart, and none of the previous section applies.\nIt is not free:\nCost Why Proxmox\u0026rsquo;s logo shows for about a quarter of a second OVMF draws it before any option ROM driver gets the screen; only a rebuild avoids that The guest sees one more PCI device, with no driver the ROM has to ride on a device, and pci-testdev is QEMU\u0026rsquo;s do-nothing one Type 0 still says Proxmox the vendor string is compiled in; set it through args with uefi=on, as above Root only it is an args line And the finding underneath it wants saying plainly. An unsigned driver ran in firmware on a guest with Microsoft\u0026rsquo;s keys enrolled and Secure Boot enforcing, because Proxmox\u0026rsquo;s OVMF does not check option ROMs. Here that is useful. It also means that on this firmware Secure Boot checks boot loaders, not whatever the VM\u0026rsquo;s hardware config attaches, and since only root writes args, that is a question about who holds root on your nodes.\nIt is a prototype. It has run on QEMU 10.2.2 with Proxmox\u0026rsquo;s firmware, not yet on a Proxmox node, and what follows is source, not a product:\ngithub.com/damo2929/RebrandPCIRom: the option ROM driver, logo converter and container build.\nThe Web Interface Is Two Packages The logo at the top left of the web interface is not in pve-manager. Workspace.js asks for a proxmoxLogoSvg component with the prefix pwt, and that lives in proxmox-widget-toolkit, which the Backup Server and the Mail Gateway share31. Only the tab icons are in pve-manager:\nWhat File Package Header logo, drawn at 200×35 /usr/share/javascript/proxmox-widget-toolkit/images/proxmox_logo.svg proxmox-widget-toolkit Browser tab icon /usr/share/pve-manager/images/favicon.ico pve-manager 128×128 icon, also the touch icon /usr/share/pve-manager/images/logo-128.png pve-manager All three are plain files, so this is a job for dpkg-divert, and here the cost from the last section does not apply at all, because a logo has no security fixes to miss and nobody\u0026rsquo;s guest is any less safe for a node still serving last month\u0026rsquo;s favicon.\ndivert() { dpkg-divert --package example-branding --add --rename --divert \u0026#34;$1.distrib\u0026#34; \u0026#34;$1\u0026#34; install -m 0644 \u0026#34;$2\u0026#34; \u0026#34;$1\u0026#34; } divert /usr/share/javascript/proxmox-widget-toolkit/images/proxmox_logo.svg /root/branding/logo.svg divert /usr/share/pve-manager/images/favicon.ico /root/branding/favicon.ico divert /usr/share/pve-manager/images/logo-128.png /root/branding/logo-128.png Tested in the same container, against real packages:\nStep Result Divert on proxmox-widget-toolkit 5.2.9, install own SVG Proxmox\u0026rsquo;s SVG renamed to proxmox_logo.svg.distrib Upgrade to 5.2.10 own SVG still in place, Proxmox\u0026rsquo;s 5.2.10 copy in .distrib dpkg --verify proxmox-widget-toolkit no complaints, dpkg knows about the diversion apt-get install --reinstall proxmox-widget-toolkit own SVG still in place dpkg-divert --remove --rename Proxmox\u0026rsquo;s original back, byte for byte Two things stay Proxmox\u0026rsquo;s. The header image is drawn in a fixed 200×35 box, so draw your SVG to that shape or it gets squeezed. And the alt text says Proxmox and the image links to https://www.proxmox.com, both written into the JavaScript component rather than a file you can divert31. Changing those means patching a minified bundle that breaks on every upgrade. Not worth it over alt text.\nWhat Proxmox\u0026rsquo;s licence and trademark ask of you Proxmox VE is licensed under the AGPL version 332, and section 13 of that licence is the network clause: \u0026ldquo;if you modify the Program, your modified version must prominently offer all users interacting with it remotely through a computer network … an opportunity to receive the Corresponding Source of your version\u0026rdquo;33. Does swapping three images count as modifying the Program? That is for a lawyer. The cheap answer is to put your images and the divert script somewhere public and link it from the login page. It costs nowt and it settles the question.\nThe trademark side is clearer. Proxmox\u0026rsquo;s media kit says \u0026ldquo;Don\u0026rsquo;t alter the logo or incorporate the logo or symbol into your logo\u0026rdquo;34. So replace theirs outright with your own, never a recoloured or reworked version of it, and keep Proxmox out of your product\u0026rsquo;s name.\nYour Name on It Means Your Maintenance Branding a VM is cheap. A handful of SMBIOS strings, a JPEG, a bitmap and three images in a web interface: an afternoon\u0026rsquo;s work, most of it spent finding out where things live.\nKeeping it is the job. It is also the part that gets skipped.\nThe SMBIOS strings and the splash look after themselves, because they live in config that no package will ever touch, and the web interface logo looks after itself because dpkg has been told about it. The firmware does not. The day you put your logo into a firmware image, you took over a piece of somebody else\u0026rsquo;s release process, and Proxmox build, test and ship new firmware with security fixes in it which their customers get with the next upgrade, and yours get when you get round to rebuilding.\nThat trade is fine made knowingly. A hosting company whose guests boot firmware two releases behind because the logo mattered more than the changelog has not built a product. It has built a sticker.\nIf your name is on the boot screen, the firmware behind it is yours to keep current, whoever wrote it. Nobody will check. That is exactly why it has to be done.\nUEFI Forum, ACPI Specification 6.6, §5.2.23 Boot Graphics Resource Table — \u0026ldquo;The Boot Graphics Resource Table (BGRT) is an optional table that provides a mechanism to indicate that an image was drawn on the screen during boot\u0026rdquo;. Archived copy; the live page returns 403 to automated requests.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft Learn, Boot screen components — \u0026ldquo;This is the standard interface that Windows uses to access the logo.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nFedora Project, Changes/FlickerFreeBoot — \u0026ldquo;a new plymouth theme which incorporates the firmware\u0026rsquo;s bootsplash image\u0026rdquo;; Fedora\u0026rsquo;s plymouth.spec makes bgrt the default theme.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDebian, dpkg-divert(1), trixie — \u0026ldquo;File diversions are a way of forcing dpkg(1) not to install a file into its location, but to a diverted location.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDMTF, DSP0134 System Management BIOS Reference Specification 3.10.0 — type 1 System Information, type 2 baseboard, type 3 \u0026ldquo;the system\u0026rsquo;s mechanical enclosure(s)\u0026rdquo;, type 11 \u0026ldquo;free-form strings defined by the OEM\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nProxmox VE, qm.conf(5) — smbios1: \u0026ldquo;Specify SMBIOS type 1 fields\u0026rdquo;; args: \u0026ldquo;Arbitrary arguments passed to kvm … this option is for experts only.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nqemu-server, src/PVE/QemuServer.pm, the smbios1 format and command-line builder — every field but uuid has the pattern [A-Za-z0-9+\\/]+={0,2}; values are decoded when base64 is set.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\npve-manager, www/manager6/Parser.js, printQemuSmbios1 — \u0026ldquo;smbios values can be arbitrary, so encode and mark config as such\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nqemu-server, src/PVE/API2/Qemu.pm, clone — \u0026ldquo;auto generate a new uuid\u0026rdquo;, the other smbios1 fields are kept.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nQEMU, System Emulation, Invocation — -smbios fields per type; -acpitable \u0026ldquo;For file=, take whole ACPI table from the specified files, including all ACPI headers\u0026rdquo;; -boot \u0026ldquo;Currently Seabios for X86 system support it.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nedk2-stable202608, OvmfPkg/SmbiosPlatformDxe/SmbiosPlatformDxe.c — OVMF adds its own type 0 only when NeedSmbiosType0 is still set after walking QEMU\u0026rsquo;s tables.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nqemu-server, src/PVE/QemuServer.pm, custom args — args are split and pushed onto the end of the command line.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nqemu-server, src/PVE/API2/Qemu.pm, config permission checks — smbios1 is a hardware-type option needing VM.Config.HWType; the fallback \u0026ldquo;catches args, lock, etc.\u0026rdquo; and dies with \u0026ldquo;only root can set\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nqemu-server, src/PVE/QemuServer.pm, the name option — format =\u0026gt; 'dns-name'.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nqemu-server, src/PVE/QemuServer.pm, vmgenid — \u0026ldquo;notify the guest operating system when the virtual machine is executed with a different configuration (e.g. snapshot execution or creation from a template)\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nqemu-server, src/PVE/QemuServer/Memory.pm — memory default =\u0026gt; 512; cores and sockets default to 1 in QemuServer.pm.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nqemu-server, src/PVE/QemuServer.pm, vm_start — lock_config at line 5457, exec_hookscript($conf, $vmid, 'pre-start', 1) at 5581, and config_to_command at 5634 with the same $conf, not reloaded in between.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nqemu-server, src/PVE/QemuServer.pm, -boot — menu=on,strict=on,reboot-timeout=1000,splash=/usr/share/qemu-server/bootsplash.jpg.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nedk2-stable202608, OvmfPkg/Library/QemuBootOrderLib/QemuBootOrderLib.c — OVMF reads etc/boot-menu-wait, the splash-time value, and nothing else from -boot.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\npve-cluster, src/pmxcfs/memdb.h — #define MEMDB_MAX_FILE_SIZE (1024 * 1024) // 1 MiB.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nSeaBIOS, src/jpeg.c — PIC writes blue first on little-endian, PIC_32 at line 946 writes red first with no endian branch; ERR_NOT_SEQUENTIAL_DCT and ERR_NOT_YCBCR_221111 are the only formats refused.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nSeaBIOS, vgasrc/svgamodes.c — mode 0x111 is 640×480 at 16 bits, 0x142 is 640×480 at 32.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\npve-edk2-firmware, debian/rules at 4.2026.08-1 — cp -a debian/Logo.bmp MdeModulePkg/Logo/Logo.bmp, the PcdFirmwareVendor string and the OVMF build flags.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nqemu-server, src/PVE/QemuServer/OVMF.pm — the hardcoded firmware table under /usr/share/pve-edk2-firmware/, the SMM choice, and the pflash0 block node.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\npve-edk2-firmware, debian/changelog — 4.2026.08-1: \u0026ldquo;Besides many bug and security fixes … fix CVE-2024-13745\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDebian, apt.conf(5), trixie — \u0026ldquo;Pre-Invoke, Post-Invoke: This is a list of shell commands to run before/after invoking dpkg(1) … should any fail APT will abort.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nQEMU v11.0.3, hw/pci/pci.c — DEFINE_PROP_STRING(\u0026quot;romfile\u0026quot;, PCIDevice, romfile), a property every PCI device has.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nedk2-stable202608, OvmfPkg/OvmfPkgIa32X64.dsc — gEfiSecurityPkgTokenSpaceGuid.PcdOptionRomImageVerificationPolicy|0x00, the platform file Proxmox builds.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nedk2-stable202608, OvmfPkg/Library/PlatformBootManagerLib/BdsPlatform.c — BootLogoEnableLogo (), called from PlatformBootManagerAfterConsole.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nedk2-stable202608, BootGraphicsResourceTableDxe.c — SetBootLogo2 at line 239 copies the image; the ReadyToBoot handler at 417 uninstalls and reinstalls the table \u0026ldquo;If BGRT data change happens\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nproxmox-widget-toolkit, src/Logo.js — proxmoxLogoSvg, 200×35, alt: 'Proxmox', linking to proxmox.com; used from pve-manager Workspace.js with prefix: 'pwt'.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nProxmox VE wiki, FAQ — \u0026ldquo;Proxmox VE code is licensed under the GNU Affero General Public License, version 3.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nGNU Affero General Public License v3, §13 Remote Network Interaction — quoted in full in the text above.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nProxmox, Media kit — \u0026ldquo;Don\u0026rsquo;t alter the logo or incorporate the logo or symbol into your logo.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/proxmox/branding-a-proxmox-vm/","summary":"A stock Proxmox VM shows Proxmox\u0026rsquo;s logo at boot, says Proxmox in its firmware vendor string and calls itself QEMU in every SMBIOS table. This replaces all of it with your product\u0026rsquo;s name: SMBIOS types 0, 1, 2, 3 and 11 with the serial number and SKU filled from the VM\u0026rsquo;s own name and size, the SeaBIOS splash and its colour trap, a branded OVMF built from Proxmox\u0026rsquo;s own tree or an option ROM driver that replaces the logo and the BGRT under Secure Boot without touching the firmware, and the web interface logo. Every change is placed so an upgrade cannot quietly undo it, and the one an upgrade can leave dangerously stale, the firmware, gets a hook that says so.","title":"Branding a Proxmox VM, and Keeping It Branded Through the Next Upgrade"},{"content":"Everything in this post was run against a NetBox 4.7.2 on my own desk, with a real estate loaded into it: one site, one cabinet, fourteen devices, cabled, powered and addressed across three customers. Every screenshot is that instance, and every error message is one I actually got.\nIt goes in the order you would meet it. What the thing is, why you would want one, the fork you will hear about within a week of searching, how to stand one up, the order you have to fill it in, and then what you get back out. The customisation and extension work is at the end, because none of it makes sense until you have seen the shape of what you are extending.\nWhat NetBox Is NetBox is a database with a very specific opinion about what a network is made of, and a web application on top of it. It is a Django application on PostgreSQL, it has been open source under Apache 2.0 since DigitalOcean released it in June 2016, and the project is stewarded today by NetBox Labs alongside a team of volunteer maintainers.12\nUnderneath that it is 149 models across ten applications, reached through 146 REST endpoints and one GraphQL endpoint. Counted from the instance I built for this, rather than read off a feature page:\nApplication Models Application Models dcim 56 virtualization 7 extras 23 tenancy 6 ipam 18 wireless 3 circuits 11 users 7 vpn 10 core 8 Fifty-six of those are DCIM, the physical layer: sites, locations, racks, device types, devices, and every kind of port, bay and cable termination a device can have. Eighteen are IPAM. The rest cover circuits, tunnels and IKE policies, virtual machines and clusters, wireless links, tenancy, and the machinery that makes the whole thing extensible.\nThe count is not the point. The joins are, and the fastest way to see that is the screen NetBox is best known for.\nStart with the cabinet, because it is the screen that sells the thing.\nThat elevation is drawn from the data, not uploaded. Each device is there because something says it occupies those rack units, facing that way, and the colours come from the role you gave it. Space utilisation reads 28.6% because NetBox worked it out. Nobody maintains that number.\nTwo things follow from it that are worth more than the picture.\nYou can ask which units are free, and get an answer you can act on. You can also reserve units before anything is installed in them, which is the difference between selling space you have and selling space you think you have.\nWhat it deliberately will not do A product that knows what it is not is rarer than one that does everything badly, and NetBox\u0026rsquo;s own documentation is blunt about it. It does not provide network monitoring, DNS service, RADIUS, configuration management or facilities management.1\nMore importantly it holds the desired state of your network, not its operational state, and the documentation says automated import of live network state is \u0026ldquo;strongly discouraged\u0026rdquo;, because every record should be vetted by a human first.1\nThat is the decision everything else rests on, and it is the one people argue with. The argument goes: surely a source of truth should be the truth, so discover the network and load it in. The answer is that a discovered network tells you what is there, and what is there includes every mistake anybody has ever made. A switch port left in the wrong VLAN in 2021 is a fact. It is not an intention.\nNetBox holds the intention. Your monitoring holds the reality. The interesting number is the difference between them, and you cannot compute a difference from one input.\nThe other tenet is stated just as plainly: given a choice between a relatively simple eighty per cent solution and a much more complex complete one, take the simple one.1 You will feel that the first time you want to model something it does not model, and there is a whole run of sections near the end about what to do when you hit it.\nNetBox does NetBox does not Record what should be there Poll what is there Hold the intended VLAN for a port Tell you the port is down Say which customer a prefix belongs to Bill them for it Render a device\u0026rsquo;s configuration from a template Push it to the device Track the circuit, the provider and the commit Monitor the circuit Say which rack a device is in, and which U Open the cabinet Read that right-hand column as a list of tools you still need. Read it wrong and you will try to make NetBox into all of them, which is how a source of truth becomes another system nobody trusts.\nWhy You Need One Ask a managed service provider where the authoritative record of a customer\u0026rsquo;s network lives, and you will get an answer. Ask two of their engineers separately and you will get two.\nOne will point at a spreadsheet. One will point at a diagram last saved by somebody who left in 2023. Somebody else will say the firewall config is the documentation, which is at least honest, because a config does describe what a box is doing. It just does not describe why, or who asked for it, or which of the four customers behind that box is paying for the rule.\nThe failure is never the day you notice the record is wrong. It is the day somebody needs it.\nThe moment What you have to produce What it costs when you cannot A renewal conversation An itemised account of what the monthly charge buys The competitor\u0026rsquo;s quote is itemised, because they went and counted An engineer hands their notice in Everything they knew, written down Six months of finding out, one ticket at a time A customer leaves A description of their own estate Three weeks to assemble it, and a reference they will give honestly An auditor asks a scope question Which systems hold personal data, and where they physically are A regulator told something that later turns out not to be true A migration needs pricing A count of what is actually there You bid on an estimate and eat the difference None of those are exotic. They are Tuesday.\nWhat you get out of it is not documentation. Documentation is a thing you write and then stop maintaining. What you get is a database that refuses to hold a contradiction, and which answers questions nobody thought to ask it in advance. There is a cabinet further down this post that turns out to be 28.6 per cent full of kit and 90.7 per cent full of power. Nobody set out to find that. It fell out of entering a wattage on a device type once.\nThe honest caveat comes with it, and the last section is about that. A record is only worth what it costs you to keep right. But the alternative is a business that cannot describe itself, and the first person to find that out is usually a customer.\nThe Other One: Nautobot You will run into this within about a week of searching, so it is worth knowing what happened.\nIn 2021, Network to Code forked NetBox and called the result Nautobot. Not a soft fork or a distribution: a hard fork that has diverged for five years. Its own v1.0 release notes describe it as \u0026ldquo;a divergent fork of NetBox 2.10\u0026rdquo;, the repository was created on 19th February 2021 and v1.0.0 landed on 26th April 2021.3\nTheir stated reasons are on their own blog and worth reading in their words rather than mine. Three things drove it. They wanted to sell enterprise support: \u0026ldquo;We need to offer high-touch support models with Service Level Agreements (SLAs) we can guarantee. We need flexibility to offer Long-term Support (LTS) for customers who can\u0026rsquo;t upgrade at the pace of a fast moving open source project.\u0026rdquo; They wanted the source of truth to sit at the centre of an automation platform rather than serve documentation. And \u0026ldquo;there became a growing divergence in our vision about what a Source of Truth for networking should look like and how to get there\u0026rdquo;.4\nThat first reason needs its date attached, because it has stopped being true. In February 2021 there was no company behind NetBox to sell you anything. NetBox Labs was not founded until 2023, spun out of NS1 after IBM acquired it, co-founded by NetBox\u0026rsquo;s own lead maintainer.5\nAnd it is not a third party who built a business on somebody else\u0026rsquo;s project. NetBox Labs is the custodian of NetBox: the project\u0026rsquo;s own documentation says \u0026ldquo;the open source project is stewarded by NetBox Labs and a team of volunteer maintainers\u0026rdquo;.1 They sell NetBox Enterprise for self-managed installs, host it for you as NetBox Cloud, and offer 24/7 support.6\nSo \u0026ldquo;you cannot buy support for NetBox\u0026rdquo; was a fair thing to say when Network to Code forked, and is not a fair thing to say now. Both projects have a commercial company behind them that will sign something, and in NetBox\u0026rsquo;s case that company is the one stewarding the project.\nThe line most people miss is the next one, and it is the reason this is not a grubby story: \u0026ldquo;the NetBox project team suggested that we should consider forking.\u0026rdquo;\nThat is about as civilised as a fork gets. Two groups wanted different things, said so over a long period, and split rather than fighting over one codebase. Both halves are still Apache 2.0. Nobody took anything they were not entitled to take.\nWhat the fork was actually for Nautobot 1.0\u0026rsquo;s release notes list what it added relative to NetBox 2.10, and the list tells you the argument better than any blog post:3\nWhat Nautobot added in 2021 Where NetBox is now GraphQL support NetBox has it Git integration as a data source NetBox has it, as synchronized data sources Single sign-on NetBox has it Secrets NetBox has it via a plugin Scripts and reports consolidated into Jobs NetBox is moving scripts out to a plugin in 4.7 Custom fields on all models NetBox has broad custom field support Data validation plugin API NetBox has custom validation rules Customisable statuses, as database objects NetBox statuses are still a Python choice set User-defined relationships between models NetBox has no equivalent UUID primary keys NetBox uses integer keys Plugin API enhancements NetBox\u0026rsquo;s plugin framework has grown a lot since Most of the top half has converged. Several things Nautobot shipped in 2021 arrived in NetBox afterwards, which is what usually happens when two projects are solving the same problems in public.\nThe bottom three have not converged, and they are architectural rather than cosmetic. I checked both codebases today rather than trusting the 2021 notes.\nNautobot\u0026rsquo;s Status is a database model, described in its own source as a \u0026ldquo;Model for database-backend enum choice objects\u0026rdquo;, so a status is a row somebody can add in the UI. NetBox\u0026rsquo;s statuses come from a Python choice set, which is why the section above adds one by editing configuration.py and restarting. Nautobot has Relationship and RelationshipAssociation models, so you can define a relationship between two existing object types without writing code. NetBox\u0026rsquo;s answer to that problem is a plugin, which is the last run of sections in this post. And Nautobot\u0026rsquo;s primary keys are UUIDs, where NetBox\u0026rsquo;s are integers.\nThey also did the thing they forked to do. As of today Nautobot is publishing 3.2.x and 2.4.x on the same day, which is a genuine long-term maintenance line running alongside current, and that was one of the three stated reasons.\nWhich one The honest numbers first. NetBox has 21,625 stars and 3,133 forks; Nautobot has 1,617 and 422.7 Both were pushed to within the last two days, both are Apache 2.0, and both now have a commercial company behind them selling support and hosting.\nThat gap is not a verdict on quality. It reflects a five year head start and the fact that most people needing a source of truth find NetBox first. But it does decide the thing that usually matters more than features, which is how many plugins, integrations, Ansible modules, forum answers and colleagues you will find for the one you pick.\nSo: if you want an inventory and a source of truth that other systems read, and you want the biggest ecosystem and the easiest hiring, the answer is NetBox, and that is what the rest of this post is about. If your reason for wanting a source of truth is specifically to drive automation from it, or you want statuses and relationships defined by your team rather than by a configuration file and a restart, go and look at Nautobot properly before you decide.\nDo not let anybody sell you either one on support alone. Both sides have that covered now, and the differences that will still be there in five years are the ones in the table above.\nWhat you should not do is pick one because somebody told you the other is dead. Neither is, and both were still shipping releases this week.\nGetting One Running Two routes, and I ran both of them for this. Containers if you want it working in twenty minutes, packages on a host if it is going to be load bearing.\nThe container stack The community maintains netbox-docker, which is the quick answer:8\ngit clone -b release https://github.com/netbox-community/netbox-docker.git cd netbox-docker tee docker-compose.override.yml \u0026lt;\u0026lt;\u0026#39;EOF\u0026#39; services: netbox: ports: - 8000:8080 EOF docker compose pull docker compose up That gives you the application, PostgreSQL, Redis and a background worker, wired together. I built the same thing by hand under podman so I could see the parts, and the parts are worth knowing because two of them catch people:\nContainer Doing what If you leave it out netbox The Django application behind gunicorn nothing works postgres The database, 15 or later nothing works redis / valkey Two databases: one for tasks, one for caching nothing works netbox-worker rqworker, which drains the task queue webhooks never fire, background jobs never run, and nothing warns you netbox-housekeeping The periodic tidy-up change log records never expire That fourth row is the one. Everything looks healthy without a worker. Event rules queue up and sit there.\nTwo things bit me on first run and neither is in an error message you would search for.\nThe first migration takes a long time. Not a minute. On this machine it was several, because NetBox 4.7 replaces django-mptt with PostgreSQL ltree and rebuilds every hierarchical table on the way through. The container just sits there applying migrations. Leave it alone.\nWithout API_TOKEN_PEPPERS you cannot create a v2 API token, and the only sign is a warning in the log:\nUserWarning: API_TOKEN_PEPPERS is not defined. v2 API tokens cannot be used. Set at least one, at least fifty characters, before you go looking for why the API rejects you.\nOn a host, from the packages The documented route is tested on Ubuntu 24.04. I ran it end to end on a clean one, and it landed on this:\nComponent What 24.04 gave me What NetBox 4.7 needs PostgreSQL 16.15 15 or later Redis 7.0.15 6.0 or later Python 3.12.3 3.12, 3.13 or 3.14 Django 6.1.1 comes with NetBox NetBox v4.7.2 Those minimums are NetBox\u0026rsquo;s own, and 4.7 raised both the PostgreSQL and the Redis floor.9\nThe whole job is five steps.\n# 1. the services sudo apt install -y postgresql redis-server sudo -u postgres psql -c \u0026#34;CREATE DATABASE netbox;\u0026#34; sudo -u postgres psql -c \u0026#34;CREATE USER netbox WITH PASSWORD \u0026#39;something-you-generated\u0026#39;;\u0026#34; sudo -u postgres psql -c \u0026#34;ALTER DATABASE netbox OWNER TO netbox;\u0026#34; sudo -u postgres psql -d netbox -c \u0026#34;GRANT CREATE ON SCHEMA public TO netbox;\u0026#34; # 2. the build dependencies sudo apt install -y python3 python3-pip python3-venv python3-dev build-essential libxml2-dev libxslt1-dev libffi-dev libpq-dev libssl-dev zlib1g-dev git # 3. the application, at a release tag rather than at main sudo mkdir -p /opt/netbox \u0026amp;\u0026amp; cd /opt/netbox sudo git clone https://github.com/netbox-community/netbox.git . sudo git checkout v4.7.2 sudo adduser --system --group netbox sudo chown -R netbox /opt/netbox/netbox/media/ /opt/netbox/netbox/scripts/ /opt/netbox/netbox/reports/ # 4. the configuration: five values, no more cd /opt/netbox/netbox/netbox/ sudo cp configuration_example.py configuration.py python3 /opt/netbox/netbox/generate_secret_key.py # run it twice sudo $EDITOR configuration.py # 5. let the upgrade script do the rest sudo /opt/netbox/upgrade.sh Those five values are ALLOWED_HOSTS, DATABASES, REDIS, SECRET_KEY and API_TOKEN_PEPPERS. Generate the last two separately and do not reuse one for the other. As such, run the key generator twice and paste the results into different lines.\nupgrade.sh builds the virtual environment, installs every Python dependency, runs the migrations, builds the documentation for offline use and collects the static files. When it finishes on a fresh box it prints a warning that looks alarming and is not:\nWARNING: No existing virtual environment was detected. A new one has been created. Update your systemd service files to reflect the new Python and gunicorn executables. (If this is a new installation, this warning can be ignored.) Then a superuser, and it runs:\nsource /opt/netbox/venv/bin/activate cd /opt/netbox/netbox \u0026amp;\u0026amp; python3 manage.py createsuperuser For anything real, put gunicorn in front of it rather than runserver. The configuration and the unit files are already in the repository, which is the detail worth knowing because people write their own:\nsudo cp /opt/netbox/contrib/gunicorn.py /opt/netbox/gunicorn.py sudo cp -v /opt/netbox/contrib/*.service /etc/systemd/system/ sudo systemctl daemon-reload sudo systemctl enable --now netbox netbox-rq netbox.service runs gunicorn on 127.0.0.1:8001, netbox-rq.service runs the worker, and nginx or Apache goes in front to terminate TLS and serve /static. Two services, and the second one is the same worker the container stack needs.\nThe Order You Have To Fill It In A fresh NetBox is an empty database with opinions, and the first hour with one is usually spent finding out what those opinions are. You go to add a device, and it will not let you.\nThose red asterisks are the whole lesson. NetBox will not record a thing before the things it hangs off exist, and it is not being awkward: a device with no type is a row that cannot answer any of the questions a device is for.\nSo rather than guess, I asked the model which foreign keys are actually mandatory, and then tried to break each rule through the API to see what comes back.\nPOST /api/dcim/device-types/ no manufacturer {\u0026#34;manufacturer\u0026#34;: [\u0026#34;This field is required.\u0026#34;]} POST /api/dcim/devices/ no device type {\u0026#34;device_type\u0026#34;: [\u0026#34;This field is required.\u0026#34;]} POST /api/dcim/interfaces/ no device {\u0026#34;device\u0026#34;: [\u0026#34;This field is required.\u0026#34;]} POST /api/ipam/aggregates/ no RIR {\u0026#34;rir\u0026#34;: [\u0026#34;This field is required.\u0026#34;]} POST /api/circuits/circuits/ no provider, no type {\u0026#34;provider\u0026#34;: [\u0026#34;This field is required.\u0026#34;], \u0026#34;type\u0026#34;: [\u0026#34;This field is required.\u0026#34;]} POST /api/virtualization/virtual-machines/ nothing at all {\u0026#34;__all__\u0026#34;: [\u0026#34;A virtual machine must be assigned to a site, cluster, or device.\u0026#34;]} Every one refused, and said exactly which thing was missing. Here is the same information as a list of prerequisites:\nTo create Required first Optional, but you want it first Tenant nothing a Tenant group Region, Site group nothing a parent of the same kind, they nest Site nothing a Region, a Site group, a Tenant Location a Site a parent Location, a Tenant Rack type a Manufacturer Rack a Site a Location, Rack group, Rack role, Rack type, Tenant Device type a Manufacturer Device role nothing a parent Device role, they nest Device a Device role, a Device type, a Site a Rack and position, a Platform, a Tenant Interface, and every other component a Device Cable two things to terminate on a Tenant Aggregate an RIR a Tenant VLAN nothing a VLAN group, a Role, a Tenant Prefix nothing a Site, a VLAN, a Role, a VRF, a Tenant IP address nothing an Interface to assign it to, a Tenant Circuit a Provider and a Circuit type a Tenant Circuit termination a Circuit a Site to land on Cluster a Cluster type a Cluster group, a Site, a Tenant Virtual machine a Site, a Cluster or a Device a Platform, a Tenant VM interface a Virtual machine The second column is the one that costs you. Prefixes, IP addresses, VLANs and tenants require nothing at all, so nothing stops you creating them on day one in any order you like. Whether they are any use is a different question, because an address with no interface behind it is a row in a list, and a device with no tenant is a device you will be editing again later.\nThe order to actually work in That gives a sequence. Work down it and nothing ever refuses you.\nFirst, the things everything else hangs off. None of these are exciting and all of them are cheap to get wrong.\nTenant groups, then tenants. Do these before anything else. Twenty-eight models take a tenant, including Site, Location and Rack, so if the customers do not exist yet you cannot stamp them on as you go and you will be bulk-editing later. Regions and site groups. Both optional, both nest inside themselves, and they are independent of each other. Regions are for geography, site groups for function, and you can use one, both or neither. Sites. The root of nearly everything. A site requires nothing, which is why it is the first thing you can actually create. Locations. These need a site, and they nest, so a hall containing rows containing pods is one model three deep. Rack roles and rack groups. Both optional. Rack groups are flat and sit alongside locations as a second axis, which is handy for rows and pods. Manufacturers, then rack types. A rack type needs a manufacturer. Skip rack types if you are not modelling the cabinets themselves. Racks. These need a site. If you also give one a location, that location has to belong to that site, and NetBox checks. Then the hardware catalogue, not the hardware. Before you can add a single device you need manufacturers, device types, device roles and, in practice, platforms.\nManufacturers. You may already have some from step 6, because rack types need them too. Module types do as well. Device types, each needing a manufacturer. Device roles, which nest, and platforms. This is the step people skip, and it is the step that decides how much typing the rest of the job takes. A device type carries its own interfaces, ports and bays as templates, so every device you create from it arrives with the right components already on it. Get the type right once and racking forty of them is forty names.\nPlatform is the odd one in that list, because a device does not strictly require one. Do it now anyway. It is what carries the config template and the NAPALM driver later, and going back to set it on an estate is the same evening you would have spent on tenants.\nThen the kit itself.\nDevices. Role, type and site are all required. Rack and position are optional, and if you give them they have to be consistent with the site. Components, if the device type did not already supply them. Cables, between the components. Then addressing, because an address wants an interface to live on and the interface only exists after step 12.\nRIRs, then aggregates. Prefix and VLAN roles, VRFs, VLAN groups then VLANs. Prefixes, then the individual IP addresses. Then the commercial layer.\nProviders, provider accounts and circuit types. Circuits, then terminate them onto sites. And the virtual estate, which mirrors the physical one.\nCluster types and cluster groups, then clusters. Virtual machines, which need a site, a cluster or a device, then their interfaces. Steps 1 to 10 are an afternoon and they feel like admin. There is nowt glamorous in any of it. They are also the afternoon that decides whether step 11 takes a morning or a fortnight.\nThe rules that bite later Required fields are the easy half, because they fail immediately and tell you why. The ones that catch people out are the consistency rules, which only fire once you have enough data to contradict yourself:\ndevice at Leeds, put in a Manchester rack {\u0026#34;rack\u0026#34;: [\u0026#34;Rack MCR1-A07 (A07) does not belong to site Leeds Edge.\u0026#34;]} device at Leeds, in a Manchester location {\u0026#34;location\u0026#34;: [\u0026#34;Location Hall 2 does not belong to site Leeds Edge.\u0026#34;]} rack at Leeds, in a Manchester location {\u0026#34;__all__\u0026#34;: [\u0026#34;Assigned location must belong to parent site (Leeds Edge).\u0026#34;]} a second device in a unit that is already taken {\u0026#34;position\u0026#34;: [\u0026#34;U39.0 is already occupied or does not have sufficient space to accommodate this device type: MX204 (1.0U)\u0026#34;]} That last one is the one I would point at if somebody asked why bother with any of this. NetBox knows the device is 1U, knows what is in the cabinet, and will not let you record two things in the same space. Your spreadsheet will let you do that all afternoon and never say a word, and you will find out when somebody is stood in the hall with a box in their hands.\nNone of these are configurable and none of them should be. They are the difference between a record and a wish.\nUsing It: What Each Screen Gives You With the estate loaded, here is what you actually get back out of it. Everything below is that same instance: one cabinet, fourteen devices, cabled and addressed across three customers.\nA device is its components Click into one of those hypervisors and the interesting tab is not the summary, it is the interfaces.\nRead across one row. The interface, its speed, what you wrote about it, the address on it, the label on the cable, and the port on the far end. That is one query, and it is the answer to the question everybody actually asks, which is \u0026ldquo;what is this plugged into\u0026rdquo;.\nNote where the address sits. It is on eno1, not on the server. That sounds like pedantry right up until you have a box with a management interface, two data interfaces and a loopback, and somebody asks which address answers for it. A model that hangs addresses off devices cannot tell you. This one can, and it can also hold the perfectly ordinary case of four addresses on one interface.\nFollow the cable Cables terminate on components, not on devices. An interface at one end, an interface at the other, or a front port, a rear port, a power outlet, a circuit termination. That is the detail the whole feature rests on, and it is why the interface list above could print the far end in its own column without being told.\nModel a cable device to device and you have drawn a picture. Terminate it on the ports and NetBox can walk it.\nOne hop here, because it is a direct attach cable. Put patch panels in the middle and it walks them, panel by panel, and tells you what is on the far end of a run through three cabinets. That is the job that otherwise involves a torch and somebody holding the other end of a tone probe.\nThe trace also prints the full location of each end. Site, hall, cabinet, face, rack unit. If you have ever been on a call trying to tell a remote hands engineer which box to look at, that line is the whole value.\nAddresses, as a tree rather than a tab IPAM is the half people arrive for.\nThe indentation is computed from the addresses themselves. You do not tell NetBox that 10.20.20.0/24 sits inside 10.20.0.0/16, it works that out, and it will keep working it out when somebody adds a /26 in the middle of it next year.\nEvery row carries the things you actually filter by: the VLAN it maps to, the role, and the customer it belongs to. Utilisation is computed too.\nOpen one and you get the addresses in it, and the gaps.\nThose green rows are the free space, shown in line with the used space. There is an API call that hands you the next free address out of a prefix, which is the one piece of IPAM automation that pays for itself immediately, because it is the thing people otherwise do by squinting at a spreadsheet and hoping.\nAn interface takes as many addresses as you like, of both families, and you nominate one of each as the device\u0026rsquo;s primary.\nTwo IPv4 and two IPv6 on one port, which is an ordinary Tuesday and something a device-centric model cannot express at all.\nEnter addresses with the mask of the network they are on, not as a /32. NetBox will accept 10.20.20.11/32 and file it under the right prefix, because containment is worked out from the host address. What it will not do is second-guess you afterwards: the mask is stored exactly as typed and handed back to whatever reads it, so a /32 on a LAN address renders a /32 into the device configuration. The place that bites is three months later in a config template, not today in the form.\nOne install, three customers Most objects in NetBox can be assigned to a tenant. An enterprise uses that for business units. If you sell managed services, you create one per customer.\nThat panel on the right is the answer to \u0026ldquo;what does this customer have\u0026rdquo;, and it assembled itself. No report, no spreadsheet, no asking the engineer who built it.\nIt is worth being careful about what tenancy means, because getting it wrong on day one is a year of unpicking later. A tenant means the object is dedicated to that customer. A router that serves only them gets their tenant. A firewall serving four of them does not belong to any of them, so it gets none, and the relationship goes somewhere else. More on that further down, because it is the point where most people find they need to add something of their own.\nPut The Numbers On The Device Type This is the step that separates an inventory from something that answers questions, and it costs about ten minutes per device type.\nA device type can carry its weight and, through its power port templates, its power draw. Put those on the type once and every device you ever create from it inherits them. Leave them off and NetBox will happily tell you a cabinet is 28.6% full and nothing else.\nI put weight on all five types here, gave each one two power port templates with a maximum and an allocated draw, gave the PDU type an inlet and twelve outlets, then created a power panel and two feeds into the cabinet and cabled the lot up: every device\u0026rsquo;s PSU1 to PDU A, every PSU2 to PDU B, and each PDU\u0026rsquo;s inlet to its feed.\nThen the rack page changes completely.\nSpace utilisation 28.6%. Power utilisation 90.7%.\nThat cabinet is one third full of kit and nearly out of power, and that is a fact about your estate that no spreadsheet will ever volunteer. It is also the fact that decides whether the next order gets racked there or somewhere else, and it fell out of data you entered once, on the types.\nTwo things that will leave it reading zero I got 0.0% at first, twice, and both causes are worth knowing because neither produces an error.\nThe outlets have to reference the inlet. A power outlet on a PDU has a power_port field pointing at the upstream port on the same device. Leave it blank and the chain is broken: NetBox has no way to know that those twelve outlets are fed by that inlet, so nothing aggregates.\nLeave the inlet\u0026rsquo;s draw fields empty. This is the counter-intuitive one. NetBox only computes a power port\u0026rsquo;s draw from what is plugged into it if both its own draw fields are blank:\nif self.allocated_draw is None and self.maximum_draw is None: ...aggregate the downstream power ports... # otherwise return {\u0026#39;allocated\u0026#39;: self.allocated_draw or 0, ...} I had helpfully set maximum_draw: 7400 on the PDU inlet, because that is what the PDU is rated at. NetBox therefore believed me, took allocated_draw as unset, and reported zero. Clear both and it works it out:\nmcr1-pdu-a INPUT -\u0026gt; allocated 5340 VA, maximum 8480 VA, across 12 outlets mcr1-pdu-b INPUT -\u0026gt; allocated 5340 VA, maximum 8480 VA, across 12 outlets feed MCR1-A07-A: available 5888 VA (230 V x 32 A x 80% max utilisation) RACK power utilisation: 90.7 % RACK weight: 166.6 kg of 900 kg So the rule is: put real numbers on the leaves, and leave the intermediate ports blank so NetBox can add them up. An administratively defined value always wins over the computed one, which is correct behaviour and exactly the wrong thing to do on a PDU.\nCooling, which is new Version 4.7 added cooling to DCIM, and it arrived with the half that is getting expensive. A rack carries a cooling capability of air, hybrid or liquid and a capacity in kilowatts; a device type carries a cooling method. Above that sit cooling sources for the chillers and CRAC units, cooling feeds representing a loop out to a rack, and intake and outflow components on the devices themselves for cold plates and manifolds.\nThe cabinet above reads Hybrid, 15.00 kW, because I told the rack so. If you are taking delivery of liquid cooled kit this year, that is a model for the thing you are currently keeping in a spreadsheet.\nOne Install, Many Customers Tenancy is the reason a provider can run one NetBox rather than one per customer, and it is worth understanding what it does and, more importantly, what it does not.\nTwenty-eight of NetBox\u0026rsquo;s models carry a tenant field. Sites, locations, racks and rack reservations. Devices, cables and virtual device contexts. Prefixes, IP addresses, ranges, aggregates, VLANs, VLAN groups, VRFs, route targets, ASNs. Circuits and circuit groups. Clusters and virtual machines. Tunnels, L2VPNs, wireless LANs and links. Power feeds and cooling feeds.\nThat is every billable noun. Set it consistently and \u0026ldquo;what does this customer have\u0026rdquo; stops being an investigation.\nBut a tenant is a label, not a lock. It says the object is dedicated to that customer. It does not stop anybody who can log in from reading the lot, which is fine inside one company and no use at all the moment a customer has an account.\nTwo things follow, and the second is the one people get wrong.\nA tenant means dedicated. A router that serves only one customer gets their tenant. A firewall serving four of them does not belong to any of them, so it gets nothing, and you record the relationship on the thing you actually sold instead. Forcing a tenant onto shared kit makes every report built on tenancy quietly wrong.\nAnd access is a separate mechanism entirely.\nPermissions are where the isolation lives NetBox\u0026rsquo;s object permissions take a JSON constraint, and the constraint is a Django ORM filter. It narrows the queryset before anything is built from it, so every view, export, API call and search is narrowed with it.\nThree fields do the work. The object types it applies to, the actions it grants, and that constraint at the bottom. Everything else is bookkeeping.\nHere is the device list as an administrator.\nAnd here is the same URL, same install, signed in as the customer.\nTwo rows instead of fourteen, and look at the left-hand menu. It has collapsed to the four things that account is allowed to touch. Nobody configured that. The navigation is built from the same permissions, so a customer never sees a link to something that would refuse them.\nThe two ways it says no This is the detail worth knowing, because the two refusals mean different things and both are deliberate.\nThe customer\u0026rsquo;s token asks for It gets /api/dcim/devices/ 200, one row of two /api/dcim/devices/1/, their own 200 /api/dcim/devices/2/, somebody else\u0026rsquo;s 404 /api/tenancy/tenants/ 403 /api/dcim/sites/, never granted 403 no token at all 403 404 means the type is yours but that row is not. The constraint removed it from the queryset, so as far as the request is concerned it does not exist. A 403 there would confirm it does, and let a curious customer count your estate by walking the IDs.\n403 means the type was never yours. Tenants and sites were never granted, so they refuse outright, and the customer cannot enumerate who else is on the platform.\nThen let them write Read-only is the easy case. The real question is whether a customer with edit rights can be trusted with them, so I granted change under the same constraint and went looking for the way out.\nAttempt Result Edit their own device 200, saved Edit another customer\u0026rsquo;s device 404 Edit their own device, moving it to the other customer\u0026rsquo;s tenant 403 Delete their own device 403, delete was never granted The third row is the one that matters. Reassigning your own device to somebody else\u0026rsquo;s tenant is the obvious escape, because the object is inside your constraint when the request arrives and outside it afterwards. NetBox evaluates the constraint against the state the object would be left in, so it refuses. I read the device back rather than trusting the status code, and the tenant had not moved.\nThat is the hole most home-grown multi-tenancy leaves open, and it is normally found by a customer rather than by a test.\nVariables That Need To Inherit Here is the distinction people get wrong, and getting it right saves a lot of editing.\nA custom field is a value on one object. You set it per object, and it stays there. Good for a fact about the thing itself: an asset tag, a support contract number, a commissioning date.\nA config context is a value attached to a characteristic, which everything matching that characteristic inherits. Good for a variable that should cascade: your NTP servers, your syslog targets, your SNMP community, your DNS domain, your management VLAN, your backup window.\nIf you find yourself setting the same custom field to the same value on forty devices, you wanted a config context.\nA context is arbitrary JSON, and it can be attached to a region, site group, site, location, device type, role, platform, cluster, cluster type, cluster group, tenant group, tenant or tag. Tenant is in that list, which for a provider means a fact true of one customer everywhere follows them onto every device you ever add for them.\nThe merge is per key, and weight decides who wins each one.\nRead the right-hand panel against the left. The region supplies four keys at weight 1000. The site supplies one key at weight 2000. The rendered context keeps the region\u0026rsquo;s domain, NTP servers and community untouched, and takes the site\u0026rsquo;s syslog server, because that is the only key anything contested.\nYou write the exception, not a fresh copy of everything with the exception in it. That is the whole value, and it is why this scales where a per-device variable file does not.\nLocal context beats everything above it The inheritance stack has a top, and it is the object itself. Local context data on a device wins over every source context that applies to it, whatever their weights.\nI set this on mcr1-core-01:\n{\u0026#34;syslog_servers\u0026#34;: [\u0026#34;10.20.10.99\u0026#34;], \u0026#34;note\u0026#34;: \u0026#34;this box logs somewhere else\u0026#34;} and its rendered context became:\n{ \u0026#34;note\u0026#34;: \u0026#34;this box logs somewhere else\u0026#34;, \u0026#34;domain\u0026#34;: \u0026#34;mcr1.example.net\u0026#34;, \u0026#34;ntp_servers\u0026#34;: [\u0026#34;172.16.10.22\u0026#34;, \u0026#34;172.16.10.33\u0026#34;], \u0026#34;snmp_community\u0026#34;: \u0026#34;n0rthwest\u0026#34;, \u0026#34;syslog_servers\u0026#34;: [\u0026#34;10.20.10.99\u0026#34;] } The local syslog server beat both the site override at weight 2000 and the region at weight 1000. Everything it said nothing about still inherited, so the domain, the NTP servers and the community came through untouched, and the new key was simply added.\nThat is the escape hatch for the one box that is genuinely different, and the panel tells you plainly that it overwrites all source contexts. If you find yourself using it on many devices rather than one, you have found a characteristic those devices share and you wanted a context scoped to it.\nThree traps An unscoped context is global. Create one and forget to assign it to anything, and it applies to every device and virtual machine you have. That is documented behaviour and occasionally what you want. It is also silent.\nKeys cannot have hyphens if you want to reach them simply. Context data is JSON so ntp-servers is perfectly legal, but Jinja variable names are not JSON keys. {{ ntp-servers }} is parsed as a subtraction and raises UndefinedError: 'ntp' is undefined. Write {{ ntp_servers }} against hyphenated data and you get an empty string, HTTP 200, and no warning anywhere. Use underscores in the keys and the obvious template works.\nA profile can stop the typos. A config context profile groups related contexts and enforces a JSON schema on their data at save time. I gave one a schema requiring syslog_servers as an array of IPv4 strings, then made the two mistakes people actually make:\n{\u0026#34;syslog-server\u0026#34;: [\u0026#34;10.1.1.1\u0026#34;]} singular, by accident 400 Data does not conform to profile schema: \u0026#39;syslog-servers\u0026#39; is a required property {\u0026#34;syslog_servers\u0026#34;: [...], \u0026#34;syslog_port\u0026#34;: 99999} 400 Data does not conform to profile schema: 99999 is greater than the maximum of 65535 Rejected where somebody made them, rather than four hundred devices later.\nRendering The Configuration Context data plus a Jinja template gives you a configuration file. The device is in scope as device, its merged context is in scope as ordinary variables, and you can walk its components.\nEverything on that page came from somewhere different, and that is the point:\nLine Where it came from host-name mcr1-core-01 the device domain-name mcr1.example.net the regional config context location \u0026quot;Manchester DC1 / MCR1-A07 / U39\u0026quot; the site, the rack and the position in it server 172.16.10.22 the regional context, weight 1000 host 10.20.10.99 any notice the device\u0026rsquo;s own local context, beating both the site and the region family inet address 10.20.10.11/24 the address on that interface, with its mask as typed The template is resolved device, then role, then platform, and the request fails if none of the three has one. So you assign a template to a platform once, and every device running that software renders from it unless its role or the device itself says otherwise. That is what a platform is for: the operating system or vendor software family, not the hardware.\nNetBox renders. It does not push. Getting the output onto the box is your automation\u0026rsquo;s job, and the day a template bug would otherwise have reconfigured four hundred devices, you will be glad those are two different systems.\nPlugins A plugin is a Django application installed alongside NetBox. It can add models, add pages, extend both APIs, inject content into existing templates, add navigation, add background job queues and load further Django apps. There is very little it cannot do, because underneath it is just Django.\nInstalling one is four commands and a restart:\nsource /opt/netbox/venv/bin/activate pip install netbox-topology-views netbox-qrcode # add the package names to PLUGINS in configuration.py, then python3 manage.py migrate python3 manage.py collectstatic --no-input sudo systemctl restart netbox netbox-rq On the container stack the same packages go in plugin_requirements.txt and you rebuild the image. Either way, pin the versions and check the compatibility matrix first, because a plugin that has not caught up with a NetBox release will refuse to start the whole application rather than disabling itself.\nThe published catalogue lists 31. These are the ones worth knowing about, with the licence each one actually ships under:\nPlugin What it does Licence DNS Zones, records and name servers as a source of truth MIT BGP Sessions, communities and routing policies Apache 2.0 Topology Views Graphical topology maps built from your cables Apache 2.0 Floorplan Graphical site and location maps LGPL 3.0 QR Code Codes on racks, devices and cables, for asset labels Apache 2.0 ACLs Access lists and rules Apache 2.0 Prometheus SD Serves Prometheus its host list straight from NetBox MIT Documents Documents attached to circuits and devices Apache 2.0 Lifecycle Hardware end of life, licences and contracts Apache 2.0 Contract Contracts and invoices MIT Reorder Rack Drag and drop rack units Apache 2.0 Branching Isolated, mergeable branches of your data NetBox Limited Use Custom Objects New object types, defined in the UI NetBox Limited Use Prometheus SD is the honest shape of the whole idea. NetBox knows what exists, so let NetBox tell the monitoring system, and stop maintaining a second list of hosts that drifts.\nOne thing to know before you build on the last two rows. NetBox itself is Apache 2.0 and has been since DigitalOcean released it in 2016, and the project is stewarded today by NetBox Labs alongside a team of volunteer maintainers.21 NetBox Branching and NetBox Custom Objects are not: they ship under the NetBox Limited Use License 1.0, which grants use \u0026ldquo;only as part of a NetBox installation obtained from NetBox Labs or a NetBox distributor authorized by NetBox Labs, and only for your own internal use\u0026rdquo;, and which does not grant the right to use the software \u0026ldquo;to provide a managed service or software products that includes, integrates with, or extends NetBox in a way that competes with any product or service of NetBox Labs\u0026rdquo;. If you installed NetBox Community from GitHub and you run it on behalf of customers, read the terms yourself before you put a schema behind them.10 NetBox\u0026rsquo;s own installation guide recommends both plugins without mentioning it.11\nCustom Fields A custom field adds an attribute to an existing model. Values are stored as JSON alongside each object, so there is no migration and no restart, and there are thirteen types including object and multi-object references to other NetBox records.\nTwo fields there, grouped under an \u0026ldquo;Asset\u0026rdquo; heading that I chose. Note the Dimensions panel to the right of them: 9.5 kg, which nobody typed on this device. It came from the device type.\nCustom fields are validated, and it is worth knowing that they are, because it is a real difference from the route below. I gave the contract field a regular expression and the API enforced it:\nPATCH {\u0026#34;custom_fields\u0026#34;: {\u0026#34;support_contract\u0026#34;: \u0026#34;nonsense\u0026#34;}} {\u0026#34;__all__\u0026#34;: [\u0026#34;Invalid value for custom field \u0026#39;support_contract\u0026#39;: Value must match regex \u0026#39;^[A-Z]{2,4}-[0-9]{6}$\u0026#39;\u0026#34;]} Two things changed in 4.7 that matter at scale. Creating a field with a default value, or deleting a field, has to rewrite the stored data of every object it applies to, so on a large table that work is handed to a background job and the field reports \u0026ldquo;provisioning\u0026rdquo; or \u0026ldquo;deleting\u0026rdquo; while it runs. A field is live only while active: during either operation it does not appear on objects, in forms, in filters or in either API. That needs a worker running, or it sits in that state indefinitely.\nUse a custom field when you are adding a fact about the thing itself. Use a config context when the value should cascade. Use the next section when the thing you need does not exist.\nCustom Select Menus, And Overriding What Ships Two different mechanisms live under this heading and they solve different problems. One is for your own fields. The other rewrites NetBox\u0026rsquo;s.\nA choice set, for your own select field A custom field of type \u0026ldquo;selection\u0026rdquo; draws its options from a choice set, which is an object you manage in the UI like anything else. So \u0026ldquo;support tier\u0026rdquo; becomes a real dropdown rather than free text somebody will spell three ways.\nIt is enforced, including through the API:\nPATCH {\u0026#34;custom_fields\u0026#34;: {\u0026#34;support_tier\u0026#34;: \u0026#34;platinum\u0026#34;}} {\u0026#34;__all__\u0026#34;: [\u0026#34;Invalid value for custom field \u0026#39;support_tier\u0026#39;: Invalid choice (platinum) for choice set Support tier.\u0026#34;]} Choice sets are shared, so one set can back the same field on several models, and changing the list in one place changes it everywhere.\nFIELD_CHOICES, for NetBox\u0026rsquo;s own fields This is the one people do not know exists. Several of NetBox\u0026rsquo;s built-in choice fields can be extended or replaced from configuration.py, including device status, site status, rack status, circuit status and plenty more.12\nAppend a plus sign to add to what ships. Leave it off to replace the list outright.\nFIELD_CHOICES = { # add to what NetBox ships with \u0026#39;dcim.Device.status+\u0026#39;: ( (\u0026#39;burn-in\u0026#39;, \u0026#39;Burn-in\u0026#39;, \u0026#39;cyan\u0026#39;), {\u0026#39;value\u0026#39;: \u0026#39;awaiting-rma\u0026#39;, \u0026#39;label\u0026#39;: \u0026#39;Awaiting RMA\u0026#39;, \u0026#39;color\u0026#39;: \u0026#39;orange\u0026#39;, \u0026#39;description\u0026#39;: \u0026#39;Faulty, with the vendor\u0026#39;}, ), # replace the stock list outright \u0026#39;dcim.Site.status\u0026#39;: ( (\u0026#39;surveyed\u0026#39;, \u0026#39;Surveyed\u0026#39;, \u0026#39;purple\u0026#39;), (\u0026#39;building\u0026#39;, \u0026#39;Building out\u0026#39;, \u0026#39;orange\u0026#39;), (\u0026#39;active\u0026#39;, \u0026#39;Active\u0026#39;, \u0026#39;green\u0026#39;), (\u0026#39;closing\u0026#39;, \u0026#39;Closing\u0026#39;, \u0026#39;red\u0026#39;), ), } I put exactly that on this instance. Device status came back with the seven stock values and my two:\noffline, active, planned, staged, failed, inventory, decommissioning, burn-in, awaiting-rma and site status came back with only mine, the stock list gone:\nsurveyed, building, active, closing A choice can be a plain tuple of value, label and colour, or a dictionary which also takes a description shown as a subtitle in the form. And they behave like native values everywhere, because as far as the rest of NetBox is concerned they are native values:\nThat badge is my status, in my colour, sorting and filtering like any other.\nThree things worth knowing before you use it.\nReplacing deletes the stock values from the menu, not from the database. Any object already holding one of them keeps it, but the value is no longer offered and will not be a valid choice next time somebody edits that object. If you replace a list, check nothing is sitting on a value you just removed.\nIt lives in the configuration file, so it needs a restart and it is not something a UI user can change. For a provider that is the right way round: the set of statuses your business recognises is a governance decision, not a Tuesday afternoon one.\nExtend before you replace. The stock values are what every plugin, script and integration expects to see. Appending costs nothing. Replacing is a decision you own forever.\nCustom Objects Here is a real gap. NetBox models clusters, virtual machines and virtual disks. It does not model the datastore those disks actually live on, and for anybody running Proxmox or VMware that is the object that connects the storage you bought to the workload that uses it.\nThe Custom Objects plugin lets you define a new object type from the UI or the API, without writing any code. So:\nSeven fields, two of which are the whole point. cluster is a single object reference to a real NetBox cluster, set to protect so nobody can delete a cluster out from under its storage. provisioned_by is a multi-object reference to the devices that actually serve it.\nThat gives you a first-class object with its own navigation entry, list view, filters, import and export:\nAnd because the references are real, the relationship shows up from the other end too. Open the cluster and the datastores are listed against it.\nA custom object type inherits most of what makes a NetBox object a NetBox object: list and detail views, a navigation entry, REST endpoints, full-text search, change logging, journaling, tags, bookmarks, import and export, event rules and notifications. None of that had to be written.\nIt is also not a JSON blob pretending to be a table. The plugin issues real DDL, and what appears in PostgreSQL is a real table with real constraints:\nTable \u0026#34;public.custom_objects_2\u0026#34; Column | Type | Nullable | Default -----------------+----------+----------+------------------ id | bigint | not null | identity name | varchar | | cluster_id | bigint | | backing | varchar | | capacity_gb | bigint | | thin_provisioned| boolean | | Indexes: \u0026#34;custom_objects_2_name_key\u0026#34; UNIQUE CONSTRAINT, btree (name) Foreign-key constraints: ... FOREIGN KEY (cluster_id) REFERENCES virtualization_cluster(id) ON DELETE RESTRICT The protect I asked for became ON DELETE RESTRICT, and it works: deleting that cluster comes back 409 naming what depends on it.\nTwo things to know before you rely on it The validation is on the form, not on the API. This is the one that would catch an automation. I declared required, a regular expression and numeric bounds on the fields of a custom object type, and then wrote to them through the REST API:\nWhat I declared What I sent Result validation_regex a value matching nothing 201 Created required: true the field omitted entirely 201 Created validation_minimum: 1 a negative number 201 Created unique: true a duplicate 400, rejected on_delete_behavior: protect delete the referenced object 409, rejected The two that held are the two that became database constraints. The rest exist only on the web form, so a person typing is constrained and a nightly sync is not. Compare that with the core custom field further up, where the same regular expression was enforced through the API. Until that changes, put anything an invoice depends on behind a real constraint.\nDeleting a type drops a table. Deleting a field drops a column. That is DDL from a web form, executed by whoever has the permission, and the documentation says as much.13 Restrict who can delete these.\nWhen to write a real plugin instead If the object matters, if automation writes to it, and if you would be upset to find rubbish in it, write the model yourself. A minimal NetBox plugin is a PluginConfig, a model, a serializer, a viewset, a table, a form, a few views and a URL map, and it comes to under two hundred lines of mostly declarations. Subclassing PrimaryModel gets you tags, custom fields, change logging, journaling, export templates and ownership for free, and every constraint you put on the model is enforced everywhere, because NetBox\u0026rsquo;s own serializer machinery runs it.\nThere is a cookiecutter template and a full plugin tutorial maintained by the community. Start from those rather than from a blank directory.\nAnd it is whatever licence you choose, which nobody can change later.\nCustom Validation, And Validation Classes Everything above is about recording what is there. This is about refusing to record what should not be.\nNetBox validates every object before it writes it, and you can add rules of your own on top. There are three mechanisms, they all live in configuration.py, and between them they cover almost anything a house standard needs.\nOne: plain rules, no code A validator can be a plain mapping of field names to conditions. No Python, and it is portable between installs because it is just data.\nCUSTOM_VALIDATORS = { \u0026#39;dcim.site\u0026#39;: ( {\u0026#39;description\u0026#39;: {\u0026#39;required\u0026#39;: True}}, ), \u0026#39;dcim.device\u0026#39;: ( {\u0026#39;name\u0026#39;: {\u0026#39;regex\u0026#39;: r\u0026#39;^[a-z0-9]+-[a-z]+-([0-9]{2}|[a-z])$\u0026#39;}}, ), } The available conditions are min, max, min_length, max_length, regex, required, prohibited, eq and neq. You can reach into a related object with a dotted path, so region.name on a site is fair game, and you can match on request.user.username although the documentation quite rightly tells you to use permissions for that instead.\nBoth of those rules fire immediately:\nPOST a site with no description {\u0026#34;__all__\u0026#34;: [\u0026#34;Custom validation failed for description: [\u0026#39;This field must not be empty.\u0026#39;]\u0026#34;]} POST a device named \u0026#34;Server1\u0026#34; {\u0026#34;__all__\u0026#34;: [\u0026#34;Custom validation failed for name: [\u0026#39;Enter a valid value.\u0026#39;]\u0026#34;]} That second one is a naming standard, enforced. Not written on a wiki page, not in somebody\u0026rsquo;s head, not a thing the new starter gets told off about in a review three weeks later. The database will not take it.\nTwo: a validator class, when the rule is a sentence Plain rules check one field against a constant. Real house rules are usually conditional: this matters only when that. For those you subclass CustomValidator, override validate(), and call fail().\nI put three in a module at /opt/netbox/netbox/house_rules.py:\nfrom extras.validators import CustomValidator class BillableKitNamesItsCustomer(CustomValidator): \u0026#34;\u0026#34;\u0026#34;Anything in a role we sell has to say whose it is.\u0026#34;\u0026#34;\u0026#34; BILLABLE_ROLES = {\u0026#39;hypervisor\u0026#39;} def validate(self, instance, request): role = getattr(instance, \u0026#39;role\u0026#39;, None) if role and role.slug in self.BILLABLE_ROLES and not instance.tenant: self.fail( f\u0026#34;A {role} is billable kit, so it must name the customer it belongs to.\u0026#34;, field=\u0026#39;tenant\u0026#39;, ) class RackedDeviceNeedsAPosition(CustomValidator): \u0026#34;\u0026#34;\u0026#34;A device in a rack with no rack unit is a device nobody can find.\u0026#34;\u0026#34;\u0026#34; def validate(self, instance, request): if instance.rack and instance.position is None: if instance.device_type and instance.device_type.u_height: self.fail( \u0026#34;A device in a rack needs a rack unit. Somebody has to find it.\u0026#34;, field=\u0026#39;position\u0026#39;, ) and wired them up by dotted path:\nCUSTOM_VALIDATORS = { \u0026#39;dcim.device\u0026#39;: ( {\u0026#39;name\u0026#39;: {\u0026#39;regex\u0026#39;: r\u0026#39;^[a-z0-9]+-[a-z]+-([0-9]{2}|[a-z])$\u0026#39;}}, \u0026#39;house_rules.BillableKitNamesItsCustomer\u0026#39;, \u0026#39;house_rules.RackedDeviceNeedsAPosition\u0026#39;, ), \u0026#39;ipam.prefix\u0026#39;: ( \u0026#39;house_rules.CustomerPrefixNeedsATenant\u0026#39;, ), } Note that a model takes a tuple of validators, plain rules and classes mixed freely, and they all run. Even a single validator has to be passed as an iterable, which is an easy five minutes to lose.\nThen the rules do what they say:\nPOST a hypervisor with no tenant {\u0026#34;tenant\u0026#34;: [\u0026#34;A Hypervisor is billable kit, so it must name the customer it belongs to.\u0026#34;]} POST a device into a rack with no position {\u0026#34;position\u0026#34;: [\u0026#34;A device in a rack needs a rack unit. Somebody has to find it.\u0026#34;]} POST a prefix with role \u0026#34;customer\u0026#34; and no tenant {\u0026#34;tenant\u0026#34;: [\u0026#34;A customer prefix must be assigned to a tenant.\u0026#34;]} POST a leaf switch with no tenant accepted, because a leaf switch is not in BILLABLE_ROLES Because fail() takes a field, the message lands on the right box in the form rather than at the top of the page, which is the difference between a rule people learn and a rule people resent.\nThat is also the answer to the gap in the custom objects section. The validation there existed only on the web form. A CustomValidator runs in the model layer, so it applies to the UI, the REST API, GraphQL mutations, bulk imports and anything a script does. There is one place to write the rule and no way round it.\nThree: protection rules, for deletion CUSTOM_VALIDATORS guards writes. PROTECTION_RULES guards deletes, and it takes exactly the same two forms.\nPROTECTION_RULES = { \u0026#39;dcim.device\u0026#39;: ( {\u0026#39;status\u0026#39;: {\u0026#39;eq\u0026#39;: \u0026#39;offline\u0026#39;}}, ), } That says a device can only be deleted when it is offline. Which produces:\nDELETE an active device {\u0026#34;detail\u0026#34;: \u0026#34;Deletion is prevented by a protection rule: [\\\u0026#34;Custom validation failed for status: [\u0026#39;Ensure this value is equal to offline.\u0026#39;]\\\u0026#34;]\u0026#34;} set it to offline, then DELETE 204 No Content Two keystrokes of friction between somebody and a live device, and it is the cheapest insurance in the whole application. Make the decommissioning process the thing that unlocks the delete button.\nThe one that will catch you Existing data is not checked until you next touch it. Adding a rule does not go back and validate what is already there. It sits quietly until somebody saves an object that breaks it, and then they get an error about a decision they had nothing to do with.\nI watched it happen. mcr1-hv-06 was created before I wrote any of this, as a hypervisor with no tenant. It sat there perfectly happily. Then I edited its description:\nPATCH {\u0026#34;description\u0026#34;: \u0026#34;touching it to trigger revalidation\u0026#34;} {\u0026#34;tenant\u0026#34;: [\u0026#34;A Hypervisor is billable kit, so it must name the customer it belongs to.\u0026#34;]} The edit had nothing to do with the tenant. The rule fired anyway, because validation runs on the whole object.\nThat is correct behaviour and it is also how a new rule turns into a support ticket. Before you turn one on, run a query for the objects that would fail it and fix them first. The API makes that easy: the rule is a filter, so ask for the devices with that role and no tenant, and you have your list.\nTurn the rule on afterwards. Then it only ever catches new mistakes, which is what it is for.\nDriving It From Ansible A source of truth nothing reads will rot. The netbox.netbox collection is how most people stop that happening, and it works in both directions: NetBox tells Ansible what exists, and Ansible tells NetBox what it built.\nIt is at version 3.23.0, licensed GPL-3.0, and has been pulled from Galaxy over 13.4 million times. It carries 91 modules and one inventory plugin.14 Everything below was run against the same instance, from a container with nothing in it but ansible-core 2.21.4, pynetbox 7.8.0 and the collection.\nThe inventory is a query, not a file This is the half that pays for itself on the first afternoon. nb_inventory builds your Ansible inventory out of NetBox directly:\nplugin: netbox.netbox.nb_inventory api_endpoint: http://netbox:8080 token: \u0026#34;{{ lookup(\u0026#39;env\u0026#39;, \u0026#39;NETBOX_TOKEN\u0026#39;) }}\u0026#34; config_context: false group_by: - sites - device_roles - tenants - racks device_query_filters: - has_primary_ip: true That produces this, with nobody maintaining a host list:\n@sites_manchester-dc1: @device_roles_hypervisor: |--mcr1-core-01 |--mcr1-hv-01 |--mcr1-core-02 |--mcr1-hv-02 |--mcr1-hv-01 |--mcr1-hv-03 |--mcr1-hv-02 |--mcr1-hv-04 ... |--mcr1-hv-05 @racks_MCR1-A07: @tenants_ravenscroft-legal: |--mcr1-core-01 |--mcr1-hv-01 |--mcr1-core-02 |--mcr1-hv-02 ... @tenants_padgate-foods: @device_roles_core-router: |--mcr1-hv-03 |--mcr1-core-01 |--mcr1-hv-04 |--mcr1-core-02 @tenants_hartley-components: |--mcr1-hv-05 Look at the right-hand column. Because tenancy is set on the devices, you get a group per customer for free, so --limit tenants_ravenscroft-legal runs a play against exactly one customer\u0026rsquo;s kit. Add a device in NetBox and it is in the group on the next run. Decommission one and it is gone. Nobody edits anything.\nEach host arrives carrying what NetBox knows about it:\nansible_host \u0026#34;2001:db8:20:20::11\u0026#34; primary_ip4 \u0026#34;10.20.20.11\u0026#34; primary_ip6 \u0026#34;2001:db8:20:20::11\u0026#34; device_roles [\u0026#34;hypervisor\u0026#34;] sites [\u0026#34;manchester-dc1\u0026#34;] racks [\u0026#34;MCR1-A07\u0026#34;] tenants [\u0026#34;ravenscroft-legal\u0026#34;] device_types [\u0026#34;sys-1029u-tn10rt\u0026#34;] manufacturers [\u0026#34;supermicro\u0026#34;] status {\u0026#34;label\u0026#34;: \u0026#34;Active\u0026#34;, \u0026#34;value\u0026#34;: \u0026#34;active\u0026#34;} Twenty-two keys in total, and one of them is worth noticing before it surprises you. ansible_host is the IPv6 address, because that device has a primary v6 set. I checked it against two others that only have v4 and they came back v4, so the rule is that the plugin prefers v6 where a primary v6 exists. Which is correct, and is also the sort of thing you would rather discover now than while wondering why a play is connecting over a path you had not thought about.\nWriting back The modules are the other direction, and the one worth having is IP allocation, because NetBox knows what is free and your playbook does not.\n- name: Rack the device netbox.netbox.netbox_device: netbox_url: \u0026#34;{{ nb_url }}\u0026#34; netbox_token: \u0026#34;{{ nb_token }}\u0026#34; data: name: mcr1-hv-07 device_type: SYS-1029U-TN10RT device_role: Hypervisor site: Manchester DC1 rack: MCR1-A07 position: 14 face: Front tenant: Hartley Components state: present - name: Let NetBox pick the next free address out of the customer prefix netbox.netbox.netbox_ip_address: netbox_url: \u0026#34;{{ nb_url }}\u0026#34; netbox_token: \u0026#34;{{ nb_token }}\u0026#34; data: prefix: 10.20.22.0/24 tenant: Hartley Components assigned_object: device: mcr1-hv-07 name: eno1 state: new Note what is not in there. No IP address. You name the prefix and NetBox hands back the next free one:\nTASK [Let NetBox pick the next free address out of the customer prefix] **** changed: [localhost] \u0026#34;device : mcr1-hv-07 (created)\u0026#34; \u0026#34;interface: eno1 (created)\u0026#34; \u0026#34;address : 10.20.22.1/24 (allocated)\u0026#34; And it is there in the UI a second later, cabled to nothing yet but racked, tenanted and addressed:\nThe trap in that playbook Run it a second time without changing a line and this happens:\n\u0026#34;device : mcr1-hv-07 (already correct)\u0026#34; \u0026#34;interface: eno1 (already correct)\u0026#34; \u0026#34;address : 10.20.22.2/24 (allocated)\u0026#34; The device and the interface are idempotent. The address is not, and it is not a bug. state: present means \u0026ldquo;make it look like this\u0026rdquo;. state: new means \u0026ldquo;give me a new one\u0026rdquo;, every single time, and that is exactly what it did: two addresses on one interface after two runs.\nThat is fine when you are genuinely provisioning something new, and it will quietly eat a prefix if you put it in a job that runs nightly. Allocate once and record the result, or use state: present with the address you already hold. The module is doing what you asked. The question is whether you asked for what you meant.\nIt is not only Ansible The collection gets the attention because Ansible is where most network teams already are, but the API is the product and plenty of things speak it.\nThing What it is Licence pynetbox The Python client the collection itself uses Apache 2.0 terraform-provider-netbox Manage NetBox objects as Terraform resources, actively maintained by e-breuninger MPL 2.0 nornir_netbox NetBox as a Nornir inventory, for people doing Python rather than YAML Apache 2.0 Prometheus SD Serves Prometheus its scrape targets from NetBox MIT go-netbox A Go client, though it has not been touched since May 2025 see repository Diode NetBox Labs\u0026rsquo; own ingestion pipeline for pushing discovered data in NetBox Limited Use The Terraform one is the interesting entry for anybody already managing infrastructure that way, because it lets a single plan create the cloud resource and the NetBox record that documents it, rather than leaving the second half to somebody\u0026rsquo;s memory.15\nDiode carries the same licence as Branching and Custom Objects, so the same reading applies before you build on it.\nWhat The Afternoon Actually Buys Everything above took me a day on one machine, and most of that day was seeding data so the screens had something in them. The install is twenty minutes either way. The decisions are the part that matters and they are all made in the first hour: tenants before anything, device types before devices, the numbers on the types rather than on the kit.\nWhat you get for that is not documentation. That is not a distinction anybody makes until they have had both: documentation is a thing you write and then stop maintaining. What you get is a database that refuses to hold a contradiction: it will not let you put two things in one rack unit, or a rack in a location that belongs to another site, or a device on a type that does not exist. Every one of those refusals is an argument you are not having in six months.\nAnd it will tell you things nobody asked it. That cabinet is 28.6% full of kit and 90.7% full of power. Nobody set out to find that. It fell out of putting a wattage on a device type once, and it is the difference between racking the next order there and finding out the hard way.\nThe honest caveat is the same one every source of truth has. It is only worth what it costs you to keep it right, and the only version of this that survives a busy quarter is the one where something breaks visibly when the data is wrong. Wire your monitoring, your provisioning or your firewall rules to read from it, and a wrong entry stops being a documentation problem somebody will get to. It becomes an outage at half nine on a Tuesday, with a name on it. That sounds like a cost. It is the whole mechanism.\nNobody is going to thank you for it either. A correct record shows up as the migration that took a fortnight instead of a quarter, the audit that took an afternoon, the customer who got a straight answer on the phone. None of that appears on a report next to your name.\nDo it anyway. The alternative is a business that cannot describe itself, and a business that cannot describe itself is not being run. It is being remembered, by fewer people every year.\nNetBox, Introduction — \u0026ldquo;Today, the open source project is stewarded by NetBox Labs and a team of volunteer maintainers\u0026rdquo;, and the origin at DigitalOcean in 2015, open sourced June 2016.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNetBox, LICENSE.txt — Apache License 2.0, copyright DigitalOcean, LLC. Origin and stewardship from the introduction.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNautobot v1.0.0 release notes — \u0026ldquo;a divergent fork of NetBox 2.10\u0026rdquo;, published 26th April 2021, and the list of what it added relative to NetBox 2.10.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNetwork to Code, \u0026ldquo;Why Did Network to Code Fork NetBox?\u0026rdquo;, 25th February 2021 — the SLA and long-term support reasoning, the divergence of vision, and \u0026ldquo;the NetBox project team suggested that we should consider forking\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNetBox Labs was founded in 2023 as a spin-out from NS1 following its acquisition by IBM, co-founded by NetBox lead maintainer Jeremy Stretch, and announced a $20m Series A in April 2023: NetBox Labs, \u0026ldquo;Let\u0026rsquo;s Go: Announcing NetBox Labs\u0026rdquo; and the Series A announcement.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNetBox Labs, NetBox Enterprise — the self-managed commercial edition, with \u0026ldquo;24/7 expert assistance from the NetBox Labs team\u0026rdquo;; NetBox Cloud is the hosted offering. Specific SLA terms are not published on that page.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nStar and fork counts, release dates and licences for both projects read from the GitHub API on 30th September 2026: netbox-community/netbox and nautobot/nautobot. Nautobot\u0026rsquo;s parallel 3.2.x and 2.4.x releases both dated 28th September 2026.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nnetbox-community/netbox-docker — the community container stack, Apache 2.0.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNetBox, Installation — the supported versions table, and the v4.7 release notes for the raised PostgreSQL and Redis minimums. Version and edition read from netbox/release.yaml.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNetBox Limited Use License 1.0 — carried by netbox-custom-objects and, identically, by netbox-branching. Both quoted clauses are verbatim.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNetBox, HTTP Server installation, \u0026ldquo;What\u0026rsquo;s Next?\u0026rdquo; — \u0026ldquo;Some of the most popular plugins include\u0026rdquo; NetBox Branching, NetBox Custom Objects, NetBox DNS and NetBox BGP, with no mention of licence terms.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNetBox, Data Validation configuration — FIELD_CHOICES, and the plus-sign suffix that extends rather than replaces: \u0026ldquo;To replace the available choices, specify the app, model, and field name separated by dots \u0026hellip; To extend the available choices, append a plus sign\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nnetboxlabs/netbox-custom-objects documentation — \u0026ldquo;Deleting a Custom Object Type drops an entire database table and should be done with caution.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nnetbox.netbox on Ansible Galaxy and netbox-community/ansible_modules — version 3.23.0, GPL-3.0, 91 modules plus the nb_inventory inventory plugin. Download count read from Galaxy on 30th September 2026. Everything shown was run with ansible-core 2.21.4 and pynetbox 7.8.0.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nLicences, activity and star counts read from each project on 30th September 2026: pynetbox (Apache 2.0), terraform-provider-netbox (MPL 2.0, v6.0.0-rc.1 on the Terraform registry), nornir_netbox (Apache 2.0), netbox-plugin-prometheus-sd (MIT), go-netbox (last pushed 9th May 2025) and Diode, which carries the same NetBox Limited Use License 1.0 as Branching and Custom Objects.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/netbox/what-is-netbox/","summary":"A hands-on walk through NetBox 4.7.2 with an estate loaded into it rather than an empty demo. What each screen holds and what you do with it: rack elevations, cable traces, the IPAM tree, dual-stack interfaces, and the cabinet that turns out to be 90.7 per cent full of power while only 28.6 per cent full of kit. How to stand one up as a container stack and on a host from the packages, both run end to end. The order you have to fill it in, derived from which foreign keys are actually mandatory. How one install is carved into customers so another customer\u0026rsquo;s device returns 404 and an ungranted object type returns 403. Config contexts and the inheritance that makes them worth more than custom fields. Adding your own values to NetBox\u0026rsquo;s own status menus. Three ways to add the object your business runs on, including a datastore modelled with no code. How to write validation rules and validator classes so the database refuses what your house standard forbids. And the story of the 2021 fork that produced Nautobot, what it was for, and what has since converged. Plus driving the whole thing from the netbox.netbox Ansible collection, where the inventory is a query rather than a file and one module argument quietly allocates a new address every time it runs.","title":"NetBox: What It Holds, And How To Make It Hold Yours"},{"content":"Every name your machine looked up today, it believed. Your bank, your updates, your mail server. An answer came back off the network and nothing anywhere checked whether it came from the organisation that owns the name or from whoever got a packet in first. There is no signature on an ordinary DNS answer, nothing to verify, and nothing that would notice if there were something wrong. That was a reasonable design in 1983 and it stopped being one a long time ago.\nDNSSEC is the fix for exactly that, and it is neither new nor expensive. The zone\u0026rsquo;s owner signs their records. The zone above them publishes a hash of their key and signs that, and so on upwards, until the chain reaches a single root key your resolver was born knowing. Anything that checks the chain can tell a real answer from a forged one. It has been a standard since 2005, the root has been signed since July 2010, and the registries publish the records for nothing.1 2\nThe top of the tree is finished. I counted the live root zone for this post and 1,351 of 1,438 top level domains are signed, every one of the 1,038 generic ones among them. Then it falls off a cliff. Of the 2,390 gov.uk domains that still resolve, thirty-nine are signed, and nine of those are parish councils. MI6 managed it. The National Cyber Security Centre did not.\nThis is what a forged answer actually costs and what signing does about it, counted from the live root zone on 27 September 2026 and down through government, the banks, the distributions you install from and the certificate authorities. One command per domain, nobody\u0026rsquo;s permission needed, and every name listed so you can argue with my choices rather than my arithmetic. Underneath it is the question I actually wanted answering. Forty-two years after Mockapetris wrote DNS down, is the reason almost nobody signs that almost nobody ever understood what DNS was promising?\nWhat DNSSEC Is, And What It Is For In one sentence: DNSSEC puts a cryptographic signature on DNS answers, so a resolver can prove an answer came from the zone\u0026rsquo;s owner rather than from whoever managed to reply first.\nIt is not encryption, it is not a firewall and it is not a filter. It is a signature and a chain of keys reaching back to a single key your resolver already trusts. That is the entire idea.\nNow the attack, because the protocol only makes sense once you have seen what it is for.\nA resolver asking for a name sends a query and waits. The answer is matched on a handful of fields: the query name, the type, the source and destination addresses and ports, and a 16 bit transaction ID. Anybody who can produce a packet matching those fields, before the real answer arrives, wins. The resolver caches the forgery and hands it to every client that asks, for as long as the attacker said to.\nDan Kaminsky\u0026rsquo;s 2008 work is the one everybody half remembers.3 The bit that matters is not that cache poisoning existed, because it was known about for years before that. It is that he found a way to retry the guess indefinitely. Ask for a name that does not exist, and each failed attempt costs nothing and lets you try again immediately, so the attacker is no longer racing once against a cached record that lasts a day. The industry response was to add randomness: random source ports on top of the random transaction ID, which took the guess from one in 65,536 to something around one in two billion.\nThat is a bigger number. It is not a proof.\nThe resolver matches on fields an attacker can guess WITHOUT SIGNING: FIRST MATCHING PACKET WINS resolver asks, then waits the real server answers in its own time off-path attacker only has to arrive first It is accepted if five fields match, and every one of them is guessable: name · type · addresses · source port · 16-bit transaction ID WITH SIGNING: THE ANSWER CARRIES PROOF A signature that chains to the one root key the resolver already had. Guessing does not produce one. The resolver matches on fields an attacker can guess. Signing replaces guessing with arithmetic And the mitigation has been eroding ever since, because every one of those fields is a guess made harder rather than a fact made checkable. Port randomisation gets undone by a NAT that rewrites ports predictably. Fragmentation attacks sidestep the matching entirely. Off-path attacks keep coming back in new forms, and each one gets its own patch.\nSigning ends the argument. A signed answer either verifies against a key the resolver can chain to the root, or it does not, and no amount of guessing gets an attacker a valid signature.\nWhat Winning That Race Actually Buys Them Be concrete about it, because \u0026ldquo;cache poisoning\u0026rdquo; sounds abstract and the consequences are not.\nWhat they forge What they get The A record for your site Traffic, and a login form that looks like yours Your MX records Every inbound email, including password resets The name a certificate authority validates against A valid certificate for a domain they do not own The name your machines fetch updates from Nothing arrives, and nobody is told Your NS delegation All of the above at once, for as long as the TTL says The third row turns a DNS problem into a certificate problem. Domain-validated issuance works by asking you to publish a record and then looking it up. If an attacker can control the answer the CA sees during that lookup, the CA issues them real paper for your name, and after that the padlock in the browser is theirs. Everything a user has been trained to check will say the site is fine.\nThe fourth row is the quiet one. A forged answer does not have to send anybody anywhere. Pointing an update service at a black hole is enough, and nothing in the stack is built to shout about updates that never arrived.\nWhat Signing Does About It The mechanism is simpler than its reputation.\nEvery signed zone holds a key pair. The zone owner signs each record set with the private half and publishes an RRSIG alongside it. The public half goes in the zone as a DNSKEY. So far this proves nothing, because an attacker who can forge an A record can forge a DNSKEY to go with it.\nWhat closes it is the DS record, and the DS lives in the parent zone, not yours. It is a hash of your key, published and signed by the zone above you. So uk vouches for your key, the root vouches for uk, and your resolver was born knowing exactly one thing: the root\u0026rsquo;s key.2\nEach parent vouches for its child, and one missing DS breaks everything below THE ONE THING A RESOLVER IS BORN KNOWING the root key shipped with the resolver Everything else is learned by asking, and proved by the level above it. DS for uk uk signed, and vouched for by the root DS for your zone damiendye.uk signs its records with RRSIG A DS is a hash of the child's key, published and signed by the parent. So the parent is the only one who can put you in the chain. Not you. AND IF ONE DS IS MISSING The chain stops there. Everything below is unverifiable, however carefully the zone underneath was signed. Each parent vouches for its child, and one missing DS breaks everything below it Three record types do the work, and you want to know which is which, because only one of them is somebody else\u0026rsquo;s job:\nRecord Lives in Says DNSKEY your zone here is my public key RRSIG your zone here is my signature over this record set DS your parent\u0026rsquo;s zone I vouch for that key You can publish the first two on your own all afternoon and it changes nothing. Until the parent publishes a DS, you are a signed zone that nobody can verify, which is the same thing as an unsigned zone with extra work in it.\n$ dig +short @127.0.0.53 uk DS 43876 8 2 A107ED2AC1BD14D924173BC7E827A1... $ dig +short @127.0.0.53 damiendye.uk DS 2371 13 2 A5B2825C57899A5A15EE9703832C8358E0D19EF29DC72DD83C691ED77C33BD7F $ dig +short @127.0.0.53 bbc.co.uk DS (nothing) That is this site\u0026rsquo;s own zone in the middle. uk vouches for it, algorithm 13 is ECDSA P-256, and a validating resolver can therefore prove every answer about it back to the root. bbc.co.uk returns nothing, so it cannot.\nOne command tells you whether a zone is in the chain of trust. If the parent publishes no DS for you, you are not signed, whatever else you have configured, and a validating resolver will treat every answer about your domain as unverifiable.\nThat one record is the whole test.\nWhat It Does Not Do Worth saying plainly, because overselling it is half of why people distrust it.\nIt does not Because Encrypt anything Every name you look up is still in the clear. That is what DNS over TLS is for, and they are not substitutes Make a compromised zone safe An attacker who owns your DNS management signs their forgeries with your key, quite happily Protect the last hop Between a validating resolver and the application, unless that hop is trusted too Stop a name being taken off you A registrar or a court can still do that, and the signature will be perfectly valid throughout Say anything about the content A signed answer is authentic, not honest. Malware can sign its zone too, and does People trip over that last row. DNSSEC proves the answer came from whoever controls the zone. It has no opinion whatsoever on whether they are decent.\nIt does exactly one thing. It makes forging an answer arithmetically infeasible rather than merely unlikely.\nHow To Check Any Zone, Including Somebody Else\u0026rsquo;s This is the part that makes the rest of the post possible, and it takes one command.\nA zone is in the chain of trust if its parent publishes a DS for it. That is the whole test, and you can run it against anybody, without permission, from any machine:\ndig +short damiendye.uk DS Something back means signed. Nothing back means not signed, whatever else the zone has configured.\nIf you want more than a yes or no, there are three more worth knowing:\nCommand Tells you dig +short \u0026lt;zone\u0026gt; DS Does the parent vouch for this zone at all delv @1.1.1.1 \u0026lt;zone\u0026gt; A Does the chain validate end to end, and if not, where it breaks resolvectl query \u0026lt;zone\u0026gt; What a validating resolver concludes, with the verdict spelled out dig +dnssec \u0026lt;zone\u0026gt; SOA The RRSIG and its expiry date, which is the thing to monitor Two traps to know before you trust your own results, because both caught me while measuring this post.\ndig +short … DS will print a CNAME chain if the name is aliased, and a hostname containing digits looks enough like a DS record to fool a naive script. A real DS is four fields: key tag, algorithm, digest type, hex digest. Match that shape, or you will count unsigned zones as signed.\nA DS under an unsigned parent means nothing. The chain has to reach the root. update.microsoft.com has a DS, and microsoft.com does not, so the branch is unverifiable regardless. Always check the whole path, not one level.\nEverything counted in the rest of this post was measured this way. No scanner, no third-party dashboard, no vendor\u0026rsquo;s league table. Ask the parent whether it vouches for the child, and insist the answer parses as a DS.\nThe Root Zone Is Finished Here is where the received wisdom is simply out of date. People still talk about DNSSEC as though the infrastructure is the problem.\nI fetched the live root zone on 27 September 2026, serial 2026092701, and counted the delegations against the DS records.4\ndelegations signed share gTLDs (.com, .org, .dev, the lot) 1,038 1,038 100% IDN TLDs 151 136 90.1% ccTLDs 248 176 71.0% arpa 1 1 100% Total 1,438 1,351 93.9% Every single generic top level domain is signed. All 1,038 of them, no exceptions, because ICANN\u0026rsquo;s Registry Agreement requires it of anything delegated under the new gTLD programme. The 87 that are not signed are almost entirely country codes, and the list is mostly small territories and a handful of states:\nae ao aq ba bb bo bs cd cf cg ck cu cv cw do eg fk gb gf gh gm gp gq gt gu hm im iq jm jo kh km kn kp mh mk mo mp mq mt mv mw mz ne ni np nr om pa pf pk pn ps qa sd sl sm so st sv sy sz td tg tj tk to va vg vi ye zw gb is in there, which is a curiosity rather than a problem, since nobody uses it for owt. So are the Vatican, North Korea, Cuba and Syria. So is tk, which was for years the largest source of free domains on the internet and a reliable source of abuse with it.\nSo the top of the tree is done. The registries did the hard part, the expensive part and the part that needed international coordination, and they finished it. As such, nothing below this point can be blamed on the infrastructure.\nNinety-four per cent at the top, one and a half at the bottom THE CHAIN IS BUILT the root signed since 2010 1,351 of 1,438 TLDs every gTLD, no exceptions the name people actually type and here it stops SHARE SIGNED, COUNTED 27 SEPTEMBER 2026 TLDs in the root 93.9% Linux and BSD distros 29% UK banks 27% Package registries 18% Microsoft zones 15% Certificate authorities 11% every gov.uk domain 1.63% 39 signed, of the 2,390 in the official register that still resolve Counted, not estimated. The chain is built all the way down to the TLD and then nobody uses it And Then It Stops Dead Below the TLD, the picture inverts completely.\nI measured the same way for every set below: ask the parent for a DS, accept only a record that actually parses as one. Everything here was counted on 27 September 2026.\nWhat Signed Share TLDs in the root zone 1,351 / 1,438 93.9% Linux and BSD distributions 19 / 96 20% UK banks and building societies 13 / 98 13% Package registries and supply chain 4 / 22 18% Microsoft\u0026rsquo;s zones 4 / 26 15% Certificate authorities 1 / 9 11% Every gov.uk domain that still resolves 39 / 2,390 1.63% Ninety-four per cent at the top. One and a half per cent at the bottom. The chain of trust is a chain with one end bolted to the wall and the other end lying on the floor.\nA note on method, because a number this bad deserves one. The gov.uk row is a census and not a sample: it is every second level domain in the government\u0026rsquo;s own published register, filtered to the 2,390 that still resolve. The other rows are curated lists of the organisations whose forgery would actually hurt, which is a judgement call, and I have listed every name in them below so you can disagree with my choices rather than with my arithmetic.\nThe Government Census: 39 Of 2,390 The official list of gov.uk domains that the government publishes runs to 3,004 second level names. Of those, 2,390 still resolve. Thirty-nine are signed.\nBefore the detail, one thing about that register. The most recent version gov.uk publishes is dated 1 October 2016.5 A decade old, for the authoritative list of the domains that the British state answers on. That is its own small finding and I will leave it there.\nNow look at which thirty-nine.\nSigned Not signed mi6.gov.uk, sis.gov.uk gchq.gov.uk, mi5.gov.uk nationalcrimeagency.gov.uk ncsc.gov.uk, cyberessentials.ncsc.gov.uk cheltenham.gov.uk, cotswold.gov.uk, somerset.gov.uk, waverley.gov.uk, southribble.gov.uk, sedgemoor.gov.uk, fdean.gov.uk, westoxon.gov.uk birmingham.gov.uk, manchester.gov.uk, leeds.gov.uk, glasgow.gov.uk, sheffield.gov.uk, liverpool.gov.uk, bristol.gov.uk, cardiff.gov.uk, edinburgh.gov.uk, belfast.gov.uk peakdistrict.gov.uk, snowdonia-npa.gov.uk, eryri-npa.gov.uk hmrc.gov.uk, dvla.gov.uk, dwp.gov.uk, nhs.uk, homeoffice.gov.uk, mod.uk, parliament.uk nine parish and town councils companieshouse.gov.uk, landregistry.gov.uk, police.uk, met.police.uk, tfl.gov.uk, ons.gov.uk Read that first column again. The Secret Intelligence Service has signed its zone. The National Cyber Security Centre has not.\nNeither has GCHQ, which is the NCSC\u0026rsquo;s own parent organisation. Neither has the Cyber Essentials scheme, which exists to certify that other people\u0026rsquo;s security is adequate. I am not going to pretend that is anything other than remarkable.\nAnd nine of the thirty-nine are parish and town councils. Abinger. Aldenham. Ashmansworth. Wheathampstead. Frampton on Severn. Places with a clerk, a part time website and a budget that would not cover a day of consultancy in Whitehall. They managed it. HMRC, which holds the tax record of every adult in the country, did not.\nCan I ask what in the process allows that. Not who: what. Because the NCSC publishes guidance telling other people to deploy DNSSEC, and the department that writes the guidance has not done the thing the guidance says. Either it is important, in which case the body that says so should have done it years ago, or it is not, in which case the guidance should say that instead.\nThe excuse on offer will be scale and legacy. It does not survive the first column. somerset.gov.uk is a unitary authority serving 580,000 people and it is signed. birmingham.gov.uk is a unitary authority serving 1.1 million and it is not. Same country, same registry, same registrars, same money available for the same suppliers. One did it.\nThe Software You Update From This is the part that should worry you more than the banks, and it is the part almost nobody measures.\nEvery machine you run fetches code from somewhere on a schedule, and it finds that somewhere by asking DNS.\nNinety-nine distributions and BSDs, ninety-six of them still resolving. Nineteen are signed.\nSigned (19) debian.org, fedoraproject.org, opensuse.org, gentoo.org, almalinux.org, artixlinux.org, cachyos.org, garudalinux.org, getsol.us, linuxmint.com, q4os.org, system76.com, tails.net, whonix.org, freebsd.org, netbsd.org, hardenedbsd.org, midnightbsd.org, opnsense.org The big commercial ones ubuntu.com, canonical.com, redhat.com, suse.com, oracle.com, rockylinux.org, centos.org, eurolinux.com Ubuntu flavours kubuntu.org, xubuntu.org, lubuntu.me, ubuntustudio.com, ubuntukylin.com Arch family archlinux.org, arcolinux.com, endeavouros.com, manjaro.org, blackarch.org Independent and minimal alpinelinux.org, voidlinux.org, nixos.org, devuan.org, slackware.com, antixlinux.com, mxlinux.org, puppylinux.com, tinycorelinux.net, slitaz.org, porteus.org, funtoo.org, calculate-linux.org Desktop-focused zorin.com, elementary.io, deepin.org, uniontech.com, bodhilinux.com, peppermintos.com, sparkylinux.org, neon.kde.org, nobaraproject.org, ultramarine-linux.org, vanillaos.org, solus-project.com Security and privacy kali.org, parrotsec.org, qubes-os.org, backbox.org, pentoo.ch, trisquel.info, pureos.net, hyperbola.info, dragora.org Regional and state altlinux.org, astralinux.ru, rosa.ru, openkylin.top, openeuler.org, opencloudos.org, openanolis.cn, mageia.org, openmandriva.org, pclinuxos.com Embedded, immutable, appliance openwrt.org, dd-wrt.com, librecmc.org, raspberrypi.com, armbian.com, flatcar.org, talos.dev, bottlerocket.dev, truenas.com, pfsense.org BSD openbsd.org, dragonflybsd.org, ghostbsd.org And the two that matter most kernel.org, gnu.org Nineteen in ninety-six. And the names in that unsigned list are not obscure: Ubuntu, Red Hat, SUSE, Oracle, Arch, Alpine, NixOS, Rocky, every Ubuntu flavour, and both of kernel.org and gnu.org.\nThe security and privacy distributions are the ones I expected to be different, and mostly they are not. Tails and Whonix have signed, which fits. Kali, Parrot, Qubes, Trisquel and PureOS have not, which does not. OpenBSD has built a twenty-five year reputation on getting precisely this class of thing right and has not signed openbsd.org, while FreeBSD, NetBSD, HardenedBSD and MidnightBSD all have.\nThe package registries are close to a clean sweep of the wrong column: npm, crates.io, RubyGems, Packagist, Docker Hub and Quay are all unsigned. pypi.org is the exception, and credit where it is due, because it is also among the most attacked.\nMicrosoft, And The Update Channel Four of twenty-six Microsoft zones are signed, and they are all one corner of the estate:\nSigned Not signed live.com, outlook.com, office.com, office365.com microsoft.com, windows.com, windowsupdate.com, update.microsoft.com azure.com, azurewebsites.net, microsoftonline.com, windows.net github.com, npmjs.com, linkedin.com, visualstudio.com, xbox.com, bing.com The Outlook and Office names are signed and nothing else is, which looks like one team\u0026rsquo;s decision rather than a company\u0026rsquo;s. Note what is in the right-hand column alongside the update service: github.com and npmjs.com, two of the largest code distribution points on the internet, both Microsoft-owned, both unsigned.\n$ dig +short @127.0.0.53 windowsupdate.com DS (nothing) $ dig +short @127.0.0.53 microsoft.com DS (nothing) $ dig +short @127.0.0.53 com DS 19718 13 2 8ACBB0CD28F41250A80A491389424... com is signed, so there is nothing technical in the way. Microsoft could publish a DS this afternoon.\nThe stock answer to why that does not matter is that DNS is only one layer. Update payloads are code signed, the client checks the signature, and the traffic runs over TLS. Break the DNS and you still cannot get code to execute.\nThat answer depends entirely on the signing being sound. It has not been.\nWhen What happened to Microsoft\u0026rsquo;s signing trust 2012 Flame forged a certificate chaining to the Microsoft Root Authority using an MD5 collision against the Terminal Services licensing enrolment path, then used it on a fake update server6 2021 Netfilter became the first rootkit found carrying a WHQL signature issued by Microsoft directly, having passed the Windows Hardware Compatibility Program while talking to a command and control server7 2021 FiveSys, another WHQL-certified driver, turned out to be a rootkit that installed its own root certificate and proxied the machine\u0026rsquo;s HTTP and HTTPS traffic7 2022 Microsoft-signed malicious drivers turned up in ransomware attacks, and Microsoft revoked the signatures and suspended the developer accounts7 2023 Storm-0558 obtained a Microsoft consumer signing key from a crash dump, after compromising an engineer\u0026rsquo;s account, and forged authentication tokens accepted for enterprise email at around 25 organisations including government agencies8 Be precise about that last row, because people overstate it. The stolen key signed identity tokens, not binaries. It belongs in the table for a different reason: the custody of Microsoft\u0026rsquo;s signing material failed, undetected for two years, through a crash dump and a compromised engineer account.\nThe rows above it are the code signing failures, and they are worse. Twice in one year, Microsoft\u0026rsquo;s own hardware certification programme put a Microsoft signature on a working rootkit and shipped it. Not a forged certificate, not a stolen key. The legitimate process, signing malware, exactly as designed.\nSo the defence that lets you treat DNS as optional has been subverted by forgery, by process abuse and by key theft, over eleven years. An attacker holding a signature the machine will accept is not a thought experiment. To a state-backed actor it is a procurement problem, not a research one.\nGive that attacker a forged DNS answer and the picture is complete: signed code they control, delivered from a server the machine believes is Microsoft\u0026rsquo;s, over a connection nothing in the stack will question. That is Flame, with better key material.\nDNSSEC is the only layer in that chain that does not care whose signing key the attacker holds. It does not validate the payload, it validates where the machine was sent, and it fails independently of every certificate and signature in play. Which is precisely what you want from a second layer, and precisely why it should not be the one that is switched off.\nAnd there is a cheaper attack that needs no key at all. Forge the answer so the update service resolves nowhere useful, and the machine simply never patches. No error the user will act on, no alarm, just a fleet quietly falling behind while the dashboard says everything is fine. If I wanted an estate ready to exploit in six months, I would not push anything to it. I would just make sure nothing arrived.\nThe Banks, In Full Not a sample this time. Ninety-eight UK banks and building societies, every one of them resolving on the day, from the high street down to societies with one branch and a Victorian name.\nThirteen are signed.\nSigned (13) lloydsbank.com, halifax.co.uk, bankofscotland.co.uk, tsb.co.uk, co-operativebank.co.uk, smile.co.uk, monzo.com, aldermore.co.uk, hampshiretrustbank.co.uk, investec.com, handelsbanken.co.uk, allica.bank, weatherbys.bank High street, unsigned hsbc.co.uk, firstdirect.com, barclays.co.uk, natwest.com, rbs.co.uk, ulsterbank.co.uk, santander.co.uk, nationwide.co.uk, virginmoney.com, clydesdalebank.co.uk, metrobankonline.co.uk, bankofireland.co.uk, aibgb.co.uk, danskebank.co.uk Digital and challenger, unsigned starlingbank.com, revolut.com, chase.co.uk, marcus.co.uk, atombank.co.uk, zopa.com, tandem.co.uk, kroo.com, monese.com, cashplus.com, anna.money, mettle.co.uk Retail, unsigned tescobank.com, sainsburysbank.co.uk, marksandspencer.com, johnlewisfinance.com SME and specialist, unsigned shawbrook.co.uk, paragonbank.co.uk, oaknorth.co.uk, recognisebank.co.uk, redwoodbank.co.uk, ccbank.co.uk, closebrothers.com, unitedtrustbank.co.uk, gbbank.co.uk, cynergybank.co.uk, securetrustbank.com Private banks, unsigned coutts.com, hoaresbank.co.uk, arbuthnotlatham.co.uk, rathbones.com, brownshipley.com, butterfieldgroup.com Building societies, unsigned every single one tested: coventrybuildingsociety.co.uk, ybs.co.uk, skipton.co.uk, leedsbuildingsociety.co.uk, principality.co.uk, westbrom.co.uk, newcastle.co.uk, thenottingham.com, cumberland.co.uk, progressivebs.co.uk, saffronbs.co.uk, newburybs.co.uk, monbs.com, furnessbs.co.uk, ipswichbuildingsociety.co.uk, leekbs.co.uk, theloughborough.co.uk, mansfieldbs.co.uk, marsdenbs.co.uk, themelton.co.uk, familybuildingsociety.co.uk, penrithbs.co.uk, scottishbs.co.uk, srbs.co.uk, swansea-bs.co.uk, teachersbs.co.uk, thetipton.co.uk, thevernon.co.uk, beverleybs.co.uk, chorleybs.co.uk, dudleybuildingsociety.co.uk, esbs.co.uk, ecology.co.uk, harpendenbs.co.uk, hrbs.co.uk, darlington.co.uk, hanley.co.uk, bathbuildingsociety.co.uk Thirty-eight building societies. Not one of them signed. Those are the institutions holding the mortgage on a great many houses.\nTwo patterns in that signed column stand out. Lloyds Banking Group has signed three of its brands, lloydsbank.com, halifax.co.uk and bankofscotland.co.uk, and the Co-operative Bank has signed both of its. So the decision is taken once, at organisation level, and then applied. No per-domain hurdle in it.\nAnd Monzo has signed while Starling has not. Two challenger banks founded within a year of each other, on comparable modern infrastructure, with the same regulator and the same registrars available. One did it.\nThe One Place Where Everybody Does It Look at the last two names in the signed column: allica.bank and weatherbys.bank.\n.bank is a restricted top level domain run by fTLD Registry Services, and its security requirements make DNSSEC mandatory, alongside TLS and email authentication, with annual re-verification of every registrant.9\nThe same institutions, two regimes Signed On an ordinary domain, where DNSSEC is optional 13 of 98 On .bank, where it is mandatory and re-checked yearly both of them Two is a small number, so take it as a demonstration rather than a statistic. The requirement is still the only thing that changed.\nThat is the answer to every excuse further down this post, and it arrives before the excuses do. The same banks, the same suppliers, the same budgets and the same skills produce 13% compliance when it is optional and 100% when somebody checks annually. Nothing technical moved. Somebody just asked.\nNobody Is Guarding The Guards One certificate authority in nine.\nSigned Not signed entrust.com letsencrypt.org, digicert.com, sectigo.com, globalsign.com, identrust.com, buypass.com, zerossl.com, certum.eu These are the organisations whose entire business is proving that something is what it claims to be. They are also the organisations that validate domain control over DNS, by asking you to publish a record and then looking it up. The lookup that decides whether you get a certificate for a domain is, at eight of these nine, a lookup nobody can verify.\nNo hypothetical, that one. It is the documented shape: forge the validation lookup, get the certificate issued to you, and now you hold valid paper for a name you do not own. DNSSEC is one of the few things that makes that materially harder, and the people it would protect most have not deployed it.\nI will tell you this for nowt. If your business model is identity, and you have not signed your own zone, the argument that it is hard is not available to you.\nAnd Who Is Actually Checking? Signing is only half of it. A signed zone protects nobody unless something on the other end verifies the signature, and this is where the picture gets worse rather than better.\nThere are two places validation can happen, and they are not equivalent:\nValidating at the resolver leaves one unauthenticated hop. Validating on the device does not TWO PLACES THE PROOF CAN BE CHECKED, AND THEY ARE NOT THE SAME THING Validation at the resolver your application believes the bit one AD bit unauthenticated the resolver checks the signatures chain verified the signed zone RRSIG and DS You get the verdict, not the proof, across the one hop this whole exercise exists to distrust. Validation on the device your application checks them itself chain verified end to end, where the answer is used the signed zone RRSIG and DS Nothing in between can lie to you, because nothing in between is being asked. Nearly all the validation in the world is the top one. Windows cannot do the bottom one at any setting. One of these hands you a verdict. The other hands you the proof Nearly all the validation in the world is the first kind. Google, Cloudflare and Quad9 all validate, and between them cover an enormous number of users. That counts for something. But it means the property most people have is my resolver says this was fine.\nWhat Each Operating System Can Do Validates on the device Default How you get it Linux with systemd-resolved Yes, fully off one line in a drop-in Linux with a local unbound or knot-resolver Yes, fully n/a install it, point the stub at it FreeBSD with local_unbound Yes, fully off service local_unbound onestart macOS Ventura and later, iOS 16 and later Yes off per app or per request, in code macOS and iOS before that API only off kDNSServiceFlagsValidate, app must ask Android not offered by the platform n/a a third-party library, in your own app Fisher-Price OS (Windows) No. Cannot n/a not available at any price Two of those deserve more than a row.\nApple quietly did the work. iOS 16 and macOS Ventura added client-side DNSSEC validation, in Apple\u0026rsquo;s own words at WWDC 2022: \u0026ldquo;iOS 16 and macOS Ventura now support client side DNSSEC validation.\u0026rdquo;10 It is opt-in rather than automatic, and it is opt-in at the right granularity, so an application that cares can ask for it per session or per request:\nlet configuration = URLSessionConfiguration.default configuration.requiresDNSSECValidation = true A real validating resolver on a phone, checking signatures on the device, and almost nobody noticed it ship. Before that, mDNSResponder exposed kDNSServiceFlagsValidate for anyone willing to use the C API.\nAndroid does not offer it. The DNS resolver has been an updatable module since Android 10 and it gained DNS over TLS in Android 9, so the platform has not been standing still on DNS. But neither the AOSP resolver documentation nor the public DnsResolver API documents DNSSEC validation, and the reason libraries like MiniDNS exist and advertise bringing \u0026ldquo;DNSSEC close to your application\u0026rdquo; is that the platform does not bring it. I could not find a primary source stating flatly that it is impossible, so I will put it no higher than: not offered, and you would be writing it yourself.\nAnd It Gets Discarded Even Where It Works One more thing, from this machine, which I did not expect to find.\nThe upstream resolver here validates and says so. systemd-resolved, with DNSSEC=no, throws that away before any application sees it:\n# ---- before: stock Fedora, DNSSEC=no ---------------------------------- $ dig @192.0.2.53 damiendye.uk A | grep flags # the upstream ;; flags: qr rd ra ad $ dig @127.0.0.53 damiendye.uk A | grep flags # the local stub ;; flags: qr rd ra # ---- the fix: two lines in a drop-in ---------------------------------- $ sudo mkdir -p /etc/systemd/resolved.conf.d $ printf \u0026#39;[Resolve]\\nDNSSEC=allow-downgrade\\n\u0026#39; \\ | sudo tee /etc/systemd/resolved.conf.d/10-dnssec.conf $ sudo systemctl restart systemd-resolved # ---- after: same query, same stub, nothing else changed --------------- $ resolvectl status | grep -m1 DNSSEC= DNSSEC=allow-downgrade/supported $ dig @127.0.0.53 damiendye.uk A | grep flags ;; flags: qr rd ra ad Same name, same answer, same second. Before the drop-in the upstream did the work, set the AD bit, and the local daemon dropped it on the floor, so the default is not simply we do not validate here, it is we do not validate here, and we will not pass on the verdict of anybody who did. After it, the bit is back and this time it is ours rather than somebody else\u0026rsquo;s claim.\nresolvectl puts the verdict in words, and it separates the two that matter:\n$ resolvectl query damiendye.uk | tail -2 -- Data is authenticated: yes; Data was acquired via local or encrypted transport: no $ resolvectl query ncsc.gov.uk | tail -2 -- Data is authenticated: no; Data was acquired via local or encrypted transport: no Note what the second one is not. It is not an error. An unsigned zone comes back unauthenticated rather than bogus, resolution succeeds, and nothing anywhere complains. That is the whole problem with 1.63% restated as a command: switching validation on does not make unsigned zones fail, it makes them visible, and only to whoever goes looking.\nThe Largest Client OS Fails The Most Basic Check Now the other end of the scale, and it is not close.\nThe Windows DNS client is, in Microsoft\u0026rsquo;s own description, \u0026ldquo;security-aware\u0026rdquo; but \u0026ldquo;non-validating\u0026rdquo;. It does not perform DNSSEC validation. It cannot be made to. What it does instead is ask its configured DNS server to validate, and then look for the AD bit in the reply.11\nThree things about that, and none of them small:\nIt never checks a signature. The strongest guarantee available to the largest client operating system in the world is a single bit set by whatever machine answered. It does not even do that by default. The client only sets the DO bit and requires AD for namespaces listed in the Name Resolution Policy Table, a Group Policy object somebody has to configure deliberately. With no rule in it, the query carries no DNSSEC expectation at all. Microsoft says the bit needs IPsec to mean anything. Their documentation is explicit that because the client is non-validating and leans on the server, \u0026ldquo;IPsec is used to establish this trust relationship\u0026rdquo;. An AD bit arriving over an unauthenticated link is a claim from whoever got there first, which is the precise attack this whole technology exists to stop. So on the desktop with the largest installed base, out of the box: no validation, no AD requirement, and no authenticated channel to the resolver. The most basic check is not merely switched off. It was never implemented.\nAnd the usual defence, that client operating systems simply do not do this sort of thing, died in 2022. Apple shipped device-side validation to every iPhone and Mac in the same year Microsoft did not. The comparison is no longer desktop against server, or mobile against fixed. It is one vendor that built it against another that has not.\nThat matters more than the same gap on Linux does, because of who is affected. A Linux server is usually run by somebody who could turn validation on this afternoon and knows what a DS record is. The machines that cannot validate at all, at any setting, are the ones sat on desks in the organisations whose zones are unsigned in the tables above. systemd-resolved ships the capability switched off, which is a decision you can reverse in a line.12 The Fisher-Price OS (Windows) ships without the capability.\nThe Service That Is Held Together By DNS Everything so far has been names on the public internet. The same hole exists inside the building, and in there it is load-bearing.\nActive Directory has no fixed address for anything. A domain-joined machine does not know where its domain controller lives, so it asks DNS. The locator process queries _ldap._tcp.dc._msdcs.\u0026lt;domain\u0026gt; for the controllers and _kerberos._tcp for the KDC, and whatever comes back is what it goes and authenticates against.13\ndig +short _ldap._tcp.dc._msdcs.corp.example SRV Forge that answer and the machine takes its authentication traffic to a host you picked. Be straight about what that is and is not: Kerberos will not hand a ticket to an impostor who does not hold the keys, so this is not a domain takeover on its own. It is a position in the path, which is what the interesting attacks are built out of. Force the fallback to NTLM and relay it. Sit in the middle of traffic that was supposed to go to a controller. Or point the whole estate at nothing and watch logons stop.\nA forged locator record does not hand over the domain, it puts the attacker in the path WHERE IS MY DOMAIN CONTROLLER? THE ANSWER IS JUST A DNS RECORD Unsigned `_msdcs` zone member server _ldap._tcp.dc SRV first answer wins a host they chose now in the path NTLM fallback, relay, or no logons Kerberos will not issue tickets to an impostor, so this is not a domain takeover. It is a position. Signed zone, validating client member server checks the signature forgery discarded your controller the signed answer Rejected before anything tries to authenticate. Windows Server can sign this zone and has been able to since 2012. The client reading it still cannot validate. Not a domain takeover. A position in the path, which is what the rest gets built on Group Policy, logon scripts, mapped drives and every trust downstream of the logon are all found the same way. It is the highest value set of names most organisations own, and it is sat there unsigned.\nMicrosoft built the server half of the fix, and built it well. An AD-integrated zone can be signed, online signing of dynamic zones landed in Windows Server 2012, and because the zone lives in the directory the private signing keys replicate to the other DNS servers through AD replication itself.14 The genuinely awkward part of DNSSEC, getting keys to the machines that need them, was solved here fourteen years ago by the thing the zone was already replicating through.\nThen it stops in the same two places as everything else:\nThe two halves What Microsoft shipped What you get out of the box Signing the _msdcs zone online signing of AD-integrated dynamic zones, keys replicated by AD itself nothing, until an administrator signs it Checking the signature the non-validating client from the last section a policy-table rule plus IPsec, or the bit means nothing So the internal zone that decides which machine gets to be your domain controller is, in most estates, every bit as unsigned as the public one. The difference is that no outsider can count it, so there is no table in this post shaming anybody. Go and look at your own before you assume.\nThe Operating System Should Do This, And It Should Be On Which brings me to the part of this I want to argue rather than count.\nValidating a DNS answer is operating system work. It sits in the same class of job as keeping the clock right, carrying a store of trusted certificates and having a TLS stack, and for the same reasons: every application needs it, almost none of them should be writing it, and the check has to happen once, in a place where it can be done properly. We settled that argument for certificates a long time ago. Nobody ships a mail client with its own private opinion about root CAs.\nThree of the four platforms above can already do it. Not one of them does it out of the box:\n$ grep DNSSEC= /usr/lib/systemd/resolved.conf # Fedora 44, systemd 259.9 #DNSSEC=no That is the compiled-in default, written out commented so an administrator can see what it is. systemd\u0026rsquo;s own manual takes a different view, recommending allow-downgrade in general and true wherever the upstream can be relied on.15 The code ships, the root trust anchor ships, and the documentation ships recommending you turn it on. The default still says no.\nPlatform Who has to act before a signature gets checked systemd-resolved an administrator, once, in a drop-in file macOS 13 and later, iOS 16 and later the author of every single application Android the author of every application, using a third-party library Windows nobody can That column is the whole problem. A requirement is one lever, and the .bank count further up shows what it does. A default is the same lever with nobody having to enforce anything, because it decides the outcome for everybody who never opens the config file, and that is very nearly everybody. Optional got us 1.63% in government and 13% in banking. People do not opt into security they cannot see, and being right about the technology has never once changed that.\nThe honest objection is that validating by default breaks users when the upstream resolver is broken rather than the zone. That is what allow-downgrade is for, and it is a real compromise rather than a free one, because a downgrade is something an attacker can provoke deliberately. Even so, shipping allow-downgrade instead of a flat no would be an enormous improvement, and it would put the breakage where it belongs, on whoever is still running a resolver that cannot handle DNSSEC in 2026.\nTurning Validation On, And Which Bit Is Fedora\u0026rsquo;s Nearly all of what follows is systemd\u0026rsquo;s rather than Fedora\u0026rsquo;s, and runs the same on Debian, Ubuntu or Arch. Two things here genuinely are the distribution\u0026rsquo;s:\nUpstream systemd Fedora 44, as installed Compiled default-dnssec allow-downgrade no Main configuration file /usr/lib/systemd/resolved.conf, every default commented out for reference also ships /etc/systemd/resolved.conf holding Cache=yes, which replaces it Upstream picks allow-downgrade in meson_options.txt, which is the compromise setting, not the brave one.16 Fedora compiles it to no and then ships a second main configuration file, and since only the first file found is used, the one that documents the defaults is no longer the one in force.15 So do not edit either of them. The vendor copy is the package\u0026rsquo;s and an update will hand your change straight back; the /etc copy is a file whose other lines are silently absent.\nUse a drop-in instead, which is what the block further up does. It overrides whichever main file won, it survives package updates, and it holds only the thing you changed. Check the result with systemd-analyze cat-config systemd/resolved.conf, which prints every file in the order applied and settles what actually won.\nThree settings, and the choice between the last two is a real one:\nDNSSEC= What it does What it costs you no validates nothing, and discards the upstream\u0026rsquo;s verdict too forgery arrives silently, as it does today allow-downgrade validates, and stands down when the upstream cannot cope an attacker can provoke the stand-down deliberately yes validates, full stop, no way back your names go with the upstream the day it breaks Start on allow-downgrade, because a broken resolver then costs you nothing. Move to yes once resolvectl status has said supported for a fortnight and you know what your upstream actually is. Per-distribution detail for everywhere that is not Fedora is in the resolved post.12\nSo What Is Actually Stopping People Right. The numbers are the numbers. Why?\nFour reasons get given, and they are not all rubbish.\nThe reason How much of it is real Key management is hard Was true. Now largely automated by the DNS provider It can take your domain off the internet True, and the real one Registrar and provider support is patchy Largely solved, and easy to check before you commit No visible benefit, no compliance stick True, and probably decisive Key management was a genuine obstacle and mostly is not any more. Signing used to mean running your own key ceremonies, remembering to re-sign before signatures expired, and rolling keys by hand on a schedule you had to track yourself. That work is now done by the DNS provider on most managed platforms, and signing is a toggle. It is not nothing, but it is no longer a project.\nThe failure mode is the honest objection, and it is the only one of the four I have real sympathy for. Get DNSSEC wrong and your domain does not degrade, it disappears. Every validating resolver refuses your records, which is what it is supposed to do, and the people who cannot reach you cannot be told why by you, because telling them requires DNS.\nThree things make it worse than an ordinary outage:\nIt fails on a clock, not on a change. Signatures carry an expiry. A zone nobody has touched since Tuesday can be gone by Sunday because a re-signing job quietly stopped running. It fails for some people and not others. Only validating resolvers reject you. Your own monitoring, if it does not validate, will report the site as perfectly healthy while a growing fraction of the internet cannot reach it. It fails at the layer you use to fix things. Remote access, your status page and your own email may all be under the name that just vanished. That fear is rational, it has taken large operators off the air, and any honest case for DNSSEC has to sit with it rather than wave it away.\nBut notice the shape of it: it is a fear of an operational discipline you do not currently have, not a fear of the technology.\nUnsigned fails quietly on your users. Misconfigured fails loudly on you TWO WAYS THIS GOES WRONG, AND ONLY ONE OF THEM GETS TALKED ABOUT Unsigned A forged answer is simply believed. Nothing logs it. Nothing alerts. Lasts as long as the attacker's TTL. Your monitoring stays green. The cost lands on whoever trusted the answer. Not on you. Signed, and broken The domain disappears outright. Fails on a clock, not on a change. Only for validating resolvers, so your own checks may look fine. The cost lands on you, loudly, with your name on the incident. Which is why the second one gets a business case and the first one does not. Nobody is ever blamed for an attack that was never detected. Unsigned fails quietly, on your users. Misconfigured fails loudly, on you And notice who pays in each column. An unsigned zone that gets forged costs your customers, silently, and nobody ever files an incident because nobody ever finds out. A signed zone that expires costs you, immediately, in public, with your name on the postmortem. Both are failures. Only one of them turns up in anybody\u0026rsquo;s objectives.\nCertificates had the same problem and solved it twice over: the failure was made visible, and then it was automated. A browser warning made an expired certificate everyone\u0026rsquo;s problem, monitoring followed, and then Let\u0026rsquo;s Encrypt made renewal something a cron job did at three in the morning. None of that made certificates easier in principle. It made forgetting them harder.\nProvider support is worth ten minutes of checking rather than assuming. Some registrars still make publishing a DS a support ticket. Plenty do not.\nAnd the last one is the real answer, which the .bank registry already proved further up this post. There is no browser padlock for DNSSEC. No customer has ever chosen a bank because its zone was signed, no auditor fails you for it, and no general regulation in the UK requires it. The benefit is entirely invisible when it works, the cost of getting it wrong is an outage with your name on it, and the person carrying that outage is not the person who would get the credit.\nPut a requirement and an annual check in front of the same institutions and compliance goes from 13% to every single one. Nothing about the technology changed between those two numbers. Nothing about the budgets, the suppliers or the skills changed either. The only variable was whether anybody was going to look.\nGiven those incentives the surprising thing is not that 1.63% of government domains are signed. It is that thirty-nine of them are.\nTurning It On Without Taking Yourself Off The Air The whole risk sits in one place, so put the effort there. Signature expiry is the failure that arrives with nobody having touched anything, so alarm on it:\ndig +dnssec damiendye.uk SOA | awk \u0026#39;/RRSIG/ {print \u0026#34;sig expires\u0026#34;, $9}\u0026#39; Treat it like a certificate. Watch the date, alarm well before it, and make the renewal automatic so the alarm is a backstop rather than a workflow.\nThen turn it on at a quiet time, on something that is not your primary domain, and leave it a fortnight before you do the one that matters. If your DNS is on a managed platform the signing itself is very likely a toggle, and the only genuinely manual step is lodging the DS with your registrar.\nIs This The Same Failure As IPv6? It is the obvious comparison, so I measured both standards in the same populations on the same day. Same organisations, same people, two decisions.\nn DNSSEC signed IPv6 reachable UK banks and building societies 98 13.3% 36.7% Linux and BSD distributions 96 19.8% 67.7% Every live gov.uk domain 2,380 1.6% 31.2% The hard migration is beating the easy one, three to one and nineteen to one THE SAME ORGANISATIONS, BOTH STANDARDS, THE SAME DAY DNSSEC signed reachable over IPv6 UK banks 98 tested 13.3% 36.7% Linux and BSD 96 tested 19.8% 67.7% Every live gov.uk 2,380 domains 1.6% 31.2% IPv6 touches every router and host. DNSSEC is one record at your registrar. The same organisations, both standards, measured the same day Nineteen times further along in government, on the same domains.\nNow sit with which is which. IPv6 touches every router, every host and every application, and wants dual stack running in parallel for years. DNSSEC signing is a toggle and one record pasted at your registrar.\nThe far harder job is beating the easy one everywhere I looked. Which rules out the comfortable explanation that this is just how infrastructure standards go, slowly and grudgingly. DNSSEC is not moving slowly. It is not moving.\nOne difference belongs to DNSSEC alone. Deploy IPv6 and you get something back: reachability, no carrier-grade NAT to buy. Sign your zone and you personally get nothing. The protection lands on your users, and only on the ones behind a validating resolver. You take on a permanent outage risk on behalf of people you will never meet, and that is a harder thing to put in front of a board than any technical obstacle in this post.\nI have written the IPv6 half of this argument elsewhere and will not repeat it here.17\nIs It That We Still Do Not Understand DNS? Forty-two years since Mockapetris wrote it down in November 1983.18 Sixteen since the root was signed.\nI think the understanding problem is real, and I think it is more specific than people not knowing how DNS works. Plenty of competent engineers can describe recursion, delegation and caching perfectly well. What is missing is one step further on, and it is the step that matters.\nAlmost nobody has internalised that DNS is an authorisation system.\nIt is treated as plumbing. A lookup table. Something that turns names into numbers and belongs to whoever runs the network, filed mentally next to DHCP. And that framing is wrong in a way that quietly decides a great deal, because in practice the DNS answer is what decides which machine your traffic goes to, which server your updates come from, and which host a certificate authority believes is yours. Whoever controls the answer controls all three.\nYou can see the misunderstanding in the pattern of who has signed. Not budget, and not skill either, and once you line it up it is hard to read as anything other than a pattern of what each of them thinks DNS is:\nWho has signed What DNS is to them Registries and TLD operators, 100% of gTLDs the product itself An intelligence service an attack surface, because their threat model has forgery in it Nine parish councils a toggle their host offered, which somebody flicked A DNS provider\u0026rsquo;s own customers a default they inherited Who has not Banks, departments, vendors, CAs plumbing, and a level below the interesting problems The people closest to DNS as a thing in itself have all signed. The people who consume it as a utility have not, almost without exception, however well resourced and however security-conscious they believe themselves to be. GCHQ has not signed. Eight of nine certificate authorities have not signed. These are not organisations short of clever people or of threat models.\nThat is also why the 2008 response to Kaminsky was to make the guess harder rather than to finish deploying the thing that makes guessing irrelevant. Making the guess harder is a plumbing fix, and plumbing is how the industry had DNS filed.\nA generation learned DNS as a directory, and never revised the entry when it quietly became the thing that decides who you are talking to.\nWe Built It, And Then We Left It The registries, the operators and the standards people did the hard part. They wrote it, argued it through the IETF for a decade, signed the root in a ceremony with witnesses, and got 100% of generic top level domains signed. That is a genuine piece of collective engineering and it is finished.\nAnd then the rest of us did not do our bit, because our bit is boring, invisible, carries all the personal downside and none of the credit, and nobody is checking.\nThat is the shape of every standard nobody enforces: the cost is carried alone, and the benefit only shows up once enough others have carried it too. So the registries carried it, and the people who publish the guidance about carrying it did not.\nExcept the job here was smaller than almost any of them. The chain was already built, and paid for, by somebody else. All that was left was one record.\nAnd If This Is The Bit We Can See One last thought, and I want to be clear it is an inference rather than a measurement, because everything else in this post is counted and this is not.\nDNSSEC is about the easiest security control there is to assess. It costs nothing, the work is an afternoon, the standard has been finished for twenty years, and anybody can check it from the outside in one command without asking permission. No audit, no questionnaire, no NDA. One dig.\nSo what does 1.6% tell you about the controls you cannot see from out here?\nThe control Can an outsider check it Costs money Visible when it works A DS record at your registrar Yes, one command no no MFA on the accounts that matter no yes no Backups restored this year, not merely taken no yes no Network segmentation no yes no Every row below the first is harder than signing a zone, costs real money, needs somebody to own it, and shares the exact property that sank DNSSEC: invisible when it works, and nobody outside checking. The only thing separating the top row from the rest is that you can check it, for free, about anybody, right now.\nIf an organisation has not done the free thing that takes an afternoon and can be verified by a stranger in one command, I am not inclined to assume it has done the expensive things that take a programme and can only be verified by someone they let in.\nNow, the obvious move is to test that against the breach record, and I tried. Of 23 UK organisations with documented major incidents, 22 are unsigned.\nThat number proves nothing and I am not going to pretend otherwise. At a base rate of 1.6%, one signed organisation in a list of 23 is precisely what chance predicts. Worse, the signed set is nine parish councils and three national parks while the unsigned set is every large department and every major city, so size drives both who gets attacked and who ends up in the news. Any comparison of breach rates between the two groups would be measuring how big an organisation is, not whether it signed.\nSo no, I cannot show you that signed organisations get breached less. Nobody can, not with data anybody can get, and anyone telling you different is kidding.\nThe Organisations That Got Hit Wrote Down What Failed You do not need my inference, because they published it themselves, and what failed is the list above.\nThe British Library is the best of them, because they wrote it down voluntarily and in detail after their 2023 ransomware attack. Their own review names the causes: entry most likely through a third-party account on a Terminal Services server with no multi-factor authentication, then legacy infrastructure and limited network segmentation letting the attackers move across the estate, with the growing complexity of third-party access having been flagged as a risk internally in 2022 and still there a year later.19 The ICO reached the same conclusions.20\nThree rows of the table above, then, confirmed from the inside by the organisation itself rather than inferred by me from out here.\nAnd before anybody uses that as a stick, the British Library deserves the opposite, and I will be emphatic about why.\nThey were open. Almost nobody else is. There was no obligation on them to write a word of it down. The standard playbook after an incident is to say as little as the law allows, put it through a communications team, decline to confirm anything specific, and wait for the news cycle to move on. That is what most of the organisations in the unsigned columns above have done when it was their turn, and it is why writing this section at all depends on one institution\u0026rsquo;s decision to behave differently.\nInstead they published an eighteen page review naming their own failures, in public, so that other institutions could learn from them. That is the behaviour you would want from every organisation in this post and get from almost none of them. The reason I can show you what actually fails inside a breached organisation is that the British Library chose to tell you.\nAnd there is a second reason the rest stay quiet, which is worse than a communications strategy. Some of them are not there any more.\nKNP Logistics had been moving freight as Knights of Old since 1865. In June 2023 the Akira group got in, encrypted the company and asked for about five million pounds. By September the group was insolvent and 730 people were out of work.21 A hundred and fifty-eight years, gone in fourteen weeks, and nobody there is writing up any lessons for you.\nSo when the unsigned columns above look quiet, that quiet is made of three different things: organisations that have not been hit yet, organisations that were hit and said as little as the law allows, and organisations that were hit and are gone. Only the first group has time left to act.\nWhich makes the next bit the honest test of my own argument rather than a cheap shot. I checked their zone:\n$ dig +short bl.uk DS (nothing) $ dig +short bl.uk DNSKEY (nothing) bl.uk is unsigned. So is britishlibrary.co.uk. Two years after a ransomware attack that shut the institution down for months, after a public review, after an ICO finding, and after the most thorough round of security attention any organisation ever gets, the free control that takes an afternoon and that a stranger can verify in one command is still not done.\nI do not read that as negligence, and I do not think it makes them worse than the unsigned organisations that have published nothing. I read it as the strongest evidence in this post for what the whole thing has been about. If DNSSEC does not get done here, at an organisation that has been through the fire, written the lessons up and had the regulator go over it, then it is not getting skipped because people are careless. It gets skipped because nothing and nobody ever puts it on the list.\nThat is the honest version of the argument. Not unsigned zones cause breaches, which is unprovable and probably false. Rather: the controls nobody outside can see are, on the published evidence of the organisations who have been through it, just as neglected as the one control everybody outside can see. DNSSEC is not the cause. It is the sample you are allowed to take.\nWhich is why the measurement is worth taking at all. Not because an unsigned zone is the end of the world on its own, but because it is one of the very few security properties an outsider can check honestly, for free, about anybody, without being let in. Treat it as a smoke alarm rather than a verdict, and then go and ask the harder questions of whoever sets it off.\nThe Only Part You Control Which leaves the only part any of us actually controls. Your own zone. Not the NCSC\u0026rsquo;s, not Microsoft\u0026rsquo;s, not your bank\u0026rsquo;s.\nGo and ask your parent whether it vouches for you:\ndig +short yourdomain.uk DS If that comes back empty, you are not signed, and on most managed DNS in 2026 the fix is a toggle and a DS record at your registrar. Turn it on, then watch the expiry the way you already watch your certificates, because that is the discipline the whole thing actually needs.\nThat is the whole job. One record, and a date in your monitoring.\nSo please, do at least this one. Not because a regulator is coming, because for most of you one is not, and not because anybody will thank you, because they will not. Do it because nobody is coming, and a standard you keep when nobody is checking is the only kind that was ever worth anything.\nThirty-nine parish clerks and a spook managed it. Frame thissen and get it done.\nRFC 4033 — \u0026ldquo;DNS Security Introduction and Requirements\u0026rdquo;, Arends et al., March 2005. The current DNSSEC specification, alongside RFC 4034 and RFC 4035.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nIANA root trust anchors — the XML IANA publishes. The first key digest carries validFrom=\u0026quot;2010-07-15\u0026quot;, the date the root was signed; the current KSK, key tag 20326, carries validFrom=\u0026quot;2017-02-02\u0026quot;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCERT VU#800113 — \u0026ldquo;Multiple DNS implementations vulnerable to cache poisoning\u0026rdquo;, the 2008 advisory covering the Kaminsky technique, and the source of the coordinated source-port randomisation response.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nThe root zone file — fetched 27 September 2026, SOA serial 2026092701. The counts in this post come from parsing the NS delegations and DS records in that file directly.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nList of gov.uk domain names — the government\u0026rsquo;s own register. The most recent file published is dated 1 October 2016 and lists 3,004 second-level domains.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft Security Response Center — Flame malware collision attack explained — \u0026ldquo;An attacker took advantage of the Terminal Services licensing system\u0026rsquo;s enrollment process for certificates that chained up to the Microsoft Root Authority which did not require internal access to Microsoft PKI\u0026rdquo;; the forged certificate \u0026ldquo;could be used to sign code that chained up to the Microsoft Root Authority and worked on all versions of Windows\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nSentinelOne — Driving Through Defenses: targeted attacks leverage signed malicious Microsoft drivers, and Bitdefender\u0026rsquo;s FiveSys analysis — malicious kernel drivers carrying signatures issued directly by Microsoft through the Windows Hardware Compatibility Program, including drivers later used in ransomware attacks.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft Security Response Center — Results of major technical investigations for Storm-0558 key acquisition — a consumer signing key leaked into a crash dump through a race condition, taken after an engineer\u0026rsquo;s corporate account was compromised, and used to forge tokens that the mail system wrongly accepted for enterprise accounts.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nfTLD Registry Services security requirements — the .bank and .insurance registry. \u0026ldquo;.BANK domain names must be signed with DNSSEC with strong cryptographic algorithms\u0026rdquo;, alongside mandatory TLS and email authentication, with annual re-verification of every registrant.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nApple, WWDC 2022 session 10079, \u0026ldquo;Improve DNS security for apps and servers\u0026rdquo; — \u0026ldquo;iOS 16 and macOS Ventura now support client side DNSSEC validation\u0026rdquo;, opted into per session or per request with requiresDNSSECValidation on URLSessionConfiguration, URLRequest or NWParameters.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft Learn — Understanding DNSSEC in Windows — the Windows DNS client \u0026ldquo;is non-validating, which means it does not perform DNSSEC validation and relies on its local DNS servers\u0026rdquo;; the AD bit expectation is driven by the Name Resolution Policy Table, and \u0026ldquo;IPsec is used to establish this trust relationship\u0026rdquo; with the DNS server.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nResolved: The Resolver You Are Already Running — the measured state of systemd-resolved, including why every mainstream distribution ships DNSSEC=no at compile time and how to change it.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMS-ADTS: DNS-Based Discovery — the Active Directory protocol specification for locating a domain controller, including the _ldap._tcp.dc._msdcs SRV query a client issues to find the controllers for a naming context.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft Learn — Sign DNS zones with DNSSEC on Windows Server and What is DNSSEC on DNS Server in Windows Server? — zone signing arrived in Windows Server 2008 R2 but barred dynamic updates, and Windows Server 2012 added online signing of dynamic zones. For an Active Directory integrated zone the private signing keys replicate to the other primary DNS servers through Active Directory replication.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nresolved.conf(5), systemd 259.9 as shipped in Fedora 44 — the manual recommends allow-downgrade, and true on systems where the upstream resolver can be relied on, while the packaged default in /usr/lib/systemd/resolved.conf is DNSSEC=no.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nsystemd meson_options.txt — option('default-dnssec', type : 'combo', choices : ['yes', 'allow-downgrade', 'no'], value : 'allow-downgrade'). Upstream\u0026rsquo;s chosen default is allow-downgrade; the #DNSSEC=no in Fedora\u0026rsquo;s vendor file is what that build set instead.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nWe Never Ran Out Of Addresses — the IPv6 version of this argument, including the 463 UK organisations holding IPv6 allocations who announce no IPv6 at all, and the point about a CGNAT having a purchase order while doing it properly has none.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 882 — \u0026ldquo;Domain Names: Concepts and Facilities\u0026rdquo;, P. Mockapetris, November 1983. The original specification, superseded by RFC 1034 and RFC 1035 in 1987.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nBritish Library, \u0026ldquo;Learning Lessons from the Cyber-Attack\u0026rdquo;, 8 March 2024 — the Library\u0026rsquo;s own review of the October 2023 ransomware attack, naming the absence of multi-factor authentication on the account used for entry, legacy infrastructure, limited network segmentation, and third-party access complexity logged as a risk in 2022.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nICO statement on the British Library\u0026rsquo;s 2023 ransomware attack, April 2025.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nThe Record — UK logistics firm blames ransomware attack for insolvency, 730 redundancies — KNP Logistics Group, parent of the 158-year-old Knights of Old, attacked by Akira in June 2023 after an employee password was brute-forced with no multi-factor authentication in place, and insolvent by September.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/dns/dnssec-the-root-is-signed-you-are-not/","summary":"The hard part of DNSSEC was finished years ago. Counted from the live root zone on 27 September 2026, 1,351 of 1,438 top level domains carry a DS record and every one of the 1,038 gTLDs is signed. Then it stops dead. A census of every gov.uk domain in the official register finds 39 signed out of 2,390 that still resolve, nine of them parish councils, while HMRC, the NHS, GCHQ and the National Cyber Security Centre are not among them. One certificate authority in nine has signed. So has one Linux distribution in three. windowsupdate.com has no DS at all. This is what a forged answer actually costs, what signing does about it, why the usual excuses do not survive contact with the numbers, and whether forty-two years after Mockapetris the real problem is that almost nobody understands what DNS actually promises.","title":"DNSSEC: Protecting Your Traffic From Forgery"},{"content":"There is a DNS resolver running on your machine right now. You did not install it, you have probably never configured it, and it is answering every name lookup the box makes. On Fedora, Ubuntu and most desktop Linux it is systemd-resolved, and it has been sat there quietly since the day the system was installed.\nIt can do three things worth having. It caches, so the same lookup does not cross the network twice. It validates DNSSEC, so a forged answer gets rejected instead of believed. And it speaks DNS over TLS, so the local network cannot read every name you ask for.\nBy default it does exactly one of those. The other two are switched off, and on some distributions they are switched off at compile time, which means the setting you would change is not even the setting that decided it. A validator that never validates is neither use nor ornament.\nThis is what the thing actually does, measured on a running Fedora 44 box with systemd 259, and what changes when you turn the other two on. Including what your distribution decided for you at compile time, and how much of that you can simply overrule.\nWhat Is Actually Answering Your Lookups Start with the honest picture, because \u0026ldquo;it uses /etc/resolv.conf\u0026rdquo; has not been true on these systems for years.\nsystemd-resolved exposes itself four different ways, and which one a program uses decides what it gets back.1 The glibc path is nss-resolve, wired in through /etc/nsswitch.conf. The native paths are D-Bus and Varlink, which carry the DNSSEC verdict and the interface scope that getaddrinfo has no way to express. Then the stub listener, a real DNS server on the loopback for anything that speaks raw DNS and knows nothing about any of the above.\nFour doors, one daemon.\nOn this machine that wiring looks like this:\n$ ls -l /etc/resolv.conf lrwxrwxrwx. 1 root root 39 May 30 13:56 /etc/resolv.conf -\u0026gt; ../run/systemd/resolve/stub-resolv.conf $ cat /etc/resolv.conf nameserver 127.0.0.53 options edns0 trust-ad search damiendye.uk $ grep ^hosts: /etc/nsswitch.conf hosts: files myhostname mdns4_minimal [NOTFOUND=return] resolve [!UNAVAIL=return] dns resolve in that line is nss-resolve. dns after it is the traditional nss-dns, sat there as a fallback that only fires if resolved is not running at all.\nFour ways in, one daemon, and a routing decision before anything leaves the box HOW A PROGRAM ASKS WHAT THE DAEMON DOES WHERE IT GOES nss-resolve glibc, no verdict D-Bus, Varlink native, full verdict 127.0.0.53 the full stub 127.0.0.54 the proxy stub systemd-resolved 1. hosts and synthetic 2. cache 3. DNSSEC validator 4. routing 5. transport the proxy stub skips 2 and 3 LLMNR on 5355 upstream DNS Four ways in, one daemon, and a routing decision before anything leaves the box There Are Two Stub Listeners, Not One Everybody knows about 127.0.0.53. Fewer people know there is a second one.\n$ ss -lntup | grep \u0026#39;:53 \u0026#39; udp UNCONN 0 0 127.0.0.54:53 0.0.0.0:* udp UNCONN 0 0 127.0.0.53%lo:53 0.0.0.0:* tcp LISTEN 0 4096 127.0.0.54:53 0.0.0.0:* tcp LISTEN 0 4096 127.0.0.53%lo:53 0.0.0.0:* They are not two addresses for the same thing. The second one deliberately does less:2\n127.0.0.53 127.0.0.54 Role The full local resolver Proxy mode only Cache Yes No DNSSEC validation Yes, when enabled Never LLMNR and Multicast DNS Yes No Synthetic names (localhost, _gateway) Yes No Upgrades to DNS over TLS Yes Yes Synthetic name for it _localdnsstub _localdnsproxy You get back What resolved decided What the upstream said The documentation is blunt about the second column. It will \u0026ldquo;pass most DNS messages relatively unmodified to the current upstream DNS servers and back, but not try to process the messages locally, and hence does not validate DNSSEC, or offer up LLMNR/MulticastDNS\u0026rdquo;.2\nThat makes 127.0.0.54 the right target for a program that wants to do its own validation, or one that needs the raw upstream response rather than resolved\u0026rsquo;s interpretation of it.\nBoth addresses have synthetic names, _localdnsstub and _localdnsproxy, which resolve without any configuration at all.\nFour Modes For /etc/resolv.conf, Not Three The mode is detected automatically from what the file is, and there are four of them.1\n/etc/resolv.conf is What it means Clients bypassing NSS symlink to run/systemd/resolve/stub-resolv.conf The recommended mode. Lists 127.0.0.53 and the live search domains Go through resolved symlink to /usr/lib/systemd/resolv.conf Static, lists 127.0.0.53, carries no search domains Go through resolved symlink to run/systemd/resolve/resolv.conf Lists the real upstream servers, kept current Bypass resolved entirely a real file managed by something else resolved reads it as a consumer, not a provider Bypass resolved entirely That third row catches people out. It looks like the tidy option, it is kept up to date, and it quietly means every program that reads resolv.conf directly is talking to your ISP\u0026rsquo;s resolver with no cache, no validation and no encryption, whatever you have configured in resolved.conf.\nThe trust-ad option in the file above matters too. Without it, glibc strips the AD bit off the answer before your program sees it, on the reasonable grounds that a random resolver\u0026rsquo;s claim to have validated something is worth nothing. With 127.0.0.53 as the nameserver the claim is coming from your own machine, so trust-ad is correct here and resolved writes it in for you.\nOn The systemd Conspiracy Theories Before any of the detail, because this is the subject where it always turns up.\nIf your objection to what follows is a theory about Lennart Poettering personally, or about Red Hat\u0026rsquo;s motives, or about systemd being a plot to take something off you, you can sling your hook, because I am not interested in hearing your fiction.\nEverything in this post came off a live machine: the source, the build flags, the shipped binary, the man pages and the measured behaviour. Every claim has a command next to it that you can run yourself and check. The build flags are public. The source is public. The one default I think is genuinely wrong was chosen by people who wrote their reasoning down where anybody can read it and argue with them, which is exactly what I do further down.\nThat is a good deal more transparency than you get from most of the software you run without a word of complaint.\nBring evidence or give over.\nThe Cache Is The Part That Just Works This is the one feature that is on by default everywhere, and it is the one that earns its keep with no configuration at all.\nThe measurement, on this box, over 32 domains. Flush the cache, query the lot, query them again, and read the query time dig reports rather than timing the process:\n$ resolvectl flush-caches $ for n in $NAMES; do dig +tries=1 @127.0.0.53 \u0026#34;$n\u0026#34; A | grep \u0026#39;Query time\u0026#39;; done median mean p90 max total for 32 names cold cache 103.5 ms 106.7 ms 184 ms 238 ms 3,414 ms warm cache 0.0 ms 0.5 ms 1 ms 3 ms 16 ms 3,414 milliseconds against 16. That is the whole argument for a local cache, and it is why this is the default.\nWorth being honest about what that number is, though. It is the saving on a run of thirty-two names that had never been looked up, against the same run repeated, and no real workload looks like either. The useful figure is the tail rather than the median: the worst lookup in that set cost 238 ms cold and 3 ms warm. A page that pulls in eight hostnames does not care about your median, it waits for your slowest one.\nThe Knobs [Resolve] Cache=yes # yes | no-negative | no CacheFromLocalhost=no # default: do not cache answers from 127.0.0.1 StaleRetentionSec=0 # serve expired records when upstream is down Setting Default What it does Cache=yes on Cache positive and negative answers Cache=no-negative Positive answers only, for when you are tired of waiting out a negative TTL CacheFromLocalhost=no on Do not cache at all when the upstream is on 127.0.0.1 StaleRetentionSec=0 off Serve expired records when the upstream stops answering DNSCacheSize=4096 4096 Records held per scope. systemd 261 and later only That third row is the one that bites. If your upstream is a dnsmasq or an unbound on loopback, resolved will not cache its answers at all, on the grounds that the thing it is talking to is already a cache. Point resolved at a filtering resolver on another host and you get two layers of caching. Point it at one on loopback and you get one.\nStaleRetentionSec= is the interesting one, and it is off by default. Set it, and when the upstream stops answering, resolved keeps serving records past their TTL rather than failing.2 It always tries the upstream first. It does not apply to NXDOMAIN, because a name not existing is a perfectly valid answer and there is nowt stale about it. For a laptop that goes in and out of coverage, or a machine that has to keep working through a DNS outage, this is worth having:\n[Resolve] StaleRetentionSec=1d systemd 261 added per-protocol cache sizing on top of this, with DNSCacheSize=, MulticastDNSCacheSize= and LLMNRCacheSize=, each defaulting to 4096 records and capped at 2^24.3 Not on this box, which runs 259, and not on any current stable distribution release. Worth knowing it is coming, because until now the cache size was not tunable at all.\nThe daemon also flushes everything on memory pressure, which is sensible and occasionally surprising when you are trying to work out why a cache hit rate looks poor.\nSplit DNS Is The Reason To Keep It If you take one thing from this, take this section, because it is the feature that genuinely does something the old stub resolver could not.\nThe traditional resolv.conf has one list of nameservers for the whole machine. One list. Bring up a VPN and something has to overwrite it, which means either your internal names work and the rest of DNS goes through the corporate resolver, or the reverse. There is no third option. The file cannot express one.\nresolved routes per-query, per-interface.1 Every link has its own servers and its own domains, and a lookup is sent to the link whose domain best matches the name, by label count.\nConfiguration Written as Effect Search domain damiendye.uk Suffix for single-label names, and routes matching queries to this link Route-only domain ~internal.example Routes matching queries to this link, never used as a suffix Catch-all route ~. Send everything not matched elsewhere to this link Default route DNSDefaultRoute=yes Take unmatched queries, without claiming ~. So a laptop with a work VPN up gets this:\nresolvectl domain tun0 \u0026#39;~corp.example\u0026#39; \u0026#39;~10.in-addr.arpa\u0026#39; resolvectl dns tun0 10.0.0.53 Names under corp.example and reverse lookups for 10.0.0.0/8 go down the tunnel. Everything else keeps going out of the local link as it did before. Nobody\u0026rsquo;s resolv.conf got rewritten, and when the tunnel drops the routes go with it.\nOne query, two links, and a routing decision made on label count THE QUERY THE DECISION THE LINK db01.corp.example 14.0.10.in-addr.arpa blogs.damiendye.uk Most labels wins Each link has its own servers and domains. tun0 ~corp.example wlp4s0 DefaultRoute A tilde routes only. No tilde also suffixes single-label names. One query, two links, and a routing decision made on label count One rule is worth stating plainly, because it is the one people reach for and get wrong. ~. on a link means prefer this link for everything. It also implicitly stops any other link being a default route. Want a link to take the leftovers without claiming the whole namespace? Set DNSDefaultRoute=yes and leave ~. alone.\nCheck what it decided rather than assuming:\n$ resolvectl status Link 3 (wlp4s0) Current Scopes: DNS LLMNR/IPv4 LLMNR/IPv6 Protocols: +DefaultRoute LLMNR=resolve -mDNS -DNSOverTLS DNSSEC=no/unsupported Current DNS Server: 192.0.2.53 DNS Servers: 192.0.2.53 2001:db8:1::53 DNS Domain: damiendye.uk Default Route: yes Every address in this post is from the documentation ranges, 192.0.2.0/24 and 2001:db8::/32, so do not paste them into a config and expect an answer.4 5 The figures and the behaviour are real, off a live machine. The addresses are stand-ins, because a global IPv6 address built the usual way carries the interface\u0026rsquo;s MAC in its bottom half, and publishing one hands out a piece of hardware inventory along with it.\nThat -DNSOverTLS and that DNSSEC=no are the next two sections.\nIt Can Validate DNSSEC The daemon is described upstream as \u0026ldquo;a caching and validating DNS/DNSSEC stub resolver\u0026rdquo;.1 The validating half is real, it is not a wrapper round something else, and it works. It is just not switched on.\nTurning it on per link needs no sudo, because resolvectl goes through polkit and an active local session is allowed:\n$ resolvectl dnssec wlp4s0 yes $ resolvectl status wlp4s0 | grep DNSSEC DNSSEC=yes/supported That /supported suffix is resolved reporting what it found when it probed the upstream, separately from what you asked for. yes/supported means you asked for validation and the server can carry it. no/unsupported on a default install means nobody asked, so nobody probed.\nOnce it is on, every lookup comes back with a verdict, and there are three of them.\nSecure, insecure and bogus are three different answers DOES THE PARENT PUBLISH A DS, AND DO THE SIGNATURES VERIFY? SECURE damiendye.uk Signed, and the chain checks out. Data returned. INSECURE systemd.io No DS. Not signed, so nothing to check. Data returned. BOGUS dnssec-failed.org Claims to be signed. The proof fails. Lookup refused. Two of the three return data. Turning validation on does not stop unsigned sites working. Secure, insecure and bogus are three different answers, and only one of them is a failure Here are all three on this machine, against real names:\n$ resolvectl query damiendye.uk -- Data is authenticated: yes; Data was acquired via local or encrypted transport: no $ resolvectl query systemd.io -- Data is authenticated: no; Data was acquired via local or encrypted transport: no $ resolvectl query dnssec-failed.org dnssec-failed.org: resolve call failed: DNSSEC validation failed: missing-key (DNSKEY Missing: no SEP matching the DS found for dnssec-failed.org.) damiendye.uk is signed, so the chain from the root validates and the data is authenticated. systemd.io is not signed, so there is nothing to check and resolved says so honestly rather than pretending. That is the systemd project\u0026rsquo;s own domain, which is a point I will come back to. dnssec-failed.org is published with deliberately broken signatures. That one fails hard, with a diagnostic naming the exact problem.\nHere is the distinction that gets lost. Insecure is not a failure. An unsigned zone returns data and tells you it could not be verified. Only a zone that claims to be signed and then cannot prove it gets rejected.\nThe chain itself is visible if you want to see it:\n$ dig +short @127.0.0.53 uk DS 43876 8 2 A107ED2AC1BD14D924173BC7E827A1... $ dig +short @127.0.0.53 damiendye.uk DS 2371 13 2 A5B2825C57899A5A15EE9703832C8358E0D19EF29DC72DD83C691ED77C33BD7F $ dig +short @127.0.0.53 damiendye.uk DNSKEY 256 3 13 oJMRESz5E4gYzS/q6XDrvU1qMPYIjCWz... 257 3 13 mdsswUyr3DPW132mOi8V9xESWE8jTo0d... Root vouches for uk, uk vouches for damiendye.uk, and the 257 key signs the 256 key that signs the records. Algorithm 13 there is ECDSA P-256, which is what you want on a new zone rather than the RSA that the uk DS above is still using.\nThrough the plain stub, a program that never heard of resolved gets the same verdict as an AD flag:\n$ dig @127.0.0.53 ietf.org A | grep flags ;; flags: qr rd ra ad; $ dig @127.0.0.53 dnssec-failed.org A | grep status ;; -\u0026gt;\u0026gt;HEADER\u0026lt;\u0026lt;- status: SERVFAIL And the daemon keeps a running tally, which is the quickest way to see whether validation is doing anything:\n$ resolvectl statistics DNSSEC Verdicts Secure: 134 Insecure: 62 Bogus: 0 Indeterminate: 2 What Validation Costs Here I am going to disappoint you, because I tried to measure this properly and could not.\nWarm, it is clean and it is nothing. The same 32 names against a full cache cost 16 ms without validation and 23 ms with it. Call it a rounding error.\nCold is where the cost should show, and a single client cannot isolate it. I ran the sweep both ways and got contradictory answers: on one ordering validation looked 74% slower, on another it looked faster, which is impossible and tells you what is really being measured. Whichever sweep runs second benefits from the upstream resolver having already fetched everything the first one asked for. The variable that dominates is not validation, it is whose cache was warm.\nSo I will not give you a number I cannot stand behind. What is true is the mechanism, and the documentation states the shape of it plainly: validation \u0026ldquo;requires retrieval of additional DNS data, and thus results in a small DNS lookup time penalty\u0026rdquo;.2 A cold validated lookup walks the delegation chain fetching DS and DNSKEY at every level before it can answer, so it costs extra round trips on names nobody has asked for yet, and nothing at all on names anybody has.\nWhich is also why the same page warns that turning the cache off \u0026ldquo;comes at a performance penalty, which is particularly high when DNSSEC is used\u0026rdquo;.2 The two features are not independent. Validation is affordable precisely because the cache means you pay for it once.\nIf you want the real figure for your own network, measure it there. One client on one connection cannot tell you.\nSo Why Is Validation Switched Off Because your distribution turned it off when it built the package, and it did that for a reason it has written down.\nUpstream systemd ships validation on. The build option says so:\noption(\u0026#39;default-dnssec\u0026#39;, type : \u0026#39;combo\u0026#39;, choices : [\u0026#39;yes\u0026#39;, \u0026#39;allow-downgrade\u0026#39;, \u0026#39;no\u0026#39;], value : \u0026#39;allow-downgrade\u0026#39;) and the upstream manual agrees: DNSSEC= \u0026ldquo;Defaults to allow-downgrade\u0026rdquo;.6\nNow look at what Fedora passes to that build:7\n-Ddefault-dnssec=no -Ddefault-dns-over-tls=no -Ddefault-mdns=no -Ddefault-llmnr=resolve And what Debian and Ubuntu pass:8\n-Ddefault-dnssec=no -Ddefault-llmnr=no -Ddefault-mdns=no -Ddns-over-tls=openssl So the file at /usr/lib/systemd/resolved.conf on a Fedora box, the one headed \u0026ldquo;Entries in this file show the compile time defaults\u0026rdquo;, reads #DNSSEC=no. Read that again, because it is doing something sly: the line looks like the software telling you its own default, and what it is actually printing back at you is Fedora\u0026rsquo;s build flag, with upstream\u0026rsquo;s real default nowhere on the page.\nSetting Upstream default Fedora build Debian/Ubuntu build Result on this box DNSSEC= allow-downgrade no no DNSSEC=no DNSOverTLS= no no no (compiled with OpenSSL) -DNSOverTLS MulticastDNS= yes no no -mDNS LLMNR= yes resolve no LLMNR=resolve Every one of those matches what resolvectl status prints on this machine. Source, build flag, running system, all three agree.\nFedora\u0026rsquo;s reasoning is in the change proposal and it is refreshingly blunt. The feature \u0026ldquo;is known to cause compatibility problems with certain network access points\u0026rdquo;, and Fedora \u0026ldquo;is not prepared to handle an influx of DNSSEC-related bug reports\u0026rdquo;, so it goes off.9\nCan I ask why the answer to a feature that breaks on bad networks is to disable it for everybody, rather than ship allow-downgrade like upstream does and let it turn itself off where it has to? Not who decided. What in the process led there. Because allow-downgrade exists precisely for the captive-portal case, it is upstream\u0026rsquo;s default for exactly that reason, and shipping no instead means a machine on a perfectly good network gets no validation either.\nTo be fair to them, allow-downgrade has its own problem, and it is a real one. The mode detects a resolver that cannot do DNSSEC and quietly stops validating. An attacker who can shape your DNS responses can make that detection fire on purpose, and the documentation says so plainly: it \u0026ldquo;makes DNSSEC validation vulnerable to \u0026lsquo;downgrade\u0026rsquo; attacks\u0026rdquo;.2 A security feature any attacker can switch off is doing less than it looks.\nSo neither default is good. no gives you nothing. allow-downgrade gives you something an attacker can take away, and yes gives you the real thing plus everything in the next section.\nHow Much Of The Web Is Signed Anyway Before spending much effort on validation it is worth knowing what fraction of your lookups it can possibly protect. So I asked resolved for the verdict on the apex of thirty-two domains that this subject actually implicates: the bodies that wrote the standard, the outfits that ship the resolver, and the infrastructure everybody resolves whether they mean to or not.\nSixteen of thirty-two. Half.\nSigned Unsigned Standards and registries ietf.org, iana.org, icann.org, rfc-editor.org, ripe.net, isc.org, nlnetlabs.nl, nic.cz, afnic.fr, verisign.com none Distributions and vendors debian.org, fedoraproject.org, opensuse.org, almalinux.org redhat.com, ubuntu.com, canonical.com, suse.com, rockylinux.org, archlinux.org systemd itself none systemd.io, freedesktop.org Infrastructure cloudflare.com, gov.uk github.com, kernel.org, google.com, wikipedia.org, mozilla.org, apache.org, gnu.org, quad9.net Read down the first column. Every single standards body and registry has signed. Ten out of ten, no exceptions: the people who wrote DNSSEC, and the people who run the registries that publish the DS records for everybody else. They did the work on their own zones.\nNow read across the third row. systemd.io has no DS. The project that wrote the validating resolver this entire post is about has not signed its own domain, and neither has freedesktop.org, where its documentation lives. Below that, quad9.net is unsigned too, which is worth sitting with for a second: a public resolver whose whole pitch is that it validates DNSSEC for you, on a zone nobody can validate.\nAnd the vendors split cleanly along one line. The community distributions signed. debian.org, fedoraproject.org, opensuse.org, almalinux.org. The companies did not. redhat.com, ubuntu.com, canonical.com, suse.com. Those are the four organisations that package and ship this resolver to most of the Linux estate.\nNo gap in the technology there. The protocol has been deployable for over fifteen years, the tooling is free, and the registries take the DS record without charging for it. I will tell you this for nowt: the people who wrote the standard have signed, and most of the people shipping the software have not.\nThere is a sharper version of the same point in gov.uk, which is signed, and then does this:\n$ dig +short @127.0.0.53 www.gov.uk CNAME www-cdn.production.govuk.service.gov.uk. $ dig +short @127.0.0.53 service.gov.uk DS (nothing) uk has a DS. gov.uk has a DS. service.gov.uk has none, so the chain stops dead there, and www.gov.uk resolves insecure even though the apex above it is signed properly. Somebody did the work on gov.uk and then pointed the actual website at an unsigned delegation. Half a job.\nAs such, if you are signing a zone, check the names people actually type. A signed apex that CNAMEs into an unsigned CDN zone buys you nowt.\nDNS Over TLS Works, With Conditions DNS over TLS is implemented, it is compiled into every mainstream build, and it works. It is off by default everywhere, including upstream, where DNSOverTLS= \u0026ldquo;Defaults to no\u0026rdquo;.6\nThe configuration is two lines, and the second one is the one people miss:\n[Resolve] DNS=9.9.9.9#dns.quad9.net 149.112.112.112#dns.quad9.net DNSOverTLS=yes That #dns.quad9.net is not decoration. It sets the name used for SNI and for validating the certificate. Leave it off and the certificate is \u0026ldquo;checked against the server\u0026rsquo;s IP\u0026rdquo; instead.2 That works with the big providers because they put IP addresses in their certificate SANs, but it is the weaker check, it breaks the moment a provider stops doing that, and it gives you no protection against being redirected to a different address that happens to hold a valid certificate for itself. Fedora\u0026rsquo;s own magazine article on this omits the hostname,10 which is a shame, because the syntax is right there in the shipped config file\u0026rsquo;s comments.\nStrict And Opportunistic Are Very Different Settings Mode On a server that supports DoT On a server that does not Authenticates the server yes Encrypted All lookups fail Yes opportunistic Encrypted Silently plaintext No no Plaintext Plaintext n/a opportunistic reads like the sensible middle ground and mostly is not. The documentation says it outright: in that mode \u0026ldquo;the resolver is not capable of authenticating the server, so it is vulnerable to \u0026lsquo;man-in-the-middle\u0026rsquo; attacks\u0026rdquo;,2 and anyone who can drop your port 853 traffic can force the downgrade. It protects you from passive observation on a network where nobody is trying. Against somebody who is, it does nowt.\nyes is the honest setting. It also means that when it cannot connect, you get no DNS at all, which is worth knowing before you put it on a machine you cannot walk up to.\nI tried strict DoT to Quad9 from this box and the query hung. No error, no timeout, nothing in the journal between the flush and my reverting it two minutes later:\nSep 27 17:12:11 systemd-resolved[40105]: wlp4s0: Bus client set DNS server list to: 9.9.9.9#dns.quad9.net, ... Sep 27 17:12:14 systemd-resolved[40105]: wlp4s0: Bus client set DNSOverTLS setting: yes Sep 27 17:12:15 systemd-resolved[40105]: Flushed all caches. Sep 27 17:14:39 systemd-resolved[40105]: wlp4s0: Bus client set DNS server list to: 192.0.2.53, ... I cannot tell you from here whether port 853 is reachable on that connection, because the shell I ran the test from could not reach port 443 either and was clearly filtered itself. What the log does show is the failure mode: strict DoT that cannot connect does not report anything useful. It waits. If you turn this on and your DNS goes quiet, check ss -tn dport = :853 before you go looking at resolved.\nTesting It Properly # does the transport actually come up resolvectl flush-caches resolvectl query ietf.org # expect: acquired via ... encrypted transport: yes # is anything still going out in the clear sudo tcpdump -ni any \u0026#39;port 53 and not host 127.0.0.53\u0026#39; The second command is the one that tells the truth. With DoT working, nothing should leave the machine on port 53, and anything that does is a program that found its way round the stub.\nDNS Over HTTPS Does Not Exist Here Straightforward answer to the question: systemd-resolved does not support DNS over HTTPS. Not partially, not behind a flag, not with a build option nobody turns on. There is no DoH in there at all.\nAnd that is not read off the documentation, which could simply be out of date. It is what the binary contains:\n$ strings /usr/lib/systemd/systemd-resolved | grep -icE \u0026#39;application/dns-message|dns-query|:443\u0026#39; 0 $ grep -oE \u0026#39;name=\u0026#34;DNSOver[A-Za-z]*\u0026#34;\u0026#39; /usr/share/dbus-1/interfaces/org.freedesktop.resolve1.Manager.xml name=\u0026#34;DNSOverTLS\u0026#34; No DoH wire format, no HTTP/2, one encryption property on the D-Bus interface and it is TLS. The complete list of resolved.conf directives upstream runs to seventeen settings and none of them mentions HTTPS.6\nIt has been asked for. Issue #8639, \u0026ldquo;Add support for DNS-over-HTTPS to systemd-resolved\u0026rdquo;, was opened on 2 April 2018 and is still open, with two pull requests attached and nothing merged.11 Eight years and counting.\nSame payload, same encryption, different port BOTH CARRY THE SAME DNS QUERY, INSIDE THE SAME TLS DNS over TLS DNS over HTTPS tcp/853 tcp/443 In systemd-resolved yes, since v239 no, not at all An observer sees the names no no A network can block it yes, its own port not without pain You can audit your own yes no Requested in systemd issue 8639, April 2018. Still open. Same payload, same encryption. The difference is which port it is on, and who can tell Is that the wrong call? Not obviously, and the argument for DoT is decent. Both carry DNS inside TLS and both stop the same passive observer. DoH\u0026rsquo;s advantage is that it hides on port 443 with everything else, so a network that wants to block encrypted DNS has to work much harder. That is genuinely useful if the network operator is the adversary.\nIt cuts the other way too. Where the operator is you, that indistinguishability means you cannot audit your own DNS either, and every application shipping its own DoH client stops using the system resolver, which is how you get a browser that ignores your split DNS, your cache and your internal zones. On your own kit, a resolver visible on a known port is a feature.\nSo: if DoT reaches your resolver, use it and you lose nothing. If port 853 is blocked, resolved has no answer and you want a DoH proxy in front of it, dnscrypt-proxy or cloudflared on loopback with resolved pointed at it. Caching and validation stay resolved\u0026rsquo;s job. Only the transport moves.\nWhat Your Distribution Actually Ships The same daemon behaves very differently depending on who packaged it. Enabling the service is the easy half. The half that gets skipped is plugging it into the networking tools the distribution actually uses, and on one of these four it is genuinely not plugged in until you do it yourself.\nInstalled Enabled nss-resolve wired in Config arrives via Supported Fedora (33+) yes yes yes, in systemd-libs NetworkManager yes Ubuntu yes yes yes netplan, then NM or networkd yes Debian (12+) separate package no no, separate package you choose and wire it yes RHEL / Rocky / Alma (9, 10) yes no yes NetworkManager, once told Technology Preview Validation And A Full Cache: The Settings Themselves These are the same four lines everywhere, because resolved.conf is resolved.conf on every distribution. What differs is only how the rest of the configuration reaches the daemon, which is the next four sections.\n# /etc/systemd/resolved.conf.d/60-local.conf [Resolve] DNSSEC=allow-downgrade Cache=yes CacheFromLocalhost=no StaleRetentionSec=1d Use a drop-in rather than editing /etc/systemd/resolved.conf, because the main file is lower precedence than any drop-in and a package update can argue with you over it. The numbering is a documented convention: vendors take 10 to 40 under /usr/, you take 60 to 90 under /etc/, so yours wins.6\nOn systemd 261 and later there is a fifth line worth adding, because until then the cache size was not tunable at all:\nDNSCacheSize=16384 # systemd 261+; default 4096, max 2^24 Apply it and confirm the daemon agrees with the file, rather than assuming it read it:\nsystemd-analyze cat-config systemd/resolved.conf # every fragment, in precedence order systemctl restart systemd-resolved resolvectl status | grep -E \u0026#39;DNSSEC|Protocols\u0026#39; resolvectl statistics # verdicts should start moving DNSSEC=allow-downgrade is the setting I would put on a machine I am not going to babysit, for the reasons in the section above. DNSSEC=yes is the honest one and it fails closed. Pick deliberately.\nPlugging It Into The Network Stack Turning the service on is one command. Getting your distribution\u0026rsquo;s own networking tools to hand it their DNS configuration is the part that varies, and it is where \u0026ldquo;enabled\u0026rdquo; and \u0026ldquo;working properly\u0026rdquo; come apart.\nWhere DNS config comes from To enable DNSSEC To size the cache Extra wiring needed Fedora NetworkManager, automatically drop-in drop-in none Ubuntu netplan, then NM or networkd drop-in, or per-link in .network drop-in none Debian whichever stack you chose drop-in, or per-link in .network drop-in libnss-resolve, plus the stack below RHEL family NetworkManager, once told drop-in drop-in dns=systemd-resolved Fedora The most complete integration of the four, and the only one where everything is already plugged in.\nNetworkManager owns the network and hands DNS to resolved without being told to. On this box there is no dns= line anywhere in /etc/NetworkManager/ and it still works, because NetworkManager detects the running daemon and uses it. /etc/resolv.conf is the stub symlink, and glibc is wired in because libnss_resolve.so.2 ships in systemd-libs, which is not optional:\n$ rpm -qf /usr/lib64/libnss_resolve.so.2 systemd-libs-259.9-1.fc44.x86_64 $ grep ^hosts: /etc/nsswitch.conf hosts: files myhostname mdns4_minimal [NOTFOUND=return] resolve [!UNAVAIL=return] dns So on Fedora the drop-in above is the whole job. Nothing else to connect.\nUbuntu Enabled by default, and nss-resolve is wired in the same way. The difference is that configuration usually arrives through netplan:\nnetwork: version: 2 ethernets: enp1s0: dhcp4: true nameservers: addresses: [9.9.9.9, 149.112.112.112] search: [example.com] Netplan hands that to its renderer, NetworkManager on desktop or systemd-networkd on server, and the renderer hands it to resolved. Note what netplan cannot express. There is no netplan key for DNSSEC=, and none for DNSOverTLS=. Those go in the drop-in above, which is where they belong anyway.\nIf you are on systemd-networkd, the per-link settings are also available natively in the .network file, and a per-link setting beats the global one:\n# /etc/systemd/network/10-lan.network [Network] DNS=9.9.9.9#dns.quad9.net DNSSEC=yes DNSOverTLS=yes Domains=~. Debian Needs Wiring In By Hand This is the one that is genuinely not plugged in, and the reason the section above exists.\nSince Debian 12 systemd-resolved is a separate package, and the release notes are explicit about the upgrade: \u0026ldquo;The new systemd-resolved package will not be installed automatically on upgrades\u0026rdquo;, and \u0026ldquo;until it has been installed, DNS resolution might no longer work since the service will not be present on the system.\u0026rdquo;12 The same notes settle the wider question: \u0026ldquo;systemd-resolved was not, and still is not, the default DNS resolver in Debian.\u0026rdquo;\nInstalling it is three packages, not one, and the second is the one everybody misses:\napt install systemd-resolved libnss-resolve systemctl enable --now systemd-resolved libnss-resolve is only Suggests:, not a dependency,13 and Suggests is the one relationship apt does not act on. So nobody installs it.\nThe reason nobody notices is more interesting than the reason nobody installs it. Leave it out and glibc still reaches resolved, because /etc/resolv.conf points at the stub and plain nss-dns talks to it. Most of the daemon keeps working:\nWithout libnss-resolve With it Cache Yes, via the stub Yes DNSSEC validation Yes, it happens in the daemon Yes DNS over TLS Yes Yes AD bit reaches glibc Yes, resolv.conf carries trust-ad Yes Secure vs insecure vs bogus No, one bit only Yes Link-scoped addresses No Yes /etc/hosts read by the daemon No Yes Nothing in the left column looks broken, so nothing gets fixed. What you lose is what the DNS wire format cannot carry, and the clearest case is an address with a scope on it:\n$ getent hosts _gateway # via nss-resolve fe80::5054:ff:fe12:3456 _gateway $ dig +short @127.0.0.53 _gateway # via the stub, as nss-dns would 192.0.2.1 Same name, same daemon, two different answers. A link-local IPv6 address is meaningless without the interface it is scoped to, and there is nowhere in a DNS response to put that, so the stub can only hand back the IPv4. The native API has somewhere to put it and gives you the address you actually wanted.\nThe verdict row is the same problem. AD is one bit, signed or not. The native API separates secure from insecure from bogus, which is the difference between \u0026ldquo;nobody signed this\u0026rdquo; and \u0026ldquo;somebody signed this and somebody else has been at it\u0026rdquo;. An application that cares cannot tell those apart over the wire.\nThat is the gap between the service running and the service being plugged in, and it stays invisible until you go looking for it.\nCheck it landed:\ngrep ^hosts: /etc/nsswitch.conf # wants \u0026#39;resolve [!UNAVAIL=return] dns\u0026#39; Four ways in on Debian, and the one that is not installed for you WHERE CONFIG COMES FROM HOW IT GETS IN THE DAEMON ifupdown, static ifupdown, DHCP systemd-networkd NetworkManager resolvconf shim = resolvectl resolved 127.0.0.53 AND THE BIT NOBODY INSTALLS glibc getaddrinfo() libnss-resolve Suggests: only. Without it glibc still works, and the verdict never reaches it. Four ways in on Debian, and the one that is not installed for you How the rest of your networking reaches it depends on which of Debian\u0026rsquo;s stacks you are on.\nThe package declares Provides: resolvconf and Conflicts: resolvconf, openresolv,13 so it takes over the resolvconf interface outright. /usr/sbin/resolvconf becomes a symlink to resolvectl, which is a multi-call binary: invoked under that name it speaks the resolvconf(8) protocol and pushes everything it is handed straight into resolved.14 Anything that already calls resolvconf -a therefore keeps working unmodified, with systemd-resolved as the only supported backend.\nNot all of the interface survives, and the gaps fail loudly rather than silently:\nresolvconf option Under systemd-resolved -a \u0026lt;iface\u0026gt; Registers per-link DNS read from stdin. The one that matters -d \u0026lt;iface\u0026gt; Unregisters, same as resolvectl revert -x Mapped to a ~. routing domain -p Marks the link as not a default route (systemd 257+) -f Makes -a and -d quiet about an interface that is not there -m Accepted and silently ignored -u, -i, -I, -l, -r, -R, -v, -V Not supported. The command fails One trap worth knowing before you debug it: the shim only writes /etc/resolv.conf when that file is a symlink to /run/systemd/resolve/resolv.conf, and not when it is a static file.14\nifupdown, the default on a Debian server install. The dns-nameservers stanza in /etc/network/interfaces was always implemented by the old resolvconf package\u0026rsquo;s hooks, and systemd-resolved conflicts with that package and ships no /etc/network/if-up.d/ hooks of its own. For a static interface, do not rely on the stanza. Set the servers on the link explicitly and let resolved own them:\n# /etc/network/interfaces auto enp1s0 iface enp1s0 inet static address 192.0.2.10/24 gateway 192.0.2.1 up /usr/sbin/resolvconf -a $IFACE \u0026lt;\u0026lt;\u0026lt; \u0026#39;nameserver 192.0.2.53\u0026#39; down /usr/sbin/resolvconf -d $IFACE DHCP on ifupdown is already handled. isc-dhcp-client ships hooks named for the daemon, /etc/dhcp/dhclient-enter-hooks.d/resolved-enter and /etc/dhcp/dhclient-exit-hooks.d/resolved,15 so a lease\u0026rsquo;s DNS servers land on the right link with nothing to configure.\nsystemd-networkd is the cleanest option and needs no shim at all, because the two halves are the same project. Per-link DNS, DNSSEC and DoT are native keys in the .network file, exactly as in the Ubuntu example above.\nNetworkManager, on a Debian desktop, needs telling once:\n# /etc/NetworkManager/conf.d/10-resolved.conf [main] dns=systemd-resolved Pick one and know which one you picked. The failure mode here is not a daemon that refuses to start, it is two stacks both believing they own /etc/resolv.conf, which reads as intermittent DNS and wastes an afternoon. resolvectl status says which link the servers landed on. If the answer is none of them, nothing is feeding the daemon, and that is the bug.\nRHEL, Rocky And Alma And here is the one worth reading twice. The package is there, nss-resolve is available, NetworkManager can drive it, and Red Hat\u0026rsquo;s documentation says this:\nsystemd-resolved is an unsupported Technology Preview.16\nThat wording has been in the RHEL 9 release notes since 9.0 and it is still there at 9.8. Technology Preview means no production SLA, and Red Hat explicitly does not recommend it for production use.\nNor is it enabled. NetworkManager owns resolv.conf on these systems, so turning it on is a NetworkManager setting rather than just enabling the unit:\n# /etc/NetworkManager/conf.d/10-resolved.conf [main] dns=systemd-resolved systemctl enable --now systemd-resolved systemctl reload NetworkManager resolvectl status # confirm NM actually handed the servers over After that the drop-in above applies unchanged, and validation and the cache behave exactly as they do on Fedora.\nSo on RHEL you have a decision that does not exist on the others. Want a validating, caching, split-DNS-capable local resolver on a supported RHEL box? Then the supported answer is not this one. It is unbound, or dnsmasq through NetworkManager\u0026rsquo;s own dns=dnsmasq, both of which Red Hat will actually stand behind.\nThe label is not a comment on the code. It is a comment on what Red Hat will put an SLA behind, and a caching validating resolver is a thing customers open tickets about. Fair enough. But if you are buying RHEL for the support, running name resolution on an unsupported component is a decision to make deliberately and write down, not one to drift into because it was in the repo.\nNothing Is Actually Crippled, And Here Is How To Prove It The premise is worth testing, because \u0026ldquo;my distro disabled it, I will have to rebuild\u0026rdquo; is the reflex and here it is wrong.\nsystemd has two entirely different kinds of build option, and they get talked about as though they were one:17\nBuild option Kind Upstream Fedora Debian and Ubuntu Changeable at runtime resolve capability on built true no, and both build it nss-resolve capability enabled built, ships in systemd-libs enabled, ships as libnss-resolve no, and both build it dns-over-tls capability auto auto, resolves to OpenSSL openssl no, and both compile it in openssl capability enabled enabled enabled no, and both enable it default-dnssec default only allow-downgrade no no yes, in a drop-in default-dns-over-tls default only no no upstream default yes, in a drop-in default-mdns default only yes no no yes, in a drop-in default-llmnr default only yes resolve no yes, in a drop-in Read the last column. Every setting this post has complained about sits in the bottom half, and every one is a default, not a capability. The code is compiled in on both. Nobody removed anything.\nThe proof is earlier in this post and needed no compiler. On the stock Fedora package, with DNSSEC=no baked in, one command turned validation on and it worked in full: a signed zone authenticated, an unsigned one reported honestly, a broken one refused with a precise diagnostic. Had it been compiled out, resolvectl status would never have said yes/supported.\nSo check your own build before you reach for a toolchain:\n# is the crypto there at all systemctl --version | tr \u0026#39; \u0026#39; \u0026#39;\\n\u0026#39; | grep -E \u0026#39;^[+-](OPENSSL|GNUTLS|GCRYPT)$\u0026#39; # ask for the feature and read back what the daemon says it can do resolvectl dnssec \u0026lt;link\u0026gt; yes resolvectl status \u0026lt;link\u0026gt; | grep DNSSEC # yes/supported means the code is there On this Fedora box that is -GCRYPT +GNUTLS +OPENSSL, which is plenty: DNSSEC needs one of them and DoT is built against OpenSSL.\nWhat The Distributions Really Do Strip There is one thing both of them take out, and it is not in the list anybody complains about.\nUpstream compiles a fallback DNS server list into the binary: Cloudflare, Google and Quad9.17 Both Fedora and Debian build with -Ddns-servers= set to nothing, and you can confirm it on the shipped binary rather than taking my word:\n$ strings /usr/lib/systemd/systemd-resolved | grep -ciE \u0026#39;quad9|one\\.one\\.one|dns\\.google\u0026#39; 0 Nothing. No hyperscaler baked into the executable.\nThat is the one place both distributions have improved on upstream, and it deserves saying plainly given how much of this post has been about defaults I think are wrong. A machine that loses its DNS configuration should fail loudly and wait for you. It should not quietly route every name you look up to a resolver in another jurisdiction that you never chose. Want a fallback? Set FallbackDNS= and pick who it is.\nIf You Genuinely Do Need A Rebuild One real case exists: a minimal or embedded build where somebody passed -Ddns-over-tls=false and the transport is genuinely absent. Check with the commands above first, because it is rare and looks identical to the setting just being off.\nIf you do need it, rebuild the package, never make install:\n# Fedora dnf download --source systemd rpmbuild --rebuild --define \u0026#39;_with_upstream 1\u0026#39; systemd-*.src.rpm # Debian and Ubuntu apt source systemd cd systemd-*/ \u0026amp;\u0026amp; editor debian/rules # change the -Ddefault-* flags dpkg-buildpackage -us -uc -b I have not run either on this machine, so treat them as the shape of the job, not a tested recipe. The principle is what matters. systemd is PID 1, and a hand-built make install over your distribution\u0026rsquo;s copy takes it out of the package manager: no more security updates, and the next upgrade fights you for the files. Building the package keeps both. For almost everybody the honest answer is that a four-line drop-in does the same job in twenty seconds, which is why this section mostly exists to talk you out of it.\nThree Configurations Worth Running Not a menu of every option. Three positions, each of which I would actually defend.\nLaptop, hostile networks Server, your own network Strict What you are defending against The cafe and the hotel portal Nothing local; the upstream already validates A resolver path you do not trust DNSSEC= allow-downgrade no yes DNSOverTLS= opportunistic no yes StaleRetentionSec= 1d 1d unset Survives a captive portal Yes n/a No Survives a filtered .arpa Yes Yes No Survives port 853 blocked Yes n/a No An attacker can downgrade it Yes, both settings n/a No A laptop on networks you do not control. Encrypt the transport, and take the downgrade risk in exchange for the thing still working behind a captive portal.\n# /etc/systemd/resolved.conf.d/60-local.conf [Resolve] DNS=9.9.9.9#dns.quad9.net 149.112.112.112#dns.quad9.net DNSOverTLS=opportunistic DNSSEC=allow-downgrade Cache=yes StaleRetentionSec=1d A server on a network you do run, validating upstream. It already happened one hop away. Do not do it twice, and do not take a dependency on a public resolver.\n[Resolve] DNSSEC=no DNSOverTLS=no Cache=yes CacheFromLocalhost=no StaleRetentionSec=1d A machine where you want the real thing. Strict on both counts, nothing for an attacker to downgrade, and you have checked the upstream can carry it.\n[Resolve] DNS=9.9.9.9#dns.quad9.net 149.112.112.112#dns.quad9.net DNSOverTLS=yes DNSSEC=yes Cache=yes That third one will break reverse DNS if anything in your path filters .arpa, it will break entirely if port 853 is blocked, and it fails closed rather than quietly. Those are the terms. Know them before you deploy it, not after, because a box that cannot resolve owt is a long drive if it is not in the next room.\nWhichever you pick, apply it and then check what the daemon actually decided rather than what you wrote:\nsystemd-analyze cat-config systemd/resolved.conf # every file, in order systemctl restart systemd-resolved resolvectl status # the state it is really in resolvectl status is the oracle here, the same way the build is the oracle for a Hugo site. A setting in a file is an intention. What status prints is what is happening.\nThe rest of the verbs are worth knowing, because between them they answer nearly every question you will have about this daemon without reading a log:\nCommand What it tells you resolvectl status Per-link servers, domains, protocols, and the live DNSSEC and DoT state resolvectl query NAME The answer, the protocol used, the DNSSEC verdict, and whether it came from cache resolvectl statistics Cache hits against misses, and the running secure/insecure/bogus tally resolvectl show-cache Everything currently cached, per scope resolvectl flush-caches Empty the cache without restarting the daemon resolvectl dns LINK ... Set servers on one link at runtime, no config file resolvectl dnssec LINK yes Turn validation on for one link, to test before committing resolvectl domain LINK ~x Set routing and search domains on one link resolvectl revert LINK Throw away every runtime change on that link resolvectl show-server-state Per-server feature probing: what each upstream was found to support resolvectl monitor Watch queries and answers live, which beats guessing Everything in that middle block is runtime only and does not survive a link coming back up, which makes it the right way to try a setting before you write it into a drop-in. It also means nmcli device reapply will quietly undo your testing, so check status after, not before.\nThe Defaults Are A Position, Not An Accident Three features in the box. One switched on.\nTempting to read that as laziness. It is not. Every one of those defaults was argued over by people who could see the bug queue on the other side of the decision, and Fedora at least wrote down honestly that it could not face the support load. That is a real constraint and I will not pretend otherwise.\nBut a default is a position, and this one says a lookup nobody can verify is an acceptable thing to build on. We have had the standard since 2005. The registries publish the records for free, the resolver on your machine implements the whole thing, and the reason it sits idle is that too much of the internet never signed owt, so switching it on makes your machine the one that looks broken. That is the shape of every standard nobody enforces. Doing it properly is a cost you carry on your own, and the benefit only turns up when enough others have carried it too.\nWhich is why the census is the part of this I would actually go and act on. Not the settings. The settings are twenty minutes. The census says the organisations publishing the guidance, shipping the resolver and hosting the world\u0026rsquo;s source code have mostly not signed their own zones, and that gov.uk signed the apex and then pointed the one website nobody can avoid using at an unsigned delegation. Somebody did the hard part and then never checked the name people type.\nYou can only fettle your own. If you run a zone, sign it, lodge the DS, then go and look up the www name the way a visitor would and confirm the chain survives to the end. It is an afternoon. Do that and the argument for validation stops being theoretical for everybody resolving your name, which is the only part of this any of us can actually mend.\nsystemd-resolved.service(8) — \u0026ldquo;implements a caching and validating DNS/DNSSEC stub resolver\u0026rdquo;; the four client interfaces, the two stub listeners, synthetic records, the routing rules and the four /etc/resolv.conf modes.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nresolved.conf(5), as shipped in systemd 259.9 on Fedora 44 — DNSSEC=, DNSOverTLS=, Cache=, CacheFromLocalhost=, StaleRetentionSec= and the 127.0.0.54 proxy stub description.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nresolved.conf(5) — DNSCacheSize=, MulticastDNSCacheSize= and LLMNRCacheSize=, \u0026ldquo;Each defaults to 4096\u0026rdquo;, added in systemd 261.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 5737 — \u0026ldquo;IPv4 Address Blocks Reserved for Documentation\u0026rdquo;: 192.0.2.0/24, 198.51.100.0/24 and 203.0.113.0/24.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 3849 — 2001:DB8::/32 reserved for documentation, \u0026ldquo;to reduce the likelihood of conflict and confusion when relating documented examples to deployed systems\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nresolved.conf(5), upstream latest — DNSSEC= \u0026ldquo;Defaults to allow-downgrade\u0026rdquo;, DNSOverTLS= \u0026ldquo;Defaults to no\u0026rdquo;, the drop-in numbering convention, and the complete directive list with no DNS over HTTPS option.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nFedora systemd.spec, rawhide — the meson build flags -Ddefault-dnssec=no, -Ddefault-dns-over-tls=no, -Ddefault-mdns=no, -Ddefault-llmnr=resolve.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDebian systemd packaging, debian/rules — -Ddefault-dnssec=no, -Ddefault-llmnr=no, -Ddefault-mdns=no, -Ddns-over-tls=openssl. Ubuntu\u0026rsquo;s systemd derives from this packaging.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nFedora Change: systemd-resolved — Fedora 33 made it the default resolver; changes the default \u0026ldquo;from the upstream default DNSSEC=allow-downgrade to DNSSEC=no\u0026rdquo; because the feature \u0026ldquo;is known to cause compatibility problems with certain network access points\u0026rdquo; and Fedora \u0026ldquo;is not prepared to handle an influx of DNSSEC-related bug reports\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nFedora Magazine — Use DNS over TLS — the recommended DNSOverTLS=yes configuration, given with bare IP addresses and no #hostname SNI form.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nsystemd issue #8639 — \u0026ldquo;Add support for DNS-over-HTTPS to systemd-resolved\u0026rdquo;, opened 2 April 2018, still open.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDebian 12 release notes, 5.2.3 — \u0026ldquo;The new systemd-resolved package will not be installed automatically on upgrades\u0026rdquo;; \u0026ldquo;until it has been installed, DNS resolution might no longer work\u0026rdquo;; \u0026ldquo;systemd-resolved was not, and still is not, the default DNS resolver in Debian.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDebian systemd packaging, debian/control — the systemd-resolved stanza: Provides: resolvconf, Conflicts: resolvconf, openresolv, Replaces: resolvconf, and libnss-resolve listed only under Suggests:.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nresolvconf(1), the resolvectl compatibility mode — resolvectl \u0026ldquo;is a multi-call binary. When invoked as \u0026lsquo;resolvconf\u0026rsquo; \u0026hellip; it is run in a limited resolvconf(8) compatibility mode\u0026rdquo;; systemd-resolved \u0026ldquo;is the only supported backend\u0026rdquo;; which options are supported, ignored or rejected; and the rule that /etc/resolv.conf is only written when it is a symlink to /run/systemd/resolve/resolv.conf.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDebian isc-dhcp-client file list — ships /etc/dhcp/dhclient-enter-hooks.d/resolved-enter and /etc/dhcp/dhclient-exit-hooks.d/resolved.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRHEL 9.4 release notes, Technology Previews — \u0026ldquo;Note that systemd-resolved is an unsupported Technology Preview.\u0026rdquo; Carried unchanged from RHEL 9.0 through 9.8.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nsystemd meson_options.txt — the split between capability options (resolve, nss-resolve, dns-over-tls, openssl) and default-value options (default-dnssec, default-dns-over-tls, default-mdns, default-llmnr), and the compiled-in dns-servers fallback list.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/dns/resolved-the-resolver-you-are-already-running/","summary":"systemd-resolved is running on most Linux desktops right now, caching every lookup and validating none of them. This walks the resolution path from nss-resolve to the two stub listeners, measures what the cache is worth on a real machine, turns DNSSEC on and shows the three verdicts it can return, and explains why validation is off by default when upstream ships it on. Then a census of how little of the web is actually signed, DNS over TLS and its strict-versus-opportunistic trap, and the DNS over HTTPS support that has been requested since 2018 and still does not exist. Ends with what Fedora, Ubuntu, Debian and the RHEL family each ship, how to wire it into each one, and why the features your distro disabled are defaults rather than missing code.","title":"Resolved: The Resolver You Are Already Running"},{"content":"There is a way for anything on your network to look up a name without your DNS server ever hearing the question. No setting changed on the machine, no administrator rights, nothing you would find in a log. The lookup leaves as an ordinary HTTPS request on port 443, goes to a resolver somewhere on the internet, and comes back with an answer your own resolver would have refused to give.\nThe blocklist never fires. The threat feed never gets the query. The log line you would have gone looking for was never written.\nIt is called DNS over HTTPS, DoH for short, and it was built for a good reason. Plain DNS travels in clear text on port 53, so the coffee shop, the airport and your internet provider can read every name you resolve and change the answers if they fancy it. DoH wraps the lookup in the same encryption as the rest of the web and hides it in the crowd. As a privacy measure for a person on a hostile network, it does exactly what it says.\nThe trouble is that the thing it defeats and the thing you rely on are the same thing.\nYour resolver is not just a lookup service. It is a control point: where a known-bad domain gets answered with nothing, where a query to a command server shows up in a log, where Protective DNS refuses the malware domain before the connection is made. DoH takes the lookup off your resolver and hands it to one you never chose, and every control you hung off that resolver goes with it.\nOrdinary DNS goes through your resolver. DoH goes round it. Same laptop, same question. One answer passes your controls, one never meets them. Ordinary DNS, port 53 DNS over HTTPS, port 443 Laptop asks for a name Your resolver blocklist and threat feed checked query logged against the client bad name answered NXDOMAIN The internet only if it passed Laptop asks for a name Your resolver never asked nothing to block nothing to log HTTPS to a resolver of its choosing Someone else's resolver answers anything The firewall sees one more encrypted web connection. It has thousands of those a minute, and on the wire this one looks like the rest. The same laptop asking the same question. On the left it goes through your resolver, which checks the name, logs it and only then asks the internet. On the right it goes straight out over HTTPS to a resolver of the client\u0026rsquo;s choosing. Your resolver never hears it, so there is nothing to block and nothing to log, and at the border it is one more encrypted web connection among thousands. This is not an argument against encrypting DNS. Encrypted DNS is right, and the last section of this post is how to run it. It is an argument about who gets to choose the resolver, because that choice is the whole game, and DoH was designed to take it away from you and give it to the browser, the app and, if you are not careful, the attacker.\nWhat DoH Actually Is Strip the branding off and DoH is a normal web request that happens to carry a DNS question.\nOrdinary DNS is a small binary message sent over UDP or TCP to port 53. DoH puts that same message, or a JSON version of it, inside an HTTPS request to a web server that speaks the protocol; the reply is an HTTPS response1. That is the entire idea.\nRFC 8484 standardised it in October 2018, and the intent was never hidden. The point, its own introduction says, is \u0026ldquo;allowing web applications to access DNS information via existing browser APIs\u0026rdquo;1.\nExisting browser APIs. A web page. Nobody found that as an awkward side effect years later. It sits in the first paragraph of the standard, written down as the goal.\nHere is a lookup done the DoH way, from a command line, against a public resolver, asking for example.com:\n$ curl -s -H \u0026#39;accept: application/dns-json\u0026#39; \\ \u0026#39;https://cloudflare-dns.com/dns-query?name=example.com\u0026amp;type=A\u0026#39; {\u0026#34;Status\u0026#34;:0,\u0026#34;TC\u0026#34;:false,\u0026#34;RD\u0026#34;:true,\u0026#34;RA\u0026#34;:true,\u0026#34;AD\u0026#34;:true,\u0026#34;CD\u0026#34;:false, \u0026#34;Question\u0026#34;:[{\u0026#34;name\u0026#34;:\u0026#34;example.com\u0026#34;,\u0026#34;type\u0026#34;:1}], \u0026#34;Answer\u0026#34;:[{\u0026#34;name\u0026#34;:\u0026#34;example.com\u0026#34;,\u0026#34;type\u0026#34;:1,\u0026#34;TTL\u0026#34;:146, \u0026#34;data\u0026#34;:\u0026#34;23.192.228.80\u0026#34;}]} No special client. No port but 443. A single HTTPS request, the same shape as fetching a web page, and a name resolved by a machine on the other side of the world that has never heard of your network or your rules. Google runs the same endpoint, so does Quad9, so do dozens of others.\nNow look at what your firewall sees. A TLS connection to a web server on 443, which it cannot read inside because that is the point of TLS, and cannot tell from the hundreds of others opening every second to content delivery networks and analytics and adverts. The DNS question is gone. It left the building dressed as web traffic and nobody at the door could see its face.\nOn the wire, a DoH lookup is a DNS question sealed inside a web request. Same question. One is written on the envelope, one is sealed three layers in. Plain DNS, port 53 DoH, port 443 DNS query name=badsite.example type=A The firewall reads this in full. It sees the name, checks it, blocks or logs it. TLS record (all the firewall sees) HTTP request to a web server DNS query (sealed) name=badsite.example type=A On port 443 the firewall sees only the outer layer. A TLS connection to a web server, the same as a page load, an advert or an analytics beacon. The name never surfaces. A plain DNS query is written on the envelope: the firewall reads the name and acts on it. A DoH query is the same question sealed inside an HTTP request inside a TLS record, and all the firewall sees is the outer layer, a connection to a web server on 443 that looks like every other one. The name never surfaces where a control could reach it. There are three ways a client can move a DNS lookup, and it is worth laying them side by side, because the difference is the whole post.\nTransport Port Your resolver sees it Refusable at the border Encrypted Plain DNS 53 (UDP/TCP) Yes, if you force it through Yes, close 53 outbound No DNS over TLS (DoT) 853 Only if it points at yours Yes, close 853 outbound Yes DNS over HTTPS (DoH) 443 Only if it points at yours No, you cannot close 443 Yes Plain DNS is readable and blockable, and that is exactly why it was easy to control and easy to spy on. DoT encrypts the lookup but keeps its own port, so you can still refuse it at the border. DoH is the one with no handle on it: encrypted like DoT, but sharing the one port you can never close, so the only lever left is which resolver the client chose, and that is the lever this post is about.\nWho Gets To Pick The Resolver The reason this matters is that a modern machine has five different things on it that can each decide where DNS goes, and you control exactly one of them by default.\nFive things on one machine can pick a resolver. You set one of them. Who decides where the lookup goes Layer Who picks the resolver Yours by default? Operating system Your network, through DHCP or router advertisements Yes Browser The vendor's default, unless your policy overrides it Only if you set the policy Application or its library Whoever wrote the code, at build time No Script on a web page Whoever runs the site, or any script it loads No Malware The operator, hardcoded, often by address not name No None of the bottom three needs administrator rights, a setting changed, or anything you would see in the OS. Five layers on one machine, each able to choose its own resolver. The operating system uses the one your network hands out, and that one is yours. The browser uses its vendor\u0026rsquo;s default unless your policy says otherwise. An application, a script on a web page and malware each choose for themselves and need nothing from you to do it. The operating system asks the resolver your network handed out over DHCP or router advertisements. That one is yours, and it is the model everything older than about 2019 assumed: one resolver, handed out by the network, which is why network-level DNS controls worked for thirty years.\nThe browser broke it. Firefox and Chrome both ship the machinery to do their own DoH, to a resolver their vendor picked, over your head, standing down only when they spot a managed network. And an application can carry its own DoH client and a hardcoded resolver in its code, decided at build time; it does not read your DHCP or ask. Plenty of legitimate software already does.\nA script on a web page is the one that should stop you, because it needs nothing installed at all. That curl command above is a single HTTPS request, and a browser makes HTTPS requests for a living. A few lines of JavaScript on any page a user opens can send lookups to a public DoH endpoint, because the big providers deliberately allow cross-origin requests so web apps can use them, the stated purpose in RFC 8484. The page you are reading could resolve names through a resolver in another country right now and you would see one more HTTPS connection.\nAnd malware picks its own resolver for the obvious reason: it does not want you to see where it is calling. It carries the resolver in its code, often reaching it by address so there is no bootstrap lookup to catch, and needs your permission for none of it.\nRead that column again. The bottom three need no administrator rights, no setting changed, and nothing that shows in the operating system. The control you spent years building, one resolver, one blocklist, one log, assumed the top row was the only row. It has not been for years.\nThe Same Trick The Browser Vendors Feared The web-page case is neither theoretical nor recent. It is how protocol helpers get abused: a web page emits bytes, something downstream acts on them, and cannot tell they came from an attacker\u0026rsquo;s page rather than a real client, because on the wire they are identical. With DoH the downstream piece is a public resolver, which answers with no way to know the JavaScript asking came from a phishing page, and your resolver, the one with the blocklist and the log, was never in the path to have an opinion. Not a hole punched through your controls, but a road built around them, paved in the same encryption you tell everyone to use.\nIt Is Already The Malware Author\u0026rsquo;s Channel You do not have to imagine how this gets used. It has been documented for years, by named researchers, on real samples, and the direction of travel is one way only: from a criminal bot in 2019 to a state intelligence tool the year after, and busier every year since, with fresh backdoors still landing in 2026.\nSample Reported Actor What DoH carried Godlua 1st July 20192 Criminal botnet The lookup for its command server\u0026rsquo;s name PsiXBot 6th September 20193 Criminal (infostealer) Command-and-control domain resolution, via Google\u0026rsquo;s DoH OilRig (APT34) Q2 20204 Iranian state-aligned Stolen data, exfiltrated over DoH to Google and Cloudflare ChamelDoH 16th June 20235 ChamelGang (APT) Its whole command channel, DNS TXT over DoH to Google and Cloudflare BRICKSTORM 4th December 20256 China-nexus backdoor C2 buried under HTTPS and nested TLS, DoH among the layers Dohdoor 26th February 20267 Undetermined (UAT-10027) C2 lookups sent to Cloudflare\u0026rsquo;s DoH on 443 Godlua, the first widely reported case, was a Linux and Windows backdoor whose write-up records it \u0026ldquo;uses DNS over HTTPS to get the C2 name to ensure secure communication between the bots, the Web Server and the C2\u0026rdquo;2. On a network watching port 53, resolving your command server is a gift to the defender. Over DoH there is nowt to watch.\nProofpoint\u0026rsquo;s line on PsiXBot is the one worth quoting, a threat-intelligence vendor saying the quiet bit out loud. Using DoH for command and control, they wrote, \u0026ldquo;should be a warning shot for the cybersecurity community as there are no simple solutions to help identify an infected host or the process of receiving instructions\u0026rdquo;3. No simple solutions, from the people whose job is finding them.\nThen OilRig took it up a tier. Kaspersky\u0026rsquo;s description is exactly the change this post is about: \u0026ldquo;instead of plain text requests to port 53, they would use port 443 in encrypted packets\u0026rdquo;, with a tool that \u0026ldquo;allows DoH queries to Google and Cloudflare services\u0026rdquo;4. A national intelligence operation, using the public DoH providers as the pipe to carry stolen data out past whatever was watching the DNS.\nAnd it did not stop there. It spread. By 2023 ChamelGang had a C++ Linux backdoor, ChamelDoH, running its whole command channel over DoH, sending DNS TXT queries to its own nameservers through Google and Cloudflare; the researcher who found it noted that both detection and prevention \u0026ldquo;become difficult\u0026rdquo;, because the encrypted transport cannot be intercepted and a malicious request cannot be told apart from a real one5. In December 2025 CISA took apart BRICKSTORM, a China-nexus backdoor that stacks HTTPS, WebSockets and nested TLS and \u0026ldquo;also uses DNS-over-HTTPS (DoH)\u0026rdquo; to bury its C2 in ordinary web traffic6. By February 2026 it was just tradecraft: Cisco Talos caught Dohdoor, which \u0026ldquo;securely sends encrypted DNS requests to Cloudflare\u0026rsquo;s DNS server over HTTPS port 443\u0026rdquo; to find its command server, phishing its way into American schools and hospitals7.\nThat is the shape of it. Not a technique that had its moment and faded, but one that gets picked up by more hands every year. The issue is getting worse, not better, and a fresh family now turns up on schedule.\nOne property did the work in every one of them: the lookup left as HTTPS on 443 and the defender\u0026rsquo;s resolver never saw it. Not a weakness somebody might find one day. A feature attackers have shipped again and again, several of them states.\nAnd be clear about why that table fills up, year on year. Every family on it walks through a door the authors of DoH were warned they were leaving open, in the standard\u0026rsquo;s own text, and left open regardless8. That was no oversight. It was short-sightedness, chosen on purpose, by people who included ICANN\u0026rsquo;s own technologist. So I will call it what it is: on this, ICANN are cybercrime enablers. Warned in the standard\u0026rsquo;s own text what would break, its people put their name to it anyway, and the table above is what walked through the gap. The crime is no surprise. It is the bill for a decision, and whose decision it was is the rest of this post.\nAnd for all that, DoH does not even buy the thing people assume it does: protection from a forged answer. Encrypting the hop to a resolver is not authenticating what the resolver hands back, and a DoH server can still return a forged record. The standard admits it, saying the rule against unconfigured servers \u0026ldquo;does not guarantee protection against invalid data\u0026rdquo;9.\nThe one mechanism that authenticates a DNS answer is DNSSEC, and DoH is not it. Worse, for almost every client DNSSEC is validated at the recursive resolver, not on the device: the client simply trusts the resolver\u0026rsquo;s word that the answer checked out. So DNSSEC only ever protected you as far as you trusted the resolver doing the checking, and DoH\u0026rsquo;s real move is to hand that trust to a remote operator you cannot see or audit.\nAs such DoH can leave you more likely to land on a forged site, not less: the local defences that would have caught a hijacked answer, the Protective DNS feed, the operator\u0026rsquo;s own filtering, are the very things you routed around, and all that is left is one remote operator\u0026rsquo;s word, taken on trust.\nDoH moves who you trust for the answer. It does not remove the need to trust someone. Someone has to be trusted for the answer. DoH just changes who, and to one you cannot audit. Your own resolver A remote DoH resolver Your device Resolver you control Validates DNSSEC on your behalf Carries your blocklist and log You can inspect and audit it If it lies, you own it and can check Your device encrypted pipe Resolver you do not control Validates DNSSEC, and you trust its word You cannot see or audit it Your local blocklist is out of the path If it forges, nothing local catches it Encryption secures the pipe, not the truth of the answer. DNSSEC is checked at the resolver, so you trust the resolver either way. DoH just makes it one you cannot see. DoH does not remove the need to trust a resolver for the answer, it just moves it. Your own resolver validates DNSSEC, carries your blocklist and can be audited; a remote DoH resolver validates on your behalf and you take its word, over an encrypted pipe that secures the hop and says nothing about whether the answer is true. What DoH is sold as protecting against Does it? Eavesdropping on the hop to the resolver Yes, the lookup is encrypted Tampering on that hop Yes A forged record from the resolver itself No, that is DNSSEC\u0026rsquo;s job, not DoH\u0026rsquo;s A malicious, coerced or compromised resolver No, you now trust it completely Malware, tracker and court-order blocking No, it routes around them And none of this is a new idea DoH tripped over by accident. Filtering malicious domains at the resolver is exactly what OpenDNS has done for the best part of twenty years, blocking phishing and malware for millions of users before the answer ever reached them, which is the model Cisco paid to own in 201510. Two decades of resolver-level security that protected people who never configured a thing, and DoH undermined the whole model in a single hit: point the app at its own resolver, and OpenDNS, or whatever your network chose, is not in the path any more.\nWhy The Old Answer Stopped Working For as long as DNS lived on port 53, the network answer was simple. Force every client to use your resolver, and block outbound port 53 everywhere else at the border. A machine reaching for an outside DNS server got refused, so it had to come through yours, so your controls applied to everything. Crude, and it worked.\nDoH kills that in one move by using 443. You cannot block outbound 443. It is the web. So the port you would close to force DNS through your resolver is the one you can never close, and the encryption that stops your ISP snooping also stops you telling a DoH lookup from a page load.\nThe two things you would use to regain control, close the port and read the traffic, are both gone by design. Defeating exactly those two, in the hands of a hostile network, is what DoH was built to do.\nAnd reading the traffic is not the escape it sounds like, because there is only one way to do it: break and inspect all of it. To see the DNS inside a 443 connection you have to man-in-the-middle every HTTPS connection on the network, put your own root certificate on every device, and decrypt and re-encrypt the lot. That is what a growing number of admins now resort to, to claw back the visibility a single firewall rule used to give them for nowt.\nAnd it is a far worse trade: you have weakened TLS for everything, put a decryption box in the path of every login and every bank session, and built one target that compromises all of it, a posture US-CERT warned against outright11. The proportionate control was broken, so the disproportionate one is what is left.\nWhich is the bind. The mechanism that shields a journalist on airport WiFi from a hostile network is the one that shields malware on your network from you, and the protocol cannot tell the two apart. A resolver being bypassed does not know whether it is a censor or a security team. It only knows it has been cut out.\nThe Correct Answer Was Always DNS Over TLS Here is the part that gives the game away. Encrypting DNS never needed any of this. It was done two years before DoH, in a way that left the person running the network able to do their job.\nDNS over TLS, RFC 7858, is from May 2016. Its abstract says what it is for in the first line: to \u0026ldquo;provide privacy for DNS\u0026rdquo;, encryption that \u0026ldquo;eliminates opportunities for eavesdropping and on-path tampering with DNS queries in the network\u0026rdquo;12. That is the whole privacy case, met: the coffee shop and the ISP shut out exactly as they are with DoH, because the lookup is encrypted end to end.\nAnd it was co-authored by Paul Hoffman, who then co-authored the DoH standard too1. Not two rival camps, then. The same people, who had already solved privacy, solving it again a different way.\nWho Hoffman is matters, because it says where this came from. He is no bystander in DNS: his name is on more than eighty RFCs, the DNS terminology standard among them, and he does the work as a technologist at ICANN13. So the standard that hid DNS on 443 was co-written from inside ICANN, the American body that decides what goes in the root.\nAnd ICANN is not the disinterested steward the word implies. I have set its record out separately: built in California, answerable to California, and willing to use its position over the namespace for its own ends. This was no privacy campaign from the edge. It was the establishment that holds the root, meeting a privacy case it had already met in 2016 a second time, in a way that removed the network operator.\nSo ask what the second way added, because privacy was not it.\nStandard Year Port Encrypts DNS Operator can still see it is DNS DNS over TLS (RFC 7858) 2016 853 Yes Yes DNS over HTTPS (RFC 8484) 2018 443 Yes No Privacy was solved in 2016 by DoT. DoH came two years later and added only bypass. Encrypted DNS: what came first, and what came after May 2016 DNS over TLS (RFC 7858), port 853 Encrypts DNS. Operator can still see it is DNS. Privacy solved. Oct 2018 DNS over HTTPS (RFC 8484), port 443 Same lead author. Same encryption. Operator can no longer see it. Jul 2019 Godlua: first malware to hide its C2 over DoH Jul 2019 ISPA brands Mozilla an \"Internet Villain\" for DoH, then withdraws it Sep 2019 PsiXBot resolves its C2 domains over Google's DoH Feb 2020 Firefox turns DoH on by default in the US, resolver Cloudflare Q2 2020 OilRig (APT34) exfiltrates data over DoH The privacy was complete in 2016. Everything below the second dot is bypass and what bypass was used for. None of it required a new privacy standard. The order is the argument. DNS over TLS solved the privacy problem in 2016, on a port the network operator can still govern. DNS over HTTPS arrived two years later from the same lead author, added nothing to the privacy but the move to port 443, and everything below that second point is bypass and what bypass was used for. The only box that changed is the last one. DoT runs on its own port, 853, so the operator can see that it is DNS and decide what happens to it: allow it to the sanctioned resolver, refuse it elsewhere. DoH puts the same encrypted lookup on 443 and mixes it into the web, where the operator cannot pick it out.\nSame privacy, same encryption. The single difference between the two standards is whether the person running the network can still see their own DNS. That is not a privacy feature. Privacy shipped in 2016. It is a bypass feature, and it is the only thing DoH adds.\nAnd they knew it. This is documented, not inferred. RFC 8484\u0026rsquo;s own Operational Considerations say it plainly: \u0026ldquo;Filtering or inspection systems that rely on unsecured transport of DNS will not function in a DNS over HTTPS environment due to the confidentiality and integrity protection provided by TLS\u0026rdquo;8.\nThose systems are the security tools already protecting users: malware and command-server blocking, Protective DNS feeds, child-safety filters, enterprise inspection. The standard names them, in its own text, as the things that stop working. Breaking tooling that was already defending people was a choice, taken with eyes open and shipped on by default. Knowing the effect and choosing it anyway is a decision they can be asked to account for.\nWhich is why the privacy framing does not survive the timeline. Once DoT exists, privacy cannot be the reason for DoH, because privacy was already done, by the same author, two years earlier. So privacy was the smokescreen.\nPull it away and look at what the second standard actually set running: an arms race. It took a control that was one firewall line and turned it into a permanent chase, resolvers against blocklists against fresh resolvers, the operator losing by default, and a handful of US firms holding the plaintext at the end of it.\nCall that a side effect of a privacy feature if you like. On the evidence, with DoT already shipped and the standard itself admitting it would break filtering, it is the point, and privacy was the word painted on the box. And it was pushed hard under that word, by the parties who gained.\nPushed and shipped DoH What else they are Mozilla Co-authored RFC 8484, then turned DoH on by default for US Firefox, resolver Cloudflare Google Ships DoH in Chrome, runs a large public DoH resolver, and is an advertising company Cloudflare The default US Firefox resolver, so it receives those queries The people selling it as privacy were the browser vendors, and the beneficiaries were not only the users.\nCan I ask why, if the aim was privacy, the answer was not the encrypted-DNS standard that already existed and kept the operator in the loop, but a second one built so the operator could not see it? I am not asking who signed it off. I am asking what part of \u0026ldquo;privacy\u0026rdquo; required cutting out the one person accountable for the network.\nBecause that is the real target, and it is the network administrator: the professional responsible for the security and the safety of the network, treated as the adversary by default. DoH does not even remove the surveillance the privacy pitch complained about. As Bert Hubert of PowerDNS set out, DNS was \u0026ldquo;typically provided by the operator of a network\u0026rdquo;, and moving it to a third party by default is \u0026ldquo;a net-negative for privacy for everyone\u0026rdquo;, because that third party \u0026ldquo;gets a complete log per device of all DNS queries, in a way that can even be tracked across IP addresses\u0026rdquo;14.\nThe lookups are not hidden. They are handed to a US resolver operator instead of the admin, and the admin is the only party removed. The vendors know it, too. The \u0026ldquo;detect a managed network and stand down\u0026rdquo; logic in the next section exists precisely because the default overrides the admin, and they shipped the off-switch rather than change the default.\nNow ask who gains from routing DNS past the operator. The poster child is the journalist on hostile airport WiFi, and that person is real. So is the advertising and tracking industry, whose domains walk past a network\u0026rsquo;s blocklist the moment an app or a browser resolves them over DoH.\nNetwork-level blocking, the Pi-hole, the corporate blocklist, the filtering resolver, works by answering a tracker\u0026rsquo;s name with nothing. DoH is how the name gets answered anyway. And the largest operator of a public DoH resolver is Google15, an advertising company that also ships DoH in its own browser16.\nA privacy feature, sold by the firm that sells the tracking, whose design defeats the tool people run to block it. Make of that what you will. I have made my mind up.\nEven the crude version of this objection got aired. In July 2019 the UK ISP trade body nominated Mozilla an \u0026ldquo;Internet Villain\u0026rdquo; for DoH that would \u0026ldquo;bypass UK filtering obligations and parental controls, undermining internet safety standards in the UK\u0026rdquo;17. They picked a clumsy target and a worse frame, got shouted down, and withdrew it.\nStrip the politics off, though, and the observation underneath was correct, and it holds whether the control being bypassed is a child-safety filter, a company blocklist or a Pi-hole in a spare room. DoH is built to get DNS past whoever runs the network.\nAnd it is not a small matter, because DoH is now the hardest of the three to block at all. DoT sits on 853, where a network can still refuse it. DoH does not, so of every way DNS can leave a network it is the one with no handle on it, which makes it the biggest single risk of the lot.\nThat reaches the law, not just security. In the UK the blocking that carries legal weight is run at the ISP by DNS: the Internet Watch Foundation\u0026rsquo;s list of child sexual abuse material, and the copyright injunctions the High Court hands down under section 97A, are enforced through DNS and IP filtering on the resolver a customer is handed18.\nRoute the DNS half around that over DoH and the block does not apply to that client. A court can order a site blocked, a network can be told to enforce it, and a browser resolving over DoH finds the address anyway, over a channel the network cannot see and cannot close. Enforcing a cyber law or a court order at the network level, which is where it has always been done, becomes very close to impossible.\nAnd this is where the bill lands on everyone, not just the network. Break the proportionate control, the surgical one that blocked a named domain at the resolver under a court order, and the state does not give up. It reaches for a blunter instrument. The UK\u0026rsquo;s Online Safety Act is that instrument: sweeping duties on platforms, mandated age checks, and Ofcom behind them with the fines, a regime that touches every user and every service in the country19.\nI am not defending the Act. I am pointing at where the pressure for it came from. When the narrow tool that enforced the law at the network stops working, you do not get less enforcement, you get cruder, broader, more intrusive enforcement aimed at everyone, because the aimed version no longer holds. That is the price of breaking a basic control, and the people who broke it are not the ones paying it.\nAnd ICANN could not care less, because from California it is not their problem. When a block fails or a platform breaches its duties, it is not ICANN that gets held up in court, it is the operator, the platform, the company, facing Ofcom and the fines. The fallout lands on everyone downstream of a decision the body that took it never has to answer for.\nBreak the surgical control and the state reaches for the sledgehammer. Everyone pays. Break the proportionate control, and a blunter one takes its place The surgical control Block one named domain at the resolver, under a court order Precise, proportionate, aimed only at the target DoH breaks it the lookup no longer passes the resolver The blunt instrument The Online Safety Act: age checks on everyone, duties on every platform, Ofcom with the fines Aimed at the whole country Break the narrow tool and the state does not give up. It reaches for the broad one, aimed at everyone, because the aimed version no longer holds. And the people who broke the narrow tool are not the ones who pay for the broad one. The surgical control blocked one named domain at the resolver under a court order. DoH breaks it, so the state reaches for the blunt instrument: the Online Safety Act, aimed at every user and every platform in the country. Break the narrow tool and you do not get less enforcement, you get a broader one nobody can aim. Line the consequences up and they are all the same shape: a thing the network used to enforce at the resolver, and no longer can.\nWhat the network enforced How it was done Under DoH Malware and command-server blocking resolver blocklist and threat feed bypassed; Godlua, PsiXBot and OilRig did exactly this Ad and tracker blocking Pi-hole or a filtering resolver bypassed; the tracker\u0026rsquo;s name resolves anyway Court orders and the IWF list ISP DNS filtering DNS-based block does not apply to that client DoT was the honest answer. It encrypts your DNS and leaves the operator able to do their job. DoH kept the encryption and added one thing on top: the operator can no longer see it. Everything else about it is downstream of that one choice.\nMade In The US, Fought Over Everywhere Else None of this stops at the English Channel, and reading it as a British problem shrinks it to a fraction of its size. The same shape runs right across Europe and out to Australia. A legal instrument in at one end, a DNS block on the network at the other.\nAustralia has its Federal Court order the ISPs to disable access to overseas infringing sites under section 115A of the Copyright Act20. Portugal skips the court entirely: an administrative memorandum under which rights holders notify a body called MAPINET, and the ISPs block by DNS inside fifteen working days21. Different statutes, one load-bearing assumption, and it is the assumption DoH kicks out from under all of them. That the client resolves through the resolver it was handed.\nSo watch what a state does when that assumption fails. It stops leaning on the ISP and starts ordering the public resolvers themselves to return the wrong answer. That is not a forecast. It is a stack of judgments, and it is getting taller.\nCountry The blocking order The public resolver What happened Italy AGCOM\u0026rsquo;s Piracy Shield, names blocked inside 30 minutes ordered to filter, 1.1.1.1 included Cloudflare refused, was fined €14,247,698, and is appealing22 France Canal+ sports-streaming injunction Google, Cloudflare and OpenDNS all named Cloudflare serves an HTTP 451, Google fails the query in silence, OpenDNS switched itself off for the country23 Belgium over 100 sports-piracy domains the same three resolvers Cloudflare complies with a 451, Google stays silent, OpenDNS left Belgium as well23 Germany Universal Music v Cloudflare ordered at first instance, 1.1.1.1 included overturned on appeal, Cologne holding the resolver \u0026ldquo;passive, automatic and neutral\u0026rdquo;24 Netherlands BREIN\u0026rsquo;s dynamic blockade ISP level, the resolver not yet Ziggo, KPN and the rest block by DNS and IP, updated on request25 Read that as one thing, not five. A few US companies rewrote DNS for the whole planet on their own authority, broke the proportionate control every one of these countries had built, and left the courts to work out what to do with the wreckage. The answers those firms give, hauled in front of a different bench every few months, tell you exactly what the rest of the world weighs with them. One appeals and cries censorship. One fails the lookup and tells the user nothing. And OpenDNS, the resolver that quietly filtered malware for twenty years and that this post has already held up as the model to copy, does not argue at all. It switches itself off for France and for Portugal rather than serve them on those terms26.\nSit with that last one. It is the good example in this whole piece walking out of the room. The service that protected people who never configured a thing now finds a whole country easier to abandon than to serve, and the reason traces straight back to a design decision taken in California by people who never had to think about a French court or a Portuguese one.\nThat is the pattern, and I will name it. This is what US platform engineering does to everywhere that is not the United States. It builds for its own market, ships the result to the world as the default, and treats every other country\u0026rsquo;s law, every regulator, every community left downstream of the change as somebody else\u0026rsquo;s mess to clear up.\nAnd the home government backs the firms to the hilt. In February 2025 the White House put its name to a memorandum casting other countries\u0026rsquo; digital laws as \u0026ldquo;overseas extortion\u0026rdquo; of American companies, naming the UK and the EU and pointing the trade representative at tariffs by way of reply27. Six months on, a US regulator wrote to a dozen American tech firms warning that obeying the UK\u0026rsquo;s Online Safety Act or the EU\u0026rsquo;s Digital Services Act might itself break US law28. Read that twice. The country whose companies broke the controls now tells those companies that complying with another democracy\u0026rsquo;s answer to the breakage is the offence.\nThat says all of it. A law passed by an elected parliament in Westminster or Brussels is recast, from Washington, as an attack on America, and the firms are told to ignore it. It is not that they weighed the rest of the world and chose against it. The rest of the world was never on the scale.\nThe Fix Is Not To Ban Encryption So the lazy conclusion is \u0026ldquo;block DoH\u0026rdquo;, and it is wrong, the same way \u0026ldquo;block ICMP\u0026rdquo; is wrong on the firewall. You do not want unencrypted DNS back. Clear-text lookups on port 53 are a real exposure and going back to them to regain visibility is trading one hole for another.\nThe thing you actually lost was not the encryption. It was the choice of resolver. So take that back and leave the encryption exactly where it is.\nKeep the encryption. Take back the choice of resolver. Encrypted all the way. One machine allowed to talk DNS to the world. Clients 53, DoT or DoH, to your resolver only DDR tells them where Your resolver serves DoH and DoT itself blocklist, threat feed, log canary answered NXDOMAIN Border only this Upstream or the roots client straight to 53, 853 or a public DoH resolver: refused Nothing here turns encryption off. The only thing taken away from the client is the right to choose somebody else's resolver. The corrected layout. Clients may use plain DNS, DNS over TLS or DNS over HTTPS, but only to your own resolver, which now serves all three. That resolver keeps the blocklist and the log and is the only machine allowed to send DNS of any kind past the border. Port 53, port 853 and known public DoH resolvers are refused from everything else. Nothing here turns encryption off. The only thing taken from the client is the right to choose somebody else\u0026rsquo;s resolver. The shape is the same one that fixes the DNS in an Active Directory domain: the argument is never \u0026ldquo;turn the feature off\u0026rdquo;, it is \u0026ldquo;decide who is allowed to drive it\u0026rdquo;. Here is what that means in parts.\nRun encrypted DNS yourself. Stand up a resolver that serves DoH and DoT, not just port 53, so a client that wants encryption gets it from you. You cannot tell clients \u0026ldquo;no DoH\u0026rdquo; and mean it while offering no encrypted resolver of their own. Give them one, and keep the blocklist, the threat feed and the logging on it as before.\nAdvertise it, so clients find it on purpose. Discovery of Designated Resolvers (RFC 9462) lets a client query a special name, learn its network\u0026rsquo;s own encrypted resolver, and upgrade to DoH or DoT against that automatically29. A DDR-aware client then encrypts its DNS and uses yours: the privacy the user wanted and the control you needed, in one move.\nAnswer the browser\u0026rsquo;s own off-switch. Firefox checks a canary domain, use-application-dns.net, before turning DoH on by default: answer it with NXDOMAIN or SERVFAIL and Firefox stands down to the system resolver30. Two honest limits, both in Mozilla\u0026rsquo;s words. The canary \u0026ldquo;only applies to users who have DoH enabled as the default option. It does not apply for users who have made the choice to turn on DoH by themselves\u0026rdquo;30. So it is a default-off signal, not a lock, and it does nothing on the app and the malware, which never asked the browser anything.\nSet the browser policy where you manage the machine. On a device you own, do not rely on the canary. Chrome\u0026rsquo;s DnsOverHttpsMode takes off, automatic and secure, and Google\u0026rsquo;s docs state that where it is unset, \u0026ldquo;for managed devices DNS-over-HTTPS queries will not be sent\u0026rdquo;31; Chrome disables its own auto-DoH the moment it sees an enterprise policy16. Point it at your resolver, or set it off and let DDR carry the encryption. Firefox has the matching Enabled, ProviderURL and Locked knobs32. Managed machine, managed decision, and the row you can actually close.\nRefuse the alternatives at the border. The rules are short, and they all say the same thing: DNS of any kind, to the outside, only from your resolver.\nRule From Effect Block outbound port 53 Everything except your resolver No plain DNS to the outside Block outbound port 853 Everything except your resolver No DoT to an outside resolver Block HTTPS to public DoH resolver addresses Everything except your resolver Cuts off the well-known DoH providers Point your resolver\u0026rsquo;s upstream at Protective DNS Your resolver Malware domains refused before connection You will not catch every DoH endpoint by address, and you cannot, because a new one is just a web server and there is no end of web servers. So sit with what that block now costs. To do for DoH what block 53 outbound did for free, you pull a list of DoH server addresses that somebody else keeps scanning for you: the maintained blocklists exist because this is the only way left, and one of them re-resolves every known public DoH domain to its current IPs every hour, by scheduled job, and ships the changes33. That is the trade the bypass forced.\nBlocking outside DNS Then, on port 53 Now, DoH on 443 The rule one line: block 53 outbound subscribe to a DoH-server IP blocklist Completeness complete the moment it was saved never complete; new endpoints keep appearing Upkeep none rescanned hourly, pulled and reapplied What it takes the firewall the firewall plus somebody else\u0026rsquo;s maintained feed A control that was one static line is now a subscription to a moving target, chasing web servers that anyone can stand up faster than a list can name them. Protective DNS closes the other end: point your resolver\u0026rsquo;s upstream at a filtering service like the NCSC\u0026rsquo;s PDNS, which \u0026ldquo;was built to hamper the use of DNS for malware distribution and operation\u0026rdquo;34, and the bad domains are refused before the connection is made, for every client that comes through your resolver, which, once these rules are in place, is all of them.\nPut those against the five rows from earlier, and every one of them has a control that closes it. That is the test of a fix: not \u0026ldquo;did we block DoH\u0026rdquo;, but \u0026ldquo;for each thing that can choose a resolver, what stops it choosing the wrong one\u0026rdquo;.\nWho picks the resolver What closes it Operating system DDR points it at your resolver; keep it that way Browser Policy on managed machines; the canary on the rest Application / library Border block on 53, 853 and public DoH resolvers Script on a web page Same border block; Protective DNS on your upstream Malware Same border block, plus Protective DNS refusing the domain None of that turns a single bit of encryption off. Clients still get DoH, they just get it from you, and the ones that try to get it from somewhere else hit a wall. The privacy is intact and the control is back. As such, the only thing anyone lost is the freedom to pick a resolver you never approved, and that was never theirs to have on a network you are responsible for.\nCan I Ask What Chose That Resolver Here is the question to put to anyone who tells you the network is fine as it is.\nWhen a machine on your network resolves a name, what decided which resolver answered? If the honest answer is \u0026ldquo;whatever the OS was handed\u0026rdquo;, good, but only if you have closed the other four rows of that table, because otherwise it is \u0026ldquo;whatever the OS was handed, unless the browser, an app, a web page or something nastier chose different, in which case I have no idea and no record\u0026rdquo;. That is not a DNS setup. That is a hope, and hope catches nowt.\nI am not asking who runs the resolver. I am asking what, on each machine, gets to choose it, and whether you could name the resolver every lookup went to yesterday. On a network where DoH is unmanaged you cannot, and the gap is every app that bundles its own resolver, every page a user opens, and every piece of malware that read the same research I just linked. Three named threat actors built their channels on this exact gap, and one answers to a government.\nA protocol that takes the choice of resolver away from the network was always going to be a gift to whoever the network was trying to watch, and pretending otherwise because the marketing said \u0026ldquo;privacy\u0026rdquo; is how a control that mattered gets switched off by default and nobody logs the day it happened.\nEncrypt your DNS. Run the resolver yourself. Advertise it, answer the canary, set the policy, close the border. Keep the encryption and keep the choice. You can have both, and if you run networks for a living, having both is the job.\nPrivacy Was The Word On The Box So credit where it is due. Thanks to ICANN, the American body that holds the root and co-wrote this from the inside, for running the US tech playbook one more time: take a problem that was already solved, wrap the fix in a virtuous word, and centralise the result onto a handful of US firms who end up holding the plaintext.\nAnd it is always the US. The body that holds the root, the resolver two browsers now default to, the advertising firm running the biggest public one, the country the plaintext lands in: American, every time. It is a pattern, not a run of bad luck.\nPrivacy was the smokescreen. DNS was encrypted, private and still governable on port 853 in 2016, and everyone who mattered knew it. What the establishment shipped on top was a protocol built to blind the one person accountable for the network, a permanent arms race of blocklists rescanned by the hour, and court orders that stop dead at the browser.\nWhether that was ineptitude or greed I will let you decide, though the record I linked points one way. Either way it beat common sense.\nAnd decisions like this one are why the calls to take the root off ICANN keep coming back, and why they land: so the international community gets a say, instead of one country\u0026rsquo;s institution deciding the DNS for everyone else and calling it stewardship.\nWe used to do all of this with one line.\nRFC 8484 — DNS Queries over HTTPS (DoH), October 2018. The introduction states the goal of \u0026ldquo;allowing web applications to access DNS information via existing browser APIs in a safe way consistent with Cross Origin Resource Sharing (CORS)\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n360 Netlab — An Analysis of Godlua Backdoor, 1st July 2019. Records that the sample \u0026ldquo;uses DNS over HTTPS to get the C2 name to ensure secure communication between the bots, the Web Server and the C2\u0026rdquo; — the first widely reported malware to abuse DoH.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nProofpoint — PsiXBot Now Using Google DNS over HTTPS, 6th September 2019. Reports the malware resolving hardcoded C2 domains through Google\u0026rsquo;s DoH service, and warns it \u0026ldquo;should be a warning shot for the cybersecurity community as there are no simple solutions to help identify an infected host or the process of receiving instructions\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nKaspersky Securelist — APT trends report Q2 2020. On OilRig (APT34): the DNSExfiltrator tool \u0026ldquo;allows the threat actor to use the DNS over HTTPS (DoH) protocol [\u0026hellip;] instead of plain text requests to port 53, they would use port 443 in encrypted packets [\u0026hellip;] which allows DoH queries to Google and Cloudflare services\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nThe Hacker News — ChamelDoH: New Linux Backdoor Utilizing DNS-over-HTTPS Tunneling for Covert CnC, 16th June 2023, reporting Stairwell\u0026rsquo;s finding (Daniel Mayer). ChamelGang\u0026rsquo;s C++ Linux implant is \u0026ldquo;a tool for communicating via DNS-over-HTTPS (DoH) tunneling\u0026rdquo;, sending DNS TXT requests to rogue nameservers through legitimate DoH providers so that \u0026ldquo;both detection and prevention become difficult\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCISA — Malware Analysis Report: BRICKSTORM Backdoor, AR25-338a, 4th December 2025. \u0026ldquo;For C2, BRICKSTORM uses multiple layers of encryption (HTTPS, WebSockets, nested Transport Layer Security [TLS]) to hide its communications with the cyber actors\u0026rsquo; C2 server. It also uses DNS-over-HTTPS (DoH)\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCisco Talos — New Dohdoor malware campaign, 26th February 2026. The backdoor, attributed to actor UAT-10027, \u0026ldquo;securely sends encrypted DNS requests to Cloudflare\u0026rsquo;s DNS server over HTTPS port 443\u0026rdquo; to resolve its command server, and targets US education and healthcare through phishing.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 8484, Section 10, Operational Considerations: \u0026ldquo;Filtering or inspection systems that rely on unsecured transport of DNS will not function in a DNS over HTTPS environment due to the confidentiality and integrity protection provided by TLS.\u0026rdquo; The effect on network filtering is stated in the standard itself.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 8484, Section 9, Security Considerations: a DoH server \u0026ldquo;can give a client invalid data in response to a DNS query\u0026rdquo;, and while the standard disallows responses from unconfigured servers, \u0026ldquo;this prohibition does not guarantee protection against invalid data, but it does reduce the risk.\u0026rdquo; DoH secures the transport, not the authenticity of the record; that is DNSSEC\u0026rsquo;s job.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nOpenDNS — a recursive DNS resolver founded by David Ulevitch in 2005/2006, providing security filtering (phishing and malware blocking) and content filtering at the resolver for home and enterprise users; acquired by Cisco in 2015 and now part of Cisco Umbrella. Resolver-level DNS security long predates DoH.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nUS-CERT / CISA — HTTPS Interception Weakens TLS Security, alert TA17-075A, 16th March 2017. Warns that intercepting HTTPS by presenting a locally-trusted certificate and decrypting traffic weakens security, because many interception products fail to properly verify certificates and downgrade the connection\u0026rsquo;s protections.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 7858 — Specification for DNS over Transport Layer Security (TLS), May 2016. The abstract: \u0026ldquo;This document describes the use of Transport Layer Security (TLS) to provide privacy for DNS. Encryption provided by TLS eliminates opportunities for eavesdropping and on-path tampering with DNS queries in the network.\u0026rdquo; Co-authored by P. Hoffman (ICANN), who also co-authored RFC 8484. DoT uses the dedicated port 853.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPaul Hoffman (engineer) — \u0026ldquo;the author or co-author of over 80 Requests for Comments (RFCs)\u0026rdquo; and \u0026ldquo;currently a technologist at ICANN\u0026rdquo;; his IETF datatracker profile lists the record, including RFC 7858 (DoT), RFC 8484 (DoH) and RFC 8499 (DNS terminology).\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nBert Hubert (PowerDNS) — Centralised DoH is bad for Privacy, in 2019 and beyond, RIPE Labs. DNS is \u0026ldquo;typically provided by the operator of a network\u0026rdquo;; centralised DoH \u0026ldquo;by default\u0026rdquo; is \u0026ldquo;a net-negative for privacy for everyone\u0026rdquo;, because the third party \u0026ldquo;gets a complete log per device of all DNS queries, in a way that can even be tracked across IP addresses\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nGoogle — DNS-over-HTTPS (DoH) — JSON API. Google Public DNS, one of the largest public resolvers, operates a DoH endpoint (dns.google) alongside Chrome\u0026rsquo;s own DoH support.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nChromium Blog — A safer and more private browsing experience with Secure DNS, May 2020. \u0026ldquo;If you are an IT administrator, Chrome will disable Secure DNS if it detects a managed environment via the presence of one or more enterprise policies.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nTechCrunch — Internet group brands Mozilla \u0026lsquo;internet villain\u0026rsquo; for supporting DNS privacy feature, 5th July 2019 (Wayback snapshot). The ISPA nomination said DoH would \u0026ldquo;bypass UK filtering obligations and parental controls, undermining internet safety standards in the UK\u0026rdquo;; the nomination and the Internet Villain category were withdrawn after backlash.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nWeb blocking in the United Kingdom — the Internet Watch Foundation child-abuse-image list and copyright injunctions under section 97A of the Copyright, Designs and Patents Act 1988 are implemented by ISPs; \u0026ldquo;The technical measures used to block sites include DNS hijacking, DNS blocking, IP address blocking, and Deep packet inspection.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nUK Government — Online Safety Act: explainer. \u0026ldquo;The Online Safety Act 2023 [\u0026hellip;] puts a range of new duties on social media companies and search services\u0026rdquo;, with the \u0026ldquo;strongest protections [\u0026hellip;] designed for children\u0026rdquo;, including requirements to prevent children accessing harmful and age-inappropriate content; enforced by Ofcom.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCopyright Act 1968 (Australia), section 115A, \u0026ldquo;Injunctions relating to online locations outside Australia\u0026rdquo; — the Federal Court may order a carriage service provider to \u0026ldquo;take reasonable steps to disable access to the online location\u0026rdquo;, the Australian form of ISP-level site blocking by DNS and IP.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nEDRi — Portugal: \u0026ldquo;Voluntary\u0026rdquo; agreement against copyright infringements. Under a 2015 memorandum of understanding, rights holders notify MAPINET, which forwards to the regulator IGAC, and \u0026ldquo;IGAC then contacts Internet Service Providers (ISPs) to restrict access to the websites through \u0026lsquo;Domain Name System (DNS) blocking\u0026rsquo;\u0026rdquo; within 15 working days, with no court order.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nTorrentFreak — Italy Fines Cloudflare €14 Million for Refusing to Filter Pirate Sites on Public 1.1.1.1 DNS. Under Italy\u0026rsquo;s Piracy Shield, AGCOM (order 49/25/CONS) ordered DNS providers including Cloudflare\u0026rsquo;s public resolver to block; Cloudflare, calling it \u0026ldquo;unreasonable and disproportionate\u0026rdquo;, refused and was fined €14,247,698, which it is appealing.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nTorrentFreak — DNS Piracy Blocking Orders: Google, Cloudflare, and OpenDNS Respond Differently, 11th May 2025. Under Canal+ sports-streaming injunctions in France and Belgium the public resolvers were ordered to block: Cloudflare returns an HTTP 451, Google silently refuses the query, and OpenDNS \u0026ldquo;pulled the plug\u0026rdquo; on both countries rather than comply.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nTorrentFreak — Cloudflare Applauds Court for Rejecting DNS Piracy Blocking Order. In Universal Music v Cloudflare (the DDL-Music case) a lower court had ordered Cloudflare to block on its 1.1.1.1 resolver; the Higher Regional Court of Cologne declined to extend the duty to the resolver, which \u0026ldquo;contributes to the connection of internet domains in a purely passive, automatic and neutral manner\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nTorrentFreak — Dutch ISPs Must Block Pirate Bay Proxies and Mirrors Again, Court Rules, 15th October 2020. BREIN holds a \u0026ldquo;dynamic\u0026rdquo; blocking order against Ziggo, KPN and others; the court treats DNS blocking as a \u0026ldquo;clear and verifiable\u0026rdquo; measure, and new domains and proxies are added to the ISP blocklist on request.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nComplete Music Update — OpenDNS pulls plug on France and Portugal after web-blocking injunctions. Cisco\u0026rsquo;s OpenDNS: \u0026ldquo;Due to a court order in France issued under the French Sport code and a court order in Portugal issued under the Portuguese Copyright Code, the OpenDNS service is not currently available to users in France [\u0026hellip;] and in Portugal.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nThe American Presidency Project — White House Fact Sheet: Directive to Prevent the Unfair Exploitation of American Innovation, 21st February 2025. The memorandum directs the US Trade Representative to consider tariffs in response to other countries\u0026rsquo; digital-services taxes and regulations, the EU\u0026rsquo;s and UK\u0026rsquo;s included, framed as foreign governments appropriating \u0026ldquo;America\u0026rsquo;s tax base\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nA\u0026amp;O Shearman — FTC chairman warns major tech companies of censorship in EU DSA and UK Online Safety Act, on letters of 21st August 2025: \u0026ldquo;Compliance with the requirements of non-US laws, such as the requirements in the EU Digital Services Act (DSA) and UK Online Safety Act, may result in companies censoring content\u0026rdquo;, which the regulator warns could conflict with US law.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 9462 — Discovery of Designated Resolvers (DDR), November 2023. Defines how a client discovers \u0026ldquo;a resolver\u0026rsquo;s encrypted DNS configuration\u0026rdquo; and upgrades to it, querying _dns.resolver.arpa for the network\u0026rsquo;s Designated Resolver.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMozilla — Canary domain — use-application-dns.net. \u0026ldquo;The canary domain only applies to users who have DoH enabled as the default option. It does not apply for users who have made the choice to turn on DoH by themselves.\u0026rdquo; A response other than NOERROR with an A/AAAA record — such as NXDOMAIN or SERVFAIL — signals Firefox to disable application DoH.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nGoogle — DnsOverHttpsMode policy. Values off, automatic and secure; \u0026ldquo;If this policy is unset, for managed devices DNS-over-HTTPS queries will not be sent.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMozilla — Policy templates: DNSOverHTTPS. Enabled turns DoH on or off, ProviderURL sets the resolver, Locked \u0026ldquo;prevents the user from changing DNS over HTTPS preferences\u0026rdquo;, ExcludedDomains and Fallback tune the rest.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\ndibdot/DoH-IP-blocklists — a maintained list of the domain names and resolved IPv4/IPv6 addresses of public DoH servers, for firewall blocking; the lookup script \u0026ldquo;runs automatically every hour via GitHub actions\u0026rdquo; and updates the address lists when they change. jameshas/Public-DoH-Lists is a second, auto-generated equivalent. That these have to exist, and be rescanned continuously, is the point.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNCSC — Protective Domain Name Service (PDNS). \u0026ldquo;PDNS was built to hamper the use of DNS for malware distribution and operation [\u0026hellip;] a recursive resolver which prevents access to domains known to be malicious.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/dns/dns-over-https-walks-past-your-controls/","summary":"DNS over HTTPS encrypts your lookups, which is good, and sends them to a resolver of the client\u0026rsquo;s choosing on port 443, which is the problem. Your own resolver never sees the query, so the blocklist it would have failed, the threat feed it would have hit and the log line it would have written all vanish. This post walks the mechanism: why one encrypted web connection among thousands is invisible at the border, who on a machine can pick a resolver you never chose (the OS, the browser, any app, any script on any web page, and malware), and the documented cases, from a page\u0026rsquo;s own JavaScript querying a public DoH endpoint that allows cross-origin requests, to Godlua and PsiXBot hiding their command channel inside DoH, to the OilRig APT exfiltrating data over it. Then the case that the privacy story is a cover: DNS over TLS already encrypted DNS in 2016, on a port the network operator can still govern, so the only thing DoH adds over it is defeating the admin, which is why the browser vendors who wrote and shipped it, and the advertising firm that runs the biggest public DoH resolver, are the ones who gained. Then the fix, which is not to ban encryption. Run your own DoH and DoT resolver, advertise it with Discovery of Designated Resolvers, refuse port 53, port 853 and known public DoH resolvers at the border, and answer the Firefox canary so the browser stands down. Keep the encryption. Take back who chooses the resolver. It closes on the fallout: the same DNS-based enforcement runs across Europe and in Australia, and courts in Italy, France, Belgium and Germany are now ordering the public resolvers themselves to block, with Cloudflare fined and appealing, Google refusing in silence and OpenDNS switching itself off for whole countries. The pattern underneath is a US design imposed on the world, and a US government that casts other countries\u0026rsquo; laws as extortion of American firms and warns them not to comply.","title":"DNS Over HTTPS Walks Straight Past Your Controls"},{"content":"You click a link to a local paper and the headline loads, then the picture, then the first three paragraphs, and you have started reading before anything else happens. Then the page goes blank. Or it turns into a heap of unstyled links with the logo stretched across the full width of the screen. Or a box slides up asking you to either accept tracking or pay £2.99 a month, with no third button.\nThe article was on your machine. The publisher\u0026rsquo;s server had already sent all of it, text and stylesheets, and your browser had already drawn it. What took it away was code the publisher chose to run afterwards, on your computer, bought from a commercial vendor whose product is taking back what was just delivered.\nI run a Pi-hole1 at home, which blocks advertising and tracking hosts for every device on the network. On a growing list of UK news sites that was enough to get the page destroyed. So on 22nd September 2026 I built a Chrome extension to stop it, called Keep The Page. The code is on GitHub at damo2929/browserplugin, MIT licensed, and this is what it does, what I found inside those pages while building it, and the fixes that made things worse before the right one turned up.\nIt is not a paywall bypass. If the server never sent the article, nowt here will conjure it up. The Times stayed shut and that is correct. What it keeps is what you were already sent.\nThe Article Arrives, Then It Goes Every one of these works the same way from the outside. The HTML arrives complete, and then a script, usually loaded from a host belonging to a vendor rather than the paper, decides you are not worth showing it to, and removes it.\nWhat it removes, and how, depends on which vendor the publisher bought. These are the four I met, each read off a live page during the build:\nPublisher What arrives What the page then does to itself Newsquest (269 titles)2 the full article builds a wall, runs an eval payload, raises a confirm() dialog notebookcheck.net the full article deletes \u0026lt;body\u0026gt; about 7 seconds in, then dialogs, then a reload loop National World (The Scotsman, Yorkshire Post)3 the article plus about 68KB of its own CSS deletes every \u0026lt;link\u0026gt; and \u0026lt;style\u0026gt; every 100ms, forever Reach plc (Mirror, Daily Record, Manchester Evening News, Liverpool Echo and others)4 the full article covers it with \u0026ldquo;accept tracking or pay £2.99/month\u0026rdquo; The 269 is not from a press release, which says \u0026ldquo;more than 200 brands\u0026rdquo;. It is the count of domains on Newsquest\u0026rsquo;s own TLS certificates, the list the extension needed to know where to run.\nThe National World one is the one that should worry publishers most, because it breaks the page for reasons that have nothing to do with advertising. The stylesheet stripper has a partner that fetches the CSS back from the vendor\u0026rsquo;s host. My Pi-hole blocks that host. So the strip ran and the restore never came, and the paper\u0026rsquo;s own layout was left permanently dependent on an ad vendor being reachable. That is not a wall. It is a publisher handing the look of its own website to a third party and not noticing.\nCan I ask why a newspaper\u0026rsquo;s stylesheet has to wait for an advertising company\u0026rsquo;s server before it is allowed to stay on the page? There is no reader-facing reason for it. None.\nWhy I Call It Malware I use the word on purpose, as a description of what the code does, and not as a legal finding about anybody. The anti-adblock payloads meet the ordinary meaning on four counts, and every one was observed directly:\nTest What I saw Runs without consent, against your interest nobody asked for it, and its job is to take away content you already have Destroys data already delivered to you the article and its CSS arrive intact, then in-page code deletes them Obfuscated to resist reading permuted string tables like o[293 * (r + 450) % e], executed through eval Evades blocking and punishes interference CNAME cloaking to dodge DNS blocklists, anti-tamper checks that escalate to a dialog or a reload loop The obfuscation is not minification. Minified code is small. This is code arranged so that you cannot grep it for what it does, and the payload is then run through eval so nothing on disk matches what executes.\nThe cloaking is a known technique with its own research literature5. The loader is fetched from what looks like a subdomain of the newspaper, and that name resolves to the vendor:\na02342.\u0026lt;publisher-domain\u0026gt; -\u0026gt; cdn-52-x.privacy-mgmt.com (Sourcepoint) fb.html-load.com -\u0026gt; adshield-fallback-dev-wskxz.b-cdn.net The subdomain is randomised per title. As such a blocklist that names hosts cannot keep up, which is the point of doing it.\nAnd the code treats being interfered with as proof of guilt. The decoded error strings in the payload that filter-list maintainers attribute to Ad-Shield6 are literally Vital API blocked and Vital API blocked (eval). Sourcepoint\u0026rsquo;s loader writes an attribute, reads it straight back, and throws if the value has changed:\nz.call(O,\u0026#39;src\u0026#39;,G), O[x](\u0026#39;src\u0026#39;) !== G \u0026amp;\u0026amp; throw E … catch (W) { try { await l(W) } catch (x) { o(W) } } // o() raises the dialog Sourcepoint have sold this openly. Their own documentation said \u0026ldquo;on average about 30% of messaged users will turn off their adblockers\u0026rdquo;7. So I am not describing a rogue script somebody slipped in. It is a product, bought and deployed on purpose by the publisher.\nThe consent-or-pay walls are a different category and I do not call them malware. They do not destroy what was delivered and they do not hide from blocklists. They are coercive in a different way, which is the next section.\nAccept 1,467 Partners Or Pay £2.99 The wall on Reach titles is Quantcast Choice, now run by InMobi8 and served from cmp.inmobi.com. It gives two choices. Accept, or pay £2.99 a month. There is no free \u0026ldquo;no\u0026rdquo;.\nAccepting shares your data with 1,467 listed partners and writes a euconsent-v2 cookie that lasts 13 months. Nobody reads a list of 1,467 companies, nobody could weigh what each of them would do with the data if they did, and the number on its own tells you what kind of consent is being asked for.\nUK GDPR says consent has to be freely given, and that when you judge whether it was, you look at whether the service was made conditional on consent to processing it does not need9. The ICO has published guidance saying consent or pay can be lawful10, guidance it now says is under review, and the European Data Protection Board has said that for large platforms offering only the two options, in most cases it will not be11. I will not pretend the regulator has banned it. It has not.\nSo here is where I land, plainly. I chose to remove the wall and to refuse the consent. That means I read the article on the free branch without paying and without handing my data to 1,467 companies. It is a decision. The extension says so in its own settings page, and I am not dressing it up as something neutral. A consent you cannot refuse is a price, and I will not pay a price that is dressed up as a question.\nRemoving the box is the easy third of it. The other two thirds are in Answering The Consent Question Properly, because a banner you delete without answering comes back every time.\nOne Day, Eight Commits The whole thing was built on one day. There were a few hours of poking at pages before anything was committed, then eight commits between 20:32 and 22:50. I built it with Claude Code, and a lot of the graft of reading obfuscated inline script line by line was done by the agent while I watched what the pages did on screen. That split worked well, and where it went wrong is below, because it went wrong in a way worth knowing about.\nTime Commit 20:32 first commit: Newsquest, notebookcheck, National World, Reach, Page Six 20:46 the vendor behind each mechanism, written into the README 20:56 50 links off the Google News front page, as an unchosen sample 21:02 answer the consent API with every purpose denied 21:11 social links defused 22:50 switchable protections, MSN and Bing, consent refusal by click, 77 tests It is a Manifest V3 extension with two permissions, declarativeNetRequest and storage, and it has no host permissions and no network access of its own, so it cannot fetch anything, rewrite a response or talk to a server on your behalf. That constraint shaped the design more than anything else, because the interesting work has to happen inside the page.\nTwo worlds, one attribute down and one event up The page's code and the extension's settings live in different worlds Extension storage chrome.storage.local protections tracing channels social settings last 50 errors read and written by the options page Isolated world bridge.js social.js portal.js can read chrome.storage cannot touch the page's own functions Page world (MAIN) walls.js guard.js portal-early.js wraps setTimeout, cookie, __tcfapi, window.adLight no chrome.* at all Down: an attribute on the html element, only what is off \u0026lt;html data-ktp-off=\"cookies,dom\"\u0026gt; default install: nothing written Up: a ktp-report event detail: a JSON string an object does not cross reliably Chrome runs extension scripts in two worlds. The page\u0026rsquo;s world can reach the page\u0026rsquo;s own JavaScript but has no extension APIs. The isolated world can read settings but cannot touch the page\u0026rsquo;s functions. Everything crosses between them as an attribute going down and an event coming up. Chrome\u0026rsquo;s MAIN world shares the page\u0026rsquo;s own JavaScript environment12. A script there can replace setTimeout, wrap document.cookie or define window.adLight before the page does, which is exactly what fighting these walls needs. In practice it cannot call chrome.storage. The isolated world can, but cannot see the page\u0026rsquo;s functions. So a small bridge reads the settings and writes them onto \u0026lt;html\u0026gt;, and errors come back up as a CustomEvent whose detail is a JSON string, because an object does not cross that boundary reliably.\nThe attribute going down lists only what you have switched off. A default install writes nothing to the page at all. That matters because a permanent marker on \u0026lt;html\u0026gt; is precisely the kind of thing these SDKs look for, and it means a storage failure fails towards defending the page rather than away from it.\nFind The Gate, Do Not Fight The Wall The Newsquest fix is two lines of reasoning and it is the one everything else was measured against.\nTheir whole wall sits behind one flag in the page:\nvar adLight = false; // line 1647 if (adLight !== true) { …} // line 2020: loader, eval payload, confirm() adLight is the subscriber \u0026ldquo;light ads\u0026rdquo; flag. If it is true, the wall never builds. So the extension defines window.adLight as true at document_start, before the page\u0026rsquo;s own script runs, with a setter that ignores whatever is written to it:\nObject.defineProperty(window, \u0026#39;adLight\u0026#39;, { configurable: false, enumerable: true, get() { return true; }, set() { /* ignore the page\u0026#39;s \u0026#34;false\u0026#34; */ } }); A var at the top of a script does not redefine a property the global object already has. It only assigns to it13. So the page\u0026rsquo;s own var adLight = false runs, lands on the setter, and does nothing. The wall\u0026rsquo;s loader never starts, the eval payload never arrives and there is no anti-tamper check to trip, because nothing was tampered with. The flag just said yes.\nThat is the only property in the whole extension that is non-configurable. It has to be, to survive the declaration. Everything else is configurable: true, so no page ever has one of its own APIs seized for good.\nIt covers 269 titles and it cannot be switched off in the settings, which say why: it runs before chrome.storage can answer, and once set it cannot be undone. A checkbox would be decoration.\nEvery Fix That Made It Worse This is the useful section, because every mistake here is the obvious thing to try.\nWhat I tried What happened Block the loader\u0026rsquo;s host the host is a random first-party CNAME per title; and a failed fetch is itself the detection signal Guard setAttribute on injected scripts tripped Sourcepoint\u0026rsquo;s read-back check, which raised the dialog it was meant to stop Answer the confirm() with Cancel in this SDK Cancel means reload, so it caused an infinite reload loop Snapshot \u0026lt;body\u0026gt; at DOMContentLoaded to restore later the wall blanks during parsing, so the snapshot was of an empty body Guard remove() and removeChild() the wall clears the page with one innerHTML = '', not node by node Guard everything at once on notebookcheck wall escalated to a dialog, then a reload loop; strictly worse than doing nothing Redirect the vendor host to a local stub a redirect rule with only declarativeNetRequest invalidated the whole ruleset, silently killing the Newsquest rule on 269 sites Run the page guards on every site patched global prototypes on every page I visited, including my bank The redirect one is worth a second look. Chrome gives a block rule implicit access and wants host permission for anything more14. Its documentation says invalid static rules are ignored15. What I saw was worse than that: the whole file went, with no error on the page, and rule 1 just stopped existing on 269 sites. The two rules for MSN now live in a ruleset of their own for that reason, so a bad edit to one cannot take the other down.\nThe reload loop taught the other hard rule. location.reload cannot be intercepted. The Location object is unforgeable in the HTML standard16, so its methods are non-writable and non-configurable, and trying gives you:\nObject.defineProperty(location,\u0026#39;reload\u0026#39;,...) -\u0026gt; TypeError: Cannot redefine property: reload Any wall whose failure path is \u0026ldquo;reload the page\u0026rdquo; cannot be stopped once it is on that path. Nothing can. The only fix is to make sure it never gets there. Which is the same lesson as adLight, learned the painful way: every attempt to fight a wall already running made things worse, and every fix that worked stopped the wall from starting.\nOne Timer Did The Damage notebookcheck has no adLight. The wall always runs. With tracing on, the extension showed what it scheduled:\ndropped setTimeout(7005ms) scheduled from eval \u0026lt;- the body.remove() dropped setTimeout(1251ms) / 105ms / 0ms x8 dropped setInterval(15000ms) One of those timers does all the damage: the one at 7 seconds that removes \u0026lt;body\u0026gt;. Everything after it is a reaction, because the blank page throws, the exception raises the dialog, the dialog reloads the page and the whole thing starts over from the top with a fresh set of timers.\nSo the fix is one rule. Drop a timer if it was scheduled from eval\u0026rsquo;d code, the page carries this SDK\u0026rsquo;s signature and the handler is a function. The stack tells you where a call came from, and eval leaves its mark in it.\nbefore: 4 page loads, 3 confirms, reload loop after: 1 page load, 0 confirms, alive 40,378ms, content intact One timer, and everything downstream of it notebookcheck: one timer does the damage, the rest is reaction As served article and CSS arrive, drawn loader from html-load.com eval payload schedules timers 7.0s: timer runs body.remove() blank page throws, confirm() raised page reloads and starts again 4 page loads, 3 dialogs, reload loop With the extension article and CSS arrive, drawn loader from html-load.com eval payload schedules timers every eval timer dropped at once 1 load, 0 dialogs, alive at 40 seconds notebookcheck without and with the extension. Nothing downstream of the 7-second timer needs fixing, because none of it happens once that timer is dropped. There were two other mechanisms in the build at that point, a redirect of the loader to a stub and a decoy element to soak up the wall\u0026rsquo;s writes. Both worked, and both were treating symptoms of that one timer, which only became obvious when the timer rule was tested on its own and turned out to be enough by itself. So both came out, and taking them out dropped a permission and every host permission with them. Always test whether the last change alone does the job before you keep the scaffolding round it.\nThe signature is the data-sdk attribute on the loader tag, and it matches a shape, l/\u0026lt;n\u0026gt;.\u0026lt;n\u0026gt;, rather than a version. Three versions turned up in one day. It latches on and never caches a negative, because the loader tag may not be parsed yet when the first timers are scheduled, and a remembered \u0026ldquo;no\u0026rdquo; would disarm it for good on a page that does carry the wall.\nThe Probe Said Fine. The Screen Did Not. The Scotsman and the Yorkshire Post were reported fixed off a probe that measured text. The page had 8,808 characters of it, 22 elements under \u0026lt;body\u0026gt;, steady at 2, 8 and 16 seconds. By that measure it was intact.\nIt was not. I was looking at it, and every stylesheet was gone, the links were a raw list, the SVG logo filled the screen and there was a horizontal scrollbar. All the words were present, which is all that probe could see. I had to point at the screen and say so.\nWhat the probe measured, and what was on the screen Yorkshire Post, before the fix: the same page, measured two ways What the probe measured innerText.length8,808 body children22 at 2s, 8s, 16sunchanged Verdict: intact What was on the screen document.styleSheets.length0 links as a raw list, logo full width, a horizontal scrollbar Verdict: broken After the stripper was dropped at schedule time: 2 stylesheets, 532 rules, the page renders The same page, measured two ways. Text length said the page was healthy. The stylesheet count said it was broken, and the stylesheet count was right. The cause was the stripper from the first table, and it did not match the timer rule, because it is a plain inline script and not eval. So its own source is the signature: a repeating job whose body calls querySelectorAll('link,style') and then remove(). Nothing legitimate does that. It is dropped at schedule time, nothing is ever stripped and nothing needs restoring. Yorkshire Post went from 0 stylesheets to 2, with 532 rules, and it rendered.\nFinding it needed a stack trace, not a guess. Patch Element.prototype.remove to log the stack whenever it removes a STYLE or LINK, and it named the inline script and the forEach on the first try.\nAfter that, document.styleSheets.length went into every check. A page with none is broken however much text it has. Measuring gets harder than you would think, and these are the traps that caught the build on the day:\nTrap What it said What was true text length as a render check \u0026ldquo;intact\u0026rdquo; unstyled markup one sample after load Reach walls \u0026ldquo;not present\u0026rdquo; on eight titles the wall lives about 600ms and was already swept empty console after navigating \u0026ldquo;the guard is not firing\u0026rdquo; load-time messages do not survive the navigation sampling at 1s and 4s notebookcheck healthy it blanks at 5 to 8 seconds cssRules counted across origins ESPN had 5 rules cross-origin sheets throw, so they read low The fix for the 600ms one is to poll every 100ms from the moment the page loads. On Wales Online the wall appeared at 425ms and was gone by 1,129ms.\nThen the checks were run for real. 52 articles across 26 domains, two per site, chosen to cover every mechanism: 52 of 52 with stylesheets, no wall left on screen, no reload loops. Then 50 links straight off the Google News UK front page, not chosen by me: 45 rendered normally, 2 were hosts my Pi-hole blocks on purpose, 1 was The Times\u0026rsquo;s real paywall, and 2 were the same article landed twice because Google rebuilds the link order on every load. None with zero stylesheets. Three National World titles I had never tested turned up in that run carrying the same SDK, and all three rendered, which is what matching a shape instead of a site list is for.\nThere are 77 unit tests, in the repository with everything else. There is no Node on this machine, so they run on gjs, and they load the real walls.js against a stand-in DOM, so the production patterns are what is under test. Every one was checked by breaking the thing it guards and watching it go red. They test decisions only. Passing them does not mean a page renders, and the Scotsman is why that sentence is in the README.\nAnswering The Consent Question Properly Deleting a consent banner leaves the question unanswered. That has two consequences, and the first is that the banner is rebuilt on every page load. And a publisher that holds its content back until the consent API replies will simply hang, while a vendor that gets no reply can treat the question as never asked. Silence is not refusal.\nSo the extension answers it, in this order:\nRefuse, answer no, store nothing, and only then remove A consent banner is answered, not just deleted A consent vendor's own container is on screen #qc-cmp2-container #onetrust-consent-sdk sp_message_container 1. Press their refusal Reject all, Decline, Only essential, Continue without accepting whole label only, 40 characters max matches Accept: the build fails 2. Answer the API: no __tcfapi every purpose, feature and vendor denied tcString \"\" tcloaded 3. Never store the record euconsent-v2 addtl_consent OptanonConsent didomi_token cookie and localStorage writes dropped Still on screen at the next tick? yes: remove it, unlock scrolling no: the refusal stands Three answers and a fallback. The refusal button is pressed if there is one, the consent API is told no to everything, the record is never stored, and only a banner still standing on the next tick is removed. It presses their refusal button first. Only buttons whose whole label is a refusal: \u0026ldquo;Reject all\u0026rdquo;, \u0026ldquo;Decline\u0026rdquo;, \u0026ldquo;Only essential\u0026rdquo;, \u0026ldquo;Continue without accepting\u0026rdquo; and a few more, each anchored, anything over 40 characters ignored. A unit test feeds it \u0026ldquo;I Accept\u0026rdquo;, \u0026ldquo;Accept All\u0026rdquo;, \u0026ldquo;Agree and close\u0026rdquo;, \u0026ldquo;Allow all\u0026rdquo;, \u0026ldquo;Got it\u0026rdquo;, \u0026ldquo;Subscribe\u0026rdquo; and \u0026ldquo;Pay £2.99/mo\u0026rdquo; and fails the build if any of them match. Clicking the wrong button would consent on your behalf, which is the one thing this project must never do. On msn.com, pressing \u0026ldquo;Reject All\u0026rdquo; got rid of Microsoft\u0026rsquo;s banner and it did not come back on reload, because Microsoft keep the refusal on their own servers.\nIt answers the consent API with no. The IAB\u0026rsquo;s framework gives every consenting page a function called __tcfapi to ask what you agreed to17. Where one exists, the extension answers it with every purpose, every special feature and every vendor denied, an empty consent string and eventStatus: 'tcloaded', which means the answer is final. It only appears where a consent framework is already on the page, so a site that never asked does not see it.\nIt never stores the record. document.cookie and localStorage silently drop euconsent-v2, addtl_consent, OptanonConsent, didomi_token and the rest, so a consent you never gave is never written and never replayed to 1,467 partners on the next page. UK law already requires consent before anything is stored on your device for this18. The extension enforces the no.\nEvery name on that list is anchored at both ends, and there is a story behind it. Sourcepoint\u0026rsquo;s cookies start _sp_. A lazy pattern of sp_ also matches Spotify\u0026rsquo;s sp_dc and sp_t, which are the login session, and would have logged me out of Spotify on every page. A test pins that too.\nRemoval is the fallback. Only for a banner still on screen after the click, and only a banner that is actually showing. Several vendors leave a permanent wrapper in the page whether or not a banner is up. euronews keeps a Didomi host at zero height, and tearing that out breaks the page for no gain. The Reach one sat three levels below \u0026lt;body\u0026gt; inside two wrappers that are not themselves fixed, so the check has to walk down to the box that is.\nShare Buttons, And A Portal That Ignores Its Own Settings Two more things on these pages exist to serve somebody other than the reader, and they needed different handling from the walls.\nShare and follow buttons. Every link to Facebook, Instagram, X, TikTok or LinkedIn is rewritten to http://localhost/removeme, with the original parked in an attribute so nothing is lost. Then a second pass deletes what is plainly furniture (an icon with no text, \u0026ldquo;Share on X\u0026rdquo;, anything in a container that calls itself share or social) and hides the rest. The marker is there so one selector shows everything about to go, and a mark-only mode stops after the first pass so you can look before trusting it on a site.\nThe naive version breaks pages three ways, and each is now a test:\nNaive version What it breaks What is done instead a[href*=\u0026quot;x.com\u0026quot;] also matches netflix.com, linux.com, phoenix.com parse the host and compare whole labels remove every social link \u0026ldquo;Continue with Facebook\u0026rdquo; goes and people are locked out login, OAuth, legal and developer links left alone remove a link in a sentence the words go with it hide by default; an option unwraps it and keeps the words That last one is a trade-off I made knowingly. A BBC sign-off reads \u0026ldquo;follow BBC Manchester on , , and .\u0026rdquo; with it on. I looked at that and chose hiding anyway. The option to keep the words is there for anyone who disagrees.\nMSN and Bing. The same web components run both, around 160 shadow roots on the front page. A plain query of the document found 1 link. Walking the shadow roots found 73, including the Facebook and X tiles. Anything that does not walk them sees almost none of it.\nMSN does have its own content settings. They are held on Microsoft\u0026rsquo;s servers against an anonymous ID, not in a cookie, so they are gone after a cookie clear, in a new profile and in incognito. They do not reach Bing at all. And with every one of them switched off and 11,700 pixels scrolled, this is what was still in the feed:\nMSN\u0026rsquo;s own switch, off Still on the page Casual Games the Games tile Shopping advertising tiles for Booking, Temu and eBay Comments 501 comment links Weather, Finance, Sports gone, until the next cookie clear A switch labelled Comments that leaves 501 comment links on the page is a claim made to the user, and a false one. So the extension takes those out itself, targeting MSN\u0026rsquo;s stable component names rather than its class names, which change every build.\nOne thing I learned there that applies well beyond MSN: removing an image from the page does not stop it loading. Chrome starts the fetch when src is set, even for an image never put in the page19. By the time a content script sees a card its thumbnail is already on the wire. So the feed data is blocked at the two endpoints that serve it, and never the image hosts, because th.bing.com also serves Bing image search.\nAnd the MSN feed still flashes on screen for a moment before it goes. Four attempts at hiding it earlier all measured as working and none of them stopped the flash, which means what paints is not what the extension hides. The settings page says so plainly. I would rather it said that than implied a clean load it does not deliver.\nWhat It Will Not Do It will not Because open a real paywall if the server withholds the article, it stays withheld guess at a site a domain goes in only after its wall is seen there; two News Corp sites assumed to match Page Six did not, and came out again touch a site with no signature on the BBC every hook is installed and nothing is logged, no timer dropped, no node touched call home no host permissions, no network access, no telemetry; the error log stays in your browser pretend the MSN flash is unsolved, and one Ad-Shield variant, wp-ls/…, does not match yet; its page rendered fine, so it is written down rather than fixed on a guess A Page You Were Sent Is Yours Once a server has sent your browser a page, that copy is on your machine. Running code on it afterwards to take it away is not a business model I recognise. It is the old trick of selling you something and then keeping a hand on it.\nNewspapers used to be a thing you bought and then owned: somebody on the street corner had a stack, you handed over the money, and it went home with you and was yours to read, fold, lend or light the fire with. Nobody came round the house an hour later to cut the second page out because you had skipped the adverts. The web version gave the paper away for nowt and then sold the reader. First to the advertisers, then to the 1,467 partners, and now to a vendor whose whole product is deciding whether you have behaved well enough to keep what you were given.\nThe part I cannot get past is the stylesheet. A paper that lets an ad company\u0026rsquo;s server decide whether its own layout survives has given away its own front page. It will not have known it had, because nobody inside was blocking the host, so the first person to find out was a reader with a Pi-hole, looking at a page full of raw links and being told by a dialog box that the fault was theirs. It was not.\nThere is a price for journalism and I have no problem paying it. What I will not do is let a script decide what I have already got. That line gets drawn in my browser, not theirs.\nPi-hole documentation — \u0026ldquo;The Pi-hole® is a DNS sinkhole that protects your devices from unwanted content, without installing any client-side software.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNewsquest — About us — \u0026ldquo;We are the leading local news publisher in the UK with a portfolio of more than 200 brands.\u0026rdquo; Brands are not domains, hence the higher count off the certificates.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDaily Business, 18th December 2024 — \u0026ldquo;National World, owner of The Scotsman and Yorkshire Post, has reached agreement on a £65.1 million takeover by Irish publisher Media Concierge.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nReach plc — About us, archived 5th September 2026 — \u0026ldquo;120+ brands, from household names like the Mirror, Express, Daily Record and Daily Star, to local titles like MyLondon, BelfastLive and the Manchester Evening News\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDimova et al., \u0026ldquo;The CNAME of the Game: Large-scale Analysis of DNS-based Tracking Evasion\u0026rdquo;, PETS 2021 — CNAME cloaking \u0026ldquo;effectively bypasses antitracking measures that rely on fixed hostname-based block lists.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nuBlockOrigin/uAssets issue #30988, filed under the maintainers\u0026rsquo; \u0026ldquo;Ad-Shield\u0026rdquo; label, with the message served from error-report.com: \u0026ldquo;Failed to load website properly since html-load.com is blocked.\u0026rdquo; See also Jacob Desforges, \u0026ldquo;Ad-Shield ad reinsertion\u0026rdquo;, 12th April 2026. The attribution is the filter-list community\u0026rsquo;s; the vendor discloses nothing.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nSourcepoint — Anti-adblock FAQs, archived 25th May 2022 — \u0026ldquo;Historical experience shows on average about 30% of messaged users will turn off their adblockers.\u0026rdquo; The same page asks whether \u0026ldquo;a CNAME applied to a 1st-party subdomain\u0026rdquo; would stop the detection script being blocked.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nAdExchanger, 16th August 2023 — \u0026ldquo;InMobi acquired Quantcast’s consent management platform, called Quantcast Choice\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nUK GDPR, Article 7 — \u0026ldquo;utmost account shall be taken of whether, inter alia, the performance of a contract, including the provision of a service, is conditional on consent\u0026rdquo;. Recital 42: consent is not freely given \u0026ldquo;if the data subject has no genuine or free choice\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nICO — Consent or pay, published 23rd January 2025 — \u0026ldquo;“Consent or pay” models can be compliant with data protection law if you can demonstrate that people can freely give their consent\u0026rdquo;. The page now says the guidance is under review following the Data (Use and Access) Act.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nEDPB Opinion 08/2024, adopted 17th April 2024 — \u0026ldquo;In most cases, it will not be possible for large online platforms to comply with the requirements for valid consent if they confront users only with a binary choice\u0026rdquo;. It addresses large online platforms, not regional papers.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nChrome for Developers — content_scripts manifest key — \u0026ldquo;Choosing the \u0026ldquo;MAIN\u0026rdquo; world means the script will share the execution environment with the host page\u0026rsquo;s JavaScript.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nECMAScript — CreateGlobalVarBinding — \u0026ldquo;If a binding already exists, it is reused and assumed to be initialized.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nChrome for Developers — declarativeNetRequest — \u0026ldquo;provides implicit access to allow , allowAllRequests and block rules\u0026rdquo;, and otherwise \u0026ldquo;you must request host permissions before you can perform any action on a host.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nChrome for Developers — declarativeNetRequest — \u0026ldquo;Errors and warnings about invalid static rules are only displayed for unpacked extensions. Invalid static rules in packed extensions are ignored.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nHTML Standard — the Location interface marks its members [LegacyUnforgeable], which in Web IDL means \u0026ldquo;the property will be non-configurable and will exist as an own property on the object itself\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nIAB Tech Lab — TCF v2 CMP API — \u0026ldquo;Every consent manager MUST provide the following API function: __tcfapi(command, version, callback, parameter)\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPECR, regulation 6 — \u0026ldquo;a person must not store information, or gain access to information stored, in the terminal equipment of a subscriber or user\u0026rdquo;, with consent as the condition in Schedule A1, as amended by the Data (Use and Access) Act 2025.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nHTML Standard — update the image data — runs \u0026ldquo;whenever that element is created or has experienced relevant mutations\u0026rdquo;, including when its src is set; being in the document is not a condition.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/random/the-news-site-sent-me-the-article-then-deleted-it/","summary":"Newsquest, National World, Reach and others send your browser the complete article with its stylesheets, then run commercial anti-adblock and consent code that blanks the page, deletes every stylesheet every 100ms or covers it with a choice between 1,467 tracking partners and £2.99 a month. This walks through what that code does, the four reasons I call it malware, and the Chrome extension I built over one day to keep the page: pinning Newsquest\u0026rsquo;s adLight flag before the wall builds, dropping the timers scheduled from eval, killing the stylesheet stripper at schedule time, and refusing consent properly by answering the TCF API with no, pressing the vendor\u0026rsquo;s own reject button and never storing the record. Then the fixes that made things worse, the probe that said a broken page was fine, share buttons and the MSN feed, and what the extension will not do.","title":"The News Site Sent Me The Article, Then Deleted It"},{"content":"Coal to nowt in fourteen years Kate Morley\u0026rsquo;s National Grid dashboard lets you set the window and watch the numbers move.1 The dataset starts in 2012, which happens to be peak coal, so the all-time column is the whole clean-up in one figure.\nAll time (2012–) Past year Past week Past day Carbon intensity 251 g/kWh 124 g/kWh 98 g/kWh 67 g/kWh Coal 12.4% 0.0% 0.0% 0.0% Gas 33.5% 27.0% 21.2% 15.9% Wind 19.5% 34.8% 46.0% 64.1% Solar 3.7% 7.4% 8.3% 11.1% Nuclear 18.3% 12.0% 12.0% 12.9% Fossil total 45.8% 27.0% 21.2% 15.9% Renewables total 24.4% 43.6% 55.8% 76.6% Demand 32.6 GW 30.9 GW 28.4 GW 25.8 GW Price £70.13/MWh £92.77/MWh £125.01/MWh £27.18/MWh Wind is now the biggest single source of electricity in Great Britain. 34.8% across a full year, against gas on 27% and nuclear on 12%. Not on a good day. Twelve months.\nCoal is zero. Not low. Zero, for a year. The last station shut on 30 September 2024, a hundred and forty two years after the world\u0026rsquo;s first one opened in London in January 1882, which means this country invented coal power and then spent the better part of a century and a half getting round to switching it off again.1 Carbon intensity has halved against the fourteen-year average, 251 down to 124. Half, in fourteen years.\nThe demand that went missing Look at the demand row again. 32.6 GW on the long average, 30.9 over the past year. It has been falling for twenty years, about 5 TWh a year since 2005.\n2005 2023 Domestic demand 126 TWh 93 TWh Industrial demand 117 TWh 86 TWh Average household 4,662 kWh (2007) 3,449 kWh Households dropped 33 TWh over that stretch while the country plugged in a million and a half electric cars and a quarter of a million heat pumps.2 EU Ecodesign rules did much of that work; the lighting regulations alone saved 81 TWh across the EU in 2020.3 Efficiency didn\u0026rsquo;t just absorb the electric cars, it swamped them.\nTwo caveats, because this one gets overclaimed. Industrial demand fell as well, 117 TWh to 86, and that is mostly factories closing rather than factories getting clever. And the decline is over. 2024 was the first year in nearly two decades that demand went back up.2\nThe cars that were supposed to break it Here is the fleet the grid absorbed, straight off the DVLA licensing tables. Battery electric vehicles licensed in Great Britain at the end of each year.4\nYear Cars Vans Buses HGVs Motorcycles Total 2015 20,472 4,786 194 314 917 26,756 2017 41,222 6,401 303 283 1,046 49,334 2019 89,581 10,479 504 269 2,790 103,724 2021 374,597 28,245 1,295 331 9,118 413,908 2023 916,576 64,988 3,188 732 14,142 1,000,092 2024 1,266,421 84,959 4,821 983 14,039 1,371,779 2025 1,708,499 111,837 7,450 1,472 13,460 1,843,395 Cars are up eighty three fold in ten years, the whole fleet sixty nine fold. The millionth electric vehicle landed at the end of 2023 on 1,000,092, which is about as close to the nose as a statistic ever gets. Nobody planned that.\nThree things in that table you don\u0026rsquo;t see in a headline.\nLorries have barely started. 1,472 electric HGVs in the whole of Great Britain, and the number actually fell between 2015 and 2020, down to 253. Cars were always the easy bit. This is the hard bit and it has not really begun.\nElectric motorcycles are going backwards. 14,142 in 2023, then 14,039, now 13,460. Two years of decline, the only category shrinking while everything else compounds.\nThe percentages are slowing and the metal is not. Car growth was 114% in 2020 and 35% in 2025. But 2025 put 442,000 cars on the road against 2024\u0026rsquo;s 350,000. A falling percentage of a big number still beats a big percentage of a small one.\nFor where this goes, NESO\u0026rsquo;s 2025 Future Energy Scenarios put Great Britain at 31 million electric vehicles by 2050 under Holistic Transition, 33.4 million under Electric Engagement and 36.1 million under Hydrogen Evolution.5 Today\u0026rsquo;s 1.84 million is about six per cent of the way there.\nOne detail worth noting from those scenarios: the flexibility available from smart charging fell this year, 16 GW down to 10 GW. Bigger batteries mean fewer, longer charges, and less to shuffle about.5\nThe demand that never showed up at all There\u0026rsquo;s a second thing pushing demand down, mind, and it isn\u0026rsquo;t efficiency.\nSelf-consumption Solar, no battery 30–40% Solar plus battery 80–90%6 Two million solar installations in the UK now, 22.3 GW between them, and better than 30% of new systems go in with a battery against 10% five years ago.6 Put storage behind the panels and most of what the roof makes never crosses the meter.\nThat energy isn\u0026rsquo;t saved. It has just stopped being counted. The meter is the only thing that changed.\nYear Solar capacity added 2021 ~0.4 GW 2022 ~0.5 GW 2023 1.9 GW 2024 2.3 GW 2025 2.6 GW Sixfold in five years, and the shape tells you why. 2021 and 2022 were flat. Then the bills landed, people did the arithmetic, and solar stopped being an environmental choice and became a financial one.\nOur house is in that table somewhere. Nine kilowatts on the roof, thirty kilowatt hours of battery under it, two cars and the aircon drinking the daytime generation. From National Grid\u0026rsquo;s point of view this address has been quietly shrinking for years. From a physics point of view it hasn\u0026rsquo;t. The energy just stopped appearing on anybody\u0026rsquo;s chart.\nThe war on the thing that worked None of the above happened by accident, and there\u0026rsquo;s a live effort to stop the rest of it.\nReform took ten English councils in May 2025 and said they would use \u0026ldquo;every lever\u0026rdquo; to block new wind, solar and battery projects. Carbon Brief put the figure at risk at about 6 GW: 5,076 MW of battery schemes, 786 MW of solar and 56 MW of wind sitting in those ten areas.7 The party\u0026rsquo;s energy spokesman wrote to developers in Lincolnshire to tell them \u0026ldquo;this is war\u0026rdquo;, and sent formal notice to the chief executives of SSE Renewables, Octopus Energy, Centrica and Equinor that a Reform government would tear up their deals.7\nTesting the farmland claim The stated reason is farmland and food security. That claim is testable, so let\u0026rsquo;s test it.\nLand use Area Share of UK Ground-mount solar, Sept 2024 21,200 ha ~0.1%8 Same, satellite-measured study 15,580–17,364 ha 0.06–0.07%9 Golf courses 125,000 ha ~0.5%10 70 GW of solar by 2035 n/a under 1% of farmland10 Solar covers about a tenth of one per cent of this country. Golf covers roughly five times as much, and in all the years anybody has been worrying aloud about British food security nobody has once written to a golf club to declare war on it. Not one letter.\nThere is a real argument buried under the noise, and it deserves saying plainly. CPRE found 59% of England\u0026rsquo;s largest solar farms sit on productive farmland, and 31% of that area is classed best and most versatile.11 Land quality is a fair point. Land quantity is not, and it is the quantity argument being made.\nAnd look again at what is actually in that 6 GW. Solar is 786 MW of it. Batteries are 5,076 MW, better than four fifths of the capacity being fought over.7 The farmland argument is aimed at the smaller number. Storage is the real target, and storage doesn\u0026rsquo;t grow owt.\nTesting the fire claim Battery schemes get refused on fire risk. Not everywhere Reform runs, to be fair. This one is broad local opposition rather than one party\u0026rsquo;s campaign, and it is working. More than 900 objections to a scheme near Allerton Bywater in Leeds. A greenbelt site near Eaglesham thrown out over lithium fire fears after 250 objections. A 49.9 MW project in Devon refused against the planning officer\u0026rsquo;s own recommendation.12\nSo put the fires next to the objections.\nCount Grid battery fires in the UK, all time 3 known13 South Korea, 2017–2019 cluster 2814 EPRI global incident database, since 2011 ~95 entries14 Objections to one Leeds scheme 900+12 One planning application in Leeds attracted more objections than there have been recorded grid battery fires anywhere on earth since 2011. Three in this country, ever. One of those was a site still under construction.13\nAnd 27 of the 30 incidents worldwide across 2018 and 2019 were in South Korea. One national cluster, bad enough to halt their storage market, and the reason the database exists at all.14 Strip those out and the global record is thinner still.\nMeanwhile the rate has collapsed. Failures per year have stayed roughly flat while deployment went from 11 GWh in 2018 to over 300 GWh in 2024. That is a 99% fall in failure rate per unit installed, because the standards caught up.14 In 2024, 0.3% of projects had a failure that led to a fire with safety concerns.14 I went through that, and the Moss Landing fire everybody cites, in the piece on rejected energy.\nThen there\u0026rsquo;s why they fail, which is the bit that ought to end the argument.\nRoot cause Share of failures Integration, assembly and construction 36%15 Operational 29% Design 21% Manufacturing defect 4% 89% of incidents don\u0026rsquo;t start with the battery at all.15 Only three in the whole database trace back to a cell or module defect. What actually goes wrong is the balance of system: DC and AC wiring, the HVAC, the fire suppression kit itself. And 72% of failures happen during construction, commissioning, or inside the first two years.15\nSo it isn\u0026rsquo;t the chemistry. It\u0026rsquo;s the fitting. Sub-standard assembly, corners cut on installation, commissioning done with the monitoring not yet live so a leak or an isolation fault has time to cascade into something that needs a fire engine before a single alarm has gone off anywhere a human being can hear it. Bad workmanship, in other words.\nThat matters because it changes what the answer is. If lithium were inherently prone to going up, you\u0026rsquo;d be right to keep it away from the village. It isn\u0026rsquo;t. This is a trade quality and inspection problem, which is the sort of thing we already know how to fix. Same as any other bit of electrical installation: proper standards, proper sign-off, somebody competent checking the work.\nWhich brings it back to the planning rules, where there is a genuine gap worth fixing. Councils have no legal obligation to consult the fire service on a BESS application, so some demand a full fire management plan and others treat safety as outside planning altogether.12\nThe objection is \u0026ldquo;these things catch fire\u0026rdquo;. The data says badly-installed things catch fire. One of those is an argument for refusing permission. The other is an argument for inspecting the build. Fix the gap and you remove the argument. Leave it broken and it keeps working as one.\nWorth saying none of it has worked especially well so far. A year on, those councils have found that blocking large solar is easier said in a press release than done in a planning committee, and several schemes went through anyway.7\nFollow the money As for where the script comes from, the funding is a matter of record.\nGroup Money in From Heartland Institute $676,000+ (1998–2007) ExxonMobil16 Heartland Institute undisclosed further sums Koch-linked foundations16 GWPF / Net Zero Watch $500,000+ a Koch-linked fund17 GWPF / Net Zero Watch $210,525 Sarah Scaife Foundation, via its US arm17 Heartland is an American climate denial outfit that opened a UK branch, with Nigel Farage as guest of honour at the launch.16 The Global Warming Policy Foundation campaigns here as Net Zero Watch, a registered charity running its campaigns through a private company.17\nAmerican fossil money, American talking points, a British party repeating them at a technology this country is demonstrably good at. I\u0026rsquo;ll let you join the dots on that one.\nThe talking points arrive with a president attached.\nThe claim What the evidence says Turbine noise causes cancer Completely unfounded. No evidence the sound harms health at all.18 Offshore wind kills whales NOAA and the National Marine Fisheries Service find no scientific evidence. Strandings are ship strikes, fishing gear and warming water.18 Manufacturing them makes \u0026ldquo;tremendous fumes\u0026rdquo; A turbine repays the energy used to build it in 5 to 8 months. Wind emits 37× less CO₂ than gas and 77× less than coal.19 That last one is worth dwelling on, because it inverts cleanly. Wind has the smallest carbon footprint of any generating technology the US Department of Energy measures.19 The IPCC medians put onshore wind at 11 g/kWh and coal at 820. I set those out in the piece on rejected energy. The thing being accused of making pollution is the thing that makes the least of it.\nThe bird one at least starts from something true, so put it against the other things that kill birds.\nCause of US bird deaths Per year Wind turbines 140,000–330,00018 Buildings ~600 million18 Cats 2 billion+18 Turbines account for something like one bird in six thousand. Nobody has ever been on the telly about cats. Cats are fine, apparently.\nAnd the animal-welfare line didn\u0026rsquo;t come from anyone who watches birds. It was coordinated by a conservative think tank funded by an industry group backed by ExxonMobil, Chevron and Marathon Oil.18 Same money as the table above, different delivery.\nWhich is the third channel. Think tanks write it, newspapers print it, and it gets moved at volume online. A Brown University study found bot accounts responsible for close to 40% of tweets calling climate science fake.20 I\u0026rsquo;d not claim to know who runs them, and that study is about climate denial generally rather than British solar in particular. But the pattern holds whichever end you pick it up from: the claims are wrong, they\u0026rsquo;re old, and they were paid for.\nSpeaking the same language Here\u0026rsquo;s the bit I think gets missed. Ask why Heartland opened a branch in London and not Lyon or Leipzig.\nBecause it works here as written. No translation, no localisation, no adapting the argument to a country that measures things in metric and heats its homes off a district network. The press release lands in English and runs the same day.\nThere is a mapped structure behind it, not just a shared vocabulary.\nThe bridge Atlas Network, Washington DC supports 450+ organisations in 90+ countries21 Funded via Donors Trust and the Charles Koch Foundation21 55 Tufton Street, Westminster GWPF, the IEA, TaxPayers\u0026rsquo; Alliance, Centre for Policy Studies, Adam Smith Institute, Civitas21 US–UK connections mapped by DeSmog ~2,00021 Climate deniers turned up in numbers at the Reform conference in September 2025.21 None of that required a single word to be translated.\nAnd the traffic runs both ways. UK far-right Telegram channels have been documented amplifying disinformation about American election integrity. The same pipe, pointed the other direction.22 Researchers describe English-language publications setting narratives that then get picked up and repeated by media in other languages, which puts us first in the queue rather than last.22\nA French or German reader gets a delay and a translator, and translation is a filter. Somebody has to decide the claim is worth carrying, and check it enough to put their own name on it. We get it raw, at the speed of a retweet, from a media market forty times our size that we already consume for entertainment.\nThat is the actual vulnerability. Not that Americans are arguing about their own grid. They can do as they like with it. It\u0026rsquo;s that we hear every word of it, in our own language, about our grid, from people who have never seen it.\nCheap carbon, dear leccy One thing hasn\u0026rsquo;t improved. Electricity averaged £92.77/MWh over the past year against £70.13 across the whole dataset. About a third dearer, while the carbon halved.1\nYou will have read that renewables are why. You will have read it a lot.\nUK national press, 2025 Editorials criticising renewables 4223 First year anti-renewable editorials outnumbered pro- since 201423 Right-leaning climate editorials rejecting climate action 81%23 Critical editorials leading on cost 86%23 So cost is the argument. Not birds, not landscape, not intermittency. Cost, in seven out of eight.\nThe dashboard settles that one, because it publishes price and mix together every half hour. Here is 20 September 2026, from teatime onwards.1\nTime Price Gas Gas share Carbon Solar Wind 16:00 −£19.03 2.85 GW 11.3% 71 g/kWh 6.33 11.30 16:30 £38.19 3.74 GW 15.1% 88 g/kWh 5.19 10.84 17:00 £103.70 4.46 GW 18.3% 109 g/kWh 4.06 10.21 17:30 £134.16 6.23 GW 25.6% 133 g/kWh 2.64 9.68 18:00 £175.57 7.96 GW 32.2% 150 g/kWh 1.53 8.96 19:30 £195.65 8.87 GW 39.9% 162 g/kWh 0.02 6.99 The sun went down. Gas tripled, 2.85 GW to 8.87. And the price went from minus nineteen quid to nearly two hundred in three and a half hours. Plus £57 in the half hour to half four, another £65 by five, another £41 by six. Carbon intensity more than doubled while it happened.\nTake the whole day rather than the interesting bit of it, and the relationship holds across all forty eight settlement periods.\n20 September 2026, all 48 half hours Gas share ranged 7.4% to 39.9% Price ranged −£19.03 to £195.65 Total swing £214.68/MWh Correlation, gas share against price r = 0.933 Half hours at a negative price 18, from 01:00 to 16:00 Nought point nine three, on one day, with the gas contracts fixed. And for nine hours of it Great Britain was paying people to take electricity off its hands. Gas down at 7.4%, solar getting on for 10 GW, price below zero.\nThen the sun set and it cost £195.65. Same wires. Same wind farms. Same country.\nThe mechanism is marginal pricing. The wholesale price is set by the running cost of the most expensive plant needed that half hour, and that is nearly always gas. During the crisis gas set the price 98% of the time while supplying around 40% of the electricity; in 2021, 97% of the time on 37% of generation.24 Highest proportion of any country in Europe.\nShare of generation Share of price-setting Gas 27.0% past year1 97–98% of periods24 Wind, solar, nuclear, hydro, biomass 73.0% the remainder Be honest about one thing in that data, mind, because somebody will spot it. The all-time average is £70.13/MWh on a dirtier mix than today\u0026rsquo;s, cheaper than the past year\u0026rsquo;s £92.77. That\u0026rsquo;s not wind putting prices up. That\u0026rsquo;s 2012 to 2020 being years of cheap gas. Across eras the gas price dominates; within a single day the weather does. Both of those point at the same culprit.\nSo a quarter of the mix prices all of it. Wind could be free at the point of generation, and much of it effectively is, and the number on your bill would not move, because the last gas plant on the stack still sets the rate.\nThat is not renewables making power expensive. That is a market design from the 1990s meeting a grid that no longer looks like the 1990s.\nThe counterfactual seals it. Take one gas price spike, of the sort we have already lived through, and model it against two different grids.24\nThe same gas spike hits a grid with… Household bills go up 2030 renewables targets met 8% No CfD-backed renewables at all 45% Same spike. Same gas. The only difference is how much wind and solar is sat there not caring what gas costs. More renewables, smaller hit. By a factor of five and a half.\nSo the wind farms are the reason the last crisis was survivable, not the reason it happened. Three separate tests, one answer: within a day the price tracks the wind inversely, across eras it tracks the gas price, and in the modelling the renewables are what blunt the spike. Gas is the cause. It isn\u0026rsquo;t close, and it isn\u0026rsquo;t really arguable.\nThe thing that would have fixed that evening Now hold that next to the eighteen half hours earlier the same day when the price sat below zero. That is the precise shape of problem a battery solves. Charge it while the grid is paying people to take power off its hands, push it back out into the teatime peak, and the last gas plant on the stack never gets called. The spread on 20 September was £214.68 a megawatt hour. Batteries exist to eat spreads like that.\nWhich brings us back to those 5,076 MW of battery schemes sitting in ten councils, and the party that promised every lever against them. Blocking storage doesn\u0026rsquo;t merely postpone some carbon reduction. It protects the gas margin, half hour by half hour, on precisely the evenings when gas is worth the most.\nWhether that\u0026rsquo;s the intention, follow the money. It\u0026rsquo;s the effect either way, and the funding behind the argument, per the table further up, belongs to the industry that collects the difference.\nThe levies, and who checked the numbers And the levies, since they get quoted as the smoking gun:\nBill component, 2025 Amount Share Policy costs on electricity £148.45 17% of the electricity bill25 Policy costs on gas £50.86 6% of the gas bill25 Seventeen per cent, and from April 2026 the government took 75% of the Renewables Obligation off bills and onto general taxation, about £92 a year of the £150 average cut.25 Real money, worth arguing about, and nowhere near the main event. The main event is the 27% of generation that prices the other 73%.\nI\u0026rsquo;m not going to tell you how many of those 42 editorials were lies, because proving intent isn\u0026rsquo;t something I can do from a desk. What can be shown is the error rate, and it\u0026rsquo;s poor. A complaint against one Daily Mail piece identified fifteen factual errors; the regulator required one correction.26 The Mail on Sunday and The Times both reported that NESO had found the cost of net zero would be £4.5 trillion by 2050. NESO found nothing of the sort, and it was the third time papers had misrepresented that same body.26\nBefore anybody quotes the low number of upheld complaints at me: IPSO\u0026rsquo;s complaints committee contains no professional scientists and does not consult experts on technical subjects, and it routinely treats a wrong number as an opinion.26 One upheld complaint out of fifteen errors measures the regulator, not the article.\nDraw your own conclusion. Mine is that the most-repeated argument against renewables in the British press is the one that collapses fastest when you put it next to the grid\u0026rsquo;s own meter.\nAnd we don\u0026rsquo;t set the gas price Here is what actually happened to gas, on the European TTF benchmark.27\nTTF gas price Pre-2021 average ~€20/MWh December 2021 €180 March 2022 €220 August 2022 peak ~€340 Early 2026 €35–45 Mid-September 2026 €83.40 Seventeen times the normal rate at the peak. And look at that last row. Gas went to €83 this month on supply and storage worries, which is exactly why a still Sunday evening on the GB grid priced at £179.70/MWh. The chain is short: European gas market moves, gas sets the British price, your bill follows.\nNote also that \u0026ldquo;back to normal\u0026rdquo; isn\u0026rsquo;t. Today\u0026rsquo;s €35–45 floor is still double pre-crisis, and it spikes on a rumour.\nNone of those decisions are made here, and our exposure is growing.28\nUK gas supply North Sea share of demand, 2025 about half Gas imported, 2025 464 TWh of which Norwegian pipeline 69% of imports of which LNG 31% of imports LNG as share of total supply now 14% LNG share by 2030 over 25% LNG share by 2035 close to 50% North Sea output is falling 12–13% a year and is projected down 78% by 2035 against 2025.28 The gap gets filled by tankers from Qatar and the United States, bought on a global spot market against buyers in Asia who can outbid us on any cold morning they like.\nSo the argument that we should keep the grid on gas for the sake of the bills has it backwards. Gas is the bit we don\u0026rsquo;t control, priced by events we don\u0026rsquo;t influence, from fields that are running out. Wind and sun are the bit that happens here for nowt once the kit is up.\nWhat it costs and what it\u0026rsquo;s priced at There is a difference between what a thing costs and what you are charged for it, and this whole business sits in that gap.\nThe cost of running this grid has fallen. Half the carbon, no coal, a third of the generation now coming off weather that arrives for free and sends no invoice. That is the cost. The price went the other way, because we kept a rule from the 1990s that lets the most expensive plant on the system set the rate for everything else on it, and then imported the fuel for that plant from a market where a cold snap in Asia moves what a pensioner in Barnsley pays to keep warm.\nNobody is hiding that. It is written down, in half-hourly settlement data, free, on a website one woman maintains.\nWhat I keep coming back to is who benefits from the confusion. Because the people telling you the wind farms did this are, when you follow the money back, funded by the thing that actually did. That is not a coincidence and it is not incompetence. It is the oldest trick there is: get the public angry at the cheapest part of the system so nobody looks at the dearest.\nAnd the bit under attack is the only bit we own outright. A gas turbine needs a tanker from Qatar and a price set in Rotterdam. A wind farm off the Humber needs maintaining. One of those is sovereignty and the other is a standing order, and we appear to be about to argue ourselves out of the first to protect the second.\nWe built the thing. It works. Somebody should tell people.\nSources Kate Morley — National Grid: Live — Great Britain\u0026rsquo;s generation mix, carbon intensity, demand and price, selectable over the past day, week, year and the whole dataset from 2012; also the closure of the last coal-fired station on 30 September 2024.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDrax — 2024, the year GB electricity demand turned a corner — two decades of falling demand, the load added by electric vehicles and heat pumps, and the 2024 reversal.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nEuropean Commission — Light sources, energy label and ecodesign — the lighting ecodesign regulations and the 81 TWh of electricity they saved across the EU in 2020.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDepartment for Transport — Vehicle licensing statistics data tables, VEH0141 — licensed plug-in vehicles at the end of each quarter by body type and fuel type. Figures above are the battery electric column for Great Britain at Q4 of each year.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNational Energy System Operator — Future Energy Scenarios — electric vehicle uptake to 2050 across the pathways, vehicle-to-grid capacity, and the smart charging flexibility revision.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMCS — UK homes installing a small-scale renewable every 90 seconds — certified installation totals, the two million milestone and installed capacity, and the share of new solar systems paired with battery storage.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCarbon Brief — Reform-led councils threaten 6GW of solar and battery schemes across England — the capacity sitting in the ten councils taken in May 2025, the \u0026ldquo;every lever\u0026rdquo; commitment, the letters sent to developers and energy company chief executives, and what has actually happened to the schemes since.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nHouse of Commons Library — Planning for solar farms — ground-mount solar covered an estimated 21,200 hectares at the end of September 2024, around 0.1% of total UK land area.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nLancaster University — Researchers use satellite imagery to shed light on UK solar farm land use — satellite measurement putting solar farm land use at 15,580 to 17,364 hectares, 0.06% to 0.07% of UK land area.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nFriends of the Earth — Fact check: British farming and renewables — land area under golf courses against solar, and the farmland share implied by the 70 GW by 2035 target.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCPRE — Two-thirds of mega solar farms built on productive farmland — 59% of England\u0026rsquo;s largest operational solar farms on productive farmland, 31% of that area classified best and most versatile.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nHouse of Commons Library — Battery energy storage systems — planning objections and refusals including the Allerton Bywater, Eaglesham and Devon schemes, thermal runaway as the fire mechanism, and the absence of any legal duty on councils to consult fire and rescue services on a BESS application.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nHouse of Commons Library — Battery energy storage systems — documented UK grid-scale BESS fires, including Liverpool in September 2020 and an Essex site under construction in February 2025, and the note that no reliable public record of incident counts exists.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nEPRI — BESS Failure Incident Database — the global grid-scale failure record since 2011, the South Korean cluster of 2017–2019, the fall in failure rate per unit installed against deployment growth, and the caveat that the database only captures publicly reported incidents.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nUtility Dive — Cells and modules not responsible for most battery energy storage system failures — EPRI\u0026rsquo;s root cause analysis: integration, assembly and construction at 36% of failures, operational 29%, design 21%, manufacturing defects 4%; 89% of incidents not originating in the battery; and the concentration of failures in construction, commissioning and the first two years of operation.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nLeft Foot Forward — What is the Heartland Institute? — the UK branch launch, attendance, and Heartland\u0026rsquo;s funding from ExxonMobil and Koch-linked foundations.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nopenDemocracy — Net Zero Watch: how dark oil money is funding influential UK climate sceptics — GWPF and Net Zero Watch funding routed through American Friends of the GWPF, including the Sarah Scaife Foundation payments.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nClimate Power — Fact check: Trump\u0026rsquo;s wind turbine claims — the cancer and whale claims against the NOAA and National Marine Fisheries Service position, US Fish and Wildlife Service estimates for bird collisions with turbines set against buildings and cats, and the origin of the animal-welfare argument in fossil-funded think tank work.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCNN — Fact check: five things Trump got wrong about wind turbines — the \u0026ldquo;tremendous fumes\u0026rdquo; and carbon footprint claim against the Department of Energy position, the five to eight month energy payback for an average turbine, and wind\u0026rsquo;s emissions compared with gas and coal.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nInstitute at Brown for Environment and Society — Shadowy Twitter bots spread climate disinformation — the share of tweets describing climate science as fake that were traced to bot accounts. Note this is a 2021 study of climate denial generally, not of UK renewables specifically.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDeSmog — 55 Tufton Street and Mapped: how a US-UK network pushes climate science denial — the Westminster cluster and its members, the Atlas Network\u0026rsquo;s reach and funding through Donors Trust and the Charles Koch Foundation, the roughly two thousand mapped transatlantic connections, and the attendance of climate denial groups at the 2025 Reform conference.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDFRLab — UK-based far-right Telegram channels amplified disinformation targeting US election integrity — documented transatlantic amplification in both directions, and the role of English-language publications in setting narratives subsequently repeated by media working in other languages.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPress Gazette — Record opposition to climate action in UK national newspapers in 2025 — the count of editorials criticising renewable energy, the first year since 2014 that they outnumbered supportive ones, the share of right-leaning climate editorials rejecting climate action, and cost as the dominant line of attack.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCarbon Brief — Q\u0026amp;A: Why does gas set the price of electricity, and is there an alternative? — marginal pricing in Great Britain, the share of settlement periods in which gas sets the price against its share of generation, and the modelled effect of a gas price spike with and without CfD-backed renewables on the system.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nHouse of Commons Library — What costs make up an electricity bill? — policy costs on electricity and gas bills in cash and as a share, the Renewables Obligation as the largest single policy cost, and the transfer of 75% of its cost to general taxation from April 2026.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCarbon Commentary — Climate misinformation and press regulation — the fifteen inaccuracies identified in a single Daily Mail article against one required correction, the repeated misreporting of NESO\u0026rsquo;s findings on the cost of net zero, and the composition and approach of IPSO\u0026rsquo;s complaints committee on technical subjects.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nTrading Economics — EU natural gas (TTF) price history — the Dutch TTF benchmark from the pre-2021 baseline through the 2021–22 spike to the August 2022 peak and current levels, including the September 2026 move on supply and storage concerns.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nICIS — Ebbing North Sea gas production to raise UK gas prices and exposure to LNG imports — North Sea output covering about half of demand in 2025, the volume and split of imports, the rate of UKCS decline, and projected LNG dependence to 2030 and 2035.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/energy/the-grid-we-did-fix/","summary":"Great Britain halved the carbon intensity of its electricity in fourteen years, ran a full year without coal, and made wind its largest single source. Meanwhile demand fell for two decades while the country plugged in 1.8 million electric vehicles. This is what the numbers show, including the bits that spoil the story.","title":"The Grid We Did Fix"},{"content":"The number everybody quotes and nobody reads Lawrence Livermore put their 2024 energy flow chart out the other day. It says America got through 94.61 quads of energy and threw 62.27 of them away.1\nRight. A quad is a quadrillion British thermal units. Not a quid. There\u0026rsquo;s no money anywhere in this article, and if your eye keeps reading it as pounds sterling then mine did too. The Americans keep their national energy books in quads and the rest of us work in watt hours, so here is the conversion that makes it mean something: one quad is about 293 TWh, a shade more than the whole UK grid delivers in a year.\nRejected energy is the polite name for those 62.27. The bit that did no useful work. Heat up a chimney, heat off a radiator, heat out of an exhaust. Nearly all of it is waste heat from burning summat.\nNow mind the units, because I had this wrong at first and so does most of the internet. 62.27 is quads, not per cent. Put it against the 32.34 quads that did something useful and the share is 65.8%. Two thirds of everything drilled, dug, piped and shipped, gone as warm air.\n2024 In British grids Energy in 94.61 quads 27,726 TWh 102 years of it Rejected 62.27 quads 18,249 TWh 67 years of it Energy services 32.34 quads 9,477 TWh 35 years of it That last column is the one that landed for me. Great Britain\u0026rsquo;s grid delivers about 271 TWh a year.2 So America\u0026rsquo;s rejected energy for 2024 alone comes to sixty-seven years of the entire British grid, thrown away as heat, in twelve months.\nPut another way: run this country\u0026rsquo;s electricity from now until 2093 and you still would not have generated what the United States wasted last year.\nThat number gets quoted all over. How the chart counts gets quoted hardly at all. That\u0026rsquo;s the interesting bit.\nOne housekeeping note, this being a piece about America written by a Yorkshireman. Petrol is gasoline. A forecourt is a gas station. Aircon is AC. And when I say nowt or summat, I mean nothing and something.\nWind and solar are counted at 100% Here\u0026rsquo;s the bit that catches folk out. Livermore uses the Energy Information Administration convention, and under it four sources go in with no conversion losses at all.3\nSource How it enters the chart Where the losses go Coal, gas, oil fuel energy in ~60–70% to rejected Nuclear thermal energy in ~67% to rejected Onshore wind electricity out nowt rejected Utility and rooftop solar electricity out nowt rejected Hydroelectric electricity out nowt rejected A gas plant is counted on the energy in the gas, so the two thirds it chucks away as heat lands in the rejected block. A turbine is counted on the electricity it delivers. There\u0026rsquo;s no \u0026ldquo;energy in the wind\u0026rdquo; line to lose against.\nSo on this chart, every terawatt hour that shifts from a thermal plant to a wind farm does two jobs at once. Adds to useful, takes away from rejected. It counts twice.\nWhich means the argument that scrapping wind and solar would cut waste runs backwards through the very diagram it gets made about. I\u0026rsquo;ve had that one put to me in earnest by people who ought to know better. It\u0026rsquo;s not a close call. It\u0026rsquo;s the sign inverted.\nWhere the waste actually is Split the chart by sector and it stops being an abstraction.1\nSector Energy in Rejected Useful Efficiency Residential 11.23 3.93 7.30 65% Commercial 9.49 3.32 6.17 65% Industrial 26.39 13.46 12.93 49% Transportation 28.29 22.35 5.94 21% The power stations lose another 19.21 quads up the cooling towers before any of it reaches those four.\nTransport is the worst by a distance. It takes the largest share of the energy and turns 21% of it into movement. Those 22.35 wasted quads are 36% of everything America throws away. More than a third of the national waste, in one sector.\nA petrol engine, gasoline to American readers, turns maybe 16–25% of the fuel into motion. The rest is heat and noise.4\nUseful at the wheels Lost as heat Petrol (gasoline) engine 16–25% 75–84% EV drivetrain (battery to wheels) 87–91% 9–13% That\u0026rsquo;s not a marginal gain. That\u0026rsquo;s the difference between a machine that mostly warms the sky and one that mostly moves the car. Same journey, different physics.\nAnd it runs through the whole chain, not just the vehicle. Getting liquid fuel to a forecourt, a gas station in American, costs energy before a drop of it is burned.\nStep Energy lost getting it there Refining crude 7–15% of input5 Tanker, pipeline, road tanker on top of that Grid transmission and distribution ~5%6 Oil tankers as share of world shipping ~28% by deadweight tonnage7 Crude and product moved by sea, yearly ~4.4 billion tonnes7 Road transport is around half of global oil demand. Electrify that and a fair slice of the tanker fleet has nowt left to carry.\nWorth being straight here, because overclaiming is how you lose an argument you were winning. Grid losses are real and they\u0026rsquo;re resistive heat. An EV pays for cabin heat in winter that an engine gets for free. But that free heat is only free because the engine already binned three quarters of the fuel. It\u0026rsquo;s free the way warmth off a house fire is free.\nThe biggest single thing America could do Light-duty vehicles are 58.5% of transport energy, the bit that electrifies without waiting for new technology.8 Run the numbers on swapping the drivetrain and leaving everything else alone.\nLight-duty road transport Quads Energy in today 16.55 Useful work it actually delivers 3.47 Same work through an EV drivetrain 4.09 of electricity …generated from gas, at CCGT efficiency 9.08 primary …generated from wind, solar or hydro 4.09 primary Energy saved 7.5 to 12.5 quads Even charging every one of them off gas turbines, you save about seven and a half quads. Off wind and solar it is twelve and a half. Eight to thirteen per cent of everything the United States burns, from one swap.\nThat is the largest efficiency gain available anywhere on the chart, it needs no invention, no breakthrough, no pilot scheme and no new physics, because every single one of the vehicles required to do it is already being built and sold today in volume. The cars exist. That\u0026rsquo;s the whole trick.\nExcept a better engine doesn\u0026rsquo;t fix the layout An EV still has to cover the distance, and this is where the other half of the problem sits.\nUnited States Europe Car miles per person, per year ~12,400 ~6,2009 Share of daily trips made by car 85% 50–65%9 Trips under a mile made by car ~70% ~30%9 Parking spaces per car ~8 not counted9 Look at the third row, because it takes the geography excuse away. About 30% of daily trips are under a mile on both sides of the Atlantic. Same errands, same distances. Americans drive seven in ten of them. Europeans walk, cycle or catch something for seven in ten of them.\nGeography didn\u0026rsquo;t do that. Weather didn\u0026rsquo;t either. Zoning did, putting the houses here and the shops three miles over there, backed by parking minimums that ended up building nearly eight spaces for every car in the country.9 Between the 1920s and the 1960s American cities were rebuilt around the motor car and much of western Europe copied them. From the late 1960s Europe stopped, and started undoing it.9\nSo the twelve and a half quads is the ceiling on electrification alone. Halve the miles as well and you halve what is left. One is an engineering job and the other is a planning job, and the planning job is the one nobody can buy their way out of in a single purchase.\nAnd the cheapest passenger-mile is a shared one There is a third lever, and America has more or less stopped pulling it.\nMode Energy per passenger-kilometre Petrol (gasoline) car 1.9 to 3.5 MJ10 Urban electric rail, busy 0.3 to 0.6 MJ10 Four to six times better, before anybody changes a drivetrain. A single-occupant car puts out 7.7 times the CO₂ per passenger-mile of a full coach.10\nNow the state of play.\nUnited States Europe Share of passenger-miles on public transit 0.40%11 many multiples of that Journeys taken by car 95%11 50 to 65% Railway track electrified 1.7% (the Americas)11 ~57% in the EU11 Nought point four per cent. Not a transport system with a transit component. A country that drives, with some buses in it.\nAnd 1.7% electrification means American rail is still, overwhelmingly, diesel. So every argument for shifting freight and passengers onto rail is being made about a network that still runs on oil. Electrify the track and you get the mode shift and the fuel switch out of the same job.\nHere is the honest bit, because it cuts the other way and someone will raise it. Transit is only efficient when it is full. Bus occupancy in the States has been falling for decades, and energy per passenger-mile on buses has gone up 63% since 1970.10 A near-empty bus on a fifty-minute loop round a subdivision is worse than the car it was meant to replace. That\u0026rsquo;s a real number and a big one.\nBut look at what makes a bus empty. Nobody within walking distance of the stop, nowhere worth walking to at the other end, and a layout that puts eight parking spaces by every door. Empty buses aren\u0026rsquo;t a fact about buses. They\u0026rsquo;re a fact about what got built around the stop.\nWhich brings you to the thing that actually shifts people out of cars, and the gap there is wider than the mode-share number suggests.\nUnited States Europe Cities with a metro system 13 6012 Cities with a tram network 30 far more12 Growth in metro route length since 2000 baseline three times faster12 Thirteen. In a country of three hundred and forty million people. Europe has sixty, and has been laying new track three times as fast since the turn of the century, so the gap is widening rather than closing.\nThe mode matters, too. Work on European cities found metros shift people out of cars in a way tram networks largely do not, which fits what you would expect: a metro is faster than driving at rush hour and a tram usually is not.12 Speed is the whole product. Build something slower than the car and you have built a subsidy for people who have no choice, not an alternative for people who do.\nThat is the one genuinely expensive item on the list. Zoning reform costs political will and a redraft. Tunnels cost billions. But it buys what the other two can\u0026rsquo;t. You move people across a dense city on electricity, at 0.3 to 0.6 MJ a passenger-kilometre, and faster than they could have driven it. At that point leaving the car at home stops being a sacrifice and starts being the obvious move. That\u0026rsquo;s when people actually do it.\nWhich means the fixes are the same fix wearing different hats. Electrification takes 7.5 to 12.5 quads off the drivetrain. Zoning takes the miles down. Density is what makes the transit worth running, and the transit is what makes the density liveable. Pull one lever and you get one lever\u0026rsquo;s worth. Pull all three and they multiply.\nAmerica is currently arguing about the first one.\nNone of it starts, though, without the cheapest step of the lot, which also looks like the hardest. Somebody has to say out loud that there\u0026rsquo;s a problem.\nThe diagnosis isn\u0026rsquo;t missing. It gets published every year by a federal laboratory, for free, on a public website, in a diagram plain enough to read in a minute. Sixty-five point eight per cent wasted. Transport a third of it. Twenty-one per cent efficiency on the largest block on the page. Nobody needs to commission a study or wait on the science. The science came out in August. People read the headline number, got it wrong, and moved on.\nThat\u0026rsquo;s the bit that ought to sting. No country spends billions tunnelling under its cities to fix a thing it reckons is fine, and America has quietly filed 22.35 wasted quads a year under fine. Not argued over and dismissed. Just never put on the table.\nThe leftovers have to go somewhere Heat is only waste if there\u0026rsquo;s nowhere to put it. That\u0026rsquo;s a planning decision, not a technology gap, and it were mostly made decades back.\nDistrict heating share of heat demand Denmark ~66%13 Sweden, Finland, Poland, the Baltics above 50% EU average ~13% United Kingdom ~3%14 United States campus, hospital and downtown schemes only Europe runs about 111,650 commercial and industrial sites on transcritical CO₂ as of 2025, roughly a third of all food retail.15 Meta\u0026rsquo;s Odense data centre has been pushing around 100,000 MWh a year into the local network since 2019. That is heat which would otherwise have gone up a dry cooler. It warms 12,000 homes instead.13\nWe\u0026rsquo;ve got about 14,000 heat networks in the UK and they still only meet 3% of heat demand.14 Fourteen thousand of the things and almost nowt to show for it, because they\u0026rsquo;re small, fragmented and mostly bolted onto social housing. Ofgem took over regulation in January and zoning starts this year. Target 7% by 2035, about a fifth of building heat by 2050.16\nDenmark isn\u0026rsquo;t doing summat clever we can\u0026rsquo;t. Denmark put pipes under the streets. We didn\u0026rsquo;t. That\u0026rsquo;s it. That\u0026rsquo;s the difference.\nHeat pumps, and the refrigerant nobody expects Same logic at the small end. A resistive heater can\u0026rsquo;t beat a coefficient of performance (COP) of 1.0. That is the definition of the thing. A heat pump moves heat rather than making it, so it does better.\nSystem COP Conditions Resistive (PTC) element 1.0 maximum any Automotive heat pump 2.0–3.2 0 to 15 °C Hyundai/Kia R290 propane unit claimed 3.8 −15 °C VW R-744 (CO₂) unit 3.1 −20 °C17 ADAC put 28 EVs through a winter test at −7 °C. Heat pump models averaged 22% less range loss than the resistive-only ones.18\nThe R-744 line is the one worth a second look. That\u0026rsquo;s carbon dioxide itself, working as the refrigerant, in the VW ID.3 and ID.4. Higher suction vapour density keeps capacity up as it gets colder, which is exactly when you want it.17\nAnd refrigerant grade CO₂ is a byproduct of ammonia, ethanol and fertiliser production, captured and cleaned up rather than vented.19 A waste stream doing the heating and cooling, at a global warming potential of 1 against R-134a\u0026rsquo;s 1,430. When that one leaks, nowt happens.\nIt\u0026rsquo;s not carbon capture and I\u0026rsquo;ll not pretend it is; the charge is under a kilo. The point is narrower and better. The working fluid is summat we already had too much of.\nWhat I actually run Kit Spec Why Solar array 9 kW roof faces the right way, so use it Battery 30 kWh shifts the day\u0026rsquo;s generation into the evening EVs MG4, Xpeng G6 charged off the array, not a forecourt Aircon self-powered runs on what the panels make Control Home Assistant drops the big loads into the cheap window Nowt exotic in that list and nowt new. Same principle as matching storage media to an IO pattern. Put the energy where it earns its keep, measure what you actually get, and stop trusting the label on the box.\nThe bit that surprised me was how much of the saving came from not moving fuel about. No tanker, no forecourt, no refinery taking its cut on the way through. The panels are thirty feet from the car.\nThe chart is a mirror Livermore have published these for years and they\u0026rsquo;re good. Honestly built, genuinely useful, free. Worth an hour of anybody\u0026rsquo;s time.1\nBut a Sankey diagram has no opinion. It shows you what a country decided to do with its energy, drawn to scale. The 65.8% isn\u0026rsquo;t a law of physics. It\u0026rsquo;s a picture of choices about engines, pipes and planning, taken one at a time over about seventy years, and it\u0026rsquo;d look different if the choices had been.\nDenmark\u0026rsquo;s chart looks different because Denmark dug.\nSources Lawrence Livermore National Laboratory — Energy Flow Charts — the annual Sankey diagrams for US energy, including the 2024 chart and its 94.6 quadrillion BTU total.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nKate Morley — National Grid: Live — Great Britain\u0026rsquo;s electricity demand, averaging 30.9 GW across the past year, which works out at about 271 TWh annually. I wrote about what that dashboard shows in The Grid We Did Fix.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nHawai\u0026rsquo;i State Energy Office — Statewide Energy Flowchart — sets out the EIA methodology Livermore uses, under which distributed solar, hydroelectric, onshore wind and utility solar are entered assuming 100% generation efficiency with no thermal losses represented.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nEVreporter — Understanding the complete efficiency picture of electric vehicles — tank-to-wheel and battery-to-wheel efficiency for combustion and electric drivetrains.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nConcawe — EU refinery energy systems and efficiency — refinery own-use energy as a share of crude intake, from 3–4% for simple distillation to 7–10% and above for full-conversion plants.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nUS Energy Information Administration — How much electricity is lost in transmission and distribution? — annual US transmission and distribution losses averaged about 5% of the electricity transmitted, 2018 to 2022.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nUNCTAD — World seaborne trade — tanker share of world shipping by deadweight tonnage, and crude and refined product volumes moved by sea.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nUS Energy Information Administration — Light-duty vehicles\u0026rsquo; share of transportation energy use — light-duty vehicles at 58.5% of US transportation energy, with medium and heavy trucks and buses at 23.9% and air the only other mode above 5%.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCar dependency — Wikipedia and CNN — This little-known rule shapes parking in America — per-capita car kilometres in the US against Europe, the share of daily and sub-one-mile trips taken by car on each side of the Atlantic, the roughly eight parking spaces per car produced by parking minimums, and the divergence in urban policy from the late 1960s.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nBureau of Transportation Statistics — Energy intensity of passenger modes and Public transport versus private cars: a passenger-kilometre energy comparison — energy per passenger-kilometre for petrol cars against busy urban electric rail, the emissions ratio between a single-occupant car and a full coach, and the rise in bus energy per passenger-mile as occupancy fell.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nTransportation in the United States — Wikipedia and Statista — Share of the rail network which is electrified in Europe — the US share of passenger-miles taken on public transit and by private vehicle, and electrified track as a share of network in the EU against the Americas.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nStreetsblog USA — Other countries are building transit while the US falls behind and Metros reduce car use in European cities but trams do not — the count of American and European cities with metro and tram networks, the rate at which metro route length has grown since 2000 on each side, and the finding that metro systems displace car journeys where tram networks largely do not.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nState of Green — Utilising excess heat to warm up Danish homes — Danish district heating share of domestic heat demand, and data centre excess heat recovery including the Odense export figures.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nGreater London Authority — Heat networks data report, February 2026 — the fragmented UK heat network estate and its present share of heat demand.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nATMOsphere — European transcritical CO₂ installations — 111,650 European commercial and industrial sites running transcritical CO₂ in 2025, covering about a third of food retail outlets.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDepartment for Energy Security and Net Zero — Heat Network Zoning: government response — the zoning framework, Ofgem regulation from January 2026, and the 2035 and 2050 targets.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNaturalRefrigerants.com — CO₂ heat pumps found to offer high efficiency at low ambient temperature in electric vehicles — R-744 automotive heat pump performance at low ambient temperature, and the VW ID.3 and ID.4 implementations.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nInsideEVs — For maximum winter EV driving range, you want a car with this feature — the ADAC winter test across 28 electric vehicles and the range loss difference at −7 °C.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNaturalRefrigerants.com — FAQs — refrigerant grade CO₂ as a recovered byproduct of ammonia, alcohol and fertiliser production, and its global warming potential of 1.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/energy/rejected-energy-what-the-livermore-chart-shows/","summary":"Lawrence Livermore publishes a Sankey diagram every year showing where American energy goes. In 2024, 62.27 of 94.61 quads did no useful work at all. What rejected energy means and what a quad is, why the chart\u0026rsquo;s own method makes renewables shrink the waste rather than add to it, why transport alone is more than a third of the national total, and the three things that would actually fix it.","title":"Rejected Energy: What the Livermore Chart Actually Shows"},{"content":"Webex puts a yellow banner across the top of the window. Offline - No internet connection. Phone services show as disconnected, nothing syncs, and you cannot join the meeting that started ninety seconds ago.\nThat is not an inconvenience when the thing is your work phone. Webex is where my calls come in, and my VoIP line runs through it, so a client that will not authenticate is a desk phone that will not ring, a meeting I am not in, and a colleague who gets voicemail. It stopped me working. Not degraded, not slow. Stopped.\nAnd none of it was mine to cause or to prevent. A working day went sideways because of quality control that was not applied, at a supplier that gets paid, on a product sold with a support contract against a platform its own requirements page says is supported. I did not misconfigure anything. I installed the vendor\u0026rsquo;s package, from the vendor\u0026rsquo;s repository, on a platform the vendor lists, and it could not open a TLS connection. Then I spent an evening of my own time finding out why, which is time the people who signed the package did not spend.\nThe machine is not offline. The browser next to it is loading pages. Your mail is arriving, your terminal is pulling from a remote, and if you ask the operating system whether it can reach the internet it says yes. Webex itself agrees, in its own log, eleven seconds before it tells you otherwise.\nWhat has actually happened is that the copy of OpenSSL Cisco ship inside Webex cannot find a single certificate authority, because it was compiled to look for them in /workspace/.conan2/p/b/cisco8ee8b59cf93de/p/ssl. That is a directory on a Cisco build container. It has never existed on your computer and it never will. Every TLS connection the application makes fails at certificate verification, the login token cannot be refreshed, a fifteen-second timer runs out, and the interface reaches for the only explanation it has a string for.\nSo the banner is wrong in a specific and unhelpful way. It names your network. The fault is a path inside their build.\nThis is version 46.8.0.35631, on Fedora 44, kernel 7.2.4. It is a supported platform. Cisco publish Linux system requirements and ship a signed .rpm, webex-46.8.0.35631-1.x86_641. What follows is how to prove it in about ten minutes, why none of the obvious fixes work, the one that does, and then the part that matters more than any of it: this is not a subtle bug. It is a build that nobody ever ran on a machine that had not built it.\nWhat The Banner Is Actually Telling You Webex runs its connectivity display off a composite state machine, seven sub-machines each with its own timers. At startup they initialise like this:\nConnectivityStateMachine::ConnectivityBanner - Initializing with state: Connected ConnectivityStateMachine::Network - Initializing with state: NoNetwork ConnectivityStateMachine::Services - Initializing with state: Connected ConnectivityStateMachine::Mercury - Initializing with state: Disconnected ConnectivityStateMachine::Authentication - Initializing with state: UserNotAuthenticated ConnectivityStateMachine::Syncing - Initializing with state: Synced ConnectivityStateMachine::Survivability - Initializing with state: SurvivabilityHide Forty milliseconds later the operating system reports back, and the application writes it down:\nNetworkManagerPowerNetworkWatcher.cpp:93 onConnectivityCheckSuccess:: The host is connected to a network, that appears to be able to reach the full Internet. That line is in the same file, in the same session, as the banner claiming there is no internet connection. The application knew. It had the answer in hand at 08:25:12.131 and showed the opposite at 08:25:27.132.\nBetween those two moments, this:\nTime What happened 08:25:12.131 Operating system confirms full internet reachability 08:25:12.218 Proxy detection: none configured, direct connection 08:25:12.241 First HTTPS request fails, errorCode: 167772294 Error in SSL handshake 08:25:12.569 CloudApps access token refresh fails, same code 08:25:12.571 Kms access token refresh fails, same code 08:25:15.684 Retry, both fail 08:25:21.745 Retry, both fail 08:25:27.132 Fifteen-second timer expires, banner switches to NoInternet 08:26:12.091 Sixty-second timers expire, services degrade to DisconnectedShortTerm The Authentication sub-machine never leaves UserNotAuthenticated, so Services derives itself as disconnected, so the banner fires. Every stage of that is correct behaviour given the input. The input is wrong, and the input is a number: 167772294, on every failed request, from the first one to the last.\nThat number is the whole post. Hold on to it.\nThree Paths, None Of Which Exist Webex does not use the system OpenSSL. It brings its own, along with its own libcurl, and the libcurl links against the bundled one rather than yours:\n$ ldd /opt/Webex/bin/libcurl.so | grep -E \u0026#39;ssl|crypto\u0026#39; libssl.so.3 =\u0026gt; /opt/Webex/bin/../lib/libssl.so.3 libcrypto.so.3 =\u0026gt; /opt/Webex/bin/../lib/libcrypto.so.3 Fine so far. Bundling a TLS library is a defensible choice and plenty of vendors do it. What matters is what got baked into it, and OpenSSL will tell you if you ask:\n$ strings /opt/Webex/lib/libcrypto.so.3 | grep -E \u0026#39;OPENSSLDIR|ENGINESDIR|MODULESDIR\u0026#39; OPENSSLDIR: \u0026#34;/workspace/.conan2/p/b/cisco8ee8b59cf93de/p/ssl\u0026#34; ENGINESDIR: \u0026#34;/workspace/.conan2/p/b/cisco8ee8b59cf93de/p/lib/engines-3\u0026#34; MODULESDIR: \u0026#34;/workspace/.conan2/p/b/cisco8ee8b59cf93de/p/lib/ossl-modules\u0026#34; Three of them. Not one slipped setting. The whole install prefix, carried through from the machine that compiled it, in a library that reaches the customer as a signed package on a supported platform.\nOPENSSLDIR is set once, at configure time, with --openssldir, and OpenSSL\u0026rsquo;s own build documentation is blunt about what it is for: \u0026ldquo;Directory for OpenSSL configuration files, and also the default certificate and key store.\u0026rdquo;2 Everything downstream of trust hangs off it. The default CA file is cert.pem inside it and the default CA directory is certs inside it3. /workspace is a Conan build cache. Conan is the C++ package manager Cisco build with, and it stores each package under a hash of its build inputs4. The hash cisco8ee8b59cf93de is a fact about a container that was probably deleted minutes after the build finished.\nAsk the library what it is, and it is not even stock OpenSSL:\nVERSION CiscoSSL 3.5.5.8.5.4 27 Jan 2026 BUILT_ON built on: Wed Feb 25 06:29:18 2026 UTC PLATFORM platform: conan-Release-Linux-x86_64-gcc-13 DIR OPENSSLDIR: \u0026#34;/workspace/.conan2/p/b/cisco8ee8b59cf93de/p/ssl\u0026#34; A maintained in-house fork, with a version scheme of its own, built in February, shipped in August, and still carrying the working directory it was made in.\nWhere each copy of OpenSSL on the machine looks for its trust anchors One machine, two copies of OpenSSL, and only one of them knows where the certificates are Neither copy is told at runtime. Each carries the answer as a string compiled into it, and that string is set by whoever ran the build. The system copy: OpenSSL 3.5.8, Fedora 44 compiled-in OPENSSLDIR /etc/pki/tls the directory is there cert.pem \u0026#8594; tls-ca-bundle.pem certs/ \u0026#8594; hashed anchors the chain is checked against real anchors Handshake completes. Every other program on the box works. Which is exactly why the user is certain the network is fine. The copy Webex ships: CiscoSSL 3.5.5.8.5.4 compiled-in OPENSSLDIR /workspace/.conan2/p/b/cisco8ee8b59cf93de/p/ssl no such directory on any customer machine cert.pem \u0026#8594; absent certs/ \u0026#8594; absent the store loads zero trust anchors Handshake fails: error 0x0A000086, decimal 167772294. The banner the user is shown says the network is down. Two copies of OpenSSL on one machine. Neither is told at runtime where trust lives. Each carries a string compiled into it, and one of those strings names a directory that only ever existed on somebody else\u0026rsquo;s build host. Do Not Take strings For An Answer. Ask The Library strings finds text in a file. It does not prove the library uses it. So load the shipped library and ask it directly, which takes about twelve lines of Python and no root:\nimport ctypes c = ctypes.CDLL(\u0026#34;/opt/Webex/lib/libcrypto.so.3\u0026#34;) for f in (\u0026#34;X509_get_default_cert_file\u0026#34;, \u0026#34;X509_get_default_cert_dir\u0026#34;, \u0026#34;X509_get_default_cert_file_env\u0026#34;, \u0026#34;X509_get_default_cert_dir_env\u0026#34;): getattr(c, f).restype = ctypes.c_char_p print(f\u0026#34;{f:34} {getattr(c, f)().decode()}\u0026#34;) X509_get_default_cert_file /workspace/.conan2/p/b/cisco8ee8b59cf93de/p/ssl/cert.pem X509_get_default_cert_dir /workspace/.conan2/p/b/cisco8ee8b59cf93de/p/ssl/certs X509_get_default_cert_file_env SSL_CERT_FILE X509_get_default_cert_dir_env SSL_CERT_DIR There it is from the library\u0026rsquo;s own mouth. When anything in Webex asks this copy of OpenSSL for the default trust store, it is handed a file and a directory that do not exist. That is what SSL_CTX_set_default_verify_paths does, and it is what almost every client does unless it has been told otherwise.\nThe last two lines are worth noting, because they are the escape hatch: the library will let SSL_CERT_FILE and SSL_CERT_DIR override both5. Remember that as well. It becomes important, and not in the way you would expect.\nReproducing The Exact Error Code Proof by inspection is not proof. Take the shipped libssl.so.3 and libcrypto.so.3, make a real handshake to a real host with nothing but the library\u0026rsquo;s own defaults, and watch what comes back.\nThe interesting run is the one where the defaults point at nothing. SSL_CERT_FILE and SSL_CERT_DIR override exactly the two values a missing OPENSSLDIR leaves dangling, so pointing them at a path that does not exist reproduces the shipped condition precisely:\n$ SSL_CERT_FILE=/nonexistent/cert.pem SSL_CERT_DIR=/nonexistent/certs python3 tls.py set_default_verify_paths -\u0026gt; 1 set_fd -\u0026gt; 1 SNI -\u0026gt; 1 SSL_connect -\u0026gt; -1 SSL_get_error -\u0026gt; 1 verify result 20: unable to get local issuer certificate err: 0xa000086 error:0A000086:SSL routines::certificate verify failed 0x0A000086 in decimal is 167772294.\nThat is the number in every failed line of the Webex log, and it did not come from Webex. It came from Cisco\u0026rsquo;s own TLS library, running outside their application, failing for exactly one reason: it had no trust anchors to check the chain against. Same library, same error, no application in the way.\nPoint the same two variables at the real Fedora bundle and the same code path completes:\n$ SSL_CERT_FILE=/etc/pki/ca-trust/extracted/pem/tls-ca-bundle.pem python3 tls.py SSL_connect -\u0026gt; 1 SSL_get_error -\u0026gt; 0 verify result 0: ok Nothing about the network changed between those two runs. One file path did.\nEvery network step succeeds. The step that fails is a local file lookup. Four steps cross the network and succeed. The fifth reads a file and does not. Run against the shipped libssl.so.3 and libcrypto.so.3, outside Webex, with nothing but the library's own defaults. TCP connect :443 ok ClientHello, SNI ok ServerHello, chain ok, chain received check the chain against trust no anchors loaded what the library reports back SSL_connect -\u0026#62; -1 verify result 20: unable to get local issuer certificate err: 0xa000086 error:0A000086:SSL routines::certificate verify failed 0x0A000086 is 167772294 in decimal, the number in every failed line of the Webex log. The application never wrote the word \"certificate\" anywhere. It logged \"Error in SSL handshake\" and a decimal integer, and the interface turned that into \"Offline - No internet connection\", which is the one thing the machine demonstrably was not. Nothing here is a network fault. The only thing missing was a directory. Four steps cross the network and succeed, including receiving the server\u0026rsquo;s full certificate chain. The step that fails opens a local file. The user is shown a message about their internet connection. Why Nothing Warned You Two things conspire to make this silent, and only one of them is Cisco\u0026rsquo;s fault.\nWhat never raises a word Whose design Why it stays quiet Loading a trust store that is not there OpenSSL\u0026rsquo;s, deliberately SSL_CTX_set_default_verify_paths returns 1 whether or not the paths are real. \u0026ldquo;A missing default location is still treated as a success\u0026rdquo;3 Having nothing to fall back to Cisco\u0026rsquo;s, and correct Certificates validated, self-signed refused, SSL retry off, so no degraded mode exists to mask a fault The first is reasonable. A program that ships its own anchors separately should not be forced to care, so the call succeeds, the store is empty, and nothing anywhere in the stack says I have loaded zero certificate authorities. The first thing that notices is a verification failure half a second later. I ran that call against the shipped library with the paths pointed at /nonexistent and it returned 1. It is in the output above.\nThe second is a policy that comes down from the service, and the log records it:\nWdm.cpp:1162 parseDeviceJson: Adding policy \u0026lt;\u0026lt; allowSelfSignedCertificate with value: false NetworkManager.cpp:1738 onConfigReady: ...httpRequestSSLRetryEnabled: 0 HttpRequestManager.cpp:1883 rawHttpRequest: {\u0026#34;validateCertificates\u0026#34;:\u0026#34;true\u0026#34;,\u0026#34;useClientCertificate\u0026#34;:\u0026#34;false\u0026#34;} Those three flags are the only reason this fault is an outage instead of something far worse, and that is worth sitting with rather than hurrying past.\nWork it through. The bundled library loads zero trust anchors, and nothing beneath the policy layer was ever going to notice, because the failure is silent by design all the way down. What turned it into a banner was three flags. Flip any one of them the way plenty of clients ship them and the same build does not fail at all. It connects, to anything holding any certificate, because it has nothing to check one against.\nThe same broken trust store, one policy flag away from a very different fault The trust path is equally broken in both columns. Only the policy decides how you find out. Common to both: the compiled-in OPENSSLDIR does not exist, so the store loads zero certificate authorities As Webex ships it validateCertificates: true allowSelfSignedCertificate: false httpRequestSSLRetryEnabled: 0 Nothing to verify against, so it refuses. Handshake fails. Banner. You lose a day. Loud, and harmless. Any one of them set the other way validateCertificates: false or self-signed permitted or retry without verification Nothing to verify against, so it proceeds. Handshake completes. Against anything. Silent, and not harmless. What separated the two outcomes was a policy value set by a different team, downstream of the defect, for unrelated reasons. The trust path was never protecting anybody. It was broken throughout, and the only open question was which way it would fail. The difference between \u0026ldquo;Webex is offline today\u0026rdquo; and \u0026ldquo;Webex trusted whatever answered\u0026rdquo; is a policy value set by a different team, downstream of the defect, for reasons that have nothing to do with it. What it surfaces as is the problem. The user gets \u0026ldquo;Offline - No internet connection\u0026rdquo;. The log gets \u0026ldquo;Error in SSL handshake\u0026rdquo; and a decimal integer. The word certificate does not appear anywhere the user or a first-line support engineer will ever look, and the one diagnostic breadcrumb is a number you have to convert to hex before it means anything.\nThe Fixes That Did Not Work Both of the obvious ones fail, and the reasons are different and both worth knowing.\nEditing the shipped openssl.cnf Webex ships a configuration file at /opt/Webex/lib/openssl.cnf, and as shipped it activates one provider:\n[provider_sect] fips = fips_sect Only FIPS. Not the default provider, which is where the ordinary TLS algorithms live6. Adding it back is a one-line edit and it changes nothing at all, because the bundled OpenSSL never reads that file. It looks for openssl.cnf inside OPENSSLDIR7, and OPENSSLDIR is the path that does not exist. The file sits in the installation directory looking authoritative. Nothing reads it.\nIt is a second defect hiding behind the first. Even if the path problem were fixed tomorrow by pointing OPENSSLDIR at /opt/Webex/lib, this configuration would then load and would activate only FIPS. And the FIPS module itself is loaded from MODULESDIR, the third dead path from the same tree, so fips.so shipping in /opt/Webex/lib cannot be found either. Three paths, one wrong prefix, and every one of them broken in a way that hides the next.\nSetting the environment variables The library honours SSL_CERT_FILE and SSL_CERT_DIR. I proved that above. The successful run is right there. So the obvious move is to set them in the desktop entry, which is what was tried:\nExec=env OPENSSL_CONF=/opt/Webex/lib/openssl.cnf \\ SSL_CERT_FILE=/etc/pki/ca-trust/extracted/pem/tls-ca-bundle.pem \\ SSL_CERT_DIR=/etc/pki/tls/certs/ /opt/Webex/bin/CiscoCollabHost %U No change. And the reason is not that Webex ignores the variables. The reason is that the process never received them:\n$ tr \u0026#39;\\0\u0026#39; \u0026#39;\\n\u0026#39; \u0026lt; /proc/20293/environ | grep -E \u0026#39;SSL|OPENSSL|CURL\u0026#39; $ tr \u0026#39;\\0\u0026#39; \u0026#39;\\n\u0026#39; \u0026lt; /proc/20293/environ | wc -l 99 Ninety-nine variables in the running Webex process, not one of them the three that were set. Because there are two desktop entries with the same name on this machine:\nFile What its Exec line runs Written by /usr/share/applications/webex.desktop the three-variable env line above, in full the .rpm, then edited by hand ~/.local/share/applications/webex.desktop /opt/Webex/bin/CiscoCollabHost %U, and no environment at all Webex\u0026rsquo;s own launcher And the specification is not ambiguous about which wins: \u0026ldquo;The base directory defined by $XDG_DATA_HOME is considered more important than any of the base directories defined by $XDG_DATA_DIRS.\u0026rdquo;8 $XDG_DATA_HOME is ~/.local/share. The user copy shadows the packaged one, every time, on every desktop that follows the spec9.\nSo Webex installs a second copy of its own launcher into your home directory, and that copy is the one your desktop runs. Change the packaged file all you like. You are editing a document nothing reads.\nTwo desktop entries with the same name, and the one that wins is the one Webex writes for itself The environment was set in the file the desktop never reads Two entries, one name. The search order is written down, and it does not favour the packaged copy. Edited by hand, and outranked /usr/share/applications/webex.desktop Exec=env OPENSSL_CONF=... SSL_CERT_FILE=... never consulted while the other file exists Written by Webex, and it wins ~/.local/share/applications/webex.desktop Exec=/opt/Webex/bin/CiscoCollabHost %U no environment set at all discarded launched XDG base directory specification: the base directory defined by $XDG_DATA_HOME is considered more important than any of the base directories defined by $XDG_DATA_DIRS. Proof, from the running process rather than from reasoning: tr '\\0' '\\n' \u0026#60; /proc/20293/environ | grep -E 'SSL|OPENSSL' \u0026#8594; no output, 99 variables, none of them these So the fix was real, the file was real, and the process it was meant for was started from somewhere else entirely. The environment was set in the file the desktop never reads. Webex writes its own entry under the user\u0026rsquo;s data directory, the specification says that one outranks the packaged one, and the proof is the running process: ninety-nine environment variables and none of the three. The Fix That Works If the library insists on a path, give it the path. Create the directory it was compiled to want and fill it with symlinks to the real thing:\n#!/bin/bash # Point the bundled CiscoSSL at the system trust store by building the # directory it was compiled to look for. Tested: Fedora 44, Webex 46.8.0.35631. set -euo pipefail OPENSSLDIR=$(strings /opt/Webex/lib/libcrypto.so.3 \\ | grep -oP \u0026#39;(?\u0026lt;=OPENSSLDIR: \u0026#34;)[^\u0026#34;]+\u0026#39;) [ -n \u0026#34;$OPENSSLDIR\u0026#34; ] || { echo \u0026#34;no OPENSSLDIR found in the shipped library\u0026#34;; exit 1; } for p in /etc/pki/ca-trust/extracted/pem/tls-ca-bundle.pem \\ /etc/ssl/certs/ca-certificates.crt \\ /etc/pki/tls/certs/ca-bundle.crt \\ /etc/ssl/cert.pem; do [ -f \u0026#34;$p\u0026#34; ] \u0026amp;\u0026amp; { CA_BUNDLE=\u0026#34;$p\u0026#34;; break; } done [ -n \u0026#34;${CA_BUNDLE:-}\u0026#34; ] || { echo \u0026#34;no system CA bundle found\u0026#34;; exit 1; } echo \u0026#34;OPENSSLDIR: $OPENSSLDIR\u0026#34; echo \u0026#34;CA bundle: $CA_BUNDLE\u0026#34; sudo mkdir -p \u0026#34;$OPENSSLDIR\u0026#34; sudo ln -sf \u0026#34;$CA_BUNDLE\u0026#34; \u0026#34;$OPENSSLDIR/cert.pem\u0026#34; sudo ln -sf \u0026#34;$(dirname \u0026#34;$CA_BUNDLE\u0026#34;)\u0026#34; \u0026#34;$OPENSSLDIR/certs\u0026#34; sudo tee \u0026#34;$OPENSSLDIR/openssl.cnf\u0026#34; \u0026gt; /dev/null \u0026lt;\u0026lt;\u0026#39;CONF\u0026#39; openssl_conf = openssl_init [openssl_init] providers = provider_sect [provider_sect] default = default_sect fips = fips_sect [default_sect] activate = 1 CONF echo \u0026#34;done. now restart Webex\u0026#34; Restart it, and the same startup sequence produces the opposite outcome. Same binary. Same log lines. Token refresh, which failed within sixty milliseconds before, now finishes in three hundred and thirty:\nAuthTokenRequester.cpp:579 Managed to fetch a new Kms access token. AuthTokenRequester.cpp:579 Managed to fetch a new CloudApps access token. AuthTokenSupervisor.cpp:310 Auth tokens refreshed. Expires in [64799 secs]. AuthenticationManager.cpp:1860 onUserAuthenticated: User authenticated. ConnectivityStateMachine::Authentication - UserNotAuthenticated -\u0026gt; UserAuthenticated Authenticated in about 750 ms, so the fifteen-second timer never fires and the banner never appears. Phone services go from Disconnected to Connecting, a state the broken session never reached in a minute of trying.\nBefore After Network check passes passes First HTTPS request 167772294 Error in SSL handshake HTTP 200 CloudApps token failed fetched Kms token failed fetched Authentication stuck at UserNotAuthenticated UserAuthenticated Banner at 15 s \u0026ldquo;Offline - No internet connection\u0026rdquo; none Phone services never attempted connecting Time to authenticate never ~750 ms Note what the fix does to your filesystem, because it should bother you. It creates a top-level directory called /workspace on your machine, a name the Filesystem Hierarchy Standard has no place for10, holding a Conan cache path and a build hash belonging to a company you bought software from. That is the shape of the remedy Cisco have left you: mount somebody else\u0026rsquo;s build environment at the root of your own.\nIt will also break. Three ways:\nWhen it breaks Why What you do A Webex update cisco8ee8b59cf93de is derived from the build inputs, so a rebuilt dependency means a new directory Re-run the script; it reads the path out of the new binary rather than assuming the old one A distribution change The symlink target is a distribution decision, not a standard11 Re-point it; the script probes four known locations A reinstall /workspace is not backed up, packaged, or owned by anything Run it again, every time, forever None of that is maintenance. It is you standing in for a step in someone else\u0026rsquo;s build pipeline, indefinitely, unpaid. Patching a defect at runtime, on every machine you own, because the vendor would not patch it once at build time.\nThey Knew The Rule And Applied It To Half The Build Here is what turns this from a bug report into an argument.\nRead the dynamic section of the shipped binaries:\n$ readelf -d /opt/Webex/bin/libcurl.so | grep RUNPATH 0x1d (RUNPATH) Library runpath: [$ORIGIN:$ORIGIN/../lib] $ readelf -d /opt/Webex/bin/CiscoCollabHost | grep RUNPATH 0x1d (RUNPATH) Library runpath: [$ORIGIN/../lib] $ORIGIN expands at load time to the directory holding the object itself12. It is the correct tool for a relocatable bundle and they used it properly. They had to: Webex ships the same tree twice, once to /opt/Webex and once to ~/.local/share/WebexLauncher/46.8.0.35631_9e6196c9-…/, and a launcher picks between them at startup. Two prefixes, one build, and the linker finds its libraries in both.\nSo the people who made this package understood the problem exactly. Absolute paths do not survive being shipped. They solved it for the code.\nThen they left the data paths as absolute strings naming the container that compiled them. Same build. Same afternoon.\nOne path in the bundle relocates itself. The other names a machine in a data centre somewhere. They knew the bundle had to move. They only applied it to the code. Both values are set at build time by the same people on the same day. One expands at load time, the other never expands at all. Where the code comes from: recorded as a relative expression RUNPATH in libcurl.so $ORIGIN:$ORIGIN/../lib /opt/Webex/lib resolved ~/.local/share/WebexLauncher/46.8.0.35631_.../lib also resolved Correct, and deliberately so. The same tree is shipped twice to two different prefixes and the linker finds it in both. Where the trust comes from: recorded as somebody's working directory OPENSSLDIR in libcrypto.so.3 /workspace/.conan2/p/b/cisco8ee.../p/ssl no such path not resolved ENGINESDIR and MODULESDIR: same tree, same outcome Three absolute paths into a build container, shipped to every customer, on a product sold with a support contract. The distance between the two halves of this picture is one person running the packaged artefact on a machine that did not build it. That is the whole defect. Not a hard bug. An untested one. The same build, the same day, the same engineers. The library search path is recorded as an expression that resolves wherever the tree lands. The trust path is recorded as somebody\u0026rsquo;s working directory. So this cannot be filed under they did not know. $ORIGIN is not something you stumble into. You reach for it because you have understood that an absolute path baked into a shipped artefact is a defect, and understood it well enough to go and fix it in the linker. Then the same build writes three absolute paths into the same libraries, and ships them.\nKnowing the rule and applying it to half the build is worse than not knowing it. Not knowing is a training problem and training has a fix. This is a package that had the right idea in it, in writing, in the ELF header where anyone could read it, and it went out of the door broken anyway. Which tells you that nothing downstream of the compiler was looking at the result. Nowt was.\nThe Container Was Already There Where /workspace came from is not a mystery, and you do not have to guess. It is in the package header:\n$ rpm -qi webex | grep -E \u0026#39;Build Host|Build Date|Vendor\u0026#39; Build Date : Sat 08 Aug 2026 20:47:43 BST Build Host : c964ea9239ae Vendor : Cisco c964ea9239ae is not a hostname anybody typed. It is twelve hexadecimal characters, which is what a container reports as its hostname when nothing sets one. So the package was built inside a container, by a company that plainly has the images, the registry and the orchestration to do it, and three separate artefacts on this machine say so independently:\nEvidence, read off the installed package Value What it proves Build Host in the RPM header c964ea9239ae a container ID, not a build machine OPENSSLDIR in libcrypto.so.3 /workspace/.conan2/… a path that only exists inside that container PLATFORM in the same library conan-Release-Linux-x86_64-gcc-13 a containerised Conan toolchain Requires in the RPM glibc \u0026gt;= 2.28 a deliberately chosen, very old ABI floor They then signed it. The build finished at 20:47:43 and the signature is dated 21:02:07 the same evening, key ID 9995e5bbb5ccde3c. Fifteen minutes. So there is a release gate, somebody or something operates it, and what it attests is who made the package, not whether the package works. A signature is a statement about provenance. It has never been a statement about fitness, and a process that has one and not the other has its priorities in the wrong order.\nBecause the missing step is the cheap one. The container is already in the pipeline. Take the artefact that just came out of it, start a clean image of each distribution you claim to support, install it, launch it, and read the first hundred lines of the log:\ndocker run --rm fedora:44 sh -c \u0026#39; dnf -y install ./webex-46.8.0.35631-1.x86_64.rpm \u0026amp;\u0026amp; timeout 25 /opt/Webex/bin/CiscoCollabHost \u0026amp; sleep 20 grep -c \u0026#34;Error in SSL handshake\u0026#34; ~/.local/share/Webex/current_log.txt\u0026#39; Non-zero, every time, on this build. The install is 1.1 GB, so call it a minute per target on a warm cache. Six distributions is six minutes of a machine that is already running, on hardware Cisco already own, in a pipeline that already exists. It was not run. Not once.\nAnd that is the answer to the thing people still say about Linux, that supporting several distributions is hard. It stopped being hard the day this tooling arrived, and the tooling is the same tooling they compile with. One base image per target. Same artefact into each. The matrix is a loop.\nA pipeline that builds in a container and never runs the result in one is not a pipeline. It is a compiler with a cron job in front of it and a signing key behind it, and whatever it produces is a guess.\nThey build inside a container and never run the result in one The container is already in the pipeline. It is only used for the half that suits the vendor. Every value below is read out of the package sat on the machine, not inferred. Build, inside a container Build Host: c964ea9239ae /workspace/.conan2/p/b/... conan-Release-Linux-x86_64-gcc-13 three separate proofs of a container Sign and publish built 20:47:43 signed 21:02:07 fifteen minutes, and a release gate that attests who, never whether Customer dnf install webex opens it \"Offline - No internet\" The step that is not in the pipeline anywhere docker run --rm fedora:44 sh -c 'dnf -y install ./webex.rpm \u0026amp;\u0026amp; CiscoCollabHost \u0026amp; sleep 20; grep -c \"Error in SSL handshake\" ~/.local/share/Webex/current_log.txt' Same infrastructure. One image per target distribution. Under a minute each, and it fails loudly on this build. Building for several distributions stopped being hard the day this tooling arrived, and it is the same tooling they compile with. A pipeline that builds in a container and never runs the result in one is a compiler with a cron job in front and a signing key behind. Three independent proofs in the shipped package that the build ran in a container, a signature applied fifteen minutes after the build, and the one step that appears nowhere: running the finished package in a clean image of each platform it is sold as supporting. Why Is A 2026 Release Built Against A 2018 Libc? The package declares what it needs, and the interesting line is the first one:\n$ rpm -q --requires webex | grep glibc glibc \u0026gt;= 2.28 glibc 2.28 was released on 1 August 201813. It is the version in Red Hat Enterprise Linux 814, a release whose full support ended in 2024. This is a 2026 product, compiled in 2026, targeting the C library of 2018. Then shipping its own 2.5 MB copy of libstdc++.so.6 into the user\u0026rsquo;s home directory, because the C++ runtime that goes with a base that old cannot carry the code.\nWhy would anybody still do that? Because they will not link statically.\nThat is the whole of it. The moment you dynamically link against the host\u0026rsquo;s C library, the oldest distribution you are willing to support becomes a constraint on the machine you compile on. You cannot use a symbol that the oldest target has never heard of, so you pin the build to an ancient base image and you stay there. Every year the gap widens. Every new language or library feature arrives with an argument about whether the floor can move. Ship your own libstdc++ to paper over the worst of it, and you are now maintaining a private runtime as well.\nThat trade made sense when a build host was a physical machine in a rack that somebody had to reinstall. It has not made sense for a decade. If you must link against a host libc, a container per target gives you a real build on a real version of each platform, and none of them constrain the others.\nBut look at what the old-target discipline actually bought here. It is ABI conservatism at considerable cost: an eight-year-old floor, a bundled C++ runtime, a support matrix frozen around it. And the client still will not start. Because what broke was a file path, and no amount of care about symbol versions protects a file path. They paid the compatibility tax and the bundling tax, and skipped the one check that costs nothing. Both bills. No product.\nThey Paid For A Self-Contained Bundle And Did Not Get One The long-standing way to ship commercial software on Linux is to depend on as little of the host as possible. Static where you can, a self-contained tree where you cannot, and no assumptions about the distribution underneath. It is not elegant and nobody pretends it is. It exists because of the alternative. A binary that needs a particular version of a particular library in a particular place turns every customer\u0026rsquo;s machine into a support case.\nCisco went the second route and bundled. Here is the bill:\nWhat the bundle holds Size or count Shared objects in /opt/Webex/lib 150 Their own libcurl, CiscoSSL, zlib-ng, ICU, Kerberos, CUPS client, hunspell and inference engine all of them Installed under /opt/Webex 1.1 GB A second copy under ~/.local/share/WebexLauncher, per user 1.1 GB On disc for a chat and calling client 2.2 GB Two gigabytes of dependencies. That is the full price of bundling: the download, the disc, the duplication, the security burden of being the only party who can patch any of it, the whole lot. Pay that and what you buy is a program that does not care what the host has installed, that behaves the same on Fedora and Debian and Arch and whatever a customer has standardised on this year, and that cannot be broken by a distribution upgrade it never heard about.\nExcept it does. It cares very much about one directory, and it is a directory on a build server.\nThey shipped their own Kerberos and their own ICU and their own spell checker, and they could not ship a working path to the certificates. The bundle\u0026rsquo;s entire purpose is to be self-contained, and it is not self-contained in the one respect that stops it working. Every one of those 1.1 GB gets found correctly through $ORIGIN. The forty-odd bytes that matter most do not.\nAnd this is the bit that gets me. Bundling is the expensive option. They took the expense, they did the hard engineering, they got the relocation right for a hundred and fifty libraries across two install prefixes. Then pointed the one part that decides whether any of it works at a machine no customer has ever owned.\nNobody Documented It Either A second failure sits alongside the first, and it is the one that would have caught the first.\nAsk the package what documentation it ships:\n$ rpm -qd webex | wc -l 0 $ rpm -qc webex | wc -l 0 No documentation files. And no configuration files. Zero entries marked %config in a package that ships an openssl.cnf. That is not cosmetic. An RPM marks a file %config so that the package manager preserves what the administrator changed, saving a .rpmsave rather than clobbering it1516. /opt/Webex/lib/openssl.cnf is shipped as an ordinary file, so the next update overwrites any edit you made to it and tells you nothing. The one file a customer might legitimately need to adjust is the one the packaging treats as disposable.\nNothing published says which trust store the client uses, which environment variables it honours, or where it reads its TLS configuration from. There is no page to check. None was written. The only way I established any of it was strings, readelf, ldd and ctypes against the shipped binaries. Reverse-engineering a supported product to answer a question its documentation should have answered in a sentence.\nAnd here is why that matters more than it sounds. Writing that documentation is itself a test. Put anyone at the vendor in front of a blank page headed where Webex for Linux reads its CA certificates from, and the first thing they have to do is go and look. The moment they look, they find /workspace/.conan2/p/b/cisco8ee8b59cf93de/p/ssl, and the next thing out of them is a question. Documented configuration paths are not paperwork for the customer\u0026rsquo;s benefit. They are the cheapest audit a vendor can run on its own build, and skipping them is how a path like that survives to a release.\nNone of that is a big ask. Say where your configuration lives. Say which environment variables you honour. Mark your configuration files as configuration so an update does not eat them. Then the customer who hits a fault has somewhere to look that is not a hex editor.\nNobody Ran It The missing step has been named enough times above. What is worth asking is why it went missing on this platform and not the others.\nNot on Fedora, anyway. On macOS and the Fisher-Price OS (Windows) this class of fault cannot show up the same way, because those platforms have a system TLS stack with a trust store the operating system manages. Linux has no such thing. OpenSSL is the trust store, and as such whoever ships it owns where it looks. So the one platform where the bundled library is load-bearing is the platform that got shipped untested, which is a decision about which customers are worth a smoke test.\nThe whole investigation took an evening: reading the log, spotting that network detection passed before the banner claimed otherwise, pulling the compiled path out of the binary, reproducing the exact error code against the shipped library. All of it done with an AI agent doing the log correlation and the ctypes harness while I worked out what to ask it. I mention that for one reason: the diagnosis a vendor never made before signing this package and putting it in their own repository is now inside an evening\u0026rsquo;s reach of any customer with the patience to look. Cisco\u0026rsquo;s engineers have Claude available to them the same as anyone else, and it would have built this properly. You cannot ask it to bake /workspace/.conan2 into a shipping artefact without it telling you what will happen when the artefact leaves the workspace. The tooling to catch this is no longer scarce, and neither is the knowledge. What is missing is anyone at the vendor whose job it was to look.\nMeanwhile the support path for the person hitting this is a banner that says their internet is down. They will restart the router. They will ring their provider. They will raise a ticket that goes nowhere, because the symptom Cisco chose to display points away from Cisco.\nThat is the complete list of what went wrong. It is worth saying what right would have looked like, because every item on it has a settled answer that predates this product.\nHow It Should Have Been Built Enough of what went wrong. Here is the standard, and none of it is novel. It is what shipping a binary to somebody else\u0026rsquo;s computer has asked of you for twenty years.\nStart with the decision Cisco got right in principle and wrong in execution: how much of the host you are willing to depend on. Linking statically is the strongest answer, and it is worth being concrete about it, because \u0026ldquo;just link it statically\u0026rdquo; gets waved about by people who have never had to and dismissed by people who have never tried.\nWhat it removes What it costs RUNPATH, and every way of getting it wrong, because there is no search at load time glibc does not link statically clean: name and user lookup go through dlopen, so the binary still reaches for the host\u0026rsquo;s NSS modules17 The glibc floor, so the oldest distribution stops dictating what you may compile on LGPL parts carry a relinking obligation, so they stay dynamic or you ship what is needed to relink Breakage from a distribution upgrade, a library rename or a stale ldconfig cache You own every patch: no distribution security update reaches your customers dlopen of a versioned .so from a directory that may not be there A bigger download, and no sharing of pages between processes And one line in the left column that the others are only supporting: the artefact your pipeline produced is the artefact the customer runs, byte for byte. Test it and you have tested the thing you shipped. That is the property this whole post is about the absence of.\nThe right column is worth being honest about. Genuine costs, and the reason people reach for musl or accept a hybrid. But look at the third row. Owning every patch applies identically to what Cisco already did: a bundled tree of 150 libraries is the same commitment with none of the load-time guarantees. They signed up to own every patch either way and got nothing back for it.\nSo the rule is simple, and it is the rule they broke: whatever you cannot link in, you must find by a path relative to the binary. $ORIGIN for the code, and the same discipline, deliberately, for every data path the library will go looking for. In this case there are four of them and every one has a documented handle on it:\nWhat the library goes looking for What shipped What it should have been Trust anchors OPENSSLDIR/cert.pem, fixed at build time shipped in the tree, or SSL_CERT_FILE set at startup5 Configuration OPENSSLDIR/openssl.cnf, same path, never found OPENSSL_CONF, pointed at the copy in the package5 Providers, including FIPS MODULESDIR, absolute, so fips.so is unreachable OPENSSL_MODULES, \u0026ldquo;the directory from which cryptographic providers are loaded\u0026rdquo;5 Engines ENGINESDIR, absolute OPENSSL_ENGINES, or nothing, since OpenSSL 4.0 removed engine support altogether5 Four paths, four environment variables, all of them in one manual page their own library ships with. If any string in your artefact starts with a / and was decided at build time, it is a defect waiting for a customer to find. There is no third option where an absolute build path is fine.\nThe Fixes For This One, And The Check That Catches It Down from the principle to this specific defect. Six changes, not one of them research:\nFix Effort Why it is right Set --openssldir=/opt/Webex/lib/ssl and ship the tree one configure flag The path then exists in the package, which is what the flag is for2 Set SSL_CERT_FILE/SSL_CERT_DIR in CiscoSSLUtils at startup, probing the known distribution paths a dozen lines Standard, documented, already supported by their own library5 Ship the CA bundle themselves inside the package packaging only Full control of trust, at the price of owning its freshness Fix the openssl.cnf to activate the default provider one line Needed regardless, and currently masked by the path fault6 Mark openssl.cnf as %config one spec line Stops an update silently eating an administrator\u0026rsquo;s change15 Publish the paths and the variables honoured one page The cheapest audit there is, and it finds this fault while being written The first is the one-flag fix and it would have shipped the product working. It is a single value in a build script, set once, that they got wrong because nothing downstream ever checked it.\nThe check is smaller than the fix:\n# in CI, on the packaged artefact, in a clean container test -d \u0026#34;$(strings lib/libcrypto.so.3 | grep -oP \u0026#39;(?\u0026lt;=OPENSSLDIR: \u0026#34;)[^\u0026#34;]+\u0026#39;)\u0026#34; \\ || { echo \u0026#34;shipping a trust store path that does not exist\u0026#34;; exit 1; } One line. It would have failed this release loudly, in February, on the machine that made it, in the container that made it, before the package was signed. The reason it is not there is not difficulty and not cost. It is that nobody was asked to write it, and so whether the packaged artefact behaved like software was not anyone\u0026rsquo;s job.\nWhat A Sloppy Build Process Costs Everybody Else Everything above is one product\u0026rsquo;s build process seen from the inside, and I will not walk you back through it. The question worth asking is whether that standard of care is likely to stop at one team.\nSo put the outside record next to it. CISA maintain a catalogue of vulnerabilities known to be exploited in the wild. Not theoretical, not scored. Observed being used against people. As of 14 September 2026 it holds 1,710 entries18:\nVendor Entries in the KEV catalogue Microsoft 388 Cisco 98 Apple 94 Adobe 81 Google 74 Oracle 46 Fortinet 30 VMware 26 Second, behind an operating system monopoly and ahead of everybody else. The most recent Cisco addition went in on 14 September 2026, the day before this was written.\nI counted the same file three weeks earlier for /en/random/is-your-msp-lying-to-you-part2/: version 2026.08.27, 1,685 entries, Cisco on 96. Two more since, in twenty-one days.\nBe fair about what that table does and does not prove on its own. A large installed base in high-value places attracts attention, and attention finds bugs, so any vendor at that scale will carry a long list. The number is a prior, not a verdict.\nThe composition is harder to wave away than the total. Twenty-five of Cisco\u0026rsquo;s ninety-eight sit on the security lines: the firewalls, the appliances, the VPN concentrators, the mail and web gateways, the identity services. And the four most recent entries added against them, in a row, are Secure Firewall Management Center twice, Secure Firewall ASA, and Secure Email Gateway, the last of those on 14 September 202618. Not the switches. Not the collaboration kit. The products sold specifically to be the thing that keeps everybody else safe.\nWhich is the sentence this whole post has been walking towards, so here it is plainly. This is a company whose business is selling security appliances, and it cannot get a desktop client to check a certificate. Not a hard case. Not a novel attack. The single most routine security operation in computing, performed by every browser on every page load, in a library they forked themselves, and they shipped it pointed at a directory that has never existed on a customer machine. Then signed it.\nWhat this post adds is a sample of the process behind it. Not a vulnerability. A plain packaging defect, the least subtle class of fault there is, on a supported platform, caught by any smoke test anyone cared to run, and shipped anyway. If a build pipeline will put that in front of a paying customer, it is not obvious what it would stop. Nothing did.\nCan I Ask Why This Was Signed? Can I ask why a package that cannot complete a TLS handshake was signed and published against a platform your own requirements page lists as supported? Not who signed it. I have no interest in a name and it is not about one. What permitted it, because something did, fifteen minutes after the build finished, and whatever it is it is still running today.\nNow the blunt version. This is not a project I pulled off a forge and took my chances with. It is paid for. There is a contract behind it, a support line, a requirements page making a claim about Linux, and a signing key asserting the package came from Cisco and is fit to install. Each of those is a statement to a customer, and on this release each was worth nothing, because the result was never opened.\nAnd the standard being missed here is not mine. It is theirs. The four variables that would have carried those paths are documented in a manual page you ship inside the bundle. Their own packaging format has a %config marker they did not use and a documentation section they left empty. Their own build ran in a container they never used to run the output. I am not asking a networking and collaboration company to invent anything. I am asking why it did not read its own manuals.\nAs such the excuse I will not accept is that this is hard. It is not hard for anybody, and it is certainly not hard for a company that does this daily, at this scale, for this money. Professionals doing this every day have no excuse, and being large is not one.\nAnd can I ask one more, because it is the one that actually matters. You sell firewalls. You sell an email gateway, a VPN concentrator, an identity service, a management centre for all of it, and the pitch on all of them is that you understand this better than your customer does. So how does a company that positions itself as an authority on network security ship a client that cannot validate a certificate? Not fails to validate it cleverly. Cannot look one up, because the directory was never there.\nThere is no version of that answer I want to hear that starts with the desktop team being separate from the appliance team. They are the same signature, the same pipeline, the same published claim about a supported platform, and the same missing step at the end, which is opening the box. If certificate validation is not checked before shipping in the product where breakage is loud and harmless, I have no reason to believe it is checked in the product where breakage is quiet and expensive.\nAnd the honest part, said plainly. None of this surprised me. I have come to expect this standard from this vendor, which is why, where the choice is mine, I do not buy their kit and I do not build on it. That is not a preference about badges. It is the same judgement I would make about any supplier: I have measured their work, more than once, and it keeps coming back the same. The catalogue above is one measurement. What it took to write the first half of this post is another.\nThe reason I was on Webex at all is that the choice was not mine, and that is worth naming, because it is the position most people reading this are in. You rarely get to avoid a vendor whose quality you have already taken the measure of. Somebody else signs the contract, the kit arrives, and the first person to find out what was skipped in the build is you, at your desk, with a meeting starting.\nShipping Something You Never Ran There will be an internal explanation. A busy sprint, a pipeline that changed hands, a platform with no owner. None of it is worth hearing, because each describes the same thing: the work was not done and nothing in the process required it to be.\nThere is an old standard in trades work that has never made it into software: you do not leave the job until you have run the thing. You fill the system and check every joint. You energise the board and test every circuit. Not because you doubt your work. Because the customer is going to use it, and finding out in front of them is not a professional outcome. Nobody watches you do it. You do it anyway. That is the entire meaning of the word.\nSoftware has spent thirty years arguing that it is different, that the build is the deliverable and the install is somebody else\u0026rsquo;s problem, that a green pipeline is the same thing as a working product, and it is not. A build that has never been executed outside the container that produced it has not been finished, it has been abandoned at the point where finishing gets boring. Everything downstream of it is a claim about work that was not done.\nThe fix is one configure flag. The check is one line of shell. The cost of neither is what is interesting; the interesting number is how many people typed their credentials into a Webex window, watched it say their internet was down, and believed it, because Cisco told them so and Cisco are a networking company. That is what the missing test actually bought: not a bug, a lie that the product tells confidently, every launch, to people who have no way to know better.\nAnd that is the line worth drawing before you close the tab, because the two halves of this post are not two subjects. A trust path that points at a build container and an authentication bypass on an edge appliance are the same failure at different stakes. Both are a value nobody checked, in an artefact nobody ran, signed by a process that attests provenance and not fitness. The one in front of me was the harmless kind. It broke loudly, on my own desk, and I was the first to know. The other kind does not do that for you.\nThe same missing check at two different stakes, and only one of them tells you One missing check, two stakes. Only one of them tells you it is there. Same underlying failure: a value nobody checked, in an artefact nobody ran, signed by a process that attests provenance The kind that breaks loudly what it was a trust store path that does not exist who found it the customer, first launch what it cost one evening, one desk who learned me, immediately Ends up written down in public. The kind that breaks nothing what it was a length nobody bounds-checked who found it whoever went looking what it cost an incident who learned the customer, from the report Ends up as 1 of Cisco's 98 in the catalogue. A process that will not catch the left-hand column was never going to catch the right-hand one. The difference is luck, not diligence. The defect in this post and the entries in that catalogue are the same failure wearing different clothes. One announced itself on my desk on the first launch. The other kind announces itself to somebody else first. Nobody at Cisco decided to ship an exploitable firewall, any more than they decided to ship a client that cannot reach the internet. That is not a defence. It is the charge. Neither needs deciding, and that is precisely the problem: both are what comes out the far end when a pipeline compiles, signs and publishes without anyone being made responsible for opening the result. A process that will not catch a directory that does not exist was never going to catch a length that is not checked, and a vendor that sells you the appliance guarding your perimeter is not entitled to that process.\nSo when a vendor\u0026rsquo;s name keeps appearing on that list, resist the comfortable explanation that they are simply big and heavily targeted. Scale explains the volume. It does not explain the kind. Look instead at what their build process does with the boring things, because the boring things are measurable from the outside, by you, today, on kit you already own.\nCheck your own bundles. strings and readelf and twenty minutes will tell you which of your vendors ship a path to a machine you will never see. What you are really measuring is not the path. It is whether anybody there was looking, and if the answer is no on something this cheap to catch, you already know what it is on the things that are not.\nCisco — Webex App system requirements — the supported platform list, Linux included.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nOpenSSL — INSTALL.md, --openssldir — \u0026ldquo;Directory for OpenSSL configuration files, and also the default certificate and key store.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nOpenSSL — SSL_CTX_load_verify_locations(3) — the default CA file is cert.pem and the default CA directory certs, both inside the default OpenSSL directory; and on return values, \u0026ldquo;A missing default location is still treated as a success.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nConan 2 — conan cache — package binaries live under a hashed path in the local cache, which is where /workspace/.conan2/p/b/\u0026lt;hash\u0026gt;/p comes from.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nOpenSSL — openssl-env(7) — SSL_CERT_DIR and SSL_CERT_FILE \u0026ldquo;specify the default directory or file containing CA certificates\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nOpenSSL — provider(7) — the default provider and what a configuration that omits it leaves unavailable.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nOpenSSL — config(5) — the configuration file OpenSSL loads at initialisation, and where it is looked for.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nfreedesktop.org — XDG Base Directory Specification — \u0026ldquo;The base directory defined by $XDG_DATA_HOME is considered more important than any of the base directories defined by $XDG_DATA_DIRS.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nfreedesktop.org — Desktop Entry Specification — where .desktop files are searched for and how one shadows another of the same name.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nFilesystem Hierarchy Standard 3.0 — the directories a root filesystem is expected to contain, /workspace not among them.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nupdate-ca-trust(8) — how the consolidated bundle under /etc/pki/ca-trust/extracted is produced on Fedora and its relatives.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nld.so(8) — $ORIGIN expands to the directory containing the program or shared object, which is what makes a bundled tree relocatable.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nglibc timeline — \u0026ldquo;2018-08-01 GLIBC 2.28 — The GNU C Library version 2.28 is now available\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nglibc — Release wiki — the version-to-distribution table; Red Hat Enterprise Linux 8 is glibc 2.28.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRPM — spec file reference — %config and what the package manager does with a file marked as configuration.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nFedora packaging guidelines — configuration files — when a shipped file must be marked %config.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nglibc FAQ — why a statically linked glibc binary still needs the host\u0026rsquo;s NSS modules at runtime.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCISA — Known Exploited Vulnerabilities Catalog — counted from the published JSON feed, catalogue version 2026.09.14, 1,710 entries.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/certificates/webex-certificates-on-a-build-server/","summary":"Webex 46.8.0.35631 on Fedora 44 sits behind a yellow banner reading \u0026ldquo;Offline - No internet connection\u0026rdquo; while every other program on the machine reaches the internet without complaint. It stopped the VoIP line dead. The network was never the problem: Cisco bundle their own fork of OpenSSL and built it with OPENSSLDIR set to /workspace/.conan2/p/b/cisco8ee8b59cf93de/p/ssl, a directory that exists on a build container and nowhere else, so it loads no trust anchors and every handshake fails. The post runs in order: what the broken state actually is, asking the shipped library where it thinks its certificates live, reproducing the exact error code outside Webex, why nothing warns you, the two fixes that do not work and why, and the one that does. Then the build process: they used $ORIGIN for the code and left three data paths absolute, the RPM header names a container ID so the container was already in the pipeline and never used to run the result, the package requires a glibc from 2018 because they will not link statically, 2.2 GB is shipped twice, and it carries no documentation and no file marked as configuration. Then how it should have been built, the six fixes and the one-line check that catches it. And finally the cross-check: Cisco hold 98 entries in CISA\u0026rsquo;s known-exploited catalogue, second only to Microsoft, and a trust path nobody checked and a bypass nobody checked are the same failure at different stakes.","title":"Webex Is Looking For Your Certificates On A Cisco Build Server"},{"content":"A VPN is not a product. It is two jobs bolted together, and your machine already has a program for each of them.\nThe first job is making a virtual link — an interface that looks like a network card, takes IP packets and hands them elsewhere. The second is carrying that link\u0026rsquo;s bytes end to end. Buy a VPN appliance and you buy both jobs with a licence stapled on; do it yourself and each is one program already installed. There are two ways to do the first job on Linux, and this post builds both: pppd, which brings a whole link protocol — negotiated addresses, a liveness check, framing — and a tap device, which brings nothing and is a hole in the kernel you push frames into. Different trades, not competitors, and the second half sets them side by side.\nThat protocol is PPP, and the reason this works is written into its first paragraph: PPP is a link protocol for point-to-point links1, and it does not specify what the link is made of. A modem and a phone line were the original answer. They were never the only one. PPTP put PPP inside GRE2. L2TP put it inside UDP3. PPPoE put it inside Ethernet frames. Every one of those is the same protocol with a different courier, and none of them needed PPP changing to do it.\nSo the question is not whether you can run PPP over a TCP or UDP socket, but which, and what it costs. The short answer, working shown: use UDP. Running a stream protocol inside a reliable stream is the one mistake here that looks fine on your desk and falls apart on a real link.\nA VPN Is Two Jobs. You Already Have Both. Strip the marketing off any tunnel and three things happen: something presents a virtual interface and turns packets into a byte stream, something carries that stream across a network that already works, and something encrypts it, or nothing does.\nEvery tunnel is a link half, a carrier and some crypto The same three parts every time. Only the carrier and the crypto change. TUNNEL LINK HALF CARRIER CRYPTO Dial-up PPP, 1994 PPP modem and a phone line none PPPoE PPP Ethernet frames none PPTP PPP GRE MPPE — broken, do not L2TP over IPsec PPP UDP IPsec This post, build one PPP over a pty UDP socket, via netcat DTLS This post, build two tun or tap device UDP socket, via socat DTLS WireGuard kernel interface UDP Noise, and nothing to negotiate Every tunnel on this list is the same shape. The differences are which carrier the bytes go into, and whether anything encrypts them. PPP has been doing the link half since 1994, which is why it turns up under three of these without being redesigned for any of them. A tun or tap device does the same half with no protocol at all, which is the second build in this post. PPP does the link half properly — address negotiation, a liveness check, header compression, several network-layer protocols over one link, an authentication step if you want one — a lot of finished work sitting in /usr/sbin doing nothing. What it does not do is care about the carrier: give it a file descriptor that moves bytes both ways and it runs over it. Netcat is a program whose entire purpose is being that file descriptor.\nThe encryption is the part nobody hands you, and it is the part most of this post is spent earning, because netcat has no answer for it and pretending otherwise is how people build things they should not.\nThe Stages, And Why They Go In This Order Everything below is built in stages, each adding one thing to the one before. That is not a teaching device. It is how you should build it, because when stage six breaks you need to know whether stage two still works, and you only know that if stage two was ever a thing you ran on its own.\nSix stages, each adding one thing you can test on its own Build it in this order, and you can always walk back down it 1 Raw, over TCP pppd with a pty, or a tap device, and a plain netcat listener. Works on a bench. Melts on a real path — two retransmission timers fighting. 2 Raw, over UDP Same link, datagram carrier. A lost datagram is one lost packet and nothing else. This is the foundation. Everything above it is optional; this is not. 3 TLS, over TCP ncat, stunnel or openssl wrapping the carrier. Verify the certificate at both ends, or you have encryption without identity and no access control at all. 4 DTLS, over UDP socat or openssl, and the shape to aim for: encrypted, authenticated, still datagrams. A lost packet stays a lost packet instead of two stacks arguing about it. 5 Compressed zstd between the interface and the carrier. Keep the window across frames where the carrier delivers in order; one self-contained frame per datagram, plus a dictionary, where it does not. 6 Both, in that order Compress, then encrypt. Never the other way round — ciphertext does not compress. And know the trade: compressing before encrypting leaks plaintext length. Fine on your own traffic. Six stages, each adding one thing to the one below it. Raw over TCP is where everyone starts, and the only rung that is a dead end: it works on a bench and melts on a real path. Raw over UDP is the foundation everything else sits on. Then encryption, then compression, then both together. Every stage is a link you can bring up, ping across and leave running, so when the top of the ladder misbehaves you can walk back down it one rung at a time. The link itself is built the same way — the reasoning behind the one-second ramp further down: a tunnel comes up raw, proves it can carry a frame, and only then gets clever. A build that turns everything on at once fails as one lump, and you spend the evening guessing which layer did it.\nWhat pppd Actually Wants Is A File Descriptor pppd was written for serial ports, so the naive reading is that it needs one. It does not. It needs a terminal device, and it will make one for itself:\npty script — Specifies that the command script is to be used to communicate rather than a specific terminal device. Pppd will allocate itself a pseudo-tty master/slave pair and use the slave as its terminal device. The script will be run in a child process with the pseudo-tty master as its standard input and output.4\nRead that again — the whole build is in it. pppd makes a pseudo-terminal, keeps the slave, and hands the master to a command as stdin and stdout. What that command does with the bytes is not pppd\u0026rsquo;s business. Run nc there and they go down a socket.\npppd, a pseudo-terminal, netcat and a socket One packet, from ppp0 to ppp0, and every hand it passes through HOST A HOST B ppp0 kernel interface pppd HDLC framing pty pair slave to pppd master to the child ncat stdin and stdout pty pair master to the child slave to pppd ncat stdin and stdout pppd HDLC framing ppp0 kernel interface the socket — TCP or UDP The two boxes in the middle are the only part you choose. Everything either side is unchanged whether the carrier is a phone line, a serial console, a TCP stream, a UDP datagram flow or a TLS session — which is why swapping the carrier later costs one word. A pseudo-terminal has no carrier-detect pin, so pppd waits for a carrier that never arrives. The `local` option is what stops it. pppd speaks async HDLC into the slave side of a pseudo-terminal it allocated itself. The command named by pty inherits the master side as stdin and stdout. Netcat copies stdin to a socket and the socket to stdout, so the two pppd instances are talking to each other through the pty pair and the network without either of them knowing a socket exists. There is a second route, notty, which uses pppd\u0026rsquo;s own stdin and stdout instead. But it spawns a character-shunt process that every byte passes through, so it \u0026ldquo;increases the latency and CPU overhead\u0026rdquo;4. Use pty unless you have a reason not to.\nThree practical facts before the first command, all checked on the box this was written on — Fedora 44, ppp-2.5.1-7.fc44:\n$ pppd --version pppd version 2.5.1 $ ls -l /usr/bin/pppd /dev/ppp -rwxr-xr-x. 1 root root 393784 Jan 17 2026 /usr/bin/pppd crw-------. 1 root root 108, 0 Sep 14 08:34 /dev/ppp $ pppd noauth nodetach pty \u0026#39;true\u0026#39; pppd: using the noauth option requires root privilege It needs root. No way round that. The binary is not setuid on Fedora, /dev/ppp is mode 0600 root-owned, and the options you need are privileged anyway. This is a thing you run as root or under a unit file, not something a user does casually. That matters later, when we get to what it means for your egress policy.\nThe kernel side is modular. ppp_generic does the interface, ppp_async the byte-stuffed framing a pty needs, and ppp_deflate/bsd_comp/ppp_mppe compression and encryption; they load on demand. If ppp_async is missing from a stripped container image the link comes up and carries nothing — a miserable half hour if you do not know to look.\nModem control lines do not exist on a pty. With no carrier-detect pin, pppd waits for a carrier that never comes. The fix is one word, local, which tells it to ignore the modem control lines4. Leave it out and nothing happens, with no error worth reading.\nNetcat Is Not One Program Before the build, the trap that eats the most time: \u0026ldquo;netcat\u0026rdquo; is at least four programs with incompatible flags, and which one you get depends on your distribution. On this machine /usr/bin/nc is a symlink to /usr/bin/ncat, Nmap\u0026rsquo;s rewrite. So a command copied off a fifteen-year-old wiki page fails for reasons that have nothing to do with PPP.\nImplementation Listen on a port UDP Stay up after a peer disconnects Ncat (Nmap) ncat -l 443 -u -k / --keep-open OpenBSD netcat nc -l 443 -u -k Traditional netcat nc -l -p 443 -u no GNU netcat nc -l -p 443 -u no As such, check which one you have before you blame pppd:\nreadlink -f \u0026#34;$(command -v nc)\u0026#34; nc --version 2\u0026gt;\u0026amp;1 | head -1 || nc -h 2\u0026gt;\u0026amp;1 | head -1 I will use ncat explicitly below. It is the one with TLS built in, which matters later, and being explicit means the commands do not silently mean something else on your box.\nStage One: The TCP Build, Because That Is What Everyone Tries First One end listens, one end connects. Addresses here are from the documentation range, so paste them straight into a lab and nothing routable is harmed. The port is 443 throughout, and that is deliberate: outbound 443 is open on almost every network by default. Wrap the carrier in TLS a few stages down and the traffic on it is indistinguishable from any HTTPS session. Binding a listener there needs root, which the far end — the attacker\u0026rsquo;s own box, or a relay — has. That is the whole egress argument in a port number, and I will come back to it.\nOn the listening end:\npppd nodetach noauth local passive \\ nodefaultroute noipdefault \\ 192.0.2.1:192.0.2.2 \\ lcp-echo-interval 10 lcp-echo-failure 3 \\ mtu 1400 mru 1400 \\ pty \u0026#39;ncat --listen --keep-open 443\u0026#39; On the connecting end:\npppd nodetach noauth local \\ nodefaultroute noipdefault \\ 192.0.2.2:192.0.2.1 \\ lcp-echo-interval 10 lcp-echo-failure 3 \\ mtu 1400 mru 1400 \\ pty \u0026#39;ncat 198.51.100.10 443\u0026#39; Every word there is doing a job, and one of them, noauth, is doing quiet damage. The tunnel has no access control, a section of its own later.\nEvery word on the pppd line, and what it does The UDP build's command line, word by word nodetach foreground, so you see its output and can put it under a supervisor noauth no peer authentication — this is the hole; the tunnel has no access control local ignore the modem control lines a pty does not have. Without it, nothing comes up passive wait for a valid LCP packet instead of exiting — server-socket behaviour nodefaultroute do not grab the default route when IPCP finishes. Add the route you want, yourself noipdefault do not offer the machine's own address as the local end 192.0.2.1:192.0.2.2 local:remote, pinned — pppd rejects any other answer during IPCP lcp-echo-interval / -failure the liveness check — its own section below mtu / mru 1400 leave room for whatever the carrier wraps around each frame pty '...' the command whose stdin and stdout become the link — here, a netcat socket The connecting end is the same line without `passive`: it initiates instead of waiting. Every option on the line, and what it does. The one to watch is noauth: it drops peer authentication, so the listener accepts whoever arrives. nodefaultroute and noipdefault keep the link from quietly rewriting your routing, local makes it come up on a pty at all, and the pinned local:remote pair stops pppd taking a different answer during IPCP. The connecting end is the same line without passive. Bring both ends up and you get a ppp0 on each, a point-to-point route to the far address, and an interface you can ping, route through and tcpdump. That is a VPN — no package installed, no daemon configured, no key exchanged, which is the problem, and I will come back to it.\nWhat It Negotiated, And How To Watch It Add debug and pppd logs the control protocol exchange, which is worth reading once even if you never read it again. The link comes up in stages, and each stage can fail on its own.\nLCP, then authentication, then one control protocol per address family Three stages, and each one fails for its own reason 1 LCP — the link itself Maximum receive unit, the control character map, a magic number to spot a looped-back line, and whether either end wants the other to authenticate. Fails here and the carrier is not passing bytes cleanly in both directions. Check the socket, not the PPP options. 2 Authentication — optional PAP sends a password in the clear. CHAP does a challenge and an MD5 response. Skipped here. Both prove who the peer is. Neither encrypts one byte of what follows. 3a IPCP The IPv4 address at each end. 3b IPV6CP The 64-bit interface identifiers. These two are independent. IPv6 can come up while IPv4 is still arguing, and a failure in one does not take the other down. The interface starts carrying traffic for a family the moment that family's control protocol finishes. LCP settles the link itself — how big a frame can be, which control characters need escaping, and a magic number that detects a looped-back line. Authentication is optional and skipped here. Then one control protocol per network layer: IPCP for IPv4, IPV6CP for IPv6. They are independent, so a link can carry IPv6 while IPv4 is still arguing, and a failure in one does not take the other down. The two useful debugging tools are already installed and nobody uses them:\n# log every frame in both directions to a file pppd ... debug record /tmp/ppp-trace # then read it back in a human-readable form pppdump -h /tmp/ppp-trace | less record writes a timestamped capture of every byte, pppdump turns it into something readable4, and with tcpdump -ni ppp0 you see both sides — the framing underneath and the packets on top.\nThe Frame On The Wire, And Why 0x7E Is Everywhere PPP over a serial line — and a pty is one as far as pppd is concerned — uses async HDLC framing, and understanding it is the difference between tuning this and guessing. Every frame begins and ends with the same byte: \u0026ldquo;Each frame begins and ends with a Flag Sequence, which is the binary sequence 01111110 (hexadecimal 0x7e)\u0026rdquo;5. If 0x7e marks a boundary it cannot appear inside a frame, so it is escaped. So is the escape byte 0x7d, and anything else either end asked for.\nThe frame, the flag byte, and what escaping costs One frame, and the two bytes that can never appear inside it 0x7E flag 0xFF address 0x03 control protocol 1 or 2 bytes payload — your IP packet up to the agreed maximum receive unit FCS 16-bit CRC 0x7E flag ESCAPING, THE ONLY RULE THERE IS Any byte that could be mistaken for a boundary is replaced by 0x7D followed by the original byte exclusive-ored with 0x20. payload 0x7E transmitted as 0x7D 0x5E payload 0x7D transmitted as 0x7D 0x5D Those two are mandatory. Everything else is decided by the async control character map — 32 bits, one per control character. Map set to zero — pppd's default, and correct on a socket Two bytes escaped out of 256. Overhead you will never measure. A socket is a clean 8-bit path and needs nothing else. All 32 flagged — the conservative modem setting On compressed or encrypted payloads about one byte in eight is under 0x20, so it doubles. Roughly 12% of the link, for nothing. The frame is flag, address, control, protocol, payload, frame check sequence, flag. Any byte inside it that could be mistaken for a boundary is replaced by 0x7D followed by the original byte XOR 0x20. The ACCM decides how many other bytes get the same treatment: zero of them on a clean 8-bit path, thirty-two of them if either end asks for the conservative default that assumes a modem eating control characters. The escaping rule is exact: \u0026ldquo;Each Flag Sequence, Control Escape octet, and any octet which is flagged in the sending Async-Control-Character-Map (ACCM), is replaced by a two octet sequence consisting of the Control Escape octet followed by the original octet exclusive-or\u0026rsquo;d with hexadecimal 0x20\u0026rdquo;5.\nThat last part is the tuning knob. The ACCM is 32 bits, one per control character, and a 1 means \u0026ldquo;escape this\u0026rdquo;. Ask for all of them and, on encrypted payloads where the bytes are effectively random, roughly one byte in eight is under 0x20 — about 12% overhead for nothing.\nModern pppd already does the right thing here. Read the man page rather than the folklore:\nIf no asyncmap option is given, the default is zero, so pppd will ask the peer not to escape any control characters.4\nSo asyncmap 0 is the default, not a trick, and a socket is a clean 8-bit path needing no escaping beyond the two mandatory bytes. Leave it alone; set bits only when something in the middle genuinely eats control characters — a terminal server, a serial concentrator, a bad console proxy. The frame check sequence at the end is a 16-bit CRC, and it is the thing that makes the UDP build work — which is next.\nTCP Over TCP Is The Wrong Carrier The build above works. On a lab bench, over loopback or a quiet LAN, it works beautifully, which is exactly why people ship it.\nThen it meets a real path with real loss. It falls over in a way that looks like everything except what it is.\nThe problem is two independent retransmission timers stacked on each other, the outer one hiding the loss from the inner. Olaf Titz wrote the definitive explanation, opening on precisely this build:\nA frequently occurring idea for IP tunneling applications is to run a protocol like PPP, which encapsulates IP packets in a format suited for a stream transport (like a modem line), over a TCP-based connection. […] Unfortunately, it doesn\u0026rsquo;t work well. Long delays and frequent connection aborts are to be expected.6\nOne lost packet, two carriers, two very different outcomes One packet goes missing. What each carrier does next. TCP CARRIER — the tunnel hides the loss 1 The carrier loses a segment. 2 The carrier retransmits it. The bytes arrive late, not missing — which is the whole problem. 3 The tunnelled TCP cannot see the carrier. It reads the delay as congestion, doubles its timer and retransmits data already in flight below it. 4 Now the carrier owes the original and a duplicate, over a link that just proved it is losing packets. The queue grows faster than either layer drains it. Throughput collapses well before the link does. UDP CARRIER — the loss reaches the layer that owns it 1 The carrier drops the packet and says nothing. 2 The PPP frame inside fails its checksum and is discarded. The next flag byte resynchronises. 3 The tunnelled TCP sees a real loss, because for once it has been told the truth about the path. 4 It halves its window and retransmits once. Congestion control does exactly the job it was designed for. One lost packet costs one packet. The carrier TCP guarantees delivery, so a lost segment is retransmitted and the bytes arrive late rather than not at all. The tunnelled TCP inside sees only the delay, decides the network is congested, backs off and retransmits the same data — which the carrier now has to deliver as well. Each layer\u0026rsquo;s timer is trying to fix a problem the other layer already owns, and the queue grows faster than either can drain it. A UDP carrier drops the packet, the inner TCP sees a real loss, and its congestion control does the job it was designed for. Here is the shape. Both TCPs set a retransmission timer from their round-trip time. When the carrier loses a segment it retransmits, so the inner connection\u0026rsquo;s data arrives late. And the inner TCP, which cannot see the carrier, reads late as congestion and retransmits too. Now the carrier has the original and a duplicate to deliver over a link already losing packets. The inner timer doubles, the outer queue grows, and throughput collapses long before the link does. Nothing in the logs says why.\nThis is not a subtle efficiency point. It is the difference between a tunnel that degrades gracefully and one that stops passing traffic at maybe 2% loss while ping across the same path still looks fine — the reason to use UDP, and to keep TCP in reserve only for a path that will pass nothing else.\nSo use UDP — as the default, not a preference.\nStage Two: The UDP Build, Which Is The Foundation Same pppd, different carrier. The only change is in the pty command.\nListening end:\npppd nodetach noauth local passive \\ nodefaultroute noipdefault \\ 192.0.2.1:192.0.2.2 \\ lcp-echo-interval 10 lcp-echo-failure 3 \\ mtu 1400 mru 1400 \\ pty \u0026#39;ncat --udp --listen 198.51.100.10 443\u0026#39; Connecting end:\npppd nodetach noauth local \\ nodefaultroute noipdefault \\ 192.0.2.2:192.0.2.1 \\ lcp-echo-interval 10 lcp-echo-failure 3 \\ mtu 1400 mru 1400 \\ pty \u0026#39;ncat --udp 198.51.100.10 443\u0026#39; Two things about a UDP listener that will catch you out, and neither is a PPP problem.\nThe listener cannot answer until it is spoken to. A UDP socket has no connection to accept, so netcat cannot know where to send replies until a datagram arrives. It locks onto the first source address and port it hears from and talks to that. This means the connecting end must send first — which pppd does by itself, because the non-passive end starts firing LCP configure-requests immediately. It also means that if the client\u0026rsquo;s source port changes, the carrier is silently talking to the wrong place. Behind NAT with a short UDP timeout, that is a tunnel that dies every few minutes for no visible reason.\n--keep-open does not do what you want on UDP. Ncat\u0026rsquo;s -k keeps a TCP listener accepting after a peer goes away. On UDP there is nothing to accept, so recovery means restarting the carrier — which is what persist and holdoff are for, further down.\nBoth are arguments for socat, which handles UDP peers more carefully, and for a supervisor rather than trusting it to stay up.\nWhy PPP Survives A Lost Datagram The obvious objection to a UDP carrier is that PPP expects a byte stream and UDP is not one: datagrams arrive whole or not at all, and can arrive out of order. It works anyway, because of two bytes the framing has carried all along.\nA lost datagram costs a frame, and the next flag byte recovers the link The boundaries do not line up, and it does not matter PPP FRAMES AS pppd WROTE THEM 7E frame A + FCS 7E frame B + FCS 7E frame C + FCS 7E frame D + FCS WHAT THE CARRIER ACTUALLY SENT — netcat reads a buffer, not a frame datagram 1 datagram 2 — lost datagram 3 datagram 4 Frame B loses its middle, so its checksum fails and it is discarded. Frame C loses its opening bytes and goes the same way. Two lost frames is two lost IP packets. The layer above retransmits them, and it retransmits them knowing the path lost something — which is precisely the signal a TCP carrier would have hidden by delivering them late instead. Netcat reads whatever is in the buffer and writes it into a datagram, so frame boundaries and datagram boundaries have nothing to do with each other. Lose a datagram and the receiver sees a frame whose bytes are missing: the frame check sequence fails and the frame is discarded, exactly as it would be on a noisy serial line. The next 0x7E resynchronises the stream. One dropped frame costs one packet, and the layer above retransmits it — which is the loss signal the inner TCP needed and never got over a TCP carrier. PPP over async HDLC was designed for a line that corrupts bytes. Every frame carries a 16-bit frame check sequence; a frame that fails it is discarded, and the next flag byte resynchronises the receiver. Reordering is rarer than loss and produces the same outcome: a bad FCS, a discarded frame, a resync.\nSo a lost datagram costs one PPP frame — one IP packet, the normal condition of every network ever built. The packet\u0026rsquo;s owner retransmits, and congestion control sees a real loss and responds correctly. That is the whole argument for UDP: it lets the traffic inside the tunnel find out the truth about the path.\nDo not let frames get big enough to need fragmenting by the carrier, though. One lost fragment then kills the whole datagram and your effective loss rate multiplies. Keep the PPP MTU well under the path MTU.\nAddresses, Routes, And Doing IPv6 Properly ppp0 is a point-to-point interface — no subnet, no ARP — so local:remote is the whole of the addressing and you route across it explicitly. For a single host reaching one network behind the far end:\n# on the client, after the link is up ip route add 203.0.113.0/24 via 192.0.2.1 dev ppp0 For the far end to forward on behalf of the client, the usual two steps:\nsysctl -w net.ipv4.ip_forward=1 nft add rule inet nat postrouting oifname \u0026#34;eth0\u0026#34; ip saddr 192.0.2.2/32 masquerade proxyarp will make the server answer ARP for the client on its own segment4, which is neat until a second client wants the same. Then do the part most people skip. Run IPv6 over it — a point-to-point link with no NAT, no broadcast domain and no address scarcity is the easiest place in your network to do IPv6 properly:\npppd nodetach noauth local \\ +ipv6 ipv6 ::1,::2 \\ nodefaultroute noipdefault \\ 192.0.2.2:192.0.2.1 \\ pty \u0026#39;ncat --udp 198.51.100.10 443\u0026#39; +ipv6 enables IPV6CP; ipv6 \u0026lt;local\u0026gt;,\u0026lt;remote\u0026gt; sets the two 64-bit interface identifiers4, the link comes up link-local, and you add a global /64 on top7. IPV6CP is independent of IPCP, so noip carries nothing but IPv6 — a reasonable thing to build in 2026, in one word.\nKeeping It Up When The Carrier Dies Quietly This is the failure that wastes an afternoon, so it gets its own section.\nA TCP carrier that dies badly — the far end powered off, a NAT table entry expired, a middlebox that stopped forwarding — does not close. There is no FIN, no RST, nothing. Netcat sits there holding a socket that will never deliver another byte, pppd sits there holding a pty that will never see another frame, and ip link cheerfully reports ppp0 as UP. A UDP carrier has no connection state at all, so it never notices anything.\nPPP has the answer built in, and it is off by default:\nlcp-echo-interval 10 lcp-echo-failure 3 That sends an LCP echo every ten seconds and tears the link down after three go unanswered4 — thirty seconds to catch a dead carrier, on exactly the case the man page names, \u0026ldquo;no hardware modem control lines\u0026rdquo;4, which is every pty ever made. Then decide what happens after:\npersist maxfail 0 holdoff 5 Detecting a dead carrier in thirty seconds, and rebuilding it A carrier can die without closing. This is how PPP notices, and recovers. Carrier dies silently far end off · NAT entry gone · middlebox stops forwarding Nothing reports it no FIN, no RST · ip link still says ppp0 UP · UDP has no state The detector lcp-echo-interval 10 lcp-echo-failure 3 echo every 10s, down after 3 missed — 30s The rebuild persist maxfail 0 holdoff 5 restart, never give up, wait 5s between tries One supervisor systemd unit, Restart=always, pppd reruns the pty command — fresh netcat and socket with it Set the echo on both ends. An echo only proves the path the reply came back on. Put `ip route` in /etc/ppp/ip-up.d/ and the IPv6 routes in /etc/ppp/ipv6-up.d/, so routing is reapplied every time the link returns — not set once by hand and lost on the first reconnect. lcp-echo-interval 10 lcp-echo-failure 3 is the detector: echo every ten seconds, link down after three missed. persist maxfail 0 holdoff 5 is the rebuild: restart, never give up, wait five seconds so a flapping path is not a fork bomb — and pppd reruns the pty command, so a fresh netcat and socket come with it. Set the echo on both ends; one supervisor, a systemd unit with Restart=always, beats two processes arguing. Put ip route in /etc/ppp/ip-up.d/ so routing returns with the link. It Has No Encryption And No Authentication Everything above is a working tunnel. It is not a secure one, and the gap is not a detail.\nThere is no encryption. Not weak encryption. None. Every packet you put through this is on the wire in clear, wrapped in an HDLC frame that any capture tool decodes on sight. tcpdump will show you the contents of somebody else\u0026rsquo;s tunnel as readily as your own.\nThere is no authentication of the peer. noauth says so. A netcat listener accepts whatever arrives on the port: the first connection or datagram from anywhere. Whoever gets there first gets a routed link into your network. PPP does have authentication (PAP in the clear, CHAP as a challenge-response8), and both prove who the peer is while encrypting not one byte of what follows: CHAP here gets you a tunnel that knows who it is talking to and still publishes the contents to anyone on the path.\nThere is an encryption option in the PPP family — MPPE9, the module sitting in ppp_mppe.ko. Do not reach for it: it is RC4 keyed off the MS-CHAPv2 exchange, broken in public for over a decade, and the reason PPTP is dead. Building a new tunnel on it in 2026 chooses a known-broken cipher over a working one that costs nothing.\nSo the honest summary so far: a routed link with no confidentiality and no access control. Fine for a lab, fine inside a link that is already encrypted, not fine anywhere else. The fix is to encrypt the carrier — the next two sections, and the reason to use ncat rather than whichever netcat your distribution shipped.\nStage Three: Wrapping The Carrier In TLS The clean answer leaves pppd exactly as it is and replaces the carrier with one that does TLS. pppd never learns anything changed.\nFirst the certificate. A self-signed one is enough as long as the client verifies it. An unverified TLS session is an encrypted conversation with somebody you have not identified, which stops nobody:\nopenssl req -x509 -newkey rsa:4096 -days 825 -nodes \\ -keyout tunnel.key -out tunnel.crt \\ -subj \u0026#34;/CN=tunnel.example.net\u0026#34; \\ -addext \u0026#34;subjectAltName=DNS:tunnel.example.net\u0026#34; Ncat Ncat has TLS built in. It also chains through a proxy, --proxy host:port --proxy-type http|socks4|socks510, so the carrier can terminate at a relay rather than the tunnel\u0026rsquo;s far end. That is the point the egress section turns on. Listening end:\npppd nodetach noauth local passive \\ nodefaultroute noipdefault 192.0.2.1:192.0.2.2 \\ lcp-echo-interval 10 lcp-echo-failure 3 mtu 1400 mru 1400 \\ pty \u0026#39;ncat --listen --keep-open --ssl --ssl-cert tunnel.crt --ssl-key tunnel.key 443\u0026#39; Connecting end:\npppd nodetach noauth local \\ nodefaultroute noipdefault 192.0.2.2:192.0.2.1 \\ lcp-echo-interval 10 lcp-echo-failure 3 mtu 1400 mru 1400 \\ pty \u0026#39;ncat --ssl --ssl-verify --ssl-trustfile tunnel.crt tunnel.example.net 443\u0026#39; --ssl-verify is the word that matters. Without it, --ssl gets you encryption against a passive listener and nothing against whoever answers the port first; with it, Ncat verifies trust and domain name against the trust file11. Ncat has no server-side client-cert check, though, so the server cannot identify the client. Pair it with CHAP, or use one of the next two.\nStunnel The traditional wrapper, and the one that does mutual authentication properly:\n[ppp] accept = 443 connect = 127.0.0.1:6001 cert = /etc/stunnel/tunnel.crt key = /etc/stunnel/tunnel.key CAfile = /etc/stunnel/clients.crt verify = 2 verify = 2 requires a client certificate signed by a CA in CAfile — the access control the netcat build never had. The client runs stunnel in client mode, and pppd\u0026rsquo;s pty command connects to the local plaintext side.\nSocat socat does the whole thing in one process per end, with verification on by default:\n# listening end pppd ... pty \u0026#39;socat - OPENSSL-LISTEN:443,reuseaddr,cert=tunnel.pem,cafile=clients.crt,verify=1\u0026#39; # connecting end pppd ... pty \u0026#39;socat - OPENSSL:tunnel.example.net:443,cafile=tunnel.crt,verify=1\u0026#39; socat is not installed on the machine this was written on, so that is from its documentation — check your own flags. It is the one I would reach for, because it is the only one of the four that also speaks DTLS.\nOpenssl, when there is nothing else openssl is on every box that has TLS at all, and s_server/s_client will carry a pipe:\n# listening end pppd ... pty \u0026#39;openssl s_server -quiet -accept 443 -cert tunnel.crt -key tunnel.key\u0026#39; # connecting end pppd ... pty \u0026#39;openssl s_client -quiet -verify_return_error -CAfile tunnel.crt -connect tunnel.example.net:443\u0026#39; -quiet suppresses the banner that would otherwise land in your PPP stream, and -verify_return_error makes a verification failure close the connection instead of warning and carrying on. These are debugging tools behaving like it, but on a box where you cannot install anything they get a link up.\nPlain, TLS over TCP, or DTLS over UDP Same link half at both ends. Only the middle changes. pppd or tap0 plain netcat, UDP ncat --udp --listen 6000 pppd or tap0 In clear on the wire, and the listener takes whoever reaches the port first. No confidentiality, no access control. pppd or tap0 TLS over TCP — ncat, stunnel or openssl ncat --ssl --ssl-verify --ssl-trustfile tunnel.crt host 6000 pppd or tap0 Protected, and the certificate says who the far end is. But the carrier is TCP again — and it is the only shape that can stream-compress. pppd or tap0 DTLS over UDP — socat or openssl socat - OPENSSL-DTLS-CLIENT:host:6000,cafile=tunnel.crt,verify=1 pppd or tap0 Encrypted, authenticated, and still datagrams. A lost packet stays a lost packet instead of two stacks arguing about it. Aim here. The same pppd at both ends throughout — only the pty command changes. Plain netcat gives you a link with nothing protecting it. TLS over TCP protects the bytes and reintroduces the carrier-TCP problem. DTLS over UDP is the shape to aim for: encrypted, authenticated, and still a datagram carrier, so a lost packet stays a lost packet instead of becoming a retransmission fight between two stacks. Stage Four: DTLS, Because The Carrier Should Still Be UDP Here is the awkward part, and the reason this section is separate. Every TLS option above runs over TCP, so wrapping the carrier in TLS undoes the argument for UDP and hands you the meltdown back. Encryption and the right transport should not be a trade. DTLS is TLS over datagrams, and it is what you want: it keeps the record-layer protection, drops the ordering and retransmission guarantees, and leaves lost packets lost, which is exactly what PPP\u0026rsquo;s frame check sequence is built to absorb.\nNcat cannot do it. socat and openssl can:\n# listening end pppd ... pty \u0026#39;socat - OPENSSL-DTLS-LISTEN:443,cert=tunnel.pem,cafile=clients.crt,verify=1\u0026#39; # connecting end pppd ... pty \u0026#39;socat - OPENSSL-DTLS-CLIENT:tunnel.example.net:443,cafile=tunnel.crt,verify=1\u0026#39; And with openssl alone, where -dtls selects any DTLS version:\n# listening end pppd ... pty \u0026#39;openssl s_server -quiet -dtls -accept 443 -cert tunnel.crt -key tunnel.key\u0026#39; # connecting end pppd ... pty \u0026#39;openssl s_client -quiet -dtls -verify_return_error -CAfile tunnel.crt -connect tunnel.example.net:443\u0026#39; Watch the frame sizing. A DTLS record cannot be fragmented the way a TLS record spreads across a TCP stream, so everything has to fit the path MTU at once.\nThe overhead stack, and the inner MTU left over Work down from the path MTU, and use what is left IP header20 (v4) / 40 (v6) UDP header8 DTLS record + tag≈ 30 PPP / Ethernet framinga few inner MTU — what the tunnel can carry: aim 1400, floor 1280 A DTLS record cannot be fragmented the way a TLS record spreads across a TCP stream, so all of this has to fit the path MTU at once. Same shape as OpenVPN since 2001 — a datagram carrier, DTLS, a virtual interface on top. Not eccentric; just unbundled. Work down from 1500: 20 bytes of IPv4 header or 40 of IPv6, 8 of UDP, roughly 30 for the DTLS record and its tag, a few for framing, and the rest is the inner MTU the tunnel can carry. Aim 1400 on an ordinary path; 1280 is the safe floor if anything in the middle is itself a tunnel. This shape — a datagram carrier, DTLS, a virtual interface on top — is near enough what OpenVPN has done since 2001. Not eccentric; just unbundled. Tuning The PPP Build: MTU, Compression, And What Actually Helps Four knobs, all pppd options — the tap build in the next section has none of them, because it has none of the machinery they configure.\nFour pppd knobs: one to set, one to leave, two to turn off One to set, one to leave, two to turn off SET IT MTU and MRU Both ends, with headroom: 1400 plain, 1280 under a tunnel. A fragmented frame loses a whole packet. LEAVE IT ACCM Already zero by default, which is right on a clean socket. Touch it only if something eats control characters. TURN OFF novj Van Jacobson header compression Saves a rounding error on a fast link, costs CPU per packet, and breaks under loss. Off above modem speed. TURN OFF nodeflate nobsdcomp The payload is already TLS/DTLS — compressing ciphertext is pure work, and across a security boundary, an attack. None of these makes the tunnel faster. PPP is not the bottleneck — PPPoE does 2 Gbit on the same daemon, because its data path stays in the kernel. The cost here is the pty: every byte crosses into userspace and back. That belongs to the pseudo-terminal, not PPP. MTU and MRU is the one that matters: set both on both ends with headroom, because a frame the carrier has to fragment loses a whole packet on any drop. The ACCM is already zero and correct on a clean socket. Turn Van Jacobson header compression off with novj above modem speed. It saves a rounding error, costs CPU per packet and breaks under loss. Turn deflate/bsdcomp off: the payload is already encrypted, and compressing across a security boundary is an attack, not a feature. None of it makes the tunnel faster, and the reason matters because PPP gets blamed for it and should not. PPP is not the bottleneck. PPPoE carries 2 Gbit on the same daemon, because its data path never leaves the kernel — ppp_generic and pppoe do the framing and forwarding, and pppd only handles the control plane. What costs you here is the pty: every byte crosses into userspace, through netcat, into a socket and back, a round trip PPPoE never makes. That belongs to the pseudo-terminal, not to PPP.\nWhich is a good moment to look at the other way of doing this, where there is no pty at all.\nThe Other Way: A Tap Device, And No PPP At All Everything so far has used PPP for the link half. There is a second way to make a virtual interface, and it needs no protocol at all.\nThe kernel\u0026rsquo;s TUN/TAP driver hands you an interface and a file descriptor tied together: write a packet into the descriptor and it appears on the interface as if off a wire; read and you get a packet the kernel wanted to send. That is the entire interface12, and it has been in Linux since 1999.\nip tuntap add dev tap0 mode tap user damien group damien ip link set tap0 mtu 1400 up ip addr add 192.0.2.1/30 dev tap0 A tap device is an interface at one end and a file descriptor at the other No daemon, no protocol, no pseudo-terminal. One file descriptor. KERNEL tap0 an ordinary interface, with an address and a route /dev/net/tun TUNSETIFF names the device you get your process socat, or fifty lines of Python holding the fd UDP socket one frame per datagram the network and the far end one read = one frame What is missing against the PPP build: no negotiated frame size, no liveness check, no addresses exchanged, no authentication. You configure both ends by hand and they never discuss any of it. Nothing is watching the link, so nothing will tell you it died. No daemon and no protocol. The kernel presents tap0 as an ordinary interface and hands the other side of it to whichever process opened /dev/net/tun. One read returns exactly one Ethernet frame; one write injects exactly one. Everything pppd was negotiating — addresses, frame sizes, liveness — you now configure by hand at both ends, and the two ends never discuss any of it. Two details in that first command are worth more than they look.\nuser damien makes the device persistent and unprivileged. Created this way it survives the process that uses it, and a named user can open it without being root. Root creates the device once; the thing that shovels frames does not need root at all. Hold that thought for the egress section.\nmode tun is the other half of the same driver, and it is the one most people actually want. Tun carries IP packets. Tap carries Ethernet frames. The difference matters enough to get its own section below.\nWhat you give up against PPP is everything PPP negotiates: no LCP so no agreed frame size and no liveness check, no IPCP/IPV6CP so both ends are configured by hand, no authentication, no header compression. A tap device is a hole in the kernel, and the protocol through it is whatever you put there. What you gain is no framing overhead at all, no escaping, no control channel, and an interface that can be bridged.\nTap Over UDP, Where One Datagram Is One Frame This is the cleanest mapping in the whole post, and it falls out of the design. A tap device is a datagram device — one read() returns exactly one frame — and a UDP socket is a datagram socket, one sendto() per datagram. So a frame goes into a datagram, arrives as a frame, and there is nothing to delimit, buffer or resynchronise. Lose a datagram and you have lost one frame, which is what a dropped packet looks like anyway.\nWith socat, both ends are one command each:\n# listening end socat TUN:192.0.2.1/30,tun-type=tap,tun-name=tap0,iff-up UDP-LISTEN:443 # connecting end socat TUN:192.0.2.2/30,tun-type=tap,tun-name=tap0,iff-up UDP:198.51.100.10:443 socat is not installed on the machine this was written on, so those two are from its documentation rather than a run here — check your own build\u0026rsquo;s address names before trusting them. It is worth having: it is the only tool in this post that does tun, tap, TLS and DTLS in one process.\nWhere you cannot install anything, the whole job is about fifty lines with no dependencies beyond the standard library. The heart of it, IFF_NO_PI turning off the four-byte header the driver would otherwise prepend:\nTUNSETIFF, IFF_TAP, IFF_NO_PI = 0x400454CA, 0x0002, 0x1000 # IFF_TUN is 0x0001 def open_tap(name): fd = os.open(\u0026#34;/dev/net/tun\u0026#34;, os.O_RDWR) fcntl.ioctl(fd, TUNSETIFF, struct.pack(\u0026#34;16sH\u0026#34;, name.encode(), IFF_TAP | IFF_NO_PI)) return fd # ... learn the peer from the first datagram (UDP has no accept), then shovel: while True: ready, _, _ = select.select([tap, sock], [], []) if tap in ready and peer: sock.sendto(os.read(tap, MTU), peer) # one read = one frame = one datagram if sock in ready: frame, src = sock.recvfrom(MTU) peer = src # last speaker wins — see the note below if frame: os.write(tap, frame) The full file — argument parsing, IPv6 via getaddrinfo, the zero-length datagram that tells a UDP listener where to answer, and the compression the next section adds — is in the bundle:\n\u0026#8615; tapcat.py — the whole thing, about a hundred lines tapcat.py · 6 kB peer = src on every datagram is the part to read twice. Whoever last sent a frame becomes the peer. Convenient behind NAT whose source port keeps moving, and an open door on an untrusted network, where anyone who can send one datagram to the port takes over the tunnel. Fine inside a DTLS session, which is where this is going; not fine on its own. The socket half of the script was exercised here over IPv6 loopback: the opener arrives, the listener learns the peer, a frame crosses and the reply comes back; the tap half needs root, the one part I could not run.\nOver TCP You Have To Invent The Framing Yourself Now swap the carrier for TCP and watch a whole problem appear that PPP quietly solved in 1994.\nTCP is a byte stream with no record boundaries and no promise about how bytes are grouped on arrival: two frames written back to back can arrive in one read, one frame in three. The receiver holds a pile of bytes with no idea where one frame ends, and a tap device accepts only whole frames.\nOne datagram per frame, or a length prefix you have to invent One read off a tap device is one frame. Keeping it that way is the carrier's job. OVER UDP — the boundary is free datagram = frame A datagram = frame B datagram = frame C datagram = frame D One read, one datagram, one write at the far end. Nothing to delimit, nothing to buffer, nothing to resynchronise after a loss. OVER TCP — the boundaries are gone and you have to put them back one byte stream — two frames can arrive in one read, one frame can arrive in three len frame A len frame B len frame C len frame D Two bytes of big-endian length in front of every frame, and a receiver that reads the length and then exactly that many bytes. It works, and it cannot recover. HDLC resynchronises on the next 0x7E because a flag is unambiguous; a length-prefixed stream that loses its place reads every length after it out of the middle of a frame. Add a marker and a checksum to fix that and you have rebuilt HDLC. Over UDP, one frame is one datagram and the boundary comes for free. Over TCP the boundaries are gone, so the sender has to add a length prefix to every frame and the receiver has to reassemble from it. PPP does not have this problem because it brings its own framing — a flag byte at each end and a checksum — which is also what lets it resynchronise after damage. A length-prefixed stream cannot: get one byte out of step and every frame after it is wrong. So over TCP you write your own framing. Two bytes of big-endian length in front of every frame is the usual answer:\n# sending sock.sendall(struct.pack(\u0026#34;!H\u0026#34;, len(frame)) + frame) # receiving def recv_exactly(sock, n): buf = b\u0026#34;\u0026#34; while len(buf) \u0026lt; n: chunk = sock.recv(n - len(buf)) if not chunk: raise ConnectionError(\u0026#34;carrier closed\u0026#34;) buf += chunk return buf length = struct.unpack(\u0026#34;!H\u0026#34;, recv_exactly(sock, 2))[0] frame = recv_exactly(sock, length) That works, and it is strictly worse than the UDP version. It is the TCP-over-TCP problem again; it adds two bytes and a reassembly loop per frame; and it has no way back from an error, because a length-prefixed stream that loses its place reads every subsequent length out of the middle of a frame. HDLC resynchronises on the next 0x7E; this cannot, short of reinventing HDLC worse. Third argument for UDP, and the strongest: over a datagram carrier there is no framing problem, because the carrier already has the only feature you needed.\nTun Or Tap: Layer 3 Unless You Genuinely Need Layer 2 The same driver gives you two devices, and people pick the wrong one constantly because a tutorial said tap.\ntun tap What crosses IP packets Ethernet frames Per-packet overhead none 14-byte Ethernet header ARP, DHCP, broadcast no yes, all of it, over the tunnel Non-IP protocols no yes Can join a bridge no yes Equivalent to a point-to-point link, like ppp0 a network cable Tun is a routed link — like the PPP interface from the first half: two addresses, a route, packets in and out. Tap is a virtual Ethernet cable, so every broadcast, ARP query and bit of multicast noise on the segment now crosses your tunnel and burns bandwidth.\nChange one line in the script to switch:\nIFF_TUN = 0x0001 # instead of IFF_TAP Use tap when you genuinely need layer 2. There are real reasons: a protocol that is not IP, a cluster heartbeat that expects to see broadcasts, a DHCP server that must reach clients across the tunnel, or bridging two segments into one.\nThat last one is the common case and the one to be careful with:\nip link add br0 type bridge ip link set tap0 master br0 ip link set eth1 master br0 ip link set br0 up Now the remote segment is part of your local one — its broadcasts, its spanning tree, its MAC churn and, if someone was careless, its DHCP server. Bridging two sites that both run 192.168.1.0/24 is a bad afternoon; one that loops back to the same segment is a bad week. Default to tun, reach for tap when you can name the layer 2 thing you need, and bridge only after checking what is broadcasting on both sides.\nThe Same Wrappers, One Process Per End The tap build has the same hole as the PPP one: the carrier is in clear and the listener accepts whoever gets there first. The fix is the same, and with socat it collapses into one command per end, because it makes the device and terminates the DTLS session in a single process:\n# listening end socat TUN:192.0.2.1/30,tun-type=tap,tun-name=tap0,iff-up \\ OPENSSL-DTLS-LISTEN:443,cert=tunnel.pem,cafile=clients.crt,verify=1 # connecting end socat TUN:192.0.2.2/30,tun-type=tap,tun-name=tap0,iff-up \\ OPENSSL-DTLS-CLIENT:tunnel.example.net:443,cafile=tunnel.crt,verify=1 That is the shortest correct build in this post. A virtual interface, a datagram carrier, mutual certificate authentication and encryption. Two commands. No daemon, no protocol negotiation.\nverify=1 is not optional — the same point as --ssl-verify, and with the peer = src behaviour above, an unidentified party is one datagram from owning your tunnel. If you are stuck with the Python script, do not bolt TLS into it: point it at a loopback port and put the wrapper in front, or use socat. A hand-rolled TLS wrapper around a hand-rolled tunnel is two chances to get the interesting part wrong.\nStage Five: Compressing The Stream With zstd There is one more thing worth putting in the backend, and unlike the compression options on pppd it can genuinely pay: compress the frames with zstd before they go into the carrier.\nThe important word there is frames, plural. Compress the stream, not each packet on its own. It is the single biggest measured effect in this post.\nNetwork traffic is repetitive in a way that only shows up across packets — the same headers, hostnames and JSON keys, over and over. A compressor that starts fresh on every 1,400-byte frame sees none of it; one that keeps its window across frames sees all of it.\nCompress before encrypt, and keep the window if the carrier lets you The order is fixed. The mode is the decision. tap0one read, one frame compresszstd, level 1 encryptDTLS, or TLS carrierUDP, or TCP Never the other way round. FLUSH_BLOCK — keep the window Emits everything so far, keeps the history. Each frame is encoded against every frame before it. No added latency: one frame in, one frame out. 5.0% on repetitive traffic · 7.5% on log lines Needs every frame delivered, in order. So: TCP, or TLS over TCP. Not UDP, not DTLS. FLUSH_FRAME — throw it away each packet Each datagram is a complete zstd frame and decodes on its own, so datagram N still works after 1 to N-1 were lost. The only mode a datagram carrier can use. 14.8% on the same traffic · 24.2% on log lines A trained dictionary wins most of that back: 5.3% and 15.1%, and it stays loss-tolerant. Compression goes between the tap device and the carrier, and before the encryption, because ciphertext does not compress. FLUSH_BLOCK is the mode that matters: it emits everything so far, so one frame in gives one frame out with no added latency, while keeping the compression history for the next frame. FLUSH_FRAME throws that history away every packet, which is what makes it safe over a lossy carrier and what makes it three times worse on the wire. On a current Fedora this needs nothing installed — Python 3.14 brought zstd into the standard library13; this box has 3.14.7 against zstd 1.5.7:\nfrom compression.zstd import ZstdCompressor, ZstdDecompressor # Python 3.14+ _c, _d = ZstdCompressor(level=1), ZstdDecompressor() # FLUSH_BLOCK: emit everything so far, keep the window for the next frame. out = _c.compress(frame, mode=ZstdCompressor.FLUSH_BLOCK) frame = _d.decompress(out) One compressor per direction, alive for the life of the link. One FLUSH_BLOCK per frame, so a frame goes out the moment it arrives with nothing buffered, and the far end hands back one frame per block.\nWhat the window is worth, measured Numbers from this machine. Same 1,400-byte frames, same level 1, the only difference being whether the compressor keeps its history:\nTraffic in the tunnel stream, FLUSH_BLOCK per packet, FLUSH_FRAME per packet with a trained dictionary Repetitive API calls and telemetry 5.0% 14.8% 5.3% Log lines 7.5% 24.2% 15.1% Plain text and config files 37.8% 51.3% 45.3% Random bytes, standing in for TLS 100%+ 100.7% — Throughput at level 1 ran 205 MB/s on text and over 1 GB/s on the repetitive traffic — the more compressible, the faster, because there is less to encode.\nLog traffic goes to 7.5% of its original size, against 24.2% per packet — three-fold, same data, same level, from one flag. And level 1 is the level: level 3 bought about one percent, level 9 another while dropping throughput from 211 MB/s to 61.\nThe catch, and it is the same catch as everything else here A shared window means every frame depends on the ones before it: lose one and the decompressor\u0026rsquo;s history no longer matches, and nothing after it decodes. So stream compression needs a carrier that delivers everything in order — TCP, or TLS over TCP, not UDP or DTLS. That is the one honest argument for the TCP carrier in this whole post. If what crosses your tunnel is genuinely compressible — syslog, database replication in the clear, telemetry, a chatty API — a stream-compressed TLS session moves a third of the bytes a datagram carrier would, which on a decent path can beat the retransmission penalty. Measure it on your own traffic.\nOver UDP, Where A Dictionary Does The Window\u0026rsquo;s Job Where the carrier is UDP or DTLS, and it should be by default, you cannot keep a window: each datagram stands on its own, which means FLUSH_FRAME and the weaker column above. A trained dictionary gives the compressor the cross-packet context a window would have, without any dependency between datagrams:\n# capture a few thousand real frames off the link first, one per file zstd --train frames/* -o tunnel.dict --maxdict=110000 ```[^zstddict] ```python from compression.zstd import ZstdCompressor, ZstdDecompressor, ZstdDict d = ZstdDict(open(\u0026#34;tunnel.dict\u0026#34;, \u0026#34;rb\u0026#34;).read()) _c = ZstdCompressor(level=1, zstd_dict=d) # load it ONCE, not per frame out = _c.compress(frame, mode=ZstdCompressor.FLUSH_FRAME) On the repetitive traffic that took per-packet compression from 14.8% to 5.3%, recovering 96% of the gap to full stream compression while staying loss-tolerant; on log lines 55% of the gap, on general text 44%.\nThree rules come with it. Both ends must load the same dictionary or nothing decodes — verified here, a dictionary-compressed frame raises ZstdError without it. Train on a capture of the real traffic, because a dictionary is a prior and a wrong one costs you: the text dictionary made different text slightly worse. And load it once into a long-lived compressor; passing it per call measured 5 MB/s, which is not a typo.\nThe header byte, and the one-second ramp Two small things that stop this being fragile.\nEvery datagram carries a one-byte header, and it names the mode rather than just saying \u0026ldquo;compressed\u0026rdquo;: 0x00 the frame as it is, 0x01 a self-contained zstd frame, 0x02 a block from a continuing stream. A receiver can then decode whatever the far end chose without being configured to match it, which is worth the byte on its own.\nIn frame mode the sender only sends the compressed form when it is actually smaller, because compression is not always a win: random bytes came out at 100.7%, and a 64-byte TCP ACK compresses to 73 — the ten-byte zstd header on a packet with nothing to squeeze. Packet counts on a real link are dominated by small packets, so without that check you would inflate the majority of your traffic to shrink the minority. In stream mode it always sends the compressed form, because skipping one would put the two windows out of step.\nAnd the link starts raw: for the first second every frame goes out uncompressed, whatever the settings say, and each direction ramps on its own:\nRAMP = 1.0 # seconds of raw frames before compression starts def pack(self, frame): if self.c is None or time.monotonic() - self.started \u0026lt; RAMP: return RAW + frame out = self.c.compress(frame, mode=ZstdCompressor.FLUSH_BLOCK) return ZSTD + out if len(out) \u0026lt; len(frame) else RAW + frame That is the staging idea from the top of the post applied to a single link. The tunnel comes up on the simplest path it has, proves it can carry a frame, and only then starts doing anything clever. When it breaks, you know which second it broke in.\nStage Six: Compressed And Encrypted, And What The Order Costs Order matters, and only one works: compress, then encrypt. Ciphertext does not compress, as the 100.7% row shows — which is also why TLS 1.3 removed its own compression, leaving your layer as the only place to do it. That order has a known problem, the same one I raised against deflate: compressing before encrypting leaks plaintext through ciphertext length, and where an attacker can inject chosen data alongside a secret and watch the sizes, that leak has been a working attack more than once — CRIME and BREACH against TLS14, and VORACLE against exactly this shape15.\nThat is not a reason never to compress. It is a reason to know which case you are in:\nA link carrying your own traffic between two boxes you own — replication, backups, logs, telemetry — has no attacker-chosen plaintext travelling with secrets. Compress it, and compress the stream. A link carrying arbitrary user browsing, where somebody else\u0026rsquo;s web content and your credentials go together, is the case VORACLE was written about. Leave it off. The reason I am comfortable putting this in the tap build and not the PPP one is not principle: here you choose a modern algorithm deliberately, for traffic you have looked at, while pppd\u0026rsquo;s deflate compresses everything by default with a 1996 one whether the case suits it or not.\nWhat The Tap Build Gains, And What It Gives Up Put the two halves side by side, because they are not competing. They are different trades.\npppd over a socket tap or tun over a socket Framing built in (HDLC, resynchronises after damage) none over UDP because none is needed; invent it over TCP Address setup negotiated by IPCP and IPV6CP configured by hand at both ends Liveness LCP echo, built in none; you add it or the link dies silently Authentication PAP or CHAP available, both weak none at all Layer 3 only 3 with tun, 2 with tap Bridging no yes, with tap Overhead flag, header and FCS per frame, plus escaping nothing, or 14 bytes with tap Compression deflate and BSD, on by default, from 1996 none, or per-frame zstd that you add and control Root needed yes, throughout to create the device; not to use it Lines of moving parts one daemon, thirty years old one file descriptor PPP gives you a negotiated, self-monitoring link and charges a protocol for it. A tap device gives you a raw hole for nothing, and you supply the missing parts or do without them. For a tunnel left running, the missing liveness check is the one that bites: pppd notices a dead carrier in thirty seconds and rebuilds it, the tap build notices nothing because nothing in it is watching. Add a keepalive, run it under something that restarts it, or use the build that already has one.\nWhen This Is The Right Tool, And When It Is Not It is never really the right tool, and I am not going to pretend otherwise. Everything above works, and none of it is what you should run in production. The honest version of a how-to includes the part where you put the tool down, and this is that part.\nWhat it is genuinely good for is showing you how a thing works. A VPN pulled apart into its parts, and, in the next section, how egress actually behaves once somebody is inside your network with root. Those are the reasons to have read this. The narrow cases below are real, but they are not why the post exists.\nReach for the PPP build when:\nThe carrier is not IP at all — a serial console, a USB gadget, a radio link, a named pipe, an SSH channel. pppd does not care what the bytes travel on, and a tap device cannot help you here. You are rescuing something: a machine with a serial console, no network, and a job that needs finishing tonight. PPP over that console is a routed link, installed at both ends with nothing to copy across. You want the link to look after itself. LCP echo, address negotiation and restart are free; writing them yourself is how the tap build becomes a small unreliable product. Reach for the tap build when:\nYou need layer 2 — a non-IP protocol, a cluster heartbeat that wants broadcasts, DHCP across the tunnel, or two segments that must be one. You want the fewest moving parts: over UDP with DTLS it is two commands, no daemon, no negotiation, nothing to escape. Root is scarce. Create the device once with user, and the process moving frames never needs privilege again. Both are the right tool when you are learning. Every layer is visible and separately swappable, and there is no better way to understand what a VPN product does than to build one out of the parts and watch each one come up.\nNeither is the right tool when you want a VPN. For that, use WireGuard. It is in the kernel, a fraction of the code, it does the crypto properly with nothing to negotiate and nothing to get wrong, and it is a datagram protocol by design. ssh -w gives you a tun device over an existing SSH session in one command, and OpenVPN is the mature, audited version of the DTLS-over-tap shape above. All three are better at this than anything built here.\nBuild these because you want to know what is inside the thing you buy. Not because it was clever.\nWhat This Really Shows: Egress From The Attacker\u0026rsquo;s Side This is the reason to read a build post you were just told not to use. Turn it around and look from inside your own network, as the person who has just landed there with root. Every real breach ends up there, by a stolen key, a container escape, an unpatched service, an insider. The question that then decides how bad the day gets is not \u0026ldquo;what can they run\u0026rdquo;, because they can run anything. It is \u0026ldquo;what can leave, and had you decided that before they arrived.\u0026rdquo;\nIf outbound is open by default, the answer is everything, and there is little you can now do. Nothing here was exotic. pppd, ncat, socat and ip are signed distribution packages already on the box; the tap shim is fifty lines of standard library. One allowed outbound port — and it is 443, the one every network opens by default — and there is a routed link from your network to somebody else\u0026rsquo;s, encrypted, authenticated, surviving restarts, carrying IPv4 and IPv6, and indistinguishable at your border from any HTTPS session your users make ten thousand times a day. Make it tap and bridge it and what left the building is not a route. It is the segment.\nThere is nothing for a scanner to catch: no malware signature because there is no malware, no odd protocol because it is a normal TLS handshake to 443, no unusual binary because your own package manager installed every one. The proxy logs a connection and a byte count, and both look like work.\nAnd the destination is not even fixed. A socket does not have to terminate where the packets end up, because the carrier can be pointed through a proxy — a relay that is nothing more than two connections and a pipe, which the next section builds in a line of shell. ncat takes --proxy with --proxy-type http, socks4 or socks5, so the TLS session your border sees ends at whatever the attacker told it to connect through — an internal jump host, a permitted SaaS endpoint that happens to forward CONNECT, a cloud relay — and the tunnel rides on from there to somewhere you never see. HTTP CONNECT and SOCKS both do this by design, because that is what a proxy is for. So an allow-list entry for a destination you trust is only ever trust in the destination and in everything it will forward to, which you do not control and cannot enumerate. The endpoint on your firewall log is the proxy. It was never the far end.\nSo the uncomfortable truth: once someone is inside with root and outbound is permissive, the tunnel is not the thing you get to prevent. The parts are installed, the exit is open, and it does not even lead where it appears to. Your one chance to make this hard was before the attacker arrived, at the border, by deciding what may leave.\nThat is default-deny egress, and it is the whole lesson. Outbound blocked by default; a short, named allow-list where a person justified each destination and port; everything else refused, logged and alerted. Not because it stops a determined attacker cold — a permitted destination is a permitted tunnel — but because the alternative is having no decision to enforce at all. An egress policy written as a list of permitted ports is a policy about port numbers. It was never a policy about what leaves, and once someone has root, port numbers are all it protects.\nThis is the same finding as the ping post, which builds the tunnel out of ICMP echo, and the protocol helpers post, where your firewall opens the holes itself. Three ways in, one conclusion: the control you thought you had was over protocols, and not one of these protocols is what it says. The border, decided in advance and default-deny, is the only control that was ever real.\nA Proxy Is Two Connections And A Pipe It is worth seeing how little a relay is, because it explains why you cannot reason about the far end from the near one. A proxy is not special software. It is one connection joined to another by a pipe. The oldest form uses a named pipe, a FIFO, to carry the return direction: one ncat listens, another connects onward, and the FIFO wires the reply path between them.\nmkfifo backpipe ncat -l 7000 0\u0026lt;backpipe | ncat farend.example.net 7100 1\u0026gt;backpipe Read it as plumbing: the listener\u0026rsquo;s output runs into the second ncat and on to the far end, and the replies come back through the FIFO to the client. Two sockets, one pipe, both directions, and the client\u0026rsquo;s connection terminates here, at the relay, while the packets carry on to farend and back. I ran exactly this on loopback with a third ncat echoing at the far end, and a line sent in came back having made the whole trip.\nA relay is two sockets joined by a pipe, so the endpoint moves One connection in, another out, a pipe between. That is a proxy. client opens a TLS session RELAY — the destination your border logs ncat -l 7000 terminates the client ncat farend originates a new hop far end the real other side stdout backpipe (FIFO) carries the replies back client → relay relay → far end The client's connection ends at the relay. The packets do not. Your firewall logged a connection to this box. Where it forwards to is decided inside the box, and chaining three of these puts the real far end three pipes away — every hop's log showing only a tidy local connection to the next, and nothing beyond it. A relay is two sockets and a pipe. The listener terminates the client\u0026rsquo;s connection. A second netcat originates a fresh connection onward, and the FIFO carries the return direction between them. The client\u0026rsquo;s TLS session ends here, at the relay, and a new hop begins — so the destination your border logged is this box, and the packets carry on to wherever it forwards. Chain three and the far end is three pipes away, each hop\u0026rsquo;s firewall seeing only a tidy local connection to the next. Ncat will do the same thing in one process, execing the onward connection for each client that arrives:\nncat -l 7000 --keep-open --sh-exec \u0026#39;ncat farend.example.net 7100\u0026#39; Same shape, fewer parts: the listener\u0026rsquo;s socket is joined to the exec\u0026rsquo;d ncat\u0026rsquo;s by the pipe the shell sets between them. Chain three and the tunnel crosses three networks, terminating and re-originating at each, every hop\u0026rsquo;s firewall logging a tidy local connection to the next and nothing beyond.\nThat is the whole trick, and it is why the border log is not evidence of a destination. Each relay is the far end as far as the box before it can tell, and the real other end is however many pipes away nobody was watching.\nNow put those relays on machines that are not the attacker\u0026rsquo;s.\nChained relays across compromised hosts launder the endpoint Every hop is somebody else's machine, and every owner sees only middle attacker the only box they own relay 1 another firm's network relay 2 a hijacked VPS relay 3 a home router destination where it was going sees: 1←→2 sees: 2←→3 sees: 3←→dest No hop can see past its own two neighbours. No origin, no destination, just middle. The traffic launders through a string of other people's systems, each running the same two-sockets-and-a-pipe relay under the attacker's control. It is why a compromised host's outbound so often leads to another victim, not the attacker — twenty years of C2. Every hop is a compromised host — another firm\u0026rsquo;s box, a hijacked VPS, a home router — running the same two-sockets-and-a-pipe relay under the attacker\u0026rsquo;s command and control. No hop can see past its own two neighbours: relay 2\u0026rsquo;s owner sees a connection from relay 1 and one to relay 3, and nothing else. The traffic launders through a string of other people\u0026rsquo;s systems, which is why a compromised host\u0026rsquo;s outbound so often leads to another victim rather than to the attacker. Each owner along that chain sees only a connection from the hop before to the hop after: no origin, no destination, just middle. This is not a novel idea I am handing anyone. It is how pivot chains and C2 networks have worked for twenty years, and why a compromised host\u0026rsquo;s outbound traffic so often leads to another victim rather than the attacker. Your logs show you talked to a box in a data centre somewhere. Whose, and what it forwards to, was never in them.\nThe defensive weight is one line: you cannot attribute or trust a destination you did not restrict in advance. By the time the traffic is leaving, the address it leaves to tells you almost nothing, because it is a relay on someone else\u0026rsquo;s machine and the real endpoint is laundered behind it.\nThe Link Was Never The Product What strikes me, having pulled this apart, is how little of it is new and how much of it is sold.\nRFC 1661 is from 1994. pppd has been in every Linux distribution for thirty years, the kernel modules are eight files in one directory, and the whole of what makes a VPN a VPN — a virtual interface, a negotiated link, an encrypted carrier, a route — is four programs and a certificate. None of it is hard or secret. They documented every byte and gave it away, and an industry grew up between you and it selling the same four parts in a box with a per-seat licence and a support contract that expires. The parts did not get better. They got wrapped.\nThat is not an argument for running this in production. I have just told you not to. It is an argument for knowing what is in the box you buy, because the day the vendor changes the licensing, gets acquired, or ends your model, the difference between a bad quarter and a bad year is whether anyone at your place knows what the thing was made of.\nTake an evening and build the tunnel out of the parts. Watch LCP negotiate, break the link and watch it come back, pull the certificate and see what stops. Then read your VPN vendor\u0026rsquo;s datasheet again, and see how much of it you recognise.\nThe same evening buys the other half. If a routed link out of your network is four installed programs and a certificate, the person who just got root on one of your boxes is not held back by how hard the tunnel is to build. It is not hard. They are held back only by what you decided, before they arrived, was allowed to leave. Build it once and you stop thinking of egress as something a product enforces, and start thinking of it as a decision you either made or did not.\nNot much of it is magic. Most of it is 1994 with a coat of paint, and there is nowt wrong with 1994. It worked, it was documented, and it still runs.\nRFC 1661 — The Point-to-Point Protocol (PPP), 1994. Defines the link, LCP, and the family of network control protocols that sit on top of it.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 2637 — Point-to-Point Tunneling Protocol (PPTP), which carries PPP inside GRE.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 3931 — Layer Two Tunneling Protocol version 3, the standards-track descendant of the protocol that carries PPP inside UDP.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\npppd(8) — the PPP daemon manual page. Source for pty, notty, local, passive, noipdefault, proxyarp, record, receive-all, the LCP echo options, and the asyncmap default: \u0026ldquo;If no asyncmap option is given, the default is zero, so pppd will ask the peer not to escape any control characters.\u0026rdquo; Quotations here were read from man pppd on ppp 2.5.1.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 1662 — PPP in HDLC-like Framing. The flag sequence, the octet-stuffing rule and the Async-Control-Character-Map: \u0026ldquo;Each frame begins and ends with a Flag Sequence, which is the binary sequence 01111110 (hexadecimal 0x7e).\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nOlaf Titz, Why TCP Over TCP Is A Bad Idea — the standard explanation of retransmission stacking, opening on this exact build. The original URL no longer serves the article; this is an Internet Archive capture.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 5072 — IP Version 6 over PPP. IPV6CP and the 64-bit interface identifiers, independent of IPCP.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 1994 — PPP Challenge Handshake Authentication Protocol (CHAP). Proves the peer\u0026rsquo;s identity; encrypts nothing.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 3079 — Deriving Keys for use with Microsoft Point-to-Point Encryption (MPPE). The RC4 construction keyed from the MS-CHAP exchange, and the reason PPTP is not a live option.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNcat Users\u0026rsquo; Guide — connecting through a proxy — --proxy and --proxy-type for HTTP CONNECT and SOCKS 4/5, so the TLS carrier terminates at the proxy, not at the tunnel\u0026rsquo;s far end.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNcat Users\u0026rsquo; Guide — the --ssl, --ssl-verify and --ssl-trustfile options, and the listen and UDP modes. Version tested here: Ncat 7.92.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nUniversal TUN/TAP device driver — the kernel\u0026rsquo;s own documentation for /dev/net/tun, TUNSETIFF, and the difference between tun and tap.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPEP 784 — Adding Zstandard to the standard library, which is why compression.zstd needs no package on Python 3.14. Measured here against Python 3.14.7 and zstd 1.5.7.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 7457 — Summarizing Known Attacks on TLS and DTLS, which covers CRIME and the general shape of a compression-before-encryption length leak.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nOpenVPN — the VORACLE attack — the same leak against a VPN that compresses before it encrypts, and the reason OpenVPN now advises against compression.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/networking/a-vpn-out-of-parts-and-what-egress-really-is/","summary":"A VPN is two jobs: something that makes a virtual link, and something that carries the bytes. PPP has done the first since 1994 and does not care what the second is — which is why PPTP, L2TP and every dial-up line you ever used are the same protocol over different carriers. Netcat is a carrier. This builds it both ways. First pppd: the pty option and what it does with a pseudo-terminal, the TCP version everyone tries first, why running a stream protocol inside TCP melts under loss, the UDP version that is the one to use, the async HDLC framing and the ACCM that decides how much bandwidth goes on escaping control characters, addressing and routing and IPV6CP, and keeping the link up when the carrier dies without telling you. Then the same tunnel with no PPP at all — a tap device, one datagram per frame over UDP, the length prefix you have to invent yourself over TCP, tun against tap, and bridging. Then the part netcat has no answer for: wrapping the carrier in TLS with ncat, stunnel and openssl, and in DTLS with socat, which is the shape you actually want. It is never really the right tool, and that is the point: it shows how egress behaves once an attacker has root inside your network and outbound access was not blocked by default, and why default-deny at the border is the only control that was ever real.","title":"A VPN Out of Parts: PPP, Tap Devices and Netcat"},{"content":"There is a feature on your firewall that reads the inside of your packets, finds an IP address and a port number written in the text, and opens an inbound hole for them. No rule. No change request. No log entry you would ever look at. It is on by default on most kit that has it, and has been for the best part of twenty-five years, while the industry that put it there spent the last twenty quietly concluding that it should be switched off — in standards documents, in kernel defaults and in four separate rounds of emergency browser patches.\nIt is called a protocol helper, or an Application Level Gateway, an ALG, a session helper, a fixup, an inspection engine, a conntrack helper. Same thing. It exists because NAT broke a handful of protocols that write addresses into their own payload, and somebody decided the least-bad fix was to teach the translator to read and rewrite that payload in flight.\nHere is the part that should stop you.\nThe helper cannot tell who wrote the text. It reads bytes off a connection and acts on them, immediately, without asking anything or anyone. It has no way to know whether the string PORT 192,168,1,29,4,0 was produced by a real FTP client doing a real transfer, or by a hidden form on a web page that your user opened by accident. Both look identical on the wire, because they are identical on the wire. So a stranger on the internet who can persuade anything on your network to emit the right bytes — a browser, a chat client, anything that will send what it is told — gets to pick which inbound port your firewall opens, and to what.\nThat is not a bug in one vendor\u0026rsquo;s parser. It is what the feature does: a security device taking configuration instructions from untrusted data and applying them immediately. The bug reports and the CVEs are detail on top of that shape.\nSamy Kamkar demonstrated the browser version in January 20101, a far better one in October 20202, and three months later Armis extended it so the port opened need not even be on the machine that clicked — it can be on your printer, your camera, or a logic controller two VLANs over3. In between, the IETF wrote down that these should be off by default4, Linux turned them off in the kernel5, and the browser vendors shipped a list of blocked ports that reads almost line for line like a directory listing of Linux conntrack modules6. Every one of those is a fix applied somewhere else, because the thing that actually needed fixing sits in a box nobody is going to patch.\nThe conclusion I have come to is that protocol helpers should be off by default everywhere, and off in fact on every network you are responsible for. Not tuned. Not restricted to trusted subnets as a first resort. Off, with the two or three genuine exceptions written down and dated, exactly as you would document any other inbound rule — because that is what they are.\nWhat follows is the mechanism, the attacks in the order they were found, the compliance position — in the UK this is a plain failure of a control you may already hold — and the commands to put it right.\nWhat A Stranger Gets Out Of It Start with the outcome, because the mechanism is easier to care about once you have seen the bill. Somebody opens a link. That is their whole contribution. The page runs JavaScript that talks to the attacker\u0026rsquo;s server, and that server shapes the conversation — padding it, measuring it, acknowledging parts of it and not others — until one segment arrives at your firewall looking exactly like the opening message of a VoIP call. Your firewall believes it, because believing it is the feature. It creates an inbound mapping, and the attacker connects straight back through.\nIn the 2020 version the port opens on the machine that clicked, so every service bound on that host is reachable from the internet for as long as the mapping lives2. File sharing. The remote desktop nobody meant to leave listening. On the Fisher-Price OS (Windows), Armis made the obvious point: reach the file-sharing port and you are one unpatched host from the path WannaCry took3.\nThe 2021 version is worse, and it is the part that should decide it for you. Because the H.323 helper handles call forwarding, one message can name a third address rather than either end of the connection, so the attacker is no longer limited to the machine that clicked and can walk your internal range instead, opening a port on each address and reading back whatever answers. Armis demonstrated exactly that: port 80 across a range, banners collected, a target chosen, then the printer\u0026rsquo;s raw print port opened and a job sent to it. Their second demonstration reached a logic controller on its unauthenticated management port and rewrote its programme3.\nNo vulnerability was exploited on any of those devices. They were all working exactly as intended, and the only thing that failed was the boundary — which failed by doing precisely what it was configured to do. Consider who is inside that boundary, too. Not just your staff: a visitor on the guest network, a contractor\u0026rsquo;s laptop, anyone with a browser. The attack needs no credentials, no foothold and no malware, because the browser is the delivery mechanism and it is already installed and already trusted.\nA firewall is for making sure nothing outside starts a conversation with anything inside unless you said so. The helper is the exception, granted on the strength of a string in a packet.\nWhat A Protocol Helper Actually Is A plain NAT is a dumb, honest thing. A packet arrives, it rewrites the source address and port in the header, writes a row so the reply can be turned back round, and forwards it. It never looks below the transport header, and does not know or care whether the bytes inside are a web request, a database query or a photograph of a dog.\nThat works until a protocol writes an address into its own payload. FTP does it, in plain text in the command stream: connect back to me at this address, on this port. SIP does it in the Via, Contact and SDP fields, H.323 does it, and IRC does it for direct transfers. All were designed when a host\u0026rsquo;s address was its address and every machine could reach every other, and under that assumption it is sensible. Put a translator in the path and the address in the payload becomes a lie.\nA plain translator reads the header. A helper reads the payload and opens a hole on what it finds. Same point in the path. One of them reads the letter as well as the envelope. A plain translator A protocol helper The packet arriving [IP hdr 192.168.1.29:51000] [ payload ] The packet arriving [IP hdr 192.168.1.29:51000] [ PORT 192,168,1,29,4,0 ] What it does Rewrites the source address and port in the header Fixes the checksums Writes one row so the reply can be turned back round Forwards it What it does All of that, and then reads the payload and finds an address and a port in it rewrites those too, and shifts the sequence numbers writes a second row: permit this inbound connection Tables it keeps Translation table only. One row, one flow, reversible. Nothing in the payload can add a row. Tables it keeps Translation table, and an expectation table. A row in the second one is an inbound firewall rule. The second table is the whole subject. It says: when a connection arrives from outside matching this pattern, let it through. Nobody wrote it, nobody approved it. Two devices at the same point in the path. The plain translator rewrites the header and never reads the payload. The helper reads the payload, rewrites the address it finds there, and adds a row to a second table permitting an inbound connection later. That second table is the subject of this post. So the helper was invented. It reads the payload, finds the address, rewrites it to the public one, and then does the bit everybody skips over: it creates a rule permitting the inbound connection that the payload just described. Netfilter calls that rule an expectation. Cisco calls it a pinhole, or a NAT door7. Palo Alto calls it a dynamic NAT pinhole8. Juniper calls it a gate. Different word, identical object: a hole punched in the boundary, on the strength of something read out of a packet.\nThe IETF named the pattern before most of these products existed. An ALG, RFC 2663 says, is an \u0026ldquo;application specific translation agent\u0026rdquo; that \u0026ldquo;may interact with NAT to set up state [\u0026hellip;] modify application specific payload and perform whatever else is necessary to get the application running across disparate address realms\u0026rdquo;9. Read that with an adversary in mind: whatever else is necessary, driven by application specific payload. And by February 2002 the IETF\u0026rsquo;s middlebox taxonomy was calling the mechanism what it is, noting that some ALGs create fragmentation problems \u0026ldquo;although in this case the problem is arguably the result of a deliberate layer violation (e.g., mucking with the application data stream of an FTP control connection by twiddling TCP segments on the fly)\u0026rdquo;10.\nA deliberate layer violation. Twiddling TCP segments on the fly. That is the mechanism, named by the people who catalogued it twenty-four years ago, and every attack here walks straight through it.\nThe Expectation Is The Whole Problem An expectation is a pre-authorised connection: a tuple of source address, source port, destination address, destination port and protocol, some fields filled in and some left as wildcards, plus a timer. When a matching packet arrives, the firewall treats it as related to an existing permitted flow rather than as a new inbound connection, lets it through, and consumes the expectation. On a healthy system the list of them is empty, which is the point — expectations are meant to be rare, short-lived and driven by something an inside host genuinely asked for.\nAn expectation is an inbound rule with blanks in it, and the blanks are the security model One object, five fields, and a timer. Which fields are blank is the entire security model. expect proto=tcp src=\u0026lt;who may connect\u0026gt; sport=\u0026lt;any\u0026gt; dst=\u0026lt;where it lands\u0026gt; dport=\u0026lt;which port\u0026gt; timeout=300 Helper Who may connect Where it lands Which port Worst case FTP the server, pinned the client, pinned from the payload any port on one host IRC any address at all the client, pinned from the payload any port, from anyone H.323 any address at all from the payload from the payload any port on any host Blue: fixed by the firewall from the connection it can see.\u0026#160;\u0026#160;Gold: supplied by the payload.\u0026#160;\u0026#160;Pink: left open, by design. 1. Which fields are wildcards? A wildcard source means the hole is not reserved for the far end of the conversation that made it. Anyone who reaches your outside interface first can take it. The IRC helper has to do this, because the protocol genuinely cannot know who is going to connect, which is a fair description of a protocol that should not be helped at all. 2. Who supplied the values? Not the firewall. The payload did, and the payload was written by whichever end of the conversation the helper was reading at the time. There is no authentication anywhere in that sentence. 3. Who can the hole point at? On most helpers, the host that created it. On H.323, whatever the payload says, because forwarding names a third party. An expectation is a firewall rule with some fields left blank and a timer on it. Three questions decide whether it is safe, and they are the right three to ask of any helper on any platform. Two of those answers are worse than people expect. The values come from the payload rather than from the firewall, so they come from whichever end of the conversation the helper happened to be reading. And on the IRC helper the source address is a wildcard by design, because the protocol cannot know who will connect: it \u0026ldquo;creates expectations whose destination address is the client address and source address is any address\u0026rdquo;11.\nHere is the sentence to carry forward. An expectation is an inbound firewall rule, created at line rate, by an untrusted party, with no record of who asked for it or why. Hold it until the Cyber Essentials section, because it is the whole of the argument there.\nAs Designed: FTP Says PORT, And The Firewall Believes It Take the simplest helper and watch it work correctly, because the attack is the same sequence with one participant swapped out. Active FTP uses two connections. The client opens a control connection to the server on port 21 and gives it commands in plain ASCII, and when a transfer is wanted it opens a listening socket and sends a PORT command naming the address and port to connect back to, at which point the server connects inbound. That is the protocol as specified, and it has worked that way since 1985. On the wire it is six numbers, four for the address and two for the port, high byte first:\nPORT 192,168,1,29,4,0 That is 192.168.1.29, port 1024, because 4 × 256 + 0 = 1024.\nActive FTP through a helper: four steps, all of them correct, and only two things ever checked Active FTP, working exactly as specified, and what got checked before the hole opened FTP client 192.168.1.29 Firewall + helper 203.0.113.10 FTP server 198.51.100.7 1\u0026#160;\u0026#160;control connection outbound to port 21 \u0026#8212; permitted, it started inside 2\u0026#160;\u0026#160;the client listens, then writes its own address into the stream PORT 192,168,1,29,4,0 3\u0026#160;\u0026#160;the helper acts rewrites the address, creates the expectation PORT 203,0,113,10,4,0 expect src=198.51.100.7 dst=192.168.1.29 dport=1024 4\u0026#160;\u0026#160;the server connects inbound, the expectation matches, the data flows Look at what the firewall verified before step four became possible. That the bytes were on a connection to port 21. That they began with the four characters PORT. That is the whole of it, because that is all there is to check. FTP carries no signature, no session key and nothing else to test. Active FTP through a helper, correct at every step. Note what the firewall verified before opening the hole: that the bytes were on port 21 and started with PORT. Nothing else, because there is nothing else to check. The helper watches that stream, spots PORT, rewrites the address from the private one to the public one while adjusting the sequence numbers because the string changed length, and creates an expectation permitting the server\u0026rsquo;s inbound connection on the port named. Genuinely useful, entirely reasonable given the constraints of 1994, and correct at every step.\nNowt else got checked, because there is nowt else to check — FTP has no signature to offer, no session key and no authentication of any kind, and the helper is reading a stream it is not a party to.\nOne more thing lives in that code, and it is a warning the authors wrote to themselves. If the address in the PORT command is not the client\u0026rsquo;s own — if the client asks the server to connect somewhere else entirely — the Linux helper refuses by default, and the comment in the source says why: \u0026ldquo;DMZ machines opening holes to internal networks, or the packet filter itself\u0026rdquo;12. Set the loose module parameter and that refusal goes away. The people who wrote the helper knew exactly what it could be made to do. They shipped the safe default and a switch, and twenty-five years later the switch is still there.\nNot As Designed: The Same Bytes, From A Web Page A helper reads a byte stream and matches a pattern. It does not, and structurally cannot, verify that the thing at the other end is the client it is pretending to be. Samy Kamkar published the consequence in January 2010 and called it NAT Pinning1. The mechanic is embarrassingly small: put a form on a web page, point it at the attacker\u0026rsquo;s server on port 6667, and arrange the body so that it contains a direct-chat request.\nPRIVMSG samy :^ADCC CHAT samy 3325256705 22^A The browser submits it, believing it is doing an HTTP POST. The router\u0026rsquo;s IRC helper, watching a connection on port 6667, sees DCC CHAT go past with an address and a port and does what it is built to do. The address there is 198.51.100.1 written as a single decimal number, which is how the protocol encodes it, and the port is 22; nothing in that string was chosen by the victim. The FTP variant is the same idea aimed at port 21, using a passive-mode response line instead1.\nNo cross-site scripting. No request forgery in the usual sense. No vulnerability in the browser. The browser did what browsers do, the firewall did what it was configured to do, and the result is a port forward to the attacker.\nTwo columns the helper cannot tell apart, because on the wire there is nothing to tell apart The protocol as designed, and a web page. The helper sees one picture. A real client A hidden form on a page What starts it A person opens an FTP or IRC client and connects The client opens a socket and names it in the stream PORT 192,168,1,29,4,0 What starts it A person opens a page. That is their whole contribution. A form posts to the attacker's server on the same port PORT 192,168,1,29,4,0 What the helper checks destination port matches\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;yes keyword at the start of the data\u0026#160;\u0026#160;\u0026#160;yes syntax parses\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;yes address and port present\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;yes What the helper checks destination port matches\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;yes keyword at the start of the data\u0026#160;\u0026#160;\u0026#160;yes syntax parses\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;yes address and port present\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;\u0026#160;yes an inbound port is opened, identically, in both columns What differs is the three things the helper cannot see. Whether a client was involved. Whether the user meant it. Who chose the address and the port. None of those leave a trace on the wire, so no amount of care in the parser recovers them. Both columns are perfectly well formed. The same helper, the same pattern match, the same expectation — and no FTP or IRC client anywhere in the picture. Every check the helper performs is satisfied identically in both columns, because on the wire there is no difference to find. That was 2010. Sixteen years ago. The browser vendors added the IRC ports to their blocked list, which closed that particular door and left the room it opened onto exactly as it was.\nBe blunt about the point, because it gets lost under the vendor names and the CVE numbers. The attacks do not exploit a flaw in the helpers. They use the helpers correctly. Every packet is well formed and every check passes honestly. The feature is doing its job, and its job is the problem.\nLining The Bytes Up Where The Helper Will Read Them There was a real obstacle between NAT Pinning and something much worse, and the way round it is the cleverest thing here. Most helpers will not match a pattern anywhere in a stream: they check that the keyword sits at the start of the data portion of a packet, which is what a real protocol message looks like and what an HTTP body never is, because a body arrives after a pile of headers the attacker does not control. Samy quotes the kernel\u0026rsquo;s own behaviour — the handler bails unless the method occurs at the start of the data2.\nSo the attacker needs the browser to emit a segment whose very first byte is the keyword. They cannot write the headers, but they can write the body and make it as long as they like, which turns the problem into arithmetic.\nMoving the segment boundary until the keyword lands where the helper reads The helper reads the first byte of a segment. So the attacker moves where the segments break. As the browser would send it segment 1 POST / HTTP/1.1 Host: ... segment 2 ...headers... PORT 192,168,... segment 3 ...1,29,4,0 padding padding The keyword sits in the middle of segment two, so the helper reads past it and does nothing. After the attacker sets the segment size, or acknowledges only part of the stream segment 1 POST / HTTP/1.1 Host: ... segment 2 ...headers and padding... segment 3 PORT 192,168,1,29,4,0 The keyword is now the first byte of a segment. The helper reads it and opens the port. The two levers, and neither is an attack on TCP. The attacker's server is one end of the connection, so it announces the maximum segment size the victim's stack will use, and it decides how much of the stream to acknowledge. Acknowledge part of it and the victim retransmits from the offset the attacker picked. Both are ordinary, standards-compliant TCP, used deliberately and on purpose. Why the padding matters. The helper only reads a protocol keyword when it is the first byte of a segment, so an attacker who controls only the body has to move the segment boundary until it lands there. Both levers are ordinary TCP, used deliberately. The browser sends a large request with a recognisable delimiter buried in the body; the attacker\u0026rsquo;s server sniffs it, measures how many bytes of headers came first, and now knows the offset. It then either advertises a segment size that puts the desired byte on a boundary, or sends a partial acknowledgement so the victim retransmits starting exactly there — and Armis adds that the TCP window can be crafted in those acknowledgements too, \u0026ldquo;in order to fully control how the TCP stream is to be segmented\u0026rdquo;3. That is the whole trick, and it uses nothing but the far end of a connection the victim\u0026rsquo;s own browser opened.\nOnce you have it, the browser is a packet generator pointed at the inside of your firewall. It only has to put the right thirty bytes at the start of a segment.\nThe 2021 work skipped most of the arithmetic. A browser relay connection over TCP carries an attacker-controlled username field, sent early, which accepts newlines and null bytes — Armis showed a capture where that field is \\r\\nPORT 192,168,1,29,4,0\\r\\n3. Worse, that path did not consult the browser\u0026rsquo;s blocked-port list at all, so the one mitigation the industry had shipped twice was simply bypassed.\nNAT Slipstreaming, End To End Put the pieces together and this is the whole attack, in order.\nClick to inbound connection in six steps, none of which is a vulnerability Six steps. Four of them are ordinary web and TCP. One is your firewall. One is the attacker. 1 The victim opens a page An advert, a link in a message, anything. This is the whole of the user's involvement. ordinary web 2 The page learns the inside address Handed over by the browser's media interface, or narrowed down by timing common gateway addresses. ordinary web 3 The attacker measures the path An oversized request with a marker in it, and a packet capture at the far end to see where the breaks fall. ordinary TCP 4 The page sends the real payload Padded so the protocol keyword lands on a segment boundary, exactly as the previous diagram shows. ordinary TCP 5 Your firewall opens the port The helper recognises the message, rewrites the address in it, and writes the expectation. your firewall 6 The attacker connects inbound To your public address, on the mapped port. The expectation matches and the firewall forwards it inside. your firewall Count the vulnerabilities exploited on the victim's machine. There are none. The browser was patched. The operating system was patched. Nothing was installed and no credential was used. Seconds elapsed, one click, and the only component that could have refused was the one built to refuse. The full chain, from click to inbound connection. Steps one to four are ordinary web and TCP behaviour. Step five is the firewall feature working exactly as documented, and it is the only component that could have refused. Total elapsed time: seconds. User interaction: one click. Vulnerabilities exploited on the victim\u0026rsquo;s machine: none.\nThe disclosure timeline is evidence in its own right. Samy published on 31 October 2020, and between 6 January and 26 February 2021 Chrome, Edge, Safari and Firefox all shipped mitigations3, tracked as CVE-2020-16043, CVE-2021-23961 and CVE-2021-1799.\nFour browser vendors shipped emergency patches for a firewall feature. Armis explain why in their own words: \u0026ldquo;While the underlying issue of this attack is the way NATs are implemented (in various ways in routers and firewalls, throughout numerous vendors and applications), the easiest and fastest way to mitigate was through a patch to browsers\u0026rdquo;3.\nThat is not a fix. It is four other industries putting sandbags round somebody else\u0026rsquo;s door, because the door itself was never going to be replaced.\nH.323 Call Forwarding Points At Anything On Your Network Most helpers bound the damage without meaning to. The FTP helper pins the expectation\u0026rsquo;s destination to the client that created it, so the worst an attacker gets is a port on the machine that clicked — bad, but survivable.\nH.323 is catastrophic, and the reason is a telephony feature.\nA phone system has to support call forwarding, and a forwarded call is by definition a call to somebody not on the line. So the signalling has to name a third endpoint, and the helper has to open a path to it, or forwarding does not work through NAT at all. Every serious implementation supports it, and the netfilter helper documents the behaviour explicitly, with a diagram13. Armis read the code and stated the consequence in one sentence: \u0026ldquo;a single H.323 packet sent over TCP port 1720 that initiates call forwarding can open a pinhole (named an expectation in the conntrack subsystem) to any TCP port of any internal IP on the network\u0026rdquo;3.\nPinned to one host, or aimed at anything on the network Most helpers can only reach the machine that clicked. One of them cannot be bounded at all. FTP, SIP, IRC H.323, with call forwarding The expectation says dst = the host whose connection made this The expectation says dst = whatever address the payload named the machine that clicked 192.168.1.29 the printer port 9100 the camera default password the controller no authentication Blast radius One machine. Every port on it, for the life of the rule, which is bad enough and is survivable. The device is managed, patched and running something that logs. Blast radius Every address on the network, one message each. Walk the range, read the banners, pick a target. None of those devices ever made a request or ran a browser. Call forwarding is the whole difference, and it is a feature, not a fault. A forwarded call is a call to somebody not on the line, so the signalling names a third party and the helper obliges. Why one helper is in a different category from the rest. Pin the expectation to the host that created it and the attacker reaches one machine. Let call forwarding name a third party, as the protocol requires, and they walk your whole address range. Any TCP port. Any internal address. From one packet, which the attacker got a browser to send.\nSo the attack stops being about the victim\u0026rsquo;s machine and becomes a port scan of your internal network, run from the internet, with the results read back over connections your own firewall is authorising. The devices it finds are the point, because they are not patched, frequently not patchable, and the entire security model for most of them is it is on the inside. Armis put a number on that: a year after publication, 97% of industrial controllers vulnerable to one set of critical flaws were still unpatched3. That is the reality \u0026ldquo;on the inside\u0026rdquo; is holding up.\nSo H.323 is the one to deal with first, on every platform, because it reaches past the victim — and it is also the one almost nobody uses, which makes it the cheapest security improvement here. Turn it off. Nobody will ring you.\nThe Helper Can Be Fired By A Message You Did Not Send One variant deserves its own section, because it breaks the assumption people reach for when they want a reason not to act: fine, but the trigger has to come from inside, so control the browsers and you control the problem.\nNo.\nIn July 2022 David Leadbeater found two faults in the Linux IRC helper14. The module looks for the string \\1DCC anywhere in the stream rather than checking it is a properly framed message in the right place, and its address check compares against the chat server\u0026rsquo;s address rather than the host behind the translator, so the publicly known address of a public server satisfies it. Put those together and an attacker sends the victim\u0026rsquo;s client a client-to-client ping — an entirely ordinary thing to receive from any other user — with a direct-transfer request inside it:\nPRIVMSG ExampleUser :^APING ^ADCC CHAT x 3325256705 22^A By the rules of the protocol, the client replies to a ping by echoing the payload back, so the victim\u0026rsquo;s client dutifully sends that string outbound, and the helper, watching the outbound stream, finds DCC in it and opens the port.\nA port opened by a message the victim never composed Nobody inside clicked anything. The victim's own client emitted the trigger, on request. The victim's client 192.168.1.29 Firewall + IRC helper reading the stream A stranger any other user, any network 1\u0026#160;\u0026#160;an ordinary ping, with a direct chat request hidden in the payload PRIVMSG ExampleUser :^APING ^ADCC CHAT x 3325256705 22^A 2\u0026#160;\u0026#160;the protocol says answer a ping by echoing the payload, so it does 3\u0026#160;\u0026#160;both of the helper's checks pass, and neither should have it looks for the keyword anywhere in the stream, not at the start of a framed message it validates the address against the chat server, not against the host behind the translator 4\u0026#160;\u0026#160;the expectation exists, and port 22 is open inbound to the victim Everything in that sequence followed its specification. The stranger sent a legal message. The client answered it the way the protocol requires. The helper matched the pattern it was written to match. No user decided anything, and nothing in any log will look the least bit unusual. The trigger does not have to come from someone you trust, or from a browser, or from anything a user did on purpose. Every piece of software here followed its specification, and the result is an open port and nothing odd in any log. Nobody inside did anything wrong and nobody clicked. The client followed the specification, and the specification and the helper together produced an inbound hole to port 22 on a machine inside the network. It became CVE-2022-2663, and the national vulnerability database description is admirably plain: \u0026ldquo;A firewall may be able to be bypassed when users are using unencrypted IRC with nf_conntrack_irc configured\u0026rdquo;15.\nNote the word unencrypted. Hold that one too.\nThe author\u0026rsquo;s own recommendation is where the rest of this post arrives from a different direction: \u0026ldquo;Potentially entirely deprecate and remove nf_conntrack_irc, it\u0026rsquo;s unclear it has much use anymore\u0026rdquo;14.\nThe Parsers Are The Other Half Of It Everything so far has been about helpers working correctly. There is a second, separate problem: they are protocol parsers written in C, in the fast path of a security device, on data supplied by strangers, with the bug rate you would expect. Nor is it a Linux problem or a cheap-router problem — it shows up in every vendor, in the code they charge the most for.\nCVE Component What a crafted packet does CVE-2018-0051 Junos SIP ALG Crashes the flow daemon on SRX and MX; also records that SIP ALG is on by default except on high-end models CVE-2018-15454 Cisco ASA / FTD SIP inspection Reloads the device or pins the CPU CVE-2022-2663 Linux nf_conntrack_irc Opens ports through the firewall, as above CVE-2023-22412 Junos SIP ALG \u0026ldquo;Specific SIP messages\u0026rdquo; crash the flow daemon, repeatably CVE-2023-22415 Junos H.323 ALG Out-of-bounds write from \u0026ldquo;specific H.323 packets\u0026rdquo; CVE-2024-21616 Junos SIP ALG One SIP packet exhausts the NAT pool, so genuine traffic stops translating CVE-2024-26851 Linux nf_conntrack_h323 Out-of-range bit shift decoding the H.323 bitmap CVE-2024-39551 Junos H.323 ALG Memory exhausted by \u0026ldquo;specific packets\u0026rdquo; until traffic stops Every one is reachable from the network with no authentication, by anyone who can get a packet to the outside interface, which is everyone. And read the wording: specific SIP messages, specific H.323 packets, a specific SIP packet. That is not a protocol failing under load. It is somebody building a packet on purpose.\nSit with the Cisco case. CVE-2018-15454 was published on 31 October 2018 at 8.6 severity, it was being exploited in the wild, and the advisory said the software update was not available yet16. Cisco\u0026rsquo;s mitigation, in their own advisory, is no inspect sip. So the vendor\u0026rsquo;s answer to an actively exploited flaw in the feature was to turn the feature off — which raises the question the rest of this post is built on. If turning it off is an acceptable answer during an incident, on what basis is it on the rest of the time?\nA Helper Only Works If You Do Not Encrypt This part should end the argument on its own, and gets the least attention. A protocol helper reads your payload and cannot read an encrypted one. So a helper is only doing anything at all on traffic you have chosen to leave in the clear, and keeping the helper working means keeping that traffic in the clear.\nEvery protocol on the list has had an encrypted mode for well over a decade. FTP has had TLS since 200517; turn it on and the PORT and PASV exchanges are invisible and the helper does nothing. SIP has had TLS since the base specification, with encrypted media alongside it. H.323 has its own security annexe. Chat has had TLS for a very long time, and the official advice for CVE-2022-2663 was, in as many words, to use it so the helper cannot see your transfer requests15.\nThe helper needs cleartext, so keeping the helper means keeping the cleartext You can have the encryption or you can have the helper. There is no third column. Control channel encrypted Helper working What the helper sees 17 03 03 01 a4 9c 2f e1 8b 44 d0 ... ciphertext No keyword. No address. No port. Nothing to match. What the helper sees REGISTER sip:example ... Contact: 192.168.1.29:5060 And so does every other device between you and them. What that costs you The helper does nothing at all, so the endpoints have to solve their own translation problem \u0026#8212; which every one of these protocols has been able to do for a decade. What that costs you Registration credentials, who called whom, and every internal address, in clear text, across every network between the two ends. None of which you control. The encrypted mode has existed this long FTP over TLS since 2005\u0026#160;\u0026#160;\u0026#183;\u0026#160;\u0026#160;SIP over TLS with encrypted media since the base standard\u0026#160;\u0026#160;\u0026#183;\u0026#160;\u0026#160;chat over TLS for decades\u0026#160;\u0026#160;\u0026#183;\u0026#160;\u0026#160;H.323 has its own security annexe So the honest form of \"we need the SIP helper\" is a sentence nobody says out loud. It is: we need our call signalling to cross untrusted networks in plain text, so a box we do not own can rewrite it. The trade nobody writes down. A helper works only on payload it can read, so the two columns are mutually exclusive. The right-hand one is what a working SIP ALG asks you to agree to. So the honest form of \u0026ldquo;we need the SIP ALG\u0026rdquo; is: we need our call signalling to travel in plain text across untrusted networks, so a middlebox we do not control can rewrite it. Say it that way in a design review and see how far it gets.\nThere is a sharper version, and it is why this is not a close call. Leaving a helper enabled is a standing incentive against encryption: the day somebody turns on SIP over TLS, calls break and the helper is why, so the change is backed out and the cleartext stays another year. Ask anyone who has tried it behind a consumer-grade firewall how that went.\nEvery other protocol on the internet has gone the other way: web traffic encrypted by default, DNS encrypted, mail transport encrypted, QUIC encrypting the transport header itself so middleboxes cannot read it. The middlebox era ended on the open internet years ago, and the last places relying on a device in the path reading the payload are the ones where somebody left a helper on.\nIPsec Is The One Where The Helper Cannot Read Anything Which raises the obvious question about the protocol that is nothing but encryption. The answer is worse than you would guess.\nESP has no port numbers, being an IP protocol in its own right, and a translator demultiplexes return traffic by port. So with two clients behind one address heading for the same gateway there is nothing to tell their inbound packets apart. RFC 3715 wrote that down in March 2004: a NAT cannot learn the mapping by inspection, and \u0026ldquo;it is possible that the NAT will deliver the incoming IPsec packets to the wrong destination\u0026rdquo;18.\nSo vendors built a helper. It watches the IKE exchange on UDP 500, whose early packets are in the clear, harvests the cookies and the security parameter index, and opens a gate so inbound ESP carrying that value reaches the right host inside. Juniper describes it plainly: \u0026ldquo;When ESP traffic hits the IKE ALG gates, sessions are created to capture subsequent ESP traffic\u0026rdquo;19. Cisco\u0026rsquo;s inspect ipsec-pass-thru does the same for ESP and AH \u0026ldquo;associated with an IKE UDP port 500 connection\u0026rdquo;, with a default map that sets no limit at all on ESP connections per client20.\nSame object, same authority, except the helper is not even matching a keyword. It cannot parse ESP, because ESP is the encrypted part; it is steering packets on a 32-bit number it watched go past in cleartext. The section of RFC 3715 that covers this is titled, without irony, \u0026ldquo;Helper Incompatibilities\u0026rdquo;, and records that cookie demultiplexing \u0026ldquo;results in problems with re-keying\u0026rdquo; and that devices parsing ISAKMP payloads \u0026ldquo;may not handle all payload ordering combinations\u0026rdquo;18. A guess standing in for a rule, and a hand-rolled parser in the packet path, written down twenty-two years ago.\nThe fix shipped ten months later, in the protocol: RFC 3947 has the two ends detect a translator during the key exchange, and RFC 3948 wraps ESP in UDP 4500 so there are ports again21. Juniper then says the quiet part out loud — \u0026ldquo;IKE NAT-T traffic on floating port 4500 is not processed in an IKE ALG\u0026rdquo;19. Do it properly and the helper is bypassed entirely, which is the same sentence as passive FTP and as ICE.\nMainline Linux never took this one. There is no ESP module among the conntrack protocols, and a 2021 patch adding SPI-based tracking went through review on the netfilter list and was never merged22. The vendors who ship it ship it out of tree, on the boxes least likely to be updated.\nIt Fails Cyber Essentials, Line By Line Up to here this has been a security argument. For anyone certifying in the United Kingdom it is also a compliance one, and there is no clever interpretation involved — it is three bullet points against three bullet points. Cyber Essentials is the UK government-backed scheme delivered through IASME, its first technical control is firewalls, and the current requirements document is version 3.3, April 2026. Here is what it obliges you to do, verbatim23:\nblock unauthenticated inbound connections by default ensure inbound firewall rules are approved and documented by an authorised person, and include the business need in the documentation remove or disable unnecessary firewall rules, when they are no longer needed Three firewall requirements, and what a helper does about each of them Cyber Essentials, control one, firewalls. Three obligations, and a helper's answer to each. What the requirements document says What a protocol helper does Result \"block unauthenticated inbound connections by default\" The first obligation, and the reason the control exists. Permits one, on the strength of a string in a packet. The party that supplied the string authenticated to nothing at all. fails \"ensure inbound firewall rules are approved and documented by an authorised person, and include the business need\" Written at line rate by a kernel module. No person approved it, no person saw it, no document, no business need recorded. fails \"remove or disable unnecessary firewall rules, when they are no longer needed\" Which presumes somebody decided they were needed. Removed by a timer. A timer is not a review, and nobody ever assessed whether the rule was necessary to begin with. fails This is the first of the five technical controls, not an edge case in the fifth. It applies, in the scheme's own words, to boundary firewalls, desktop computers, laptops, routers and servers. The three firewall requirements in the Cyber Essentials technical controls, and what a protocol helper does about each one. Three requirements, three failures, on the first of the five controls. An expectation exists precisely to permit an inbound connection that would otherwise be blocked, and the party whose data caused it authenticated to nothing, so the first bullet fails outright. The rule was written at line rate by a kernel module, so there is no document, no business need recorded and no authorised individual anywhere in the chain — ask for the approval record for the rule that let a connection reach port 9100 on your printer and you have none, and cannot make one, because it existed for ninety seconds eighteen months ago. And a helper\u0026rsquo;s rules are removed by a timer, which is not a review.\nThree requirements, three failures, on the first control of five, applying in the scheme\u0026rsquo;s own words to \u0026ldquo;boundary firewalls, desktop computers, laptops, routers, servers\u0026rdquo;23, which is to say everything you own.\nBe fair about what that means, because I am not the certification body. An assessor works from the question set and the evidence you give them, and that set asks whether you block unauthenticated inbound connections by default and whether your inbound rules are documented and approved. Answer yes with a helper enabled and the answer is not true. You may well pass anyway. Passing and complying are not the same thing, and the gap surfaces after an incident.\nIt is only unusually plainly worded, too: a card-industry standard, a customer\u0026rsquo;s questionnaire and your insurer\u0026rsquo;s proposal form all ask the same thing in other words. So this is the section to take to whoever signs the certificate. Not the attacks and not the CVEs. Three bullets, and the honest answer to each.\nThe Industry Already Decided This, Twenty Years Ago None of this is new or contested. What is remarkable is how long the decision has been made while the defaults carried on regardless.\nTwenty-five years of the same finding, and the one row that never appears The conclusion was reached in 2007. The defaults did not move. When What happened Who could act Jan 2001 RFC 3027 catalogues every protocol that NAT breaks, and what an ALG must do about each standards body Feb 2002 RFC 3234 calls the mechanism \"a deliberate layer violation\" and warns of extra points of attack standards body Jan 2007 RFC 4787, a Best Current Practice: NAT ALGs for UDP-based protocols SHOULD be turned off standards body Jan 2010 NAT Pinning: a hidden form on a web page opens an inbound port on the visitor's machine a researcher 2012 netfilter gains an explicit attach mechanism and a switch to stop automatic assignment the kernel Apr 2016 Linux changes the default: helpers do nothing without an explicit rule. Ships in 4.7 the kernel Oct 2020 NAT Slipstreaming, then the January 2021 variant that reaches every device on the network researchers Nov 2020 Four browser vendors ship mitigations; the web platform standard gains a forbidden port list the browsers Aug 2022 The IRC helper turns out to fire on a message the victim never composed a researcher 2023\u0026#8211;24 Four more ALG vulnerabilities in one vendor's flagship firewall line researchers Now read the \"who could act\" column, and notice who is never in it. In twenty-five years, not one row of this is a firewall vendor pushing an update that turns the feature off on kit already in the field. The standards body asked, the kernel changed upstream, and neither could reach the boxes. Twenty-five years of the same conclusion, reached by people who could not fix the thing that needed fixing. The one row that never appears is a firewall vendor turning the feature off on kit already in the field. RFC 3027 catalogued every protocol NAT breaks in January 200124. RFC 3234 put ALGs in the middlebox taxonomy a year later and was blunt about the cost of adding boxes to a path: it \u0026ldquo;creates extra points of attack, reduces or eliminates the ability to perform end to end encryption, and complicates trust models\u0026rdquo;10. Then, in January 2007, RFC 4787 — a Best Current Practice rather than a suggestion — set out how NAT is required to behave, and requirement ten said this:\nREQ-10: To eliminate interference with UNSAF NAT traversal mechanisms and allow integrity protection of UDP communications, NAT ALGs for UDP-based protocols SHOULD be turned off.4\nOff. Nineteen years ago, with the reason given: helpers get in the way of the mechanisms that actually work, and they stop you protecting your own traffic. The same section notes, wearily, that some products have ALGs \u0026ldquo;turned on permanently\u0026rdquo;4.\nThree years later NAT Pinning showed a web page opening a port1. Netfilter answered in 2012 with the CT target, which attaches a helper to a named flow by an explicit rule, and a switch to stop automatic assignment altogether11. Then on 25 April 2016 the kernel changed its default, in a commit message worth reading in full for how tired it sounds:\nFour years ago we introduced a new sysctl knob to disable automatic helper assignment [\u0026hellip;] This knob kept this behaviour enabled by default to remain conservative. This measure was introduced to provide a secure way to configure iptables and connection tracking helpers through explicit rules. Give the time we have waited for this, let\u0026rsquo;s turn off this by default now, worse case users still have a chance to recover the former behaviour by explicitly enabling this back through sysctl.5\nThat shipped in Linux 4.7, and ever since, a box with those modules loaded does nothing with them until you write a rule attaching one to a flow. Everything after it is in the diagram above, down to the web platform standard itself gaining the port list25.\nNow look at the list for what is missing. In twenty-five years not one row of it is a firewall vendor pushing a firmware update that turns these off on kit already deployed. The standards body asked. The kernel did it upstream. The browsers paid for it. The boxes carried on.\nAnd the blocked-port list is the tell. Ports 69, 137, 161, 554, 1719, 1720, 1723, 5060, 5061, 6566 and 10080 are all on it6 — TFTP, NetBIOS, SNMP, RTSP, H.323 twice, PPTP, SIP twice, the scanner protocol and the Amanda backup protocol. List the helper modules in a Linux kernel tree and you will find you have read the same list twice. Not a coincidence, and not a security control: it is one industry maintaining a permanent, growing denylist of ports because another will not change a default.\nUnder Most Of The Badges, It Is Linux The netfilter detail is the important part, not a Linux-shaped digression. A very large share of the boxes doing NAT on this planet are Linux with netfilter underneath a vendor interface: every OpenWrt derivative, which is most of the consumer and small-business router market, most ISP-supplied home routers, a good deal of carrier-grade NAT kit, and plenty of commercial appliances whose web interface gives no hint of what is beneath it. The helper parsing your SIP is very often the same nf_conntrack_sip.c that ships in the mainline kernel, compiled by somebody else and driven by a menu.\nThe evidence is in the research itself. When Samy went looking for the SIP ALG in a Netgear router he extracted the firmware and found a kernel module containing ftp_decode and sip_decode2. Armis\u0026rsquo;s tested list included OpenWrt and VyOS plus a category they simply called \u0026ldquo;various consumer grade Linux routers\u0026rdquo;, and the H.323 analysis that produced the any-internal-host finding came from reading the netfilter source, then confirmed on commercial firewalls from three vendors3.\nThe kernel default does not reach you. Linux 4.7 turned automatic helper assignment off in 2016, but only if the kernel is new enough and nobody turned it back on. Armis found VyOS setting nf_conntrack_helper back to 1 explicitly, and noted that plenty of Linux-based products re-enable it \u0026ldquo;as it is still useful for many users\u0026rdquo;3. A 2014 kernel in a 2026 product gets the 2014 behaviour, and a current kernel with the switch flipped gets the same. Neither shows up on a datasheet.\nKnowing the netfilter model tells you what to ask of every other box. The three questions from the expectation section are not Linux questions. They are the questions. Every vendor has the same object under a different name, the documentation almost never gives you the answers, and knowing what the reference implementation does is how you work out what to test.\nDo the Linux work properly, then read every other vendor against it.\nThe Linux Helpers, Properly Start with what is actually loaded:\nlsmod | grep -E \u0026#39;nf_conntrack|nf_nat\u0026#39; The helper modules are the ones named after protocols — nf_conntrack_ftp, _sip, _h323, _irc, _tftp, _pptp, _snmp, _amanda, _sane, _netbios_ns and _talk — each with a matching nf_nat_* where addresses get rewritten. Then check automatic assignment, the switch deciding whether a loaded module does anything by itself:\nsysctl net.netfilter.nf_conntrack_helper Zero is what you want, and the default from Linux 4.7 onward5. One means every loaded helper is live on every flow matching its port, from any address — the 2015 behaviour, and the one the attacks in this post assume.\nThen watch conntrack -L expect on a live firewall. With helpers off it stays empty; with them on, put a capture beside it and watch rows appear as people use the network. Worth doing once: nothing makes the point faster than watching an inbound permission you did not write appear and disappear in front of you.\nIf automatic assignment is off and you still want a helper on a specific flow, the sanctioned way is an explicit rule, which bounds it to one destination and one port:\niptables -t raw -A PREROUTING -p tcp --dport 21 -d 192.0.2.10 -j CT --helper ftp The nftables equivalent declares a ct helper object and attaches it with ct helper set in the prerouting chain, which is the same discipline with better syntax.\nNote what that rule is: a documented, approved, business-justified inbound exception, written by a person, where an auditor can read it. Which the automatic version could never be.\nIf you want them gone rather than dormant, and on a firewall you should, stop the modules loading at all:\nfor m in ftp sip h323 irc tftp pptp snmp amanda sane netbios_ns talk; do echo \u0026#34;install nf_conntrack_$m /bin/false\u0026#34; done \u0026gt; /etc/modprobe.d/no-conntrack-helpers.conf install ... /bin/false rather than blacklist is deliberate: blacklist only stops loading by alias, and anything asking for the module by name still gets it. And if a distribution\u0026rsquo;s firewall layer loads them for you — firewalld does when a zone has the FTP or TFTP service enabled — that is the layer to fix, because it will helpfully put them back.\nEvery Other Badge, And How To Turn It Off Check your own version rather than trusting anything on the internet, mine included; these defaults move between releases and models.\npfSense and OPNsense are the proof that the argument is over. There is no SIP ALG to disable, because there has never been one to enable. They are among the most widely deployed firewall distributions there are, they run telephony for a great many organisations, and if a helper were genuinely required for modern VoIP that would not be possible. The FTP proxy went the same way: Netgate took it out of the base system in January 2015.\nOpenBSD\u0026rsquo;s design is the one everyone else should have copied. pf does not rewrite payloads in the forwarding path at all. If you want FTP helped you run ftp-proxy, a separate user-space daemon, and write an explicit divert-to rule sending the control connection to it26. Three properties fall out at once: it is off unless you deliberately turn it on, it only ever sees traffic you named in a rule, and a bug in it crashes a userland process rather than the packet path. That is what opt-in looks like when somebody designs it rather than bolting it on.\nOpenWrt does not ship the ALG modules, and automatic assignment stays off if you install them.\nCisco ASA and FTD carry inspection engines in the default global policy, and no inspect sip is Cisco\u0026rsquo;s own advice during an incident:\npolicy-map global_policy class inspection_default no inspect sip no inspect h323 h225 no inspect h323 ras no inspect skinny On FTD it is configure inspection sip disable from the device CLI16. IPsec pass-through is the one Cisco got right: inspect ipsec-pass-thru is not in the default policy at all, so unless somebody added it deliberately there is nothing to remove20.\nCisco IOS and IOS XE have it on by default too — \u0026ldquo;NAT support for SIP is enabled by default on port 5060\u0026rdquo;, in Cisco\u0026rsquo;s own words7, and the same for H.323:\nno ip nat service sip tcp port 5060 no ip nat service sip udp port 5060 no ip nat service h225 Juniper SRX enables SIP and H.323 on branch models and not on the high-end ones, which on its own tells you what Juniper\u0026rsquo;s engineers think of them. Find out where you stand with show security alg status, then:\nset security alg h323 disable set security alg sip disable set security alg ftp disable set security alg ike-esp-nat disable FortiGate inspects VoIP by default through the VoIP profile, with a kernel session helper underneath it. Fortinet\u0026rsquo;s documented sequence removes the helper first27:\nconfig system session-helper show delete \u0026lt;the sip entry\u0026gt; end config system settings set default-voip-alg-mode kernel-helper-based end Read your own show output for the entry number rather than copying one; it moves between models and releases, and Fortinet warn a reboot is often needed.\nCheck Point is the awkward one to audit, because there is no single switch. The helper is a property of the service object in the rule, so the predefined SIP service gets you the protocol handler and everything it does. Avoiding it means your own plain UDP or TCP service on port 5060 with the protocol type set to none, placed above anything still using the built-ins. \u0026ldquo;Is the ALG on?\u0026rdquo; is therefore not a question a settings page can answer. As such, be careful accepting anybody\u0026rsquo;s word that it is disabled.\nPalo Alto gives you a per-application toggle, and their own documentation says the SIP ALG \u0026ldquo;creates dynamic NAT pinholes\u0026rdquo;8. Objects, Applications, find sip, tick Disable ALG, commit.\nMikroTik ships ten helpers under /ip firewall service-port — SIP, H.323, FTP, IRC, TFTP, PPTP, RTSP and more — documented in a line each, with no security warning on the page. List first, then disable what you find:\n/ip firewall service-port print /ip firewall service-port set [find name=sip] disabled=yes /ip firewall service-port set [find name=h323] disabled=yes /ip firewall service-port set [find name=ftp] disabled=yes Consumer and ISP-supplied routers. Look for \u0026ldquo;SIP ALG\u0026rdquo;, \u0026ldquo;SIP helper\u0026rdquo;, \u0026ldquo;VoIP passthrough\u0026rdquo;, \u0026ldquo;IPsec passthrough\u0026rdquo; or \u0026ldquo;application layer gateway\u0026rdquo;, usually under advanced NAT. On many ISP-supplied ones there is no setting at all — which tells you whether that box belongs on a network you are responsible for.\nWhatever the platform, finish the same way: prove it. Put a capture on the outside interface, send a crafted PORT or REGISTER line at the relevant port from outside, and confirm nothing opens. A setting you have not tested is a belief.\nMost Of What Breaks Is Already Dead Be honest about the cost, because it is the basis for asking anyone to do this. Turning off the SIP helper on a network with badly configured phones can break calls, usually one-way audio or registrations that drop. Turning off the FTP helper breaks active-mode FTP outbound. Turning off the H.323 helper breaks H.323, if you still have any. Expect at least one of those if you do it in a single change on a network nobody has looked at in years.\nNow read the list back, and notice that almost every protocol a helper serves is one the rest of the industry already buried. H.323 lost to SIP twenty years ago. PPTP has been indefensible since 1998, a case I have made at length in IPsec Was a Good Idea. IRC direct transfers belong to a decade nobody is nostalgic for. NetBIOS name service, the community-string versions of SNMP, the scanner discovery protocol and the old backup protocol are local-network relics never meant to cross a boundary. Plain FTP is gone from the browsers entirely — Firefox removed it in July 2021 and Chrome that October, both on the grounds that usage was negligible and the security was not worth the maintenance28.\nSo \u0026ldquo;we cannot turn the helper off, something will break\u0026rdquo; is usually an argument for keeping a dead protocol on life support to justify a feature that opens ports for strangers. Turning it off does not break your network. It exposes the one thing on it that should have been retired years ago, which is information you wanted anyway.\nSIP is the genuine exception and the only one. Everything else on that list is an argument you should be glad to lose, and none of it is unfixable, because every protocol involved solved its own translation problem years ago, in the protocol, where it belongs.\nWhat To Do Instead The middlebox answer and the endpoint answer, side by side The same problem, answered twice. One answer put the decision in the middle. A box in the path decides The two ends decide How it works A device reads the payload as it passes It rewrites the address it finds written there It opens an inbound hole for the connection described Nobody at either end knows any of this happened How it works Each end asks a server what it looks like from outside Each offers every path it has: local, translated, relayed The two ends test the paths against each other They keep the one that works open with their own traffic What it needs from you The payload readable, so no encryption on the control channel. That specific device in that specific path. One translator only, so a carrier-grade one downstream breaks it. And trust in whoever wrote the text, which is the part nobody thought about until 2010. What it needs from you Outbound access, and nothing else at all. It works through translators you do not own and cannot see, through two of them stacked, through a mobile network, and with the control channel encrypted end to end, because nothing in the middle reads it. For file transfer it is simpler still: passive mode, where the client opens both connections outbound and there is nothing left for a helper to do. The right-hand column is not a proposal. It is what your browser already does. Every video call made in a browser is negotiated that way, through every kind of translator, with no ALG in the path. The same problem, solved twice. One answer put the decision in a box in the middle. The other let the two ends work it out between themselves — which is what every browser already does for every video call. FTP. Passive mode, in the specification since 1985 and the default in every client for twenty years: both connections open outbound and there is nothing left for a helper to do. And if you are moving files between organisations in 2026, FTP is not the protocol for it — SFTP and FTPS are encrypted, and no helper touches either.\nSIP and everything else real-time. The endpoint asks a server on the internet what its public address and port look like from outside, offers every candidate path it has, and the two ends test them against each other and keep one that works, with a relay where no direct path exists. That is STUN, TURN and ICE, and it is what every browser on earth does for every video call, through every kind of translator, with no ALG in the path. If your phone system cannot do it in 2026, the problem is your phone system.\nIPsec. NAT traversal, which is to say RFC 3947 and RFC 3948: the two ends notice the translator during the key exchange and wrap ESP in UDP 4500 for the rest of the session21, with no gate on any box in between.\nH.323. Retire it. SIP won that argument around 2005, so there is no migration to plan, only a deletion. Direct transfers, TFTP, SNMP, NetBIOS, scanner discovery and the backup protocol go the same way: none of them has any business crossing a boundary.\nIPv6. None of this exists there, because there is no translation and so nothing for a helper to rewrite. A host has its own address, the address in the payload is true, and a stateful firewall permits what you told it to and nothing else. Every problem in this post descends from address translation, and that descends from not deploying IPv6 — an argument I have made elsewhere and will not repeat.\nNow the part I am not going to be diplomatic about.\nIf somebody tells you to turn these on — a supplier, a telephony installer, a managed service, an integrator on a framework — they are not a network engineer and they are not a security specialist. They may be very good at the thing they actually do, and this will not be it. The correct answer to a phone system that needs a firewall to rewrite its signalling is to fix the phone system, and anyone telling you to open your boundary to a string-matching parser instead either does not know what an expectation is or does not care. We are not amateur hour. There has been a right way to do this in standards track documents since 2007, and \u0026ldquo;just enable the SIP helper\u0026rdquo; is the sound of somebody reaching for whatever closes the ticket today.\nAsk them, in the room, what the expectation\u0026rsquo;s source address wildcard is set to. If the question lands as a surprise, you have your answer, and it was never about the protocol.\nIf You Are Still Running Them, Can You Call Yourself A Professional? That is a genuine question and it deserves a genuine answer, so here are three, because there are three cases. It turns on whether you know, and knowing is not something that happens to you. Making sure you know is the job.\nIf there is a SIP or H.323 or FTP helper enabled on a boundary you are responsible for, and you cannot say without looking it up what an expectation is, which of its fields are wildcards, who supplies the values, and what your certification submission claims about inbound rules — then no. Not on this. You have not chosen a configuration, you have inherited a default and never read it. The failing is not the gap; everybody has gaps and I have had this one. It is building a boundary across a gap you never closed, then signing something saying the boundary is sound.\nIf you know exactly what it does, and it is on because a regulator names the protocol, or a partner\u0026rsquo;s kit terminates nothing else, or the phone system is replaced in March and this has to hold until then — yes, obviously, and you are doing the job properly. Those are real constraints and I have worked round worse. What makes it professional rather than negligent is that you wrote down which helper, on which interface, for which flow, why, and the date it comes off. Which is the paperwork the firewall control was asking for anyway.\nThe indefensible case is the middle one. Knowing enough to be uneasy, and leaving it because nobody has made you justify it. That is not engineering. It is habit with a change number attached, and it is how a feature a Best Current Practice told you to turn off in January 2007 is still on in 2026.\nThis matters most when you are paying somebody for judgement, because you cannot inspect judgement on delivery. So inspect it before you sign. Ask what their standard build does with protocol helpers and why, ask what happens if somebody on the guest network opens a link, and ask which internal devices would be reachable with the H.323 helper left on — then watch whether they say \u0026ldquo;only the machine that clicked\u0026rdquo;, because that is the wrong answer a competent-sounding person gives. You will know inside two minutes whether you are being told something or read to, and two minutes is a much cheaper test than an incident.\nIf the answer comes back as a shrug and you sign anyway, that is a decision as well. It has just stopped being theirs and become yours.\nNobody Was Ever Made To Justify It I want to be fair to the people who built these things, because they deserve it. In 1994 the helper was a reasonable answer to a real problem. Addresses were running short, NAT was the pragmatic fix, a handful of important protocols did not survive it, and the options were to change every FTP client on earth or teach the box to read. They taught the box to read, shipped the safest defaults they could think of, and wrote warnings into the source about what it could be made to do. Those warnings are still there. I quoted one.\nWhat went wrong afterwards is not a technical failure. It is that nothing in this industry ever required anybody to look at it again. The IETF said turn them off and had no power to make anyone. The kernel changed its default and could not reach the devices already shipped. Researchers proved it four times in sixteen years, and each time the fix landed somewhere other than the firewall. Meanwhile the default stayed on — not because anyone defended it, but because a default nobody argues about survives indefinitely, and no department anywhere has ending things as its job.\nThat is the pattern, and it is bigger than one firewall feature. This trade is excellent at maintaining and hopeless at stopping. Maintenance is budgeted, staffed, billable and safe; retirement needs one person to put their name on a change with no upside if it goes well and their name all over it if it does not. So the thing stays, and one day somebody finds that it opens ports to your printer.\nThe tell, for me, is that phrase in Cisco\u0026rsquo;s own documentation: the ALG creates a NAT door. Not a filter. Not a control. A door, in the wall you paid for, opened by whoever gets a packet to it with the right words at the front — and the settled response has been to ask the people walking past not to try the handle.\nYou can close yours this afternoon, and that is the part worth ending on. Not the attacks; the attacks are only what happens when nobody does. Go and look at what your boundary permits inbound that you never wrote down, decide whether you meant it, and take out the ones you did not — putting a date next to whatever you kept.\nThat is all a boundary has ever been: a list of things somebody chose to allow, and could say why. Anything on it nobody chose is not security. It is furniture.\nSamy Kamkar — NAT Pinning, 5 January 2010. The original browser-to-ALG attack, using a hidden form to make a browser emit an IRC DCC CHAT or an FTP 227 response line so that the router\u0026rsquo;s helper opens an inbound port. \u0026ldquo;No XSS or CSRF required.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nSamy Kamkar — NAT Slipstreaming, 31 October 2020, updated January 2021. Summarised by the author as letting \u0026ldquo;an attacker to remotely access any TCP/UDP service bound to a victim machine, bypassing the victim\u0026rsquo;s NAT/firewall (arbitrary firewall pinhole control), just by the victim visiting a website\u0026rdquo;. Carries the segment-boundary technique, the note that the SIP handler \u0026ldquo;will bail unless the method (eg, REGISTER) occurs at the start of the data portion of the packet\u0026rdquo;, and the Netgear firmware analysis.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nBen Seri and Gregory Vishnepolsky, Armis — NAT Slipstreaming v2.0, 26 January 2021. The H.323 call-forwarding primitive, the relay bypass of the browsers\u0026rsquo; restricted-port list, the tested product list (OpenWrt, VyOS, consumer Linux routers, FortiGate, Cisco ASAv and csr1000v, HPE vsr1000, SonicWall TZ300), the disclosure timeline, and the conclusion that \u0026ldquo;resolving the issue will require a fundamental change of their implementations by various router/firewall vendors\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 4787 — Network Address Translation (NAT) Behavioral Requirements for Unicast UDP, January 2007, BCP 127. Section 7 carries REQ-10, and the observation that \u0026ldquo;Certain NATs have these ALGs turned on permanently, others have them turned on by default but allow them to be turned off\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPablo Neira Ayuso — netfilter: nf_ct_helper: disable automatic helper assignment, commit 3bb398d9, 25 April 2016, shipped in Linux 4.7. Changes the default of nf_conntrack_helper from enabled to disabled.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nChromium source — net/base/port_util.cc. Its kRestrictedPorts array includes 69, 137, 139, 161, 554, 1719, 1720, 1723, 5060, 5061, 6566 and 10080. Adam Rice\u0026rsquo;s announcement of the SIP ports, 5 November 2020: \u0026ldquo;a carefully-crafted HTTP request to port 5060 on an attacker\u0026rsquo;s server can fool some NAT devices into treating it as a SIP packet and setting up port forwarding.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCisco — SIP ALG Hardening for NAT and Firewall, IP Addressing Configuration Guide, Cisco IOS XE 17.x. \u0026ldquo;SIP ALG creates a firewall pinhole or a Network Address Translation (NAT) door based on the first value in the Via header field.\u0026rdquo; NAT support for SIP is \u0026ldquo;enabled by default on port 5060\u0026rdquo;. See also Using Application-Level Gateways with NAT, which states SIP and H.323 are on by default.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPalo Alto Networks — Disable the SIP Application-level Gateway (ALG). \u0026ldquo;SIP ALG creates dynamic NAT pinholes but may interfere with VoIP applications that have NAT traversal capabilities, causing communication failures.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 2663 — IP Network Address Translator (NAT) Terminology and Considerations, August 1999. Section 2.9 is where the Application Level Gateway is defined.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 3234 — Middleboxes: Taxonomy and Issues, February 2002. The layer-violation line is section 2.11; the cost of extra boxes in the path is section 5.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nEric Leblond, Pablo Neira Ayuso, Patrick McHardy, Jan Engelhardt and Mr Dash Four — Secure use of iptables and connection tracking helpers. \u0026ldquo;This system relies on parsing of data coming either from the user or the server. It is therefore vulnerable to attack.\u0026rdquo; Source of the IRC wildcard quote; documents the nf_conntrack_helper sysctl and the CT --helper target.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nLinux kernel source — net/netfilter/nf_conntrack_ftp.c. The loose module parameter defaults to false, guarding the case where the PORT address is not the client\u0026rsquo;s own; the comment names the risk as \u0026ldquo;DMZ machines opening holes to internal networks, or the packet filter itself\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nnetfilter — H.323 conntrack/NAT helper, by the author of the module. Includes the call-forwarding scenario that lets a session refer to a third-party address.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDavid Leadbeater — NAT-Again: IRC NAT helper flaws, August 2022. Demonstrates the ping-echo trigger, notes it can also scan, unmask cloaked users and disconnect them by naming port 0, and recommends: \u0026ldquo;Potentially entirely deprecate and remove nf_conntrack_irc, it\u0026rsquo;s unclear it has much use anymore.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCVE-2022-2663 — \u0026ldquo;An issue was found in the Linux kernel in nf_conntrack_irc where the message handling can be confused and incorrectly matches the message. A firewall may be able to be bypassed when users are using unencrypted IRC with nf_conntrack_irc configured.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCisco — advisory for CVE-2018-15454, 31 October 2018. The mitigation given is no inspect sip on ASA and configure inspection sip disable on FTD; NVD records that at publication, \u0026ldquo;Software updates that address this vulnerability are not yet available.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 4217 — Securing FTP with TLS, October 2005.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 3715 — IPsec-Network Address Translation (NAT) Compatibility Requirements, March 2004. Section 2.1 item (f) covers SPI selection against NAT; section 2.3 is titled Helper Incompatibilities and carries the quoted lines about IKE cookie de-multiplexing and ISAKMP payload parsing.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nJuniper — IKE and ESP ALG, Application Layer Gateways User Guide. Source of the gate description, the note that NAT-T traffic on port 4500 is not processed by the ALG, and the warning that where two clients share a translated address the device \u0026ldquo;will be unable to distinguish and route return traffic properly\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCisco — IPsec Pass Through Inspection, ASA Firewall CLI Configuration Guide 9.20. \u0026ldquo;IPsec Pass Through application inspection provides convenient traversal of ESP (IP protocol 50) and AH (IP protocol 51) traffic associated with an IKE UDP port 500 connection.\u0026rdquo; Not in the default policy; the supplied _default_ipsec_passthru_map \u0026ldquo;sets no maximum limit on ESP connections per client\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 3947 — Negotiation of NAT-Traversal in the IKE, and RFC 3948 — UDP Encapsulation of IPsec ESP Packets, both January 2005.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCole Dishington — netfilter: nf_conntrack: Add conntrack helper for ESP/IPsec, May 2021, third revision. Reviewed on netfilter-devel and not merged; net/netfilter in mainline still carries no nf_conntrack_proto_esp.c.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNCSC and IASME — Cyber Essentials: Requirements for IT Infrastructure v3.3, April 2026. Control 1, Firewalls, applies to \u0026ldquo;boundary firewalls, desktop computers, laptops, routers, servers, IaaS, PaaS, SaaS\u0026rdquo;; the three requirements quoted are its own words.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 3027 — Protocol Complications with the IP Network Address Translator, January 2001. \u0026ldquo;The purpose of this document is to identify the protocols and applications that break with NAT enroute.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nWHATWG Fetch — pull request 1109, the standards change that added the bad-port entries across browsers.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nOpenBSD — ftp-proxy(8). Control connections reach it only because you sent them there: \u0026ldquo;FTP control connections should be redirected into the proxy using the pf(4) divert-to command, after which the proxy connects to the server on behalf of the client.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nFortinet — Technical Tip: Disabling VoIP Inspection. Documents removing the SIP entry from config system session-helper, set default-voip-alg-mode kernel-helper-based, and notes that re-enabling requires a restart.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMozilla — Stopping FTP support in Firefox 90, 20 July 2021, and Google — Deprecations and removals in Chrome 95, October 2021: \u0026ldquo;Use of FTP in the browser is sufficiently low that it is no longer viable to invest in improving the existing FTP client.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/networking/protocol-helpers-turn-them-off/","summary":"A protocol helper — SIP ALG, FTP helper, H.323 ALG, conntrack helper, call it what your vendor calls it — reads the payload of a connection, finds an address and a port written in the text, and opens an inbound pinhole for them. It cannot tell whether that text came from a real FTP client or from a hidden form on a web page, because there is nothing in it to tell. Samy Kamkar proved the browser case in 2010. NAT Slipstreaming proved it again in 2020, and Armis extended it in 2021 to reach any device on your network, not just the machine that clicked. The IETF asked for these off by default in 2007, Linux turned them off in 2016, and the browser vendors ended up shipping a blocked-port list that reads like a directory of conntrack modules. This post walks the mechanism diagram by diagram — the expectation table, the segment-alignment trick, H.323 call forwarding, the IRC helper that fires on somebody else\u0026rsquo;s message — takes in the IPsec pass-through helper, which cannot read ESP at all and steers inbound packets on an SPI it watched go past in the clear, sets the lot against the Cyber Essentials firewall control it plainly fails, and gives the commands to turn it all off on Linux, Cisco, Juniper, FortiGate and MikroTik.","title":"Your Firewall Takes Instructions From Strangers. Turn Off the Protocol Helpers."},{"content":"IPsec was a good idea. Put the encryption at the network layer, underneath everything, and every protocol running over IP inherits confidentiality and integrity without being told. No library to link. No per-application certificate. No rewriting the thing you already shipped. The packet leaves protected and arrives protected, and the routers in between carry it without knowing or caring what is inside.\nThat design rests on one assumption, and it is stated in the standard rather than implied: a packet\u0026rsquo;s address identifies the machine it came from. A security association is looked up on the destination address, the protocol number and the SPI1. The Authentication Header goes further and signs the IP header itself, source and destination included2. The address is not routing metadata to IPsec. It is part of the identity and part of the integrity check.\nThen this industry spent thirty years taking the address away.\nFirst NAT, to make an office fit behind one line. Then carrier-grade NAT, to make several hundred houses fit behind one address, because turning on IPv6 was work and buying a box was procurement — which is the whole subject of We Never Ran Out of Addresses. We Ran Out of Effort. and I will not go back over it here. What matters for this post is the consequence. The one thing IPsec was built on is the one thing the modern access network no longer provides.\nSo IPsec got patched. Wrap the encrypted packet in UDP so a translator has a port to rewrite. Zero the checksum so nothing revalidates it. Send a one-byte packet every twenty seconds, for ever, so a table in somebody else\u0026rsquo;s kit does not forget you exist. Move the identity off the address and onto a name. Abandon the Authentication Header, because it cannot survive a rewritten header by construction. When a hotel blocks UDP, wrap the whole lot in TCP as well3.\nEvery one of those is a real, standardised, vendor-supported fix. Together they are a protocol held in place by its own scaffolding. And the scaffolding is the argument: you do not spend a quarter of a century propping something up because it is fundamentally sound.\nThe conclusion I have come to is that IPsec should be retired in full. Not tuned, not re-proposed with better ciphers, not kept for site-to-site because that bit still works. Retired, with dates, the way PPTP should have been retired a decade before anyone got round to it. What follows is the evidence, the diagrams, the vendor documentation that says all of this in the vendors\u0026rsquo; own words, and — because most people reading this still have to keep the things running on Monday — a working method for diagnosing IPsec faults in the meantime.\nWhat It Actually Costs You, In Tickets Before the standards, here is the bill, in the order you will meet it.\nThe tunnel drops on a clock. Every hour, or every eight, or after twenty minutes of no traffic. It comes back when somebody opens a file, so half the users never report it and the other half report it as \u0026ldquo;the VPN is slow\u0026rdquo;. Nobody has changed anything.\nTwo people in the same house cannot both connect. The second one comes up, the first one goes down. They ring the service desk separately, so the tickets never meet and nobody spots the pattern for a fortnight.\nSmall things work and big things hang. Login works. Teams works. Ping works. Copying a file stalls at the same place every time, and opening a large page in an internal application sits there until it times out.\nThe tunnel is up and no traffic passes. Both ends say established. Both ends are pleased with themselves. Nothing moves.\nNothing can reach in. Site-to-site to the branch that moved onto a fibre altnet no longer establishes in the direction it used to, and nobody can say why, only that \u0026ldquo;their IP changed\u0026rdquo;.\nTurning on QoS broke encryption. Somebody prioritised voice, and now the far end is discarding packets as replays.\nNot one of those is a misconfiguration in the ordinary sense. Every one of them is IPsec meeting the network as it now is. The rest of this post is why, in order, and how to prove which one you have got.\nIPsec Is Not Carried By The Network. It Is The Network. Start with what it was, because the design is genuinely good and the failures only make sense against it.\nESP is not a protocol running over TCP or UDP. It is a transport protocol, IP protocol number 50, sat directly on IP in the same slot TCP and UDP sit in. AH is protocol 51. Neither has a port field, because neither needs one: on the internet IPsec was designed for, the destination address already names exactly one machine, and the SPI in the ESP header names which security association on that machine. Address plus protocol plus SPI. That triple is the lookup1.\nThe design as specified: the address names the machine, so no ports are needed and either end can start IPsec as specified: the address is the identity Host A 203.0.113.10 Routers forward only Host B 198.51.100.7 either end can start What leaves the machine outer IP header src 203.0.113.10 to 198.51.100.7 protocol 50 = ESP SPI 4 B sequence 4 B encrypted payload your packet, unreadable in transit ICV 16 B Association lookup = destination address + protocol + SPI. No port field anywhere, because none is needed. Three things this design assumes, all of them true in 1995: The address on the packet is the machine. Nothing in the path rewrites a header. Either end may be dialled. The Authentication Header goes further still and signs the outer header itself, source and destination included, so a receiver can prove the addresses were not touched on the way. Not to scale. ESP overhead varies with the cipher; the figures here are AES-GCM in tunnel mode over IPv4. The design as specified. ESP sits directly on IP as protocol 50, with no ports, because it does not need them — the destination address names the host and the SPI names the association on it. AH signs the IP header itself. Both ends hold a real, reachable address, either end can start the conversation, and nothing in the middle needs to understand the payload. Look at what that buys. There is no handshake port to expose, no session layer to get wrong, no application that has to opt in. Either end can start. The middle of the network is dumb, which is exactly what the middle of a network should be. A router forwards protocol 50 the same way it forwards protocol 6, and the fact that it cannot read the payload is the point rather than a limitation.\nIt is a clean design. It is also, in 2026, a description of an internet that most people reading this cannot buy.\nThen Somebody Put A Translator In Every Path A NAT rewrites the source address, and often the source port, so several machines can share one address. That is all it does. Against IPsec it is close to a complete demolition, and the IETF was straight enough about it to publish an entire document listing the pieces: RFC 3715, IPsec-Network Address Translation (NAT) Compatibility Requirements4. Sixteen separate incompatibilities. Here are the ones that matter.\nAH is finished, by construction. In RFC 3715\u0026rsquo;s own words: \u0026ldquo;Since the AH header incorporates the IP source and destination addresses in the keyed message integrity check, NAT or reverse NAT devices making changes to address fields will invalidate the message integrity check.\u0026rdquo;4 There is no fix for that and there was never going to be. A protocol that signs the header cannot cross a box whose whole job is rewriting the header. AH did not get worked around. It got abandoned.\nThere are no ports to translate. A NAT doing port translation needs a port. ESP has not got one. Cisco put it plainly in the Catalyst configuration guide: \u0026ldquo;If PAT found a legislative IP address and port, it would drop the Encapsulating Security Payload (ESP) packet.\u0026rdquo;5 The packet is not rejected on policy. It is dropped because the box has nowhere to write the thing it needs to write.\nThe identity stops matching the packet. Again from RFC 3715: \u0026ldquo;Where IP addresses are used as identifiers in Internet Key Exchange Protocol (IKE) Phase 1 or Phase 2, modification of the IP source or destination addresses by NATs or reverse NATs will result in a mismatch between the identifiers and the addresses in the IP header.\u0026rdquo;4 An identity scheme that names machines by address cannot survive a device that renames machines for a living.\nTwo machines can pick the same SPI. The SPI is chosen by the receiver and only has to be unique to it. Put two hosts behind one address and the translator has two associations with nothing to tell them apart.\nNothing can come in first. A NAT builds its table from outbound packets. There is no outbound packet until someone starts, and on a NATted link only the inside can start. Half the protocol\u0026rsquo;s symmetry is gone.\nOne translator in the path breaks five things at once, and every fix for them costs something One translator in the path, and five things break at once Host 192.0.2.20 Translator rewrites source to 203.0.113.5 Gateway 198.51.100.7 nothing can start from this side What the rewrite destroys 1 The signed header no longer matches. AH covers the source and destination addresses, so the integrity check fails by construction. There is no fix. AH is simply not usable across a translator. 2 There is no port to rewrite. ESP is IP protocol 50 and carries no port field, so a translator doing port translation has nothing to work with and drops the packet. 3 The identity stops matching the packet. The key exchange named the peer by its address. The address in the header is now somebody else's. 4 Two hosts can choose the same SPI. The receiver picks it and only guarantees it is unique to itself, so the translator has two associations it cannot tell apart. 5 Half the symmetry is gone. The table is built from outbound packets, so only the inside can start. The compromise, and what each part of it costs Wrap the whole ESP packet in UDP on port 4500, so the translator has something it understands. Abandon the Authentication Header entirely, because nothing can save it. Zero the UDP checksum, because a checksum over rewritten addresses would only fail. Move the identity off the address and onto a name the peer asserts instead. Send one byte every twenty seconds, for ever, so a table in somebody else's kit does not forget you. It works. That is not in dispute. The argument is that nothing sound needs five concessions to cross one box. The same tunnel with one translator in the path. Five things break at once: the signed header no longer matches, there is no port for the translator to rewrite, the identity no longer matches the source address, two hosts can choose the same SPI, and nothing outside can start a conversation. The row underneath is what the industry did about it — and each fix is something given up. The Fix Was Real, And Every Part Of It Cost Something NAT traversal works. That is not in dispute and I am not going to pretend otherwise. I have run plenty of tunnels through plenty of NATs. What I want on the record is the bill, because it is paid every day by everyone and almost nobody itemises it.\nThe mechanism is RFC 3948: detect a translator during the key exchange, then put the whole ESP packet inside a UDP datagram on port 4500 so the translator has something it understands6. Cisco\u0026rsquo;s description of the wire format is exact: after encryption, \u0026ldquo;a UDP header and a non-IKE marker (which is 8 bytes in length) are inserted between the original IP header and ESP header\u0026rdquo;5. Juniper\u0026rsquo;s is shorter and says the same thing: \u0026ldquo;NAT-T encapsulates both IKE and ESP traffic within UDP with port 4500 used as both the source and destination port.\u0026rdquo;7\nNow the itemised bill.\nThe Authentication Header is gone. Not deprecated politely. Unusable. RFC 8221 now says plainly that using ESP and AH together is NOT RECOMMENDED8, and the honest reason is that the one thing AH did that ESP does not is the one thing NAT destroys.\nThe UDP checksum is deliberately zeroed. RFC 3948 requires it: with the addresses rewritten, a checksum computed over them would fail, so the standard\u0026rsquo;s answer is to stop computing it. \u0026ldquo;If the protocol header after the ESP header is a UDP header, set the checksum field to zero in the UDP header.\u0026rdquo;6 A layer of error detection removed to keep a layer of address rewriting working.\nYou now send packets to keep a table warm. RFC 3948 defines a keepalive as one byte, 0xFF, sent when nothing else has gone out for a configurable interval whose default is twenty seconds6. Juniper states the reason without decoration: \u0026ldquo;Because NAT devices age out stale UDP translations, keepalive messages are required between the peers.\u0026rdquo;7 So a laptop on a train, doing nothing at all, transmits every twenty seconds, for ever, because a table in a box neither end owns will otherwise forget it exists. Multiply by a fleet. That is the radio waking up, the battery going down, and the mobile network carrying traffic whose entire purpose is to prevent forgetting.\nYou now firewall two ports instead of one protocol, and Cisco\u0026rsquo;s own restriction list requires static translation rules for both 500 and 4500 to be in place before any of it works5.\nAnd read the rest of that restriction list, because it is the vendor telling you the shape of the thing. Dynamic NAT policies: unsupported. IPv6 traffic: incompatible with the feature. IPsec and NAT on the same device: cannot both operate5. That is not a configuration guide. That is a list of places the scaffolding does not reach.\nNowt about that is elegant, and none of it is anybody\u0026rsquo;s fault in particular. It is what happens when you keep a layer-3 protocol alive on a network that stopped honouring layer 3.\nCarrier-Grade NAT Removed What Was Left Ordinary NAT took the address off the machine and gave it to the site. You still owned the translator, so you could forward a port, pin a mapping, lengthen a timer, or put the concentrator in front of it.\nCarrier-grade NAT takes the address off the site and gives it to several hundred strangers, and the translator belongs to your ISP. Everything you used to be able to do about it, you now cannot.\nAnd be clear about what is actually in the path, because this is the bit people get wrong. The router in the house is still doing NAT. It has not been switched off. It still translates your laptop on 192.168.1.20 to whatever address the line holds — except the line now holds 100.64.12.7, which is shared space, not a public address. The carrier then translates that again. So the packet crosses two translators before it reaches the internet, and that is the ordinary case, not an unusual one.\nCarrier-grade NAT means two translations, two timers, and no reachable address at either of them Carrier-grade NAT is two translations, not one Laptop 192.168.1.20 Router in the house translation one The line 100.64.12.7 Carrier translator translation two, to 203.0.113.9 yours, and the only table you can see theirs, shared with several hundred homes, invisible to you nothing from the internet can reach either of those addresses Two tables, two timers, and the shorter one wins. Your keepalive has to beat whichever of the two ages out first, and you can only read one of them. The one that matters is the one you cannot see. Port forwarding still works and achieves nothing. The router happily forwards a port from an address the internet cannot reach. UPnP and PCP report success and open a door onto a corridor. This is why \"I have forwarded 500 and 4500 and it still will not come up\" is such a common and such a misleading ticket. The responder role is gone. Two branches on consumer fibre cannot dial each other at all. Something in the middle has to introduce them, and now you depend on a company you never chose. Identity cannot be the address. Several subscribers arrive at the far end as one address, so whatever tells them apart, it is not the field IPsec was designed to use. And there is a published ceiling: on the larger SRX platforms Juniper state no more than 1,000 tunnels from any one translated address. The quickest confirmation costs nothing: read the WAN address in the router. If it starts 100.64, that is the shared space set aside in RFC 6598, you are behind carrier-grade NAT, and half the settings page is decorative. A business line with a real static address has one translation and a reachable endpoint. That is what the extra monthly charge is actually selling you: the thing every machine used to have for nowt. The packet is translated twice: once by the router in the house, once by the carrier. Neither address it holds is reachable from outside. That means two mapping tables with two independent timers where only one is visible to you, port forwarding that succeeds and achieves nothing, no responder role at all, and a peer identity that can no longer be an address. You are double-NATted, and only one of the tables is yours. Two translations means two mapping tables, two ageing timers, and two chances of the mapping going away. Your keepalive has to beat whichever expires first, and you can read exactly one of them. Worse, the two interact: the router in the house may have its own idea about IPsec and try to help with an application gateway, so the source port your client thinks it is using is not the one leaving the house, and is not the one leaving the carrier either.\nPort forwarding still works, and does nothing at all. This is the ticket that eats the most time. Somebody forwards UDP 500 and 4500 on the home router, the router accepts it, the settings page says the rule is active — and nothing can use it, because it forwards from an address the internet cannot reach. UPnP and PCP behave the same way: the client asks for a mapping, the router grants it, and the port opens onto a corridor. Everything reports success and nothing works. The box will tell you owt you want to hear.\nThe quickest way to settle it costs nothing. Read the WAN address in the router. If it begins 100.64, that is the shared space set aside in RFC 6598, you are behind carrier-grade NAT, and half that settings page is decorative.\nThe responder role no longer exists. A device behind CGNAT cannot be the end anybody dials. Site-to-site between two branches on consumer fibre — normal, cheap, and exactly what a small business wants — requires at least one end to hold a real address, or a third party in the middle to introduce them. That third party is a company you now depend on because your ISP would not give you an address.\nThe timer belongs to someone else. RFC 4787 tells NAT operators a UDP mapping \u0026ldquo;MUST NOT expire in less than two minutes\u0026rdquo; and recommends five or more9. That is the floor and the advice, not a promise, and you cannot inspect what your carrier actually does. As such the keepalive stops being a tuning option. It is load-bearing, on a timer you are guessing at.\nYour peer\u0026rsquo;s identity cannot be its address. Several subscribers reach the far end as one address. Whatever the concentrator uses to tell them apart, it is not the IP header — which is the thing IPsec was designed to use.\nThere is a hard ceiling, and the vendors publish it. Juniper document that on SRX5400, SRX5600 and SRX5800, \u0026ldquo;the total number of tunnels from a given public translated IP cannot exceed 1000 tunnels\u0026rdquo;7. Read that as an operator rather than as a spec line. Your VPN concentrator has a per-shared-address limit, the sharing is done by a carrier you have no contract with, and how close you are to that limit depends on how many of your users happen to sit behind the same one. There is no counter you can look at. There is only the day it starts failing for some people and not others.\nEverything in this section is downstream of one decision this country made and kept making. I have written that case out in full elsewhere and I am not repeating it. The point here is narrower: IPsec\u0026rsquo;s core assumption was deleted by the access network, and IPsec has been living on workarounds ever since.\nAnd The IPv6-Only Network Does Not Save It Either Here is where I have to be honest against my own argument, because the obvious reply to everything above is: fine, so give every machine a real IPv6 address and IPsec works as specified again.\nIt does, between two ends that both have one. That is not the network most people are on.\nThe IPv6-only access networks that actually exist, mobile in particular, reach the IPv4 internet through NAT64, which is not a NAT at all in the ordinary sense. It is a protocol translator, rewriting an IPv6 packet into an IPv4 one. And RFC 6146 names what it will carry, and names what it will not, without any ambiguity at all:\n\u0026ldquo;The current specification only defines how stateful NAT64 translates unicast packets carrying TCP, UDP, and ICMP traffic. Multicast packets and other protocols, including the Stream Control Transmission Protocol (SCTP), the Datagram Congestion Control Protocol (DCCP), and IPsec, are out of the scope of this specification.\u0026rdquo;10\nIPsec, by name, out of scope. And packets carrying anything outside that list \u0026ldquo;SHOULD be discarded\u0026rdquo;10.\nSo on an IPv6-only phone or laptop trying to reach an IPv4 concentrator, and that is most concentrators, native ESP does not get dropped by a firewall or mangled by a translator. It is never carried in the first place. The translator is doing precisely what its standard tells it to do.\nThe industry\u0026rsquo;s answer to that is, inevitably, another layer: 464XLAT, which gives the device a local IPv4 stack and translates twice, out of IPv4 and back, so that things NAT64 cannot carry work anyway11. It is already on the list of bolt-ons in the IPv6 post and I am not going to re-argue it. What is worth saying here is the shape: a protocol that broke on NAT in 2004 also breaks on the translation built for the IPv6 transition in 2011, and both are answered by wrapping it in something else.\nThat is the test a protocol has to pass to belong in 2026. Does it work on the network people actually have — behind a carrier\u0026rsquo;s shared address, on an IPv6-only mobile network, through a hotel that only passes TCP 443? Anything built on a UDP port passes all three without being told. IPsec needs a different workaround for each, and on the third one it needs TCP encapsulation as well3.\nNo Ports Also Means No Second Link Here is a failure that has nothing to do with NAT, is purely modern, and gets almost no airtime.\nRouters and switches spread traffic over parallel paths. Link aggregation, equal-cost multipath: both work the same way, by hashing the five-tuple. Source address, destination address, protocol, source port, destination port. ESP has no ports. So every packet of a tunnel between the same two addresses hashes identically, and the whole tunnel lands on one member link, regardless of how many you bought.\nTwo 10G links and one IPsec tunnel gives you 10G. Four gives you 10G. The kit is working exactly as designed.\nThe workaround is the usual shape. Some silicon can hash on the SPI instead, since each SPI names one association and therefore one flow, but that is a feature you have to have bought rather than something you can assume of a path you do not own. There is an active IETF draft whose entire purpose is to wrap ESP in yet another UDP header so that ordinary routers can hash it, and it says why in one sentence: \u0026ldquo;Although the ESP SPI field within the IPsec packets can be used as the load-balancing key, but it cannot be used by legacy switches and routers.\u0026rdquo;12 Its problem statement is equally blunt about what people do instead: \u0026ldquo;Many cloud service providers allow customers to establish multiple IPsec VPN tunnels in parallel to enable ECMP and increase aggregate bandwidth. However, this approach is not ideal, as each tunnel typically requires its own public IP address, leading to higher public IP consumption and increased operational overhead.\u0026rdquo;12\nSit with that for a second. The recommended way to make an encrypted link go faster is to build several of them, each burning a public IPv4 address — during an address shortage — because the protocol has no port number to hash on. Meanwhile a UDP-based tunnel gets multipath for free, from hardware that shipped fifteen years ago, because it has a port like everything else on the modern internet.\nThat is not a legacy problem waiting to age out. It is a live limit on new builds, today, at exactly the speeds people are now buying.\nThe MTU Tax, And Who Pays It Every tunnel costs bytes. IPsec costs more than most, and the way it fails when it runs out is the worst kind of fault: intermittent, size-dependent, and invisible to every test anybody runs first.\nWhat each tunnel spends per packet before any of your data goes in Bytes gone before your data starts, to scale headers, per packet total Native ESP tunnel, AES-GCM IP 20 ESP 8 IV 8 trailer + tag 18 54 B ESP in UDP port 4500, for NAT IP 20 UDP 8 ESP 8 IV 8 trailer + tag 18 62 B L2TP/IPsec AES-CBC, NAT-T IP 20 UDP 8 ESP 8 IV 16 UDP 8 L2TP 6 PPP 4 trailer + tag 18 88 B PPTP GRE, protocol 47 IP 20 GRE 16 PPP 4 no integrity tag at all 40 B WireGuard one UDP port IP 20 UDP 8 header 16 tag 16 60 B The shaded blocks are the price of somebody else's translator. In the L2TP row, three of them sit inside the encryption: a second UDP header, a session layer, and the framing from a dial-up modem. PPTP is cheapest because it protects nothing. The 40 bytes buy no integrity check, and MS-CHAPv2 was reduced to a single DES operation in 2012. Cheap is not the measure. What the bytes buy is the measure. One worked configuration each, over IPv4. Exact figures move with the cipher, the mode and the address family. Bytes spent on every packet before any of your data goes in, for one worked configuration each. Native ESP is lean. Wrapping it for NAT costs eight more. L2TP/IPsec carries a UDP header, an L2TP header and a PPP header inside the encryption — dial-up framing, encrypted, in 2026. PPTP looks cheap because it does not carry an authentication tag at all, which is the whole problem with it. Work it through for one configuration rather than hand-waving. ESP in tunnel mode over IPv4 with AES-GCM: 20 bytes of outer IP header, 8 of ESP header, 8 of nonce, 2 minimum of trailer, 16 of integrity tag. 54 bytes before any of your packet goes in. Wrap it for NAT traversal and the UDP header makes it 62. On a 1500-byte path that leaves 1438, and the moment anything upstream is running PPPoE at 1492 you are down again.\nThen the bit that makes it a fault rather than an arithmetic problem. A sender finds out a packet was too big only by receiving an ICMP error back — Fragmentation Needed on IPv4, Packet Too Big on IPv6. If anything on the path drops those errors, the sender never learns, and it keeps sending packets that keep dying. That is the classic black hole: the handshake completes because handshakes are small, and the transfer hangs because transfers are not.\nCisco have maintained a whole document about this since the GRE and IPsec era, and it is still one of the better explanations of the interaction in anyone\u0026rsquo;s library13. The reason it needs maintaining is that people keep blocking ICMP wholesale and then wondering why tunnels behave oddly.\nI have written the two halves of that argument out already and they hold here without restating: which ICMP messages are load-bearing and which one is not, in Ping: The Diagnostic Tool That Opens a Whole Lot More, and how to find the exact hop that is eating your traffic, in The Firewall Is Eleven Hops Away. The short version for this post: the errors are the mechanism, echo is not, and a border policy that drops all of ICMP has broken your VPN in a way that will be blamed on the VPN.\nAnd Your Own Quality Of Service Can Break It One more, because it catches good engineers doing the right thing.\nESP carries a sequence number and the receiver keeps a replay window of 64 packets by default on Cisco platforms14. Now prioritise voice on the sending router. Low-latency queueing does what you asked and reorders packets relative to the sequence they were encrypted in. If a packet falls outside the window by the time it arrives, the far end discards it as a replay, and the counter that goes up is a security counter.\nCisco\u0026rsquo;s own words: \u0026ldquo;Certain QoS features, such as Low Latency Queueing (LLQ), could cause IPsec packet delivery to become out-of-order and dropped by the receiving endpoint due to a replay check failure.\u0026rdquo;14 Their answer is to widen the window to 1024 where the platform supports it, or to adopt a multiple-sequence-number-space extension that maps QoS classes to separate sequence spaces within one association14.\nSo: turning on a standard feature of your own network breaks your own tunnel, and the remedy is another protocol extension. It is the same shape as everything above.\nL2TP: Dial-Up Framing, Encrypted, In 2026 L2TP is not a security protocol and never claimed to be. RFC 2661 is a tunnelling protocol for carrying PPP sessions, published in 1999, with no confidentiality of its own — its own security section points you at IPsec for packet-level protection15. That pairing is RFC 3193, and the result is the stack in the diagram above: an outer IP header, a UDP header for NAT traversal, ESP, then inside the encryption another UDP header on port 1701, an L2TP header, and a PPP header.\nPPP. The framing from dial-up modems, carried inside an encrypted tunnel, over the internet, in 2026, because that is what L2TP was written to carry.\nCount what that costs on a worked example — outer IP 20, UDP 8, ESP header and IV 24, inner UDP 8, L2TP 6, PPP 4, trailer and tag 18 — and you are spending about 88 bytes per packet to move data that native ESP moves for 54. Thirty-four bytes, every packet, for a session layer that adds nothing you wanted and a link layer designed for a telephone line.\nThe overhead is the least of it.\nIt multiplies the NAT problem rather than dividing it. L2TP/IPsec conventionally uses ESP in transport mode, which is the mode NAT hurts most, and the well-known result is that many implementations cannot support two clients behind one address at all. Two people in one house, or forty in one office, or several hundred behind a carrier\u0026rsquo;s shared address. The far end sees one address and cannot tell the sessions apart, so the second connection replaces the first. That is the ticket from the top of this post, and it is not a bug in anybody\u0026rsquo;s product — it is the identity-by-address problem arriving in the place where the most people meet it.\nAnd in the field it is usually deployed with one shared secret for everybody. Because the secret is configured on the client profile and handed out with the setup instructions, the pre-shared key is in the onboarding document, the wiki page, the email to new starters, and every laptop that ever left. It is not a second factor. It is a password that authenticates the gateway to nobody in particular and has never once been rotated.\nL2TP is not a protocol that has aged badly. It is a protocol that was carrying the wrong thing on day one, and got bolted to IPsec to make up for what it could not do at all.\nPPTP Was Never Secure, And Is Still On Sale PPTP deserves two paragraphs, not a section, and the reason it gets them is that people still ship it.\nIt was never a standard. RFC 2637 is Informational. A vendor protocol written up, not something the IETF ever recommended. It carries PPP inside GRE, IP protocol 47, which like ESP has no ports, so it needs its own special-case handling in every NAT on the path — the \u0026ldquo;PPTP passthrough\u0026rdquo; checkbox, which on a great many home routers supports exactly one session at a time.\nThe security ended publicly in 2012. Marlinspike and Hulton showed that MS-CHAPv2\u0026rsquo;s security reduces to a single DES operation regardless of password length, built chapcrack to extract the handshake, and wired it to a cracking service that returned the key within a day for twenty dollars — a 100% success rate, not a probability16. Their conclusion was that PPTP traffic should be considered unencrypted. Apple agreed with their feet and removed PPTP from the built-in client in macOS Sierra and iOS 10 in 2016, and still publish the warning17.\nTen years on from that, PPTP is still a menu item on routers being sold this year, still in vendor how-to pages, still the thing somebody turns on because it is the one that works first time. It works first time because it is not doing the job.\nEverything That Calls Itself An IPsec VPN \u0026ldquo;IPsec VPN\u0026rdquo; is not a protocol. It is a family, and the length of the list below is the argument, because no two products implement the same subset of it, and the gaps between those subsets are where every interop job you have ever hated actually lives.\nPiece What it adds Where it stands ESP, IP protocol 50 The encryption and integrity itself RFC 4303 — current18 AH, IP protocol 51 Integrity over the IP header too RFC 4302 — unusable through NAT2 IKEv1 The original key exchange Deprecated, RFCs moved to Historic19 IKEv2 The current key exchange RFC 729620 IPComp, IP protocol 108 Compresses before encrypting, with its own associations RFC 317321 PF_KEY v2 A kernel API so a daemon can load the keys RFC 236722 NAT traversal Wraps ESP in UDP 4500 so a translator can cope RFC 3947 / 39486 TCP encapsulation For networks that block UDP as well RFC 9329, replacing RFC 82293 IKEv2 fragmentation Because the key exchange itself outgrew the MTU RFC 738323 MOBIKE So the tunnel survives the address changing RFC 455524 Dead peer detection A heartbeat, because nothing else tells you RFC 370625 XAUTH User authentication — password, token, RADIUS Never an RFC. Expired draft, 200126 Mode-Config Hands the client an address, DNS and routes Never an RFC either26 L2TP/IPsec Carries PPP inside the tunnel RFC 2661 + RFC 319327 GRE or VTI over IPsec Gives you a routable interface to run a protocol on Vendor architecture on top of ESP DMVPN mGRE plus NHRP plus IPsec, so spokes find each other Vendor architecture, NHRP RFC 233228 GETVPN Group keys with no pairwise tunnel at all GDOI, RFC 640729 PPTP The thing IPsec was supposed to replace RFC 2637 — Informational, never a standard30 Now look at the two rows in bold, because they are the ones that should stop you.\nFor the best part of two decades, the standard way to log a user into a corporate IPsec VPN was XAUTH — your username and password, your token, your RADIUS server — with Mode-Config handing the client its address, DNS servers and routes. Between them they are the entire remote-access experience. Every \u0026ldquo;Cisco IPsec\u0026rdquo; client, every VPN icon in a tray, every set of joining instructions.\nNeither of them is a standard. XAUTH was an individual internet draft that expired in 2001 and was archived without ever becoming an RFC26. The stated reason is worth reading, because it is a committee explaining why it would not do its job: the draft records that the IPSRA working group would not accept any protocol extending ISAKMP or IKE, and the IPsec working group refused anything dealing with remote access26. So the most widely deployed part of the most widely deployed VPN protocol was left homeless, implemented anyway by every vendor from their own reading of an expired draft, and shipped to millions of users.\nIKEv2 eventually fixed both, and it is worth saying so: authentication moved to EAP and the client configuration payloads went into the core specification20. But read the dates. The most-used feature of the most-used VPN protocol ran on an expired draft for about a decade before it had a standard at all, and the installed base carried on running the draft version for years after that. A protocol family does not get credit for eventually standardising the part everybody was already using.\nThat is the family you are running. Some of it is standards-track and current. Some of it is Historic. Some of it never got as far as being anything. And a product that says \u0026ldquo;IPsec VPN\u0026rdquo; on the datasheet has told you approximately nothing about which of these eighteen things it does, which is why connecting two of them together is a fortnight, a spreadsheet of proposals, and a phone call to somebody who has done it before.\nThe Fisher-Price OS (Windows) Has Never Really Interoperated This is the part where somebody says the problem is really Linux, so let us do it with sources.\nBy default, the Windows client will not connect to an IPsec server that sits behind NAT at all. Not \u0026ldquo;will struggle\u0026rdquo;. Will refuse. The fix is a registry value called AssumeUDPEncapsulationContextOnSendRule under HKEY_LOCAL_MACHINE\\SYSTEM\\CurrentControlSet\\Services\\PolicyAgent, set to 1 if the server is behind a translator or 2 if both ends are, on every client and on the server, followed by a reboot31. Microsoft\u0026rsquo;s own words: \u0026ldquo;By default, Windows Vista and Windows Server 2008 don\u0026rsquo;t support Internet Protocol security (IPsec) network address translation (NAT) Traversal (NAT-T) security associations to servers that are located behind a NAT device.\u0026rdquo;31\nTwo things about that. The first is that Microsoft co-authored the NAT traversal standard it is refusing to use. Their name is on RFC 3947 and RFC 39486. The second is that the page telling you to edit the registry was last revised in February 202631. Twenty years on, a registry hack on every endpoint is still the answer, and it is still being maintained as the answer.\nAnd read the sentence Microsoft puts just above it: \u0026ldquo;If you must use IPsec for communication, use public IP addresses for all servers that you can connect to from the Internet.\u0026rdquo;31\nThat is this entire post, in the vendor\u0026rsquo;s own documentation. Do not put IPsec behind NAT. Give every machine a real address. They wrote it down, and then the industry spent two decades doing the opposite and billing for the difference.\nIt does not stop at NAT. Go and read what an open-source gateway has to document to accept a Windows client.\nThe gateway certificate needs an extended key usage that exists for this and nothing else. serverAuth, OID 1.3.6.1.5.5.7.3.1, plus IP Security IKE Intermediate, OID 1.3.6.1.5.5.8.2.232. Your CA has to be told to emit an OID most tooling has never heard of, or the connection fails with a policy error and no useful message. A second registry value is needed before the client will offer decent cryptography. strongSwan documents adding NegotiateDH2048_AES256 under Rasman\\Parameters to get AES-256-CBC and a 2048-bit group33. Read that the other way round, which is the way that matters: without a registry edit, the default proposal is weaker than that. Server-initiated rekeying is rejected by clients behind NAT, and the documented workaround is to disable rekeying on the gateway and let the client start it33. Standard IKEv2 extensions are simply absent — no IKE redirection, no multiple authentication rounds33. None of those is a Linux bug. Every one is an open-source project writing down what it has to do to accommodate one vendor\u0026rsquo;s reading of a standard that vendor helped write.\nThat is the pattern, and it is thirty years old. PPTP was Microsoft\u0026rsquo;s protocol, written up as Informational and never standardised30. MS-CHAPv2 and MPPE were Microsoft\u0026rsquo;s authentication and encryption, and both were broken in public16. SSTP is a Microsoft tunnel nobody else terminates. DirectAccess was IPsec, and it was Windows at both ends by design — and it has now been deprecated and is being removed, with customers pushed onto Always On VPN instead34. Not one of those was ever a protocol the rest of us could meet in the middle. They were a protocol you joined, and if you did not run the right operating system on both ends you got the registry hack, the odd OID and the workaround page.\nSo let me be plain, since it is my blog. The Fisher-Price OS (Windows) has never been a genuine peer in an open protocol stack, because being a peer was never what it was for. It hides the machine from the person using it as a design goal, and a stack you cannot see is a stack you cannot make interoperate. If it is the only operating system you have ever administered, the sections above about reading kernel counters and watching the wire will have read like another language, and that is the gap — not a preference, a gap.\nNone of which need cost the reader anything. Every diagnostic in this post runs from any Unix on the network, pointed at whatever is broken, and it does not care in the slightest what the far end is running. And if that operating system is the only one you have, the diagnostic section below carries its own tools for it as well — the capture, the two cmdlets, and the error codes with what each one is really telling you. A broken tunnel still has to be fixed on Monday. Then the replacement I argue for at the end has one client, behaving the same way, on every platform including that one, which is the first time in thirty years that has been true of a VPN.\nThe Deprecation Record Reads Like An Obituary Set the workarounds aside and just read what the standards bodies have done to this family over the years. Not opinion. Requirement levels, in published RFCs.\nWhat Where it stands now Source IKEv1 Deprecated; RFCs 2407, 2408 and 2409 moved to Historic RFC 9395, 202319 DES in ESP MUST NOT RFC 82218 3DES in ESP SHOULD NOT RFC 82218 HMAC-MD5-96 MUST NOT RFC 82218 ESP together with AH NOT RECOMMENDED RFC 82218 Encryption-only ESP Shown insecure, and broken in practice in 2007 Degabriele and Paterson35 IPsec on an IPv6 node Downgraded from MUST to SHOULD RFC 6434, 201136 PPTP Never a standard; Informational only RFC 263730 The last-but-one row is the one I would put in front of anyone who tells me IPsec is fine and the network is the problem. IPv6 originally mandated IPsec — it was the security story, written into the node requirements. In 2011 the IETF changed its mind: \u0026ldquo;Previously, IPv6 mandated implementation of IPsec and recommended the key management approach of IKE. This document updates that recommendation by making support of the IPsec Architecture a SHOULD for all IPv6 nodes.\u0026rdquo;36\nEven the address family that would have given IPsec back everything NAT took away stopped requiring it fifteen years ago. That is not the network failing IPsec. That is the people who designed the network deciding it had not earned the mandate.\nThe Complexity Was Flagged In 1999, In Writing None of this is hindsight, and that is what makes it worth writing down.\nIn 1999 Niels Ferguson and Bruce Schneier were commissioned to evaluate IPsec. Their report is short, plain and worth reading whole. It opens with \u0026ldquo;IPsec was a great disappointment to us. Given the quality of the people that worked on it and the time that was spent on it, we expected a much better result.\u0026rdquo;37 It names the cause: \u0026ldquo;Our main criticism of IPsec is its complexity. IPsec contains too many options and too much flexibility; there are often several ways of doing the same or similar things. This is a typical committee effect.\u0026rdquo;37\nAnd it made three recommendations that read now like a list of things that happened anyway, twenty years late and the hard way:\nEliminate transport mode. \u0026ldquo;We therefore recommend that transport mode be eliminated.\u0026rdquo;37 Transport mode is the mode L2TP/IPsec uses, and it is the mode NAT hurts worst. Eliminate AH. \u0026ldquo;We conclude that eliminating transport mode allows the elimination of the AH protocol as well, without loss of functionality.\u0026rdquo;37 NAT eliminated it instead, for a worse reason. Never allow encryption without authentication. They warned that administrators \u0026ldquo;will be quite likely to configure ESP for only encryption, believing that it provides security.\u0026rdquo;37 Eight years later that last one stopped being a warning. Degabriele and Paterson published attacks that \u0026ldquo;break any RFC-compliant implementation of IPsec making use of encryption-only ESP\u0026rdquo; — ciphertext-only, needing nothing more than the ability to watch traffic and inject packets35. The prediction was in the public record for eight years and the standard still permitted the configuration.\nThe verdict of that 1999 report is the sentence I keep coming back to: \u0026ldquo;We have found serious security weaknesses in all major components of IPsec. As always in security, there is no prize for getting 90% right; you have to get everything right.\u0026rdquo;37\nTwo more data points, and then I will leave it.\nLogjam, 2015. The team behind it scanned a 1% sample of IPv4 for IKE and found that 86.1% of IKEv1 and 91.0% of IKEv2 servers supported the 1024-bit Oakley Group 2, and that 66.1% of profiled IKEv1 servers preferred it. Their conclusion: precomputation against a second 1024-bit group \u0026ldquo;would allow decryption of traffic to 66% of IPsec VPNs\u0026rdquo;, and the published intelligence documents on VPN exploitation are \u0026ldquo;consistent with having achieved such a break\u0026rdquo;38. Cryptographic agility, the thing IPsec has most of, is what let almost everybody sit on the same weak group for fifteen years.\nCVE-2016-1287. A buffer overflow in the IKEv1 and IKEv2 code on Cisco ASA, reachable by sending crafted UDP packets, giving remote code execution before authentication39. Think about where that box sits. It is the device you deliberately exposed to the entire internet, running the most option-laden protocol in the estate, with a fragment reassembler in front of the parser, holding the keys to everything inside. The complexity Ferguson and Schneier warned about is not an abstraction. It is attack surface, on the one machine you cannot put behind anything.\nDiagnosing It While You Still Run It You cannot turn all this off this afternoon, so here is how to work on it. This is the method, in the order that costs least, with what each result actually means.\nSix ways it gives up, placed on the path where each one happens Where each failure actually lives Client policy and routes Home router translation one Carrier translation two The internet filters and MTU Gateway selectors and identity 1 2 3 4 5 6 1 Nothing on the wire at all. The traffic never reached IPsec. Stop looking at crypto. ip xfrm policy and the route, and whatever the host firewall is doing. 2 Port 500 both ways, 4500 never appears. NAT traversal was not negotiated. tcpdump -ni eth0 'udp port 500 or udp port 4500 or ip proto 50' Bare ESP will not survive it. 3 Dies after idle, revives on first use. A mapping aged out at one of the two translators. Your keepalive is losing to a timer you cannot read. Shorten it, and stop trusting the default. 4 Outbound counter rising, inbound flat. ESP is dying in one direction, in transit. ip -s xfrm state at both ends, then find the hop that is eating it. 5 Login fine, large transfers hang. Path MTU, and the ICMP errors are not coming back. ping -M do -s 1400 stepping down, then clamp the segment size and fix the ICMP rule. 6 Both ends say up, nothing passes. Traffic selectors, policy or routing \u0026#8212; not the keys. And if a second user knocked the first one off the tunnel, it is peer identity, not capacity. One capture at the border for thirty seconds answers the first three. Do that before you open either console. The six failures in the order to test for them, with the symptom, the check and what the answer means. Each row is a different layer of the stack giving up, and the first three are answered by watching the wire for thirty seconds — which is why that is the first thing to do, not the last. Watch The Wire First, Not The Console Both consoles will tell you what they believe. The wire tells you what happened. One capture at the border, thirty seconds, answers the first three questions at once:\ntcpdump -ni eth0 \u0026#39;udp port 500 or udp port 4500 or ip proto 50 or ip6 proto 50\u0026#39; Nothing outbound at all — the problem is in front of IPsec: routing, policy, or a host firewall. Stop looking at crypto. Outbound only, nothing back — your packets are leaving and their answers are not arriving. Filtering in transit, a dead peer, or the far end rejecting silently. UDP 500 both ways but 4500 never appears — NAT traversal was not negotiated. Either one end has it disabled or the detection failed. Protocol 50 on the wire while one end is behind NAT — the negotiation decided there was no translator when there is one. It will never come back. Reading The Key Exchange Failure IKEv2 tells you why it refused, and the notify names are specific enough to diagnose from alone. It helps to have the shape of the whole exchange in front of you first, because every notify below belongs to a particular rung of it.\nThe IKEv2 exchange step by step, and which fault lives at each step Every step of the exchange, and the fault that lives on it Client behind a translator Gateway real address What fails here IKE_SA_INIT request \u0026#8212; UDP 500 nothing back \u0026#8212; 809 ERROR_VPN_TIMEOUT IKE_SA_INIT response \u0026#8212; UDP 500 NO_PROPOSAL_CHOSEN, INVALID_KE_PAYLOAD NAT detected, both ends float to 4500 no float \u0026#8212; proto 50 dies at the translator IKE_AUTH \u0026#8212; UDP 4500, encrypted too big, fragment lost, retransmit, timeout IKE_AUTH response \u0026#8212; CHILD_SA created AUTHENTICATION_FAILED \u0026#8212; 13801, 13806 ESP in UDP 4500 \u0026#8212; your traffic TS_UNACCEPTABLE \u0026#8212; up, and nothing moves keepalive \u0026#8212; 1 byte every 20 s, for ever no keepalive \u0026#8212; mapping expires on idle The float is the hinge. Above it everything is UDP 500. Below it everything is UDP 4500. If the float never happens you are putting protocol 50 into a translator that has nothing to rewrite, and it will never come back. The size fault lives in one rung. IKE_AUTH carries the certificate chain, so it is the only large message on this ladder. A pre-shared key that connects where a certificate does not is a dropped fragment, not a bad certificate. Created is not the same as working. The child association can exist and still carry nothing, if the two ends disagree about which traffic it covers. Every fault in the right-hand column is reported at the client as a timeout, whatever it actually was. The exchange from the first packet to steady state, with the fault that lives on each step. The float from 500 to 4500 is the hinge: everything above it is one port, everything below it is another, and if the float never happens you are putting protocol 50 into a translator that cannot carry it. IKE_AUTH is the only large message here, which is why a pre-shared key can connect where a certificate does not — that is a dropped fragment, not a bad certificate. Notify What it actually means Where to look NO_PROPOSAL_CHOSEN Not one of the offered cipher/integrity/DH/PRF combinations is acceptable to the far end Both proposal lists; expect a deprecated algorithm on one side INVALID_KE_PAYLOAD Diffie-Hellman group mismatch — you offered one group, it wants another The DH group, first thing in the proposal AUTHENTICATION_FAILED Key or certificate or identity wrong — a pre-shared key mismatch, an expired certificate, or an ID that is not what the peer expects The identity, not just the secret TS_UNACCEPTABLE The traffic selectors do not overlap — you asked to protect subnets the peer will not protect Both ends\u0026rsquo; selector configuration INVALID_SPI A packet arrived for an association that no longer exists, usually after a one-sided restart Whether one end has rekeyed or rebooted Juniper\u0026rsquo;s Phase 2 guidance says the same thing about the commonest of them: \u0026ldquo;no proposal chosen\u0026rdquo; means \u0026ldquo;the device did not accept any of the IKE Phase 2 proposals that the peer sent\u0026rdquo;, and the fix is a mutually acceptable proposal rather than a repeated restart40.\nOn strongSwan, the state of everything in one command:\nswanctl --list-sas # what is established, and what it negotiated swanctl --log # the negotiation as it happens On Cisco, show crypto ikev2 sa and show crypto ipsec sa, with debug crypto ikev2 when it will not come up41. On Junos, show security ike security-associations and show security ipsec security-associations, with the negotiation in show log kmd-logs42.\nWas NAT Traversal Actually Negotiated? This is the check people skip, and it explains a large share of \u0026ldquo;it works from the office and not from home\u0026rdquo;.\nEach end sends hashes of the addresses and ports it believes are in play. If the hash the far end computes from the packet it received does not match the one you sent, there is a translator between you, and both ends move to UDP 4500. If detection fails — one side has traversal disabled, or something in the middle is mangling the exchange — both ends carry on with bare ESP, which will not survive the translator.\nSo: see port 4500 in the capture, from both directions, or there is no traversal happening. Do not take the console\u0026rsquo;s word for it.\nThen check the keepalive is actually running and its interval is below whatever your carrier ages mappings at. The default is twenty seconds6; the standard\u0026rsquo;s floor for NAT operators is two minutes9; what your particular carrier does, you cannot see. If the tunnel dies after idle and revives on traffic, this is your fault every time.\nThe Kernel Counters Almost Nobody Reads On Linux the transform layer keeps a full error breakdown, and it is the fastest way to turn \u0026ldquo;it does not work\u0026rdquo; into a specific cause. The counters are documented by the kernel itself43.\ncat /proc/net/xfrm_stat # error counters, by cause ip -s xfrm state # per-SA packet and byte counters ip xfrm policy # what should be protected, and in which direction It reads better as a path than as a list. The packet passes through five stages, and each one has its own counter:\nFive stages, and the counter that names the one your packet died at Follow one packet, and let the counter name the stage it died at Your policy does this traffic get protected? XfrmOutPolBlock you are dropping it yourself, on purpose, in policy Your association encrypt, seal, number it XfrmOutNoStates policy matched and there is no association to carry it The path translator, MTU, filters nothing increments, at either end this is where NAT, MTU and filtering live, and no counter can see any of it Their association dest + protocol + SPI XfrmInNoStates \u0026#183; XfrmInStateProtoError \u0026#183; XfrmInStateSeqError it arrived and the crypto did not work out: wrong SPI, wrong key, out of window Their policy was this meant to be protected? XfrmInTmplMismatch \u0026#183; XfrmInNoPols the crypto was fine and the policy disagreed about whether it should have been Stage three is the one with no counter. Everything a quarter of a century of translation did to this protocol happens there, and the transform layer at neither end can see a single packet of it. That is the whole reason a capture comes before a console. Stages four and five are opposite faults. Four means the crypto did not work out. Five means it did, and something disagreed about whether it should have. They get looked at in the wrong order almost every time, because five looks like a crypto fault and is not. Counter names are Linux. On Junos the same stages read out of show security ipsec statistics; on the Fisher-Price OS (Windows), out of Get-NetIPsecQuickModeSA and the Windows Firewall with Advanced Security log. Five stages, and the counter that names the one your packet died at. Stage three is the only one with no counter at all — the translator, the MTU and every filter in between live there, and the transform layer at neither end can see a single packet of it. That is the whole argument for reaching for a capture before a console. Stages four and five are opposite faults and get looked at in the wrong order almost every time. Watch which one moves while the fault happens:\nCounter Kernel\u0026rsquo;s description What it means on the day XfrmInNoStates \u0026ldquo;No state is found i.e. Either inbound SPI, address, or IPsec protocol at SA is wrong\u0026rdquo; Their packets are arriving for an association you do not have — usually a one-sided rekey or restart XfrmInStateSeqError \u0026ldquo;Sequence error i.e. Sequence number is out of window\u0026rdquo; Reordering or replay-window trouble; see the QoS interaction above XfrmInStateProtoError \u0026ldquo;Transformation protocol specific error e.g. SA key is wrong\u0026rdquo; Keys disagree — the association survived a rekey on one side only XfrmInTmplMismatch \u0026ldquo;No matching template for states e.g. Inbound SAs are correct but SP rule is wrong\u0026rdquo; The association is right and the policy is wrong XfrmInNoPols \u0026ldquo;No policy is found for states e.g. Inbound SAs are correct but no SP is found\u0026rdquo; Protected traffic arriving that nothing asked to protect XfrmOutPolBlock \u0026ldquo;Policy discards\u0026rdquo; You are dropping it yourself, on purpose, in policy XfrmOutNoStates \u0026ldquo;No state is found\u0026rdquo; Traffic matched a policy with no association to carry it — the tunnel never came up XfrmInTmplMismatch and XfrmInNoPols are the two worth knowing by sight, because both mean the crypto is fine and the policy is not, which is the opposite of where everybody looks first.\nIt Says Up And Nothing Moves Both ends established, no traffic. Read the per-association counters in both directions:\nip -s xfrm state Outbound bytes rising, inbound flat — you are encrypting and sending, and nothing is coming back. Either your ESP is not reaching them or theirs is not reaching you. Ask the far end for their outbound counter; if it is rising too, the packets are dying in transit and the next question is where, which is a TTL question rather than a crypto one. Both flat — nothing is being offered to the tunnel. Routing or policy, not IPsec. On a route-based setup check the route actually points at the tunnel interface; on a policy-based one check the selectors. Both rising, applications still broken — it is not the tunnel. Go and look at what is on the far side. Juniper\u0026rsquo;s guidance for this case is the same instinct in their idiom: if only the outbound packet counter on the session is incrementing, confirm with the peer whether the traffic is being received at all44.\nSmall Things Work, Big Things Hang MTU. It is always MTU. The test takes ten seconds:\nping -M do -s 1400 10.0.0.1 # inside the tunnel, do-not-fragment set ping -M do -s 1300 10.0.0.1 # step down until it succeeds Where it starts succeeding tells you the real usable size. Then make TCP find out for itself, by clamping the advertised segment size to the path rather than hoping every ICMP error survives the journey:\nnft add rule inet filter forward tcp flags syn tcp option maxseg size set rt mtu And fix the underlying cause too, which is nearly always an over-broad ICMP rule at a border somewhere. Drop echo if you want. I have argued elsewhere that it deserves dropping. But keep Fragmentation Needed and Packet Too Big. They are the mechanism, not a nicety.\nIt Drops On A Clock Time the failures. The interval names the cause on its own.\nA fixed period matching a configured lifetime — rekey. The association expires and the replacement negotiation is failing or racing. Check both ends\u0026rsquo; lifetimes; mismatched values are normal and fine, but a hard lifetime on one end shorter than the other\u0026rsquo;s soft lifetime produces exactly this. After a period of no traffic, back on first use — a NAT mapping expired. Keepalive interval, or the absence of one. Dead-peer detection tearing it down while the link is fine — the probes are being lost rather than the peer being dead, often because the probes are the only traffic and the mapping has already gone. The Second User Kills The First Two peers reaching the concentrator from one address, authenticating as the same identity. The gateway has a choice between keeping the old association and replacing it, and a common default is to replace. So the second connection wins and the first one silently dies.\nGive every peer a genuinely unique identity rather than an address or a shared name, and set the gateway to keep multiple associations from one address rather than assuming one per peer. Then test it the only way that counts: two clients, one address, at the same time. If your acceptance test has never had two users behind one NAT, you have not tested the case that most of your users are in.\nThe Same Faults, On The Fisher-Price OS (Windows) Every check above runs from a Unix box, and if you have one on that network then use it, because it will tell you the truth faster and it does not care what the far end runs. But a lot of readers have a client that will not connect, an operating system that hides the machine from them on purpose, and nothing else to look at. So here is the same method again, in the same order, with the tools that operating system actually ships.\nWatch the wire. There is no tcpdump, but there is a capture. Run it elevated, reproduce the failure, stop it:\nnetsh wfp capture start cab=on file=ipsec netsh wfp capture stop That writes a .cab. Inside it is the trace of what the filtering platform and the key exchange actually did while the fault was happening, which is more than either console will admit to. For a live look rather than an archive, the same command set prints to the console with file=-: netsh wfp show state \u0026ldquo;Displays the current state of WFP and IPsec\u0026rdquo;, and netsh wfp show ikeevents \u0026ldquo;Displays recent Internet Key Exchange (IKE) epoch events matching the specified parameters\u0026rdquo;, filtered to one peer45.\nnetsh wfp show state file=- netsh wfp show ikeevents remoteaddr=203.0.113.5 file=- show ikeevents is the nearest thing here to the exchange log every other stack writes without being asked. Worth knowing it is there. Otherwise the interface hands you a three-digit number and nothing else, and you are diagnosing a key exchange by guesswork.\nRead the associations, and read both of them. The equivalents of ip -s xfrm state and swanctl --list-sas are two cmdlets, and the split between them is the whole diagnosis:\nGet-NetIPsecMainModeSA Get-NetIPsecQuickModeSA Main mode is the key exchange. Quick mode carries the packets. Microsoft puts the relationship plainly: \u0026ldquo;There is only one main mode SA between a pair of computers, but there can be many quick mode SAs\u0026rdquo;46. So main mode present with quick mode empty is the same fault as an IKE_SA up with no CHILD_SA under it, and it means the same thing: the two ends agreed on how to talk and then failed to agree on what to protect. Look at the traffic selectors, not the ciphers.\nRead the logs, in the two places they hide. Connection-level failures land in the Application log against the RasClient source, and Microsoft\u0026rsquo;s own note on reading them is the useful bit: \u0026ldquo;All error messages return the error code at the end of the message\u0026rdquo;47. That number is the diagnosis, and the next section is what the numbers mean. Policy and filtering decisions land somewhere else entirely, under Applications and Services Logs, in the Windows Firewall with Advanced Security channels. Two logs, two teams, one fault.\nCollect it properly when you have to escalate. The supported bundle is TSS — TSS.ps1 -Scenario NET_VPN on the client, TSS.ps1 -Scenario NET_RAS on the server, started before you reproduce the fault and stopped after48. Learn it before somebody asks you for it.\nWhen It Sits On Connecting, And Then Times Out This is the fault that fills the tickets, and the word in the error is a lie. Start by naming the code. The code is specific even when the message is useless.\nA connect timeout is one of three things, and none of them is a clock What the client calls a timeout, and what it actually was The client says: timed out 809, 718, 828, 930, 638 \u0026#8212; five codes, one word, and not one of them is a clock It is one of three things, and one test tells you which Nothing arrived the silence case Test: capture at the client \u0026#8212; does UDP 500 leave and nothing come back? Filtering in the path, or NAT traversal never negotiated, so 4500 was never sent at all. It arrived, too big the fragment case Test: does a pre-shared key connect where a certificate does not? IKE_AUTH carries the chain, so it fragments, and the fragment is dropped in the path. Never the tunnel the wrong-layer case Test: does the gateway log an authentication failure at the same second? 930 is RADIUS. 812 is the authentication method. 13801 and 13806 are certificates. None of the three is a timer. Lengthening the timeout is the one thing the interface invites you to do, and the one thing that has never once fixed any of them. The word is there because the layer reporting the failure cannot see far enough to say more. Test in that order. The first costs thirty seconds of capture, the second costs one connection attempt, the third costs a log. Five error codes carry the word timeout and not one of them is a clock. It is one of three things and a single test separates them: a capture says whether anything came back at all, one connection attempt with a pre-shared key says whether the certificate exchange was simply too big to arrive, and the gateway\u0026rsquo;s own log says whether the tunnel was ever the problem. Test them in that order, because that is also the order of what they cost. Code Name in raserror.h What actually happened 809 ERROR_VPN_TIMEOUT Nothing came back at all. Microsoft\u0026rsquo;s stated cause: \u0026ldquo;the UDP 500 or 4500 ports on the VPN server or firewall are blocked\u0026rdquo;47 — but blocked covers three different things and only one is a deny rule. Usually 4500 was never sent, because NAT traversal was never negotiated, or the helper that used to carry it has been switched off. Read on before you ask for a firewall change 789 ERROR_OAKLEY_GENERAL_PROCESSING \u0026ldquo;The L2TP connection attempt failed because the security layer encountered a processing error during initial negotiations\u0026rdquo; — credentials or certificates, not the network 718 ERROR_PPP_TIMEOUT The IPsec part worked. PPP inside the L2TP tunnel got no answer 828 ERROR_IDLE_TIMEOUT \u0026ldquo;The connection was terminated because of idle timeout\u0026rdquo; — a server-side setting, deliberately 930 ERROR_AUTH_SERVER_TIMEOUT RADIUS did not answer in time. Nothing to do with IPsec whatsoever 638 ERROR_REQUEST_TIMEOUT The generic one. Treat it as no information and go to the wire Every one of those is sourced from Microsoft\u0026rsquo;s own error list49. Now read 809 again. It is called ERROR_VPN_TIMEOUT, and the documented cause is a blocked port. The client is not timing out because the far end is slow. It is timing out because the far end is silent, and silence is the only failure mode a protocol with no ports and no handshake visibility can report.\nBe careful with that word blocked, though, because it is carrying more than it can hold. Three different things wear it, and only one of them is a deny rule.\nNothing ever sent 4500 in the first place. NAT traversal was not negotiated, so the client carried on speaking ESP, and ESP gives a translator nothing to rewrite. The default on that platform is reason enough on its own: without the registry value, it will not form a NAT-T association to a server behind a translator at all31. So 4500 is not blocked. It was never tried.\nThe helper that used to paper over it has been switched off. Firewalls and home routers carry per-protocol helpers — IPsec passthrough on consumer kit, and the whole application-layer gateway family behind it — that read a protocol the NAT cannot handle and open the return path for it. On Linux the automatic form of that was turned off by default at kernel 4.7 \u0026ldquo;for security reasons\u0026rdquo;, with the guidance since being to attach a helper deliberately with a rule or not at all; the helpers it covers include the one for PPTP50.\nAnd switching them off was the right call. A helper is a piece of your firewall that parses a payload and then punches a hole based on what it read. NAT Slipstreaming is the bill for that: a browser visiting a page, traffic shaped so the router\u0026rsquo;s SIP or H.323 helper reads it as a call, and a pinhole opened through the NAT — in the 2021 version, to any internal address, not just the machine that loaded the page51. Turning helpers off closes that. It also stops your IPsec working. Both are true at once, and the second is not an argument for undoing the first.\nWhich is the post in miniature, again. The protocol only worked because boxes in the middle were reading traffic that was not theirs and opening holes on its behalf, and the industry has spent the last decade correctly deciding to stop doing that.\nSo work the causes in this order, cheapest first.\nOne: nothing is coming back. Capture at the client, or ask the gateway. If UDP 500 leaves and nothing returns, it is filtering or reachability and no timer will fix it. If 500 completes both ways and 4500 never appears, NAT traversal did not negotiate. And if either end is behind a translator, you are back at the registry value from earlier in this post: AssumeUDPEncapsulationContextOnSendRule, 1 or 2, on both machines, then a reboot31. Without it the client refuses by design and reports it as a timeout.\nTwo: the answer is too big to arrive. This one wastes whole afternoons. Certificate authentication makes the second exchange large, because it carries a chain, and a large exchange fragments. Fragments get dropped by the same middleboxes as everything else in this post, the client retransmits into the same hole, and then it gives up and says timeout. The tell is diagnostic on its own: a pre-shared key connects and a certificate does not. That is not a certificate fault. That is a size fault, because the pre-shared key exchange is small enough to fit. Standard IKEv2 fragmentation exists precisely for this23, and strongSwan records when this platform got it: \u0026ldquo;IKEv2 fragmentation is supported since the v1803 release of Windows 10 and Windows Server\u0026rdquo;33. Anything older, or a gateway with fragmentation disabled, and you are relying on a path that will carry a fragmented UDP datagram. Plenty will not.\nThree: it connects, then drops on a clock. Time it. If it dies after a fixed idle period, that is 828 and it is configuration, not a fault — -IdleDisconnectSeconds on the server, with -SALifeTimeSeconds, -MMSALifeTimeSeconds and -SADataSizeForRenegotiationKilobytes as the other three clocks that can end a session52. That last one ends it on volume rather than time, which is why a session can die reliably during a large file copy and never during a day of email. And if it dies at rekey with the client behind NAT, that is the documented interop fault from earlier: the client rejects a server-initiated rekey with Microsoft error 13863, and the gateway-side answer is to stop initiating and let the client do it33.\nFour: it is not the tunnel at all. 930 is RADIUS. 812 is an authentication method the server would not accept47. 13801 and 13806 are certificates — wrong extended key usage, expired, missing root, or a server name that does not match the certificate\u0026rsquo;s subject47. The tunnel negotiation was fine in all four cases, and if you spend the afternoon on ciphers you will not find any of them.\nHere is the thing worth taking away from that table. Six error codes, five of them with the word timeout in the name, and not one of them is a timeout in fact. They are a blocked port, a lost fragment, a policy decision and a RADIUS server, all wearing the same word, because the layer reporting the failure cannot see enough of what happened to say anything more useful. Lengthening the timer fixes none of them, and lengthening the timer is what the interface invites you to do.\nWhat To Run Instead I am not going to pretend the replacement is exotic. It is in the kernel and it has been for years.\nWireGuard is one UDP port, one key per peer, no cipher negotiation and no protocol agility at all. The author\u0026rsquo;s position on that is deliberate and stated: \u0026ldquo;It intentionally lacks cipher and protocol agility. If holes are found in the underlying primitives, all endpoints will be required to update. As shown by the continuing torrent of SSL/TLS vulnerabilities, cipher agility increases complexity monumentally.\u0026rdquo;53 That single decision deletes NO_PROPOSAL_CHOSEN, INVALID_KE_PAYLOAD, downgrade attacks and the Logjam finding in one go, because there is nothing to negotiate and nothing to downgrade.\nIt is honest about the trade too, and so am I: no agility means that when a primitive does fall, you update the entire fleet rather than flipping a config line. That is a real operational cost and it is the right one to pay.\nThe rest lines up against the list above almost item for item. Having a UDP port means NAT and CGNAT treat it like any other flow, and ECMP and LAG hash it like any other flow. Roaming is built in rather than bolted on — an authenticated packet from a new address moves the peer\u0026rsquo;s endpoint, so a phone going from WiFi to mobile does not renegotiate anything. It answers nothing to an unauthenticated packet, so a scanner finds a closed port where IPsec would find a concentrator to talk to. And it is under 4,000 lines of code53 against a stack that needs a document listing its sixteen incompatibilities with one middlebox.\nOn performance it simply measured faster than both IPsec configurations it was benchmarked against, at 1,011 Mbit/s against 881 and 825, with lower latency53. I would not retire a protocol over a benchmark. I mention it because the last argument standing for IPsec is usually performance, and it is not true either.\nFor getting a person to an application — rather than a network to a network — the answer is not a tunnel at all. Identity at the front door, the application published through it, nothing routed. I have built that out with Proxmox and Cloudflare Access in Zero Trust VDI Without the Cloud Bill, and the relevant point for this post is that a user who needs three internal applications does not need a route to your entire estate.\nAnd say the quiet part about hardware. Yes, there are NICs and ASICs with ESP offload, and that is a genuine argument for IPsec on specific kit at specific speeds. It is an argument about silicon someone already sold you, not about the protocol being right. As such it dates, and it dates fast. PPTP survived in exactly the same way, for exactly as long as the checkbox existed.\nRetiring It Properly Retirement is a plan with dates, not a feeling. Here is the one I would put my name to.\nStop new IPsec deployments now. Not \u0026ldquo;prefer alternatives\u0026rdquo;. Stop. Every new tunnel is a tunnel someone has to migrate later, and the ones being built today will still be running in 2035 unless somebody says no this week.\nDo remote access first, because that is where every failure in this post lands hardest: the CGNAT, the shared address, the keepalives, the second user in the same house, the MTU black hole on a random hotel network. It is also the easiest to move, because the endpoints are managed and the change is a client.\nThen site-to-site over the public internet, which is the same set of problems with fewer endpoints and a maintenance window.\nLeave for last the tunnels you do not own both ends of: a partner, a regulator, a carrier\u0026rsquo;s managed service. Those move when the contract moves, and the way to make them move is the next point.\nStop buying kit whose only tunnel is IPsec. Put it in the tender. A line asking for a modern, port-based, roaming-capable tunnel is a line a vendor either answers or does not, and it is how the installed-base argument finally loses. That argument is the only one keeping any of this alive, and it is only ever defeated by purchasing.\nWrite down which tunnels are left and why, and put a date on each. A protocol nobody has owned for ten years is how PPTP got to 2026. An inventory with dates is the difference between retiring something and merely disliking it.\nIf You Are Still Deploying It, Can You Call Yourself A Professional? That is a genuine question and I am going to answer it honestly, because it is the one the rest of this post has been walking towards.\nIt depends on whether you know. And knowing is not something that happens to you. Making sure you know is the job.\nIf you are standing up a new IPsec remote-access service this year and you cannot say, without looking any of it up, why ESP has no ports, what a residential line under carrier-grade NAT does to it, why a NAT64 network will not carry it at all, or what that one-byte packet every twenty seconds is actually for — then no. Not on this, not yet. You are not choosing a protocol. You are repeating a shape, because the last one looked like that and nobody in the room asked why — including you. The failing is not the gap. Everybody has gaps. The failing is building across one you never went and closed. As such the person who inherits it in 2035 gets a decade of tickets that were all avoidable on the day you drew it.\nLet me put the right name on that, because the polite version has been in circulation for twenty years and it has changed nothing. Somebody who deploys a protocol they cannot explain is not an engineer. They are a follower. They read the cue cards — the vendor\u0026rsquo;s reference architecture, the last change request, a diagram somebody drew in 2014 and nobody has opened since — and they read them out with real conviction, and there is no understanding anywhere behind the performance. From the outside it looks exactly like competence. It goes on looking exactly like competence right up to the first failure the cards do not cover, and from that moment it is the only thing in the room that matters.\nThe cue cards are also the reason this protocol is still here. Nobody stood at a whiteboard in 2026 and argued for IPsec on the merits. It got deployed again because it was on the card, and the card was written by a vendor whose interest is that you keep buying the box that terminates it. That is how something outlives its own obituary by twenty years — not by being defended, but by never once being made to justify itself in a room where somebody could tell the difference.\nAnd this gets called out. Out loud, in the room, at the time — not muttered in the corridor afterwards. If somebody puts the word professional next to their name, the word arrives with an invitation to be asked, and asking is not rude. Being asked and having an answer is the entire difference between the word and a business card.\nIt matters most when you are paying for it. A consultancy, an MSP, a vendor\u0026rsquo;s professional services arm, the integrator on the framework — what you are buying is judgement, and judgement is the one thing you cannot inspect on delivery. So inspect it beforehand. Ask why this protocol and not another one. Ask what happens to it on a carrier-grade NAT line, on an IPv6-only mobile network, on a path that quietly drops fragments. Ask which of those they have personally hit and what they did about it. You will know inside two minutes whether you are being told something or being read to, and two minutes is a considerably cheaper test than four years of tickets. And if the answer is cue cards and you sign anyway, that is a decision as well — it has just become yours rather than theirs.\nIf you can say all of that, and you are deploying it anyway because a regulator names it by protocol, because the partner\u0026rsquo;s kit terminates nothing else, or because the replacement is in next year\u0026rsquo;s budget and this has to work in March — then yes, obviously, and you are doing the job properly. Constraints are real and I have built round worse. What separates the two is not the protocol on the diagram. It is whether you wrote down why, and whether there is a date next to it.\nThe indefensible position is the middle one. Knowing enough to be uneasy, and building it anyway because nobody made you justify it. Not engineering. Habit, with a change number attached — and it is exactly how L2TP ended up on a datasheet printed this year.\nSo ask it of yourself before the tender closes rather than after. Professional is not a word about which protocols you know. It is a word about whether you can defend the one you picked, out loud, to somebody who knows the failure modes. If you can, deploy whatever the constraints demand and sleep fine. If you cannot, you have just found what to go and read tonight, and that is not an insult. Everybody was a follower once, without exception, me included. The part that is not defensible is choosing to stay one and calling it a career.\nA Good Idea Is Allowed To Be Over I want to be fair to IPsec, because it deserves it.\nIt was the right instinct. Security belongs low in the stack, where everything inherits it and no application has to be trusted to get it right. Binding the association to the address was not a mistake in 1995 — the address was the machine, and building on that was correct. The people who wrote it were serious people solving a real problem, and the WireGuard paper, of all documents, is the one that puts it best: the layering of IPsec is sound, everything is in the right place, to academic perfection53.\nAnd then the ground moved. Not because IPsec did anything wrong, but because this industry decided that addresses were a cost to be managed rather than a thing every machine gets, and built twenty-five years of translation to avoid the alternative. IPsec\u0026rsquo;s foundation was quietly withdrawn, and rather than admit it, we shimmed it. UDP encapsulation. Then keepalives. Then TCP encapsulation for the networks that block UDP. Then another UDP header so a router can find something to hash on. Each fix reasonable on its own; the stack of them is a protocol being kept upright by people who are paid to keep it upright.\nThat is the part worth being angry about, and it is not really about IPsec at all. Our industry is very good at maintaining things and very bad at ending them. Maintenance is billable, budgeted, staffed and safe. Retirement is a decision somebody has to sign, with their name on it, and no immediate reward. So PPTP lived fourteen years past the proof it was worthless, L2TP still ships on boxes sold this year, and IKEv1 kept negotiating tunnels for a decade after it went Historic. Not because anyone defended them. Because nobody was ever required to kill them.\nThere is no prize for getting 90% right. Ferguson and Schneier wrote that about IPsec in 1999, and what has happened since is thirty years of the industry getting the last 10% wrong in a new way each time and calling the patch a solution.\nA good idea is still allowed to be over. Knowing when to stop maintaining something is a skill, and it is the one this trade is worst at. Somebody has to be the one who says a protocol has had its life, writes the date down, and takes the consequences of being the one who said it. Otherwise we will still be sending a one-byte packet every twenty seconds in 2040, to keep a table warm in a box we do not own, on a network that took our addresses off us and charged us for the privilege.\nRFC 4301 — Security Architecture for the Internet Protocol, December 2005. Defines the security association and the triple it is looked up on: destination address, security protocol and SPI.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 4302 — IP Authentication Header, December 2005. AH\u0026rsquo;s integrity check covers the immutable fields of the IP header, source and destination addresses included.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 9329 — TCP Encapsulation of Internet Key Exchange Protocol (IKE) and IPsec Packets, November 2022, replacing RFC 8229. Exists because middleboxes on public networks block UDP.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 3715 — IPsec-Network Address Translation (NAT) Compatibility Requirements, March 2004. Sixteen enumerated incompatibilities, including: \u0026ldquo;Since the AH header incorporates the IP source and destination addresses in the keyed message integrity check, NAT or reverse NAT devices making changes to address fields will invalidate the message integrity check,\u0026rdquo; and \u0026ldquo;Where IP addresses are used as identifiers in Internet Key Exchange Protocol (IKE) Phase 1 or Phase 2, modification of the IP source or destination addresses by NATs or reverse NATs will result in a mismatch between the identifiers and the addresses in the IP header.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCisco — Configuring IPsec NAT-Traversal, Security Configuration Guide, Cisco IOS XE 17.15.x (Catalyst 9300 Switches). \u0026ldquo;If PAT found a legislative IP address and port, it would drop the Encapsulating Security Payload (ESP) packet\u0026rdquo;; the UDP header and an 8-byte non-IKE marker are inserted between the outer IP header and the ESP header; the restriction list includes static rules for both port 500 and 4500, no support for dynamic NAT policies, no IPv6, and that IPsec and NAT cannot both operate on the same device.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 3948 — UDP Encapsulation of IPsec ESP Packets, January 2005, with the negotiation in RFC 3947. Defines the keepalive as \u0026ldquo;a one-octet-long payload with the value 0xFF\u0026rdquo;, sent \u0026ldquo;if no other packet to the peer has been sent in M seconds. M is a locally configurable parameter with a default value of 20 seconds,\u0026rdquo; and requires the UDP checksum be zeroed: \u0026ldquo;If the protocol header after the ESP header is a UDP header, set the checksum field to zero in the UDP header.\u0026rdquo; Microsoft, Cisco, F-Secure, Nortel and SafeNet are all on the author lists of RFC 3947 and RFC 3948.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nJuniper — Route-Based and Policy-Based VPNs with NAT-T, Junos OS IPsec VPN User Guide. \u0026ldquo;NAT-T encapsulates both IKE and ESP traffic within UDP with port 4500 used as both the source and destination port\u0026rdquo;; \u0026ldquo;Because NAT devices age out stale UDP translations, keepalive messages are required between the peers\u0026rdquo;; and on SRX5400, SRX5600 and SRX5800, \u0026ldquo;the total number of tunnels from a given public translated IP cannot exceed 1000 tunnels.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 8221 — Cryptographic Algorithm Implementation Requirements and Usage Guidance for ESP and AH, October 2017. ENCR_DES MUST NOT, ENCR_3DES SHOULD NOT, AUTH_HMAC_MD5_96 MUST NOT; ESP combined with AH is NOT RECOMMENDED.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 4787 — Network Address Translation (NAT) Behavioral Requirements for Unicast UDP, January 2007. REQ-5: \u0026ldquo;A NAT UDP mapping timer MUST NOT expire in less than two minutes,\u0026rdquo; with \u0026ldquo;a default value of five minutes or more for the NAT UDP mapping timer is RECOMMENDED\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 6146 — Stateful NAT64: Network Address and Protocol Translation from IPv6 Clients to IPv4 Servers, April 2011. \u0026ldquo;The current specification only defines how stateful NAT64 translates unicast packets carrying TCP, UDP, and ICMP traffic. Multicast packets and other protocols, including the Stream Control Transmission Protocol (SCTP), the Datagram Congestion Control Protocol (DCCP), and IPsec, are out of the scope of this specification.\u0026rdquo; Packets carrying anything else \u0026ldquo;SHOULD be discarded\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 6877 — 464XLAT: Combination of Stateful and Stateless Translation, April 2013. Gives an IPv6-only device a local IPv4 stack so that traffic NAT64 cannot carry works anyway.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nIETF — draft-xu-ipsecme-esp-in-udp-lb, Encapsulating IPsec ESP in UDP for Load-balancing. \u0026ldquo;Although the ESP SPI field within the IPsec packets can be used as the load-balancing key, but it cannot be used by legacy switches and routers\u0026rdquo;; \u0026ldquo;Many cloud service providers allow customers to establish multiple IPsec VPN tunnels in parallel to enable ECMP and increase aggregate bandwidth. However, this approach is not ideal, as each tunnel typically requires its own public IP address, leading to higher public IP consumption and increased operational overhead.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCisco — Resolve IPv4 Fragmentation, MTU, MSS, and PMTUD Issues with GRE and IPsec. The long-standing reference on tunnel overhead, path MTU discovery and what breaks when the ICMP errors do not come back.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCisco — Troubleshoot IPsec Anti-Replay Check Failures. \u0026ldquo;Certain QoS features, such as Low Latency Queueing (LLQ), could cause IPsec packet delivery to become out-of-order and dropped by the receiving endpoint due to a replay check failure.\u0026rdquo; Default window 64 packets; 1024 on later platforms; the alternative remedy is multiple sequence number spaces per association.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 2661 — Layer Two Tunneling Protocol \u0026ldquo;L2TP\u0026rdquo;, August 1999. A tunnelling protocol for PPP with no packet-level confidentiality of its own; its security section points at IPsec.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMoxie Marlinspike and David Hulton — Divide and Conquer: Cracking MS-CHAPv2 with a 100% Success Rate, 2012; tool at github.com/moxie0/chapcrack. MS-CHAPv2\u0026rsquo;s security reduces to a single DES operation regardless of password length; contemporary write-up at The Register. Their conclusion was that PPTP traffic should be considered unencrypted.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nApple — If you see a \u0026ldquo;VPN Using PPTP May Not Be Secure\u0026rdquo; alert. PPTP was removed from the built-in client in macOS Sierra and iOS 10, 2016.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 4303 — IP Encapsulating Security Payload (ESP), December 2005. IP protocol 50; the ESP header carries the SPI and sequence number and has no port field.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 9395 — Deprecation of the Internet Key Exchange Version 1 (IKEv1) Protocol and Obsoleted Algorithms, April 2023. \u0026ldquo;Internet Key Exchange Version 1 (IKEv1) has been deprecated, and RFCs 2407, 2408, and 2409 have been moved to Historic status.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 7296 — Internet Key Exchange Protocol Version 2 (IKEv2), October 2014. The current key exchange, and the source of the notify names in the diagnostics table.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 3173 — IP Payload Compression Protocol (IPComp), September 2001. Its own IP protocol number and its own associations, negotiated alongside ESP.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 2367 — PF_KEY Key Management API, Version 2, July 1998. The kernel interface a key daemon uses to install associations.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 7383 — IKEv2 Message Fragmentation, November 2014. \u0026ldquo;This document describes a way to avoid IP fragmentation of large Internet Key Exchange Protocol version 2 (IKEv2) messages.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 4555 — IKEv2 Mobility and Multihoming Protocol (MOBIKE), June 2006. Lets an established tunnel survive a change of address.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 3706 — A Traffic-Based Method of Detecting Dead Internet Key Exchange (IKE) Peers, February 2004.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\ndraft-beaulieu-ike-xauth — Extended Authentication within IKE (XAUTH). Last revision 02, October 2001, status Expired, \u0026ldquo;Expired \u0026amp; archived\u0026rdquo;, never published as an RFC. The draft records that it was offered as Informational because \u0026ldquo;the IPSRA working group will not accept any protocol which extends ISAKMP or IKE, and the IPsec working group refuses to accept any protocols that deal with remote access.\u0026rdquo; Mode-Config shared the same fate.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 3193 — Securing L2TP using IPsec, November 2001. Co-authored at Microsoft.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 2332 — NBMA Next Hop Resolution Protocol (NHRP), April 1998. The piece that lets DMVPN spokes find each other.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 6407 — The Group Domain of Interpretation, October 2011. Group keying, as used by GETVPN.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 2637 — Point-to-Point Tunneling Protocol (PPTP), July 1999. Category: Informational. A vendor protocol written up, never a standard.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft — Configure L2TP/IPsec server behind NAT-T device, KB 926179, last revised 12 February 2026. \u0026ldquo;By default, Windows Vista and Windows Server 2008 don\u0026rsquo;t support Internet Protocol security (IPsec) network address translation (NAT) Traversal (NAT-T) security associations to servers that are located behind a NAT device\u0026rdquo;; the AssumeUDPEncapsulationContextOnSendRule DWORD under HKEY_LOCAL_MACHINE\\SYSTEM\\CurrentControlSet\\Services\\PolicyAgent takes 0 (default, cannot), 1 (server behind NAT) or 2 (both ends behind NAT), and the machine must be restarted. Also: \u0026ldquo;If you must use IPsec for communication, use public IP addresses for all servers that you can connect to from the Internet.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nstrongSwan — Windows Certificate Requirements. The gateway certificate needs the serverAuth EKU, OID 1.3.6.1.5.5.7.3.1, and the IP Security IKE Intermediate EKU, OID 1.3.6.1.5.5.8.2.2.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nstrongSwan — Windows Clients. Documents adding the NegotiateDH2048_AES256 DWORD under Rasman\\Parameters to get AES-256-CBC and MODP-2048; the rekeying workaround for clients behind NAT (rekey_time = 0 on the gateway, letting the client initiate); and that the Windows client \u0026ldquo;does not currently support IKE redirection (RFC 5685) and multiple authentication rounds (RFC 4739).\u0026rdquo; Also: \u0026ldquo;IKEv2 fragmentation is supported since the v1803 release of Windows 10 and Windows Server\u0026rdquo;, and clients behind NAT refuse a server-initiated CHILD_SA rekey with Microsoft error 13863.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft — DirectAccess and Remote Access Always On VPN migration overview. DirectAccess is deprecated and will be removed in a future release of Windows Server; customers are directed to Always On VPN.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nJean Paul Degabriele and Kenneth G. Paterson — Attacking the IPsec Standards in Encryption-only Configurations, IEEE Symposium on Security and Privacy, 2007. Attacks that \u0026ldquo;break any RFC-compliant implementation of IPsec making use of encryption-only ESP\u0026rdquo;, ciphertext-only, needing only the ability to eavesdrop and to inject traffic.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 6434 — IPv6 Node Requirements, December 2011. \u0026ldquo;Previously, IPv6 mandated implementation of IPsec and recommended the key management approach of IKE. This document updates that recommendation by making support of the IPsec Architecture [RFC4301] a SHOULD for all IPv6 nodes.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNiels Ferguson and Bruce Schneier — A Cryptographic Evaluation of IPsec, Counterpane Internet Security, 1999. \u0026ldquo;IPsec was a great disappointment to us\u0026rdquo;; \u0026ldquo;Our main criticism of IPsec is its complexity\u0026rdquo;; \u0026ldquo;We therefore recommend that transport mode be eliminated\u0026rdquo;; \u0026ldquo;We conclude that eliminating transport mode allows the elimination of the AH protocol as well, without loss of functionality\u0026rdquo;; and \u0026ldquo;We have found serious security weaknesses in all major components of IPsec. As always in security, there is no prize for getting 90% right; you have to get everything right.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDavid Adrian et al. — Imperfect Forward Secrecy: How Diffie-Hellman Fails in Practice, ACM CCS 2015. Precomputation for a second 1024-bit group \u0026ldquo;would allow decryption of traffic to 66% of IPsec VPNs\u0026rdquo;; 86.1% of scanned IKEv1 and 91.0% of IKEv2 servers supported Oakley Group 2, and 66.1% of profiled IKEv1 servers preferred it.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCVE-2016-1287 and Cisco\u0026rsquo;s advisory, Cisco ASA Software IKEv1 and IKEv2 Buffer Overflow Vulnerability. Remote code execution before authentication, reached by crafted UDP packets to the IKE service.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nJuniper — How to Analyze IKE Phase 2 VPN Status Messages. \u0026ldquo;No proposal chosen\u0026rdquo; means the device \u0026ldquo;did not accept any of the IKE Phase 2 proposals that the peer sent\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCisco — Understand and Use Debug Commands to Troubleshoot IPsec, and Troubleshoot Common L2L and Remote Access IPsec VPN Issues.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nJuniper — Troubleshoot a VPN Tunnel That is Down.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nLinux kernel — XFRM proc counters. The descriptions in the table are the kernel documentation\u0026rsquo;s own.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nJuniper — Troubleshoot a VPN That Is Up But Not Passing Traffic. \u0026ldquo;If only the pkts counter in the out direction of the session is incrementing, then validate with the VPN peer that the traffic is being received.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft — netsh wfp, Windows Commands. netsh wfp capture start \u0026ldquo;Starts a capture session for network events processed by WFP\u0026rdquo; and defaults to writing wfpdiag.cab; netsh wfp show state \u0026ldquo;Displays the current state of WFP and IPsec\u0026rdquo;; netsh wfp show ikeevents \u0026ldquo;Displays recent Internet Key Exchange (IKE) epoch events matching the specified parameters\u0026rdquo; and takes a remoteaddr= filter. Setting file=- on either show prints to the console instead of writing XML.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft — Get-NetIPsecQuickModeSA, NetSecurity module. \u0026ldquo;There is only one main mode SA between a pair of computers, but there can be many quick mode SAs,\u0026rdquo; and monitoring them \u0026ldquo;can provide information about which peers are currently connected to this computer, and which protection suite is protecting the data exchanged between them.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft — Troubleshoot Always On VPN, last revised 12 February 2026. On reading client logs: \u0026ldquo;look for events labeled RasClient. All error messages return the error code at the end of the message.\u0026rdquo; Error 809\u0026rsquo;s cause: \u0026ldquo;You can encounter this issue when the UDP 500 or 4500 ports on the VPN server or firewall are blocked.\u0026rdquo; Error 812 is an authentication method mismatch between the server\u0026rsquo;s policy and the client\u0026rsquo;s profile. Error 13801\u0026rsquo;s four listed causes are a machine certificate without Server Authentication under Enhanced Key Usage, an expired RAS machine certificate, a client missing the root certificate, and a client whose \u0026ldquo;VPN server name doesn\u0026rsquo;t match the subjectName value on the server certificate\u0026rdquo;; 13806 is \u0026ldquo;IKE can\u0026rsquo;t find a valid machine certificate\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft — Guidance for troubleshooting Remote Access (VPN and AOVPN). The supported collection is TSS, run elevated: TSS.ps1 -Scenario NET_VPN on the client and TSS.ps1 -Scenario NET_RAS on the server, reproducing the fault between start and stop, with the traces written to C:\\MS_DATA.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft — Routing and Remote Access Error Codes, the codes defined in raserror.h. 638 ERROR_REQUEST_TIMEOUT; 718 ERROR_PPP_TIMEOUT; 789 ERROR_OAKLEY_GENERAL_PROCESSING, \u0026ldquo;The L2TP connection attempt failed because the security layer encountered a processing error during initial negotiations with the remote computer\u0026rdquo;; 809 ERROR_VPN_TIMEOUT, \u0026ldquo;The network connection between your computer and the VPN server could not be established because the remote server is not responding\u0026rdquo;; 828 ERROR_IDLE_TIMEOUT, \u0026ldquo;The connection was terminated because of idle timeout\u0026rdquo;; 930 ERROR_AUTH_SERVER_TIMEOUT, \u0026ldquo;The authentication server did not respond to authentication requests in a timely fashion\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nfirewalld — Automatic Helper Assignment. \u0026ldquo;With kernel 4.7 and up the automatic helper assignment in kernel has been turned off by default\u0026rdquo;, controlled by the sysctl at /proc/sys/net/netfilter/nf_conntrack_helper, and \u0026ldquo;for the secure use of iptables and connection tracking helpers it is recommended to turn AutomaticHelpers off\u0026rdquo;. The kernel\u0026rsquo;s own message on the change names the reason and the replacement: default automatic helper assignment \u0026ldquo;has been turned off for security reasons\u0026rdquo;, use the CT target to attach helpers instead. The helpers covered include ftp, irc, sip, h323, tftp, snmp and pptp.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nSamy Kamkar — NAT Slipstreaming, v1 31 October 2020, v2 26 January 2021 with Ben Seri and Gregory Vishnipolsky of Armis. The attack abuses \u0026ldquo;the Application Level Gateway (ALG) connection tracking mechanism built into NATs, routers, and firewalls\u0026rdquo; to \u0026ldquo;bypass victim NAT and connect directly back to any port on any machine on the network, exposing previously protected/hidden services and systems\u0026rdquo;. v1 used the SIP gateway on port 5060; v2 used H.323 on 1720, which let the pinhole be aimed at any internal host rather than only the machine that loaded the page.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft — Set-VpnServerConfiguration, RemoteAccess module. -IdleDisconnectSeconds \u0026ldquo;Specifies the time, in seconds, after which an idle connection is terminated\u0026rdquo;; -SALifeTimeSeconds and -MMSALifeTimeSeconds set the quick mode and main mode lifetimes; -SADataSizeForRenegotiationKilobytes \u0026ldquo;Specifies the number of kilobytes that are allowed to transfer using a security association (SA), after which the SA will be renegotiated.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nJason A. Donenfeld — WireGuard: Next Generation Kernel Network Tunnel, NDSS 2017. \u0026ldquo;It intentionally lacks cipher and protocol agility. If holes are found in the underlying primitives, all endpoints will be required to update. As shown by the continuing torrent of SSL/TLS vulnerabilities, cipher agility increases complexity monumentally\u0026rdquo;; \u0026ldquo;implemented for Linux in less than 4,000 lines of code\u0026rdquo;; \u0026ldquo;it is important to stress, however, that the layering of IPsec is correct and sound; everything is in the right place with IPsec, to academic perfection.\u0026rdquo; Benchmarks: 1,011 Mbit/s against 881 and 825 for two IPsec cipher suites, and 0.403 ms ping against 0.501 and 0.508.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/networking/ipsec-was-a-good-idea-turn-it-off/","summary":"IPsec was right in 1995: encrypt below the application, bind the security association to the IP address, let every protocol inherit it. Then NAT arrived, carrier-grade NAT finished the job, and the fix was to wrap the whole thing in UDP and keep a timer running so a translation table would not forget you. This post shows how it falls down, diagram by diagram — the security association that cannot survive a rewritten header, the two translators every CGNAT line now has, the NAT64 standard that names IPsec as out of scope, the tunnel that cannot use a second link because ESP has no ports, the MTU nobody owns, and L2TP and PPTP as the two protocols that were never fit to be here. It carries the vendor documentation from Cisco, Juniper and Microsoft that admits every one of those, the eighteen pieces that call themselves an IPsec VPN including the two that were never standards at all, why the Fisher-Price OS has never truly interoperated with an open stack, a working method for diagnosing IPsec while you still run it, and the case for retiring the lot with dates.","title":"IPsec Was a Good Idea. It Is Time to Turn It Off."},{"content":"What Azure Virtual Desktop Actually Costs Azure Virtual Desktop delivers a Windows desktop in a browser. That is genuinely what it does, and it works. The question is what it costs to keep it working.\nThe sticker price is not the price. Microsoft\u0026rsquo;s licensing stack for AVD runs roughly like this1:\nA Microsoft 365 licence that includes the Windows Enterprise entitlement — E3 or E5, or the equivalent Business Premium Azure compute underneath the desktop — a VM billed by the hour, or a reserved instance billed by the month Azure storage for the OS disk and the user profile Azure networking for the traffic between the desktop and everywhere else Optionally, Microsoft Entra ID P1 or P2 for conditional access policies Windows 365 Cloud PC simplifies the billing into a flat per-user-per-month number, but the number is not small. A 2 vCPU / 8 GB / 128 GB configuration — which is a modest office desktop — is £35.60 per user per month2. Add a GPU and it jumps to £269.40 for the GPU Standard tier. Multiply by headcount and it is a real line item, every month, forever.\nThat is the product this post replaces.\nBrowser any device no RDP client no VPN HTTPS Cloudflare Access\u0026#160;+\u0026#160;IdP authenticates the user renders RDP in browser free ≤\u0026#160;50 users no ports exposed no VPN required zero trust policy tunnel LXC cloudflared own VLAN firewalled to RDP targets only RDP Windows\u0026#160;VM GPU VF\u0026#160;0\u0026#160;·\u0026#160;RDP Windows\u0026#160;VM GPU VF\u0026#160;1\u0026#160;·\u0026#160;RDP Windows\u0026#160;VM GPU VF\u0026#160;2\u0026#160;·\u0026#160;RDP Proxmox host\u0026#160;·\u0026#160;Intel Arc Pro SR-IOV Arc Pro B-series PF → host VF 0 → VM VF 1 → VM VF 2 → VM A browser, an identity provider and a tunnel — no VPN, no exposed ports, no per-seat cloud bill The whole path: a browser, Cloudflare Access for identity, a tunnel into a firewalled LXC, and RDP to Windows VMs with Intel Arc Pro virtual functions. No VPN, no exposed ports, no per-seat cloud bill. What the Replacement Stack Looks Like Three layers, each independent, none of them billed per seat:\nThe GPU layer is the Intel Arc Pro SR-IOV build from the previous post. A single Arc Pro card splits into hardware virtual functions through standard PCIe SR-IOV — no vGPU licence, no NVIDIA subscription. Each Windows VM gets its own VF and its own GPU-accelerated desktop. The card sets the seat count, not a licence server.\nThe access layer is Cloudflare Access with browser-rendered RDP. The user opens a URL in any browser, authenticates against your identity provider, and Cloudflare renders the RDP session directly in the browser tab. No RDP client installed. No VPN. No ports exposed to the internet. Access federates to any OAuth or OIDC provider — that includes on-prem providers like Keycloak or Authentik stacked on your existing Active Directory, whether that is Samba 4 AD or a Windows domain controller. The directory keeps doing what it already does: user accounts, group policy, domain-joined VMs. Keycloak or Authentik federates against it and adds the OAuth/OIDC and MFA layer that Cloudflare Access needs. The identity plane stays on your own hardware. No Entra ID subscription required. Cloudflare Access handles what happens after that: it controls what the authenticated user can reach on your network — which applications, which protocols, which hosts. The desktop is behind a Cloudflare tunnel and unreachable from anywhere except through the Access policy, and the policy decides both identity and scope.\nThe tunnel layer is a cloudflared process running in an LXC container on Proxmox, on its own VLAN — a /30 for IPv4 with its own dedicated IPv6 prefix, nothing else in the broadcast domain. The Proxmox firewall controls what the LXC can reach, and the answer is short: TCP 3389 to the desktop VMs and nowt else. If the tunnel endpoint is compromised, the blast radius is one container on an otherwise empty VLAN, and the only thing it can talk to is what the firewall already permits. That is a much smaller surface than a VPN concentrator that hands out a routed subnet.\nCloudflare Zero Trust is free for up to 50 users3. You pay from user 51. Azure Virtual Desktop charges from seat one.\nThe Architecture User any device any browser HTTPS Cloudflare Access OAuth IdP does identity\u0026#160;+\u0026#160;MFA Access controls network\u0026#160;scope renders RDP tunnel Your premises LXC cloudflared own\u0026#160;VLAN /30\u0026#160;IPv4 Firewall Proxmox TCP\u0026#160;3389\u0026#160;only RDP Windows VMs GPU virtual functions Arc\u0026#160;Pro\u0026#160;SR-IOV The session path: user authenticates at Cloudflare\u0026rsquo;s edge, the tunnel lands in a firewalled LXC on its own VLAN, and the firewall permits only RDP to the desktop VMs. The path a desktop session takes:\nThe user opens https://vdi.example.com in any browser, on any device Cloudflare Access intercepts the request and redirects to your identity provider — Keycloak or Authentik federated against your Active Directory The OAuth provider validates the user\u0026rsquo;s identity against AD and handles MFA Access evaluates the policy — the provider confirmed who they are, Access decides what they can reach on your network Cloudflare establishes an RDP session through the tunnel to the target VM The RDP session is rendered in the browser — no client, no plugin, no download The tunnel terminates in an LXC container on the Proxmox host, on a locked-down VLAN The LXC forwards RDP to the Windows VM, which has a GPU virtual function from the Arc Pro card At no point does the Windows VM have a public IP. At no point is port 3389 open to the internet. The only thing listening on the public internet is Cloudflare, and the only thing that gets past Cloudflare is a user who passed the Access policy.\nThe Tunnel Endpoint: An LXC on Its Own VLAN The cloudflared process needs to run somewhere, and where you put it is a security decision.\nRunning it directly on the Proxmox host is the simplest option and the worst one. A tunnel endpoint that shares the host\u0026rsquo;s network namespace can reach everything the host can reach, which on a hypervisor is everything. A compromised tunnel becomes a pivot into the management plane.\nRunning it in a full VM is clean but heavy. A tunnel relay is a single Go binary that uses almost no CPU and a few hundred megabytes of RAM. Giving it a full kernel and a virtual disk is overkill.\nAn LXC container is the right shape. It gets its own network namespace, its own VLAN — a /30 IPv4 and a dedicated IPv6 prefix, nothing else sharing the broadcast domain — and its own firewall rules in the Proxmox firewall. It shares the host\u0026rsquo;s kernel but not its network stack. The firewall permits TCP 3389 to the desktop VMs, DNS to resolve them, and HTTPS outbound to Cloudflare\u0026rsquo;s edge. As such, even a compromised tunnel endpoint can only talk to the things it was already supposed to talk to — and the VLAN it sits on has no other residents to reach.\nThe LXC configuration, the Proxmox firewall rules for the VLAN, and the cloudflared tunnel setup are all in the next post.\nHow Cloudflare Access Fits Together There are two halves to this. The OAuth provider — Keycloak or Authentik, federated against your Active Directory (Samba 4 or Windows) — handles identity validation and MFA. It proves the user is who they claim to be, using the same directory the VMs are domain-joined to. Cloudflare Access handles everything after that: what the authenticated user is allowed to reach, through which protocol, and how the session is rendered.\nCloudflare Access renders the RDP session directly in the browser. This is not a download or a plugin — Cloudflare\u0026rsquo;s edge runs a headless RDP client and streams the result as a canvas in the browser tab. The user sees a Windows desktop. The browser sees HTTPS to Cloudflare. The Windows VM sees an RDP connection from the cloudflared tunnel.\nThe Access application is a self-hosted app pointed at the tunnel\u0026rsquo;s RDP ingress. The Access policy is where the two halves meet. The OAuth provider has already confirmed the identity and passed the MFA challenge. Access takes that token and decides what to do with it — which application the user can reach, whether their device posture passes, whether their location is permitted. The provider says who. Access says what.\nThe tunnel creation, the ingress rules, the LXC configuration, the Proxmox firewall rules and the Access policy itself are all in the next post — this one sets out what the stack is and why it exists. The next one builds it.\nWhat the User Sees Cloudflare Access includes an App Launcher — a resource portal that lists every application the authenticated user is allowed to reach. After the user logs in through the OAuth provider, the portal shows them their available desktops, internal web apps, and any other tunnelled resources, all in one place. One URL, one login, and a tile for each thing they have access to. It is the landing page for the whole stack, not just RDP.\nThe user clicks a desktop tile, and gets a Windows desktop in a browser tab. No client, no plugin, no download. It works on whatever device they already own — a company laptop, a personal machine, a Chromebook, a tablet. That is the point. The system is built for businesses that let people use their own kit.\nClipboard, audio and multi-monitor are not supported through the browser-rendered session. That is by design, not by accident. Every one of those channels is a data exfiltration path. A clipboard that crosses the boundary moves files out. Audio capture moves conversations out. Multi-monitor with a local desktop beside the remote one makes drag-and-drop trivial. Cutting those channels means a user can work inside the desktop but cannot pull work out of it through a side channel. The Access policy controls who gets in. The browser rendering controls what gets out.\nIf the business needs clipboard or audio for a specific workflow, Cloudflare Access also supports a native RDP client through the tunnel — the same tunnel, the same policy, the same identity check. The native path gives full RDP features to the users who need them and keeps the browser path locked down for everyone else. Two access methods, one policy engine, one tunnel.\nThe Cost Comparison A concrete example. One server, 42 GPU-accelerated desktops:\nA 64-core CPU with hyperthreading gives you 128 threads. Each VM gets 8 vCPUs — a proper desktop allocation, not a thin client. 42 × 8 = 336 vCPUs, which is well past 128 threads on paper. But this is VDI. Office desktops are idle most of the time. A user reading a document or typing an email is not loading 8 cores. Oversubscription is not a risk here, it is the design. Proxmox lets you allocate more vCPUs than physical threads because the scheduler knows most of them are sleeping. The CPU is sized for the peak, and the peak is a handful of users compiling or rendering at once, not all 42. This is not a shortcut — the cloud hyperscalers oversubscribe the same way. Every Azure VM you rent shares physical cores with other tenants on the same host. The performance model is identical. The difference is who owns the host. 12 GB of RAM per VM is a solid office desktop. 42 × 12 GB = 504 GB, plus 4 GB for Proxmox itself = 508 GB on paper. In practice KSM collapses the identical pages across those 42 Windows images, so the physical RAM needed is substantially less — but budget for 512 GB of DIMMs and let KSM hand you the headroom. Three dual Intel Arc Pro B60 cards in a board like the Supermicro H13SSL-NT. Each physical card presents two GPUs to the OS, so three cards give you 6 GPUs. Each GPU supports 7 SR-IOV virtual functions4. That is 42 GPU-accelerated desktops from three PCIe slots. The B60 lists at around $599–799 per card. The B70 is the bigger option at $949 launch for 32 GB and 4 virtual functions per GPU5 — fewer seats but more VRAM per seat. Windows 365 Cloud PC for 42 users2:\nEven without a GPU, the numbers are not small. The Basic tier — 2 vCPU, 4 GB, 128 GB — is £26.90/user/month. The Standard tier — 2 vCPU, 8 GB, 128 GB, which is a modest office desktop — is £35.60/user/month. For 42 users on Standard, that is £1,495.20/month, £17,942.40/year, before any of the extras below.\nWith a GPU it gets worse. The GPU Standard tier is £269.40/user/month. For 42 users, that is £11,314.80/month, £135,777.60/year.\nBoth tiers then add:\nMicrosoft 365 licensing if you do not already have it Azure networking — ingress and egress are billed separately, and a desktop that streams to a browser is not light on egress Azure Backup or a third-party backup solution — the VM snapshots and profile storage are not backed up for free Public IPv4 addresses — Azure charges for every public IP attached to a resource, and the price has only gone up Azure\u0026rsquo;s underlying pricing is in USD, so GBP costs move with the exchange rate — a weak pound makes every line item more expensive and you have no control over either side of that The bill never stops. Year five costs the same as year one. This stack for 42 users — the build cost:\nSupermicro H13SSL-NT motherboard: ~£700 AMD EPYC 64-core CPU: ~£1,500 512 GB DDR5 ECC RDIMM (8 × 64 GB): ~£5,200 Three dual B60 cards: ~£1,800 Chassis, PSU, boot SSDs: ~£800 Total hardware: roughly £10,000 Amortised over a five-year life, that is £2,000/year in capital cost. Add:\nCloudflare Zero Trust: free for up to 50 users Windows 11 Enterprise VDI rights — Software Assurance on Pro upgrades to Enterprise which includes VDI access for up to four VMs per user, or licence through Microsoft 365 E3/E56 Electricity — and this is worth putting a number on The power budget at peak draw: a 64-core EPYC at 360W TDP, three dual B60 cards at 400W each (1,200W), 512 GB of RAM at roughly 80W, plus storage, fans and PSU losses at around 150W. That is approximately 1,800W at the wall under full load. VDI desktops are not under full load — office use is mostly idle CPU and light GPU, so a realistic average is closer to 1,000–1,200W. Call it 1,100W.\nAt 35p per kWh, 1.1 kW running 24/7 is:\n1.1 × 24 × 365 = 9,636 kWh/year 9,636 × £0.35 = £3,373/year in electricity So the total annual cost of this stack is roughly £5,400/year — £2,000 in amortised hardware and £3,400 in electricity, before Windows licensing and internet. Set that next to the Azure bill: £5,400 versus £135,778 for the GPU tier, or £17,942 for basic desktops without a GPU. Even on the cheapest Azure tier, this stack costs less than a third. On the GPU tier, it is 4% of the Azure bill.\nWhy VDI Fits Proxmox VDI is one of the workloads Proxmox is quietly very good at, for two reasons that have nowt to do with the hypervisor itself.\nKSM. Proxmox enables KSMd — the kernel same-page merging daemon — out of the box. Twenty Windows VMs built from the same image share vast amounts of identical memory pages: the OS, the base libraries, the unchanged parts of the user profile. KSMd finds those duplicates and collapses them into a single physical page, copy-on-write. The result is that twenty desktops fit in the memory you would otherwise need for eight or ten. On a VDI host where every VM runs the same image, KSM is not a marginal optimisation. It is what makes the density affordable.\nbcache. If your storage is HDD-backed Ceph with Optane bcache, VDI is the best-case workload for it. Boot storms and login storms are read-heavy and repetitive — exactly the pattern a cache absorbs. Once the working set is warm, the desktops read from Optane at NVMe latency and the spindles barely move. The writes are user profile changes and temp files, which are small and sequential enough that bcache\u0026rsquo;s writeback handles them without ever bottlenecking the HDD.\nProfiles — two options, both on Ceph. User profiles are the other half of VDI storage, and there are two clean ways to handle them without leaving the cluster.\nFSLogix profile containers are the standard way to roam a Windows desktop profile, and they work on S3-compatible object storage. Proxmox Ceph exposes an S3 gateway through the RADOS Gateway, so the profile store lives on the same cluster as the VM disks. No separate file server, no Azure Files bill, no external dependency.\nThe other option is to skip FSLogix entirely and use Windows folder redirection to SMB shares served by clustered Samba on CephFS. The desktops redirect Documents, Desktop, AppData and the rest to a Samba share backed by CephFS, and the Samba cluster handles failover. No profile container at all — the files live on the filesystem as plain files, and CephFS handles the replication. This is simpler to manage, simpler to back up, and sidesteps FSLogix licensing entirely — useful if the licence cost is an issue or you just want fewer moving parts. Either way, the profiles sit on your own storage, backed by the same Ceph pool, and the cost is the disk you already bought.\nBetween KSM recovering memory, bcache recovering storage latency and profiles landing on Ceph — whether through FSLogix on S3 or folder redirection on CephFS — a single Proxmox host with HDDs and a modest amount of RAM serves more desktops than the spec sheet suggests. Azure charges for every gigabyte of all three. Here, the infrastructure does the work for free.\nBackup, Resilience and Compliance Keeping everything inside VDI is not just a cost decision. It is a compliance and resilience decision.\nWhen the desktop lives on the server, the data lives on the server. Nothing lands on the user\u0026rsquo;s device. A stolen laptop is a lost screen, not a lost dataset. There is no local disk to encrypt, no local copy to exfiltrate, no endpoint to forensically image after a breach. The data never left the infrastructure you control.\nThat makes backup straightforward. The VM disks and the profile stores are on Ceph, and Ceph snapshots are atomic and instant. One snapshot policy covers every desktop and every profile. Restoring a desktop to yesterday\u0026rsquo;s state is a snapshot rollback, not a rebuild. Restoring a profile is the same operation on a different pool. The backup target is the cluster, not twenty scattered endpoints.\nIt also simplifies compliance. Data residency is easy to prove when the data is on hardware you own, in a rack you can point at, in a jurisdiction you chose. Audit trails sit on your own logs. Access is gated by Cloudflare policies you wrote, authenticated by an IdP you run, and recorded by systems you control. There is no third-party cloud provider between you and the evidence an auditor asks for.\nResilience follows the same line. A dead desktop VM is a new VM from the golden image with the profile reattached. The user logs in again and the desktop is back. There is no endpoint to rebuild, no OS to reimage, no hardware to ship. The recovery unit is the VM, and spinning one up takes minutes.\nWhat You Give Up This is not free in every sense. The things Azure Virtual Desktop handles that this stack does not:\nMicrosoft manages the patching and the updates. Here, you do. Intune and Endpoint Manager integrate natively with AVD. Here, you are managing the Windows VMs yourself or through whatever tooling you choose. Azure\u0026rsquo;s network is Azure\u0026rsquo;s problem. Here, your internet connection is the path to the desktop. If it goes down, the desktops are unreachable until it comes back. Scaling is a credit card away on Azure. Here, scaling means buying another card or another host. AVD gives the user a full RDP feature set by default. Here, the browser-rendered path strips clipboard, audio and multi-monitor on purpose. Users who need those features get a native RDP client through the same tunnel and the same policy — but the default is the locked-down browser, and that is the right default for a BYOD workforce. None of that is trivial. Whether it matters depends on what you have: if you already run Proxmox, already manage Windows, and already have someone who can look after a hypervisor, then all of those are things you are already doing. If you do not, then Azure is selling you the staff you do not have, and that is genuinely worth something.\nThe question is whether it is worth £18,000 a year for 42 basic desktops — or £136,000 for GPU ones — every year, plus networking, backup, IPv4 and licensing on top, with the price set by someone else, the hardware belonging to someone else, and the bill denominated in a currency you do not control. Or whether you spend £5,400 a year on hardware and electricity and keep the rest.\nThat is what this post sets out. The next one builds it — the cloudflared tunnel, the LXC, the Proxmox firewall rules, the Cloudflare Access policy, and the working desktop in a browser tab.\nReferences Azure Virtual Desktop pricing — compute, storage and networking billed separately on top of the Microsoft 365 entitlement.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nWindows 365 plans and pricing — flat per-user-per-month, GPU configurations in the Enterprise tier.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCloudflare Zero Trust pricing — free for up to 50 users, pay-as-you-go from user 51.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nIntel Arc Pro B60 specifications — 24 GB GDDR6, 20 Xe2-cores, PCIe 5.0 x8, SR-IOV with up to 7 virtual functions.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nIntel Arc Pro B70 specifications — 32 GB GDDR6, 32 Xe2-cores, PCIe 5.0 x16, SR-IOV with up to 4 virtual functions.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nWindows 11 Licensing for Virtual Desktops — Software Assurance on a qualifying OS (e.g. Pro) upgrades to Enterprise, granting VDI rights for up to four VMs per user on your own on-premises server.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/networking/zero-trust-vdi-cloudflare-access/","summary":"Part one: the architecture and the cost argument for replacing Azure Virtual Desktop with Proxmox, Intel Arc Pro SR-IOV and Cloudflare Access. On-prem OAuth for identity, browser-rendered RDP, a firewalled LXC tunnel on its own VLAN, KSM for memory density, and profiles on Ceph. The next post builds it.","title":"Zero Trust VDI Without the Cloud Bill — Proxmox, Intel Arc Pro and Cloudflare Access"},{"content":"Ping is not a diagnostic. It is a tunnel. ICMP echo (the request a machine sends and the reply it gets back) will carry any bytes you put in it, in both directions, over a protocol most firewalls pass without inspecting and most logging counts rather than reads. A network that \u0026ldquo;only allows ping\u0026rdquo; already has a full, encryptable VPN out of it, and anyone who can reach that network can drive one.\nThis is old, and it is documented. Loki1 published the technique in Phrack 49 in 1996. Ptunnel2 has carried whole TCP sessions inside ping for twenty years and is a package away on any Linux box. No zero-day. The protocol behaving exactly as specified.\nThat is the cause: the standard, not a bug in anyone\u0026rsquo;s product. Every host on earth is obliged to take the data in an echo request and hand it straight back, unchanged. Arbitrary data in, the same data out: that is the whole of what a tunnel needs, and it has been mandated since 1981.\nThe fix is one narrow change. Drop ICMP echo, request and reply, on IPv4 and IPv6, at the border, and keep every other ICMP message. The errors — Time Exceeded, Packet Too Big, Destination Unreachable — are load-bearing; block those and you break path MTU discovery and traceroute for nothing. So this is not \u0026ldquo;block ICMP\u0026rdquo;. It is drop the one message that is a liability, and keep the ones that carry the truth.\nWhat Your Egress Policy Actually Allowed If you run a network where outbound is filtered (and if you do not, that is its own conversation), then this is the bill for one line in it. The policy that says \u0026ldquo;block everything outbound, allow ICMP because we need to ping things\u0026rdquo; is not an egress policy. It is a full tunnel with the paperwork filed under diagnostics. Anything on the inside that can send an echo request and read the reply can move data to anywhere on the outside that answers one, at whatever rate the link will bear, and past every content control you bought. In the logs it is somebody checking whether the internet is up, and nowt else. Malware has shipped this for years for that exact reason. Quiet, standard, and usually already allowed.\nThis is why it is a first-choice channel for anyone who is already inside and should not be. It is the tenancy, not the smash-and-grab. Someone after a way in and out that lasts avoids the port that trips an alert. They stand a tunnel up on echo and leave it running for months. A host that pings a lot is a host nobody watches. Two-way networking into the estate: command in, data out, over the one protocol nobody rate-limits, alerts on, or reads. And it is encrypted, the way any real tool encrypts it, so on the wire the payload is the random-looking bytes a ping carries anyway and content inspection has nothing to read. Size and timing can still give it away to someone actually looking. Peering inside the packet cannot.\nAnd the machine driving it need not belong to anyone who works there. All it takes is something that can reach your network and put a ping on the internet, and reaching your network is easier than anyone likes to admit. The WiFi carries past the walls, into the corridor, the car park, the flat upstairs. The ethernet port in the meeting-room wall, or reception, or the empty desk by the window, is very often live and will talk to anything you plug in, with no 802.1X asking who you are. A signal in range, or a socket nobody locked down. That is the entry fee. No badge, no account, no invitation.\nPicture the visitor you did invite. They are in the meeting, pleasant, taking notes, contributing. Their laptop is not. The moment it reached your network, over the air or through the port under the table, it was inside the wall, and if that network can emit an echo request, it has a way out. Nothing was dropped on your estate. No account, no privilege on anything you own, nothing for your endpoint agents to catch, because your agents are not on their machine. The person is across the table. The traffic is leaving through your front door, encrypted, indistinguishable from a laptop that pings more than it should.\nA visitor in signal range, or at an open port, gets a tunnel out through a firewall that sees only pings A visitor in range, or at an open port, gets a route out inside the wall — your network WiFi reaches the corridor, car park, flat upstairs visitor's laptop no badge, no account a live wall port no 802.1X asking who you are border firewall sees only pings server on the internet echo request\u0026#160;\u0026#8594; \u0026#8592;\u0026#160;echo reply any IP traffic, encrypted, riding inside the pings The entry fee is access to the network — a WiFi signal in range, or a live wall port with no 802.1X — and the exit is echo through the border. To the firewall it is a host that pings; inside those pings is any IP traffic, encrypted, bound for a server on the internet. The two things you will reach for do not help. NAT is a translation table, not a filter: an echo request from an inside host opens a mapping keyed on the ICMP id, which NAT treats exactly as it treats a port, and the reply comes back through it like any other flow. I ran a client behind my own home NAT and it never noticed the NAT was there. A separate guest VLAN is no better. Segregation keeps the visitor off your servers, but it does nothing to keep them off the internet, and the internet, reached with a ping, is the whole requirement. One control touches this: does that network let echo out?\nYou do not inspect this away, and you do not segregate it away. You close the door. Everything after this is why the standard leaves it open, three ways to build the tunnel so the claim stands on more than my say-so, and the one change that shuts it.\nThe Part of ICMP That Owes You Nothing Useful Start with what the standard actually says, because the whole argument rests on one sentence and people wave it away without reading it.\nAn echo request carries a data field. In IPv4, RFC 7923 states plainly that \u0026ldquo;the data received in the echo message must be returned in the echo reply message\u0026rdquo;. IPv6 tightened the wording rather than loosening it: RFC 44434 defines the field as \u0026ldquo;zero or more octets of arbitrary data\u0026rdquo; and then requires that it \u0026ldquo;MUST be returned entirely and unmodified in the ICMPv6 Echo Reply message\u0026rdquo;.\nRead that as an operator and it is a keepalive. Read it as somebody moving data and it is a gift. The standard obliges every reachable host to accept a block of bytes you pick and send them straight back to you, unchanged, on demand. Arbitrary length. Arbitrary content. No handshake, no port, no application on the far end that has to agree to anything. The kernel does it, before any userland process gets a look in.\nAnd it is both families. Moving to IPv6 does not tidy this away. It opens it wider. RFC 792 said the data \u0026ldquo;must be returned\u0026rdquo;; RFC 4443 says it \u0026ldquo;MUST be returned entirely and unmodified\u0026rdquo;, which is the same door with a stronger lock holding it open. As such, the channel exists on ICMP echo and on ICMPv6 echo alike, and a rule that shuts it on one family and not the other has shut nothing. The tunnel just moves to the family you left open. Whatever you do about this, do it on ip and ip6 together.\nNothing else in the protocol does this. A Time Exceeded carries the header of the packet that died and no more. A Destination Unreachable is a report about something that already happened. Those messages tell you facts about the network. Echo carries whatever you put in it, in both directions, and calls it a diagnostic.\nErrors Are Load-Bearing. Echo Is Not. This is the distinction the \u0026ldquo;just block ICMP\u0026rdquo; crowd never make, and it is the whole point, so I will make it once and properly.\nBlocking ICMP errors breaks the network quietly. Filter Packet Too Big and you kill path MTU discovery: the handshake completes, small transfers work, and anything carrying a full-size packet hangs forever with nothing in the logs. On IPv6 that is not even a matter of taste. Routers do not fragment, so RFC 48905 lists Packet Too Big among the messages a firewall \u0026ldquo;must not drop\u0026rdquo; and warns that without it \u0026ldquo;parts of the Internet will become inaccessible\u0026rdquo;. Filter Time Exceeded and you break traceroute, which RFC 18126 names as the reason the message is mandatory in the first place. These are not optional. They are the feedback that makes the network self-correcting, and I covered the cost of losing them at length last time.\nNow drop echo and go looking for what broke. ping across the boundary stops working. That is the list. The whole list.\nPath MTU discovery does not care, because it runs on Packet Too Big, which is an error. Traceroute does not care, because traceroute -T and -U walk TCP and UDP and read the errors that come back. Not one of them sends an echo. Neighbour discovery on IPv6 does not care, because that is types 133 to 137, and you keep those or the segment dies. The reverse-TTL trick from the last post works on any reply, and a TCP handshake gives you one. Everything that made the last post\u0026rsquo;s diagnostics work keeps working, because not one measurement in it sent an echo request.\nSo the two halves of the protocol could not be less alike. One half is the network telling you the truth about itself, and you break it at your peril. The other half is a mandated payload-return service that happens to be called a diagnostic, and the strongest thing anyone can say for keeping it open is that ping is handy. It is handy. It is also the only part of ICMP an attacker can drive, and I can show you exactly what they drive it with.\nKeep the ICMP errors, rate-limited; drop only echo, both directions and both families Two halves of one protocol, and only one of them is a liability KEEP rate-limit, never block Destination Unreachable carries why a packet could not be delivered Packet Too Big path MTU discovery depends on it Time Exceeded traceroute depends on it Parameter Problem reports a malformed header Neighbour Discovery 133\u0026#8211;137 IPv6: the local segment depends on it Break these and the network breaks quietly. DROP both directions, both families Echo request type 8 (IPv4) · type 128 (IPv6) Echo reply type 0 (IPv4) · type 129 (IPv6) Nothing needs these but the ping command. Everything an attacker drives through ICMP runs through them. A tunnel left open on one family is not shut. The ICMP errors are load-bearing: path MTU discovery, traceroute and IPv6 neighbour discovery all depend on them, and blocking them breaks the network quietly. Echo is the only part nothing depends on but ping, and the only part an attacker can drive. Drop that, keep the rest, on both families. Proving It: A VPN Made of Ping The tool is Hans7, written by Friedrich Schöller. Its own description is one line: it \u0026ldquo;makes it possible to tunnel IPv4 through ICMP echo packets, so you could call it a ping tunnel.\u0026rdquo; It brings up a tun interface at each end, gives them addresses, and moves every packet between them inside ICMP echo. To the network in the middle it is somebody pinging a server and the server answering. To me it is a route.\nAn entire IP packet travels inside the data field of a single ping A whole packet rides inside one ping The packet your SSH session sends IP header src / dst TCP header port 22 application data your keystrokes, your file Hans copies the whole packet into the echo data field The packet that leaves on the wire IP header you to server ICMP echo request type 8 · id · seq echo data field the entire packet above, unchanged To the border firewall this is one echo request, and the log records a ping. Inside the data field is a packet bound for anywhere the far end can reach. The standard requires the far end to send that data straight back, so the reply is a packet too. Every packet the tunnel carries is copied into the data field of an ICMP echo request. The standard obliges the far end to return that data unchanged, so the reply carries a packet too. The firewall counts a ping; the payload goes anywhere the far end can route. I ran it across my own line, a server on a public address and a client behind my home NAT, and pushed real traffic through it. Not synthetic pings with a flag set. An SSH login and a file pull, riding inside echo.\nIt is a package away where I ran the server, and it builds from Schöller\u0026rsquo;s source anywhere else. The server has to be Linux and needs root, because opening a tun device and a raw ICMP socket both do. The syntax is deliberately small:\n# on the server (public IP), pick the tunnel network and a password hans -s 10.8.0.0 -p \u0026#39;\u0026lt;password\u0026gt;\u0026#39; # server takes 10.8.0.1; clients are handed 10.8.0.2 and up # on the client, point it at the server\u0026#39;s public address hans -c \u0026lt;server-public-ip\u0026gt; -p \u0026#39;\u0026lt;password\u0026gt;\u0026#39; Once both ends are up there is a new interface at each side with an address on the tunnel network, and it behaves like any other point-to-point link:\nNothing about that session knows it is inside ping. SSH opens a TCP connection to 10.8.0.1, the kernel routes it out tun0, Hans wraps each packet as the payload of an echo request, and the far end unwraps it and feeds it to its own tun0. The reply comes back as the payload of an echo reply. As far as SSH is concerned it is talking over an ordinary link. As far as the firewall is concerned, nobody opened anything. A host is being pinged.\nWhat It Looks Like on the Wire This is the part that ends the argument, so watch the firewall\u0026rsquo;s-eye view rather than mine.\nWhat is really happening, and what the firewall logs, are not the same picture One tunnel, two pictures What is really happening your host tun0 · 10.8.0.2 the server tun0 · 10.8.0.1 anywhere it can reach SSH, file pulls routed out any IP traffic, through the tunnel as ordinary packets What the border firewall sees and logs your host 198.51.100.9 the server 192.0.2.7 echo request echo reply icmp echo request 198.51.100.9 \u0026gt; 192.0.2.7 icmp echo reply 192.0.2.7 \u0026gt; 198.51.100.9 icmp echo request 198.51.100.9 \u0026gt; 192.0.2.7 ... counted as pings, content unread Same wire. The firewall never sees the packets inside. The tunnel moves real traffic between two hosts and then out to anywhere the server can reach. At the border it is only echo request and echo reply — the firewall logs pings and never sees the packets carried inside them. Sit on the outside interface with tcpdump and capture only ICMP while the SSH session runs. No TCP to port 22 crosses the boundary. What crosses is echo request and echo reply, back and forth, each one fatter than a real ping because it is carrying a slice of a TCP segment in its payload:\n# on the boundary, watch only ICMP echo while traffic runs over the tunnel tcpdump -ni \u0026lt;wan-iface\u0026gt; \u0026#39;icmp[icmptype] = icmp-echo or icmp[icmptype] = icmp-echoreply\u0026#39; Read the two things that matter in that capture. First, the payload lengths: a normal ping sends 56 bytes and every line is the same size, while these vary and run large, because the size of the thing you are moving leaks into the size of the ping. Second, the rate: a diagnostic ping is one a second, and this is a flood, because it is moving a file. Neither is hidden. Both are sitting in plain sight on a protocol nobody is looking at.\nAnd that is the whole point. Not that this is clever, or hard to spot once you look. Almost nobody looks, because the box is configured to allow ICMP, the logs count it as pings, and the alerting was tuned for the ports somebody remembered to worry about. The traffic leaves looking like a health check and the health check is a route to anywhere the server can reach.\nA Proper Full-Tunnel VPN, Built by Hand Hans proves the channel is there, but it does the interface and the addressing for you and hands back a route, so it does not quite show you what you have built. To see that this is a VPN in the full sense — the whole machine\u0026rsquo;s traffic leaving through ping, not a link between two named hosts — put it together by hand. icmptunnel8, by Dhaval Kapil, is the one for that, and its own description is a single line: \u0026ldquo;Transparently tunnel your IP traffic through ICMP echo and reply packets.\u0026rdquo; Same idea, tun device and echo payloads, but you lay the plumbing yourself and nothing hides inside a binary.\nOn the server you start the tunnel, bring the interface up, and then do the thing that gives the whole game away: tell the kernel to stop answering pings itself, so its own echo replies do not fight the ones the tunnel is sending.\nsudo ./icmptunnel -s 10.0.1.1 # server mode; creates tun0, then blocks # from a second shell, bring the interface up (iproute2, not the net-tools the repo ships) sudo ip addr add 10.0.1.1/24 dev tun0 sudo ip addr add 2001:db8:1::1/64 dev tun0 sudo ip link set tun0 mtu 1472 up # 1500 − 20 (IP) − 8 (ICMP); see below # stop the kernel replying to pings — echo now belongs to the tunnel sudo sysctl -w net.ipv4.icmp_echo_ignore_all=1 sudo sysctl -w net.ipv6.icmp.echo_ignore_all=1 # let the server route the client\u0026#39;s packets onward sudo sysctl -w net.ipv4.ip_forward=1 Read the icmp_echo_ignore_all line again, because it says more than it looks. That knob is the kernel\u0026rsquo;s own, and it ships as a defence: set it and, in the kernel\u0026rsquo;s words, it \u0026ldquo;will ignore all ICMP ECHO requests sent to it\u0026rdquo;9, so an operator can take a host off the ping radar entirely. It is on both families now. net.ipv4.icmp_echo_ignore_all for years, and net.ipv6.icmp.echo_ignore_all added later to match.10 The tunnel flips that defensive switch on for the opposite reason: with the kernel no longer answering echo itself, its two ends are free to use echo as pure transport. Which is the tell. The people who wrote the stack already treat echo as something a host may reasonably refuse — the rule in this post makes that same call once, at the border, for every host behind it.\nOn the client you bring the interface up and then point the default route down it. Not one host reached through the tunnel. Everything.\nsudo ./icmptunnel -c \u0026lt;server-public-ip\u0026gt; # client mode; creates tun0 sudo ip addr add 10.0.1.2/24 dev tun0 sudo ip addr add 2001:db8:1::2/64 dev tun0 sudo ip link set tun0 mtu 1472 up # keep the route to the server itself OUT of the tunnel... sudo ip route add \u0026lt;server-public-ip\u0026gt; via \u0026lt;gateway\u0026gt; dev \u0026lt;iface\u0026gt; # ...then send everything else down it sudo ip route replace default dev tun0 The project\u0026rsquo;s own client.sh and server.sh still reach for net-tools ifconfig and route. The ip commands above are the iproute2 equivalents and do the same job. That route to the server is the line people forget: keep it out of the tunnel, or the echo packets carrying the tunnel try to travel down the tunnel, and nothing leaves. Everything else now goes through tun0, wrapped in echo, and reaches the server. Whether it goes further is a routing choice. A single masquerade rule on the server would put it on the public internet under the server\u0026rsquo;s own address, and it is deliberately not here, because the tunnel is the thing being shown and it does not need it.\nThat is a full-tunnel VPN, built out of ping in a handful of commands. Every packet the client sends — web, DNS, SSH, all of it — is captured into tun0 and leaves as an echo request to the server, and the answers come back as echo replies. A machine on a network that \u0026ldquo;only allows ICMP\u0026rdquo; has just handed its entire outbound traffic to a box outside, and the border logged a host that likes to ping.\nRolling Your Own in Python Here is the part that should trouble anyone hoping to defend against this by spotting a tool. Neither Hans nor icmptunnel is doing anything you could not write yourself in an afternoon. Open a tun device, wrap each packet as the payload of an echo request, unwrap the ones that come back. That is the whole mechanism, and the bare tunnel is about sixty lines of standard-library Python with nothing to install at all. The version below adds one thing on top: it encrypts the payload, and that is the only part that pulls in a dependency.\n#!/usr/bin/env python3 # pingvpn.py — a VPN tunnel over ICMP echo, in one short file. # # Not a product. It exists to show the channel is trivial to rebuild, so a # defence that hunts for a known tool is chasing the wrong thing entirely. # Linux, needs root (a tun device and a raw ICMP socket both do), and the # `cryptography` package for the AES (pip install cryptography). # # server: sudo python3 pingvpn.py --server # client: sudo python3 pingvpn.py --client \u0026lt;server-public-ip\u0026gt; # # The payload is encrypted with AES-128-GCM under a pre-shared key before it # goes on the wire, so what a firewall sees in the echo data is random bytes — # the same as a real ping\u0026#39;s padding, and nothing for content inspection to read. import argparse import fcntl import hashlib import os import select import socket import struct import sys from cryptography.hazmat.primitives.ciphers.aead import AESGCM TUNSETIFF = 0x400454CA IFF_TUN = 0x0001 IFF_NO_PI = 0x1000 MAGIC = 0x4954 # \u0026#39;IT\u0026#39; in the id field, so we ignore real pings ECHO_REQUEST = 8 ECHO_REPLY = 0 PSK = b\u0026#34;change-me-to-a-shared-secret\u0026#34; # pre-shared secret, both ends KEY = hashlib.sha256(PSK).digest()[:16] # 128-bit key -\u0026gt; AES-128-GCM AEAD = AESGCM(KEY) def open_tun(name=b\u0026#34;tun0\u0026#34;): fd = os.open(\u0026#34;/dev/net/tun\u0026#34;, os.O_RDWR) fcntl.ioctl(fd, TUNSETIFF, struct.pack(\u0026#34;16sH\u0026#34;, name, IFF_TUN | IFF_NO_PI)) return fd def encrypt(data): # -\u0026gt; nonce || ciphertext+tag nonce = os.urandom(12) return nonce + AEAD.encrypt(nonce, data, None) def decrypt(blob): # raises on a packet that is not ours return AEAD.decrypt(blob[:12], blob[12:], None) def checksum(data): if len(data) % 2: data += b\u0026#34;\\x00\u0026#34; total = sum(struct.unpack(\u0026#34;!%dH\u0026#34; % (len(data) // 2), data)) total = (total \u0026gt;\u0026gt; 16) + (total \u0026amp; 0xFFFF) total += total \u0026gt;\u0026gt; 16 return ~total \u0026amp; 0xFFFF def build_echo(icmp_type, payload): head = struct.pack(\u0026#34;!BBHHH\u0026#34;, icmp_type, 0, 0, MAGIC, 0) csum = checksum(head + payload) return struct.pack(\u0026#34;!BBHHH\u0026#34;, icmp_type, 0, csum, MAGIC, 0) + payload def main(): ap = argparse.ArgumentParser(description=\u0026#34;a VPN tunnel over ICMP echo\u0026#34;) group = ap.add_mutually_exclusive_group(required=True) group.add_argument(\u0026#34;--server\u0026#34;, action=\u0026#34;store_true\u0026#34;) group.add_argument(\u0026#34;--client\u0026#34;, metavar=\u0026#34;SERVER_IP\u0026#34;) args = ap.parse_args() out_type = ECHO_REQUEST if args.client else ECHO_REPLY in_type = ECHO_REPLY if args.client else ECHO_REQUEST peer = args.client # None on the server until a client is seen tun = open_tun() sock = socket.socket(socket.AF_INET, socket.SOCK_RAW, socket.IPPROTO_ICMP) print(\u0026#34;tun0 created. bring it up with an address and route, then send traffic.\u0026#34;, file=sys.stderr) while True: readable, _, _ = select.select([tun, sock], [], []) if tun in readable: # a packet wants to leave this host packet = os.read(tun, 65535) if peer: sock.sendto(build_echo(out_type, encrypt(packet)), (peer, 0)) if sock in readable: # something arrived over ICMP data, _ = sock.recvfrom(65535) ihl = (data[0] \u0026amp; 0x0F) * 4 # skip the IP header the kernel adds icmp = data[ihl:] if len(icmp) \u0026lt; 8 or icmp[0] != in_type or icmp[4:6] != struct.pack(\u0026#34;!H\u0026#34;, MAGIC): continue try: packet = decrypt(icmp[8:]) # wrong key or a real ping -\u0026gt; skip except Exception: continue if args.server: peer = socket.inet_ntoa(data[12:16]) # reply to whoever sent os.write(tun, packet) # hand the carried packet to the stack if __name__ == \u0026#34;__main__\u0026#34;: main() That is the entire tunnel. It opens tun0 and a raw ICMP socket and shuttles packets between them: what leaves the host is wrapped as an echo request, or an echo reply on the server, and sent to the far end; what arrives over ICMP is unwrapped and handed back to the stack. The id field is pinned to one value so it steps over real pings, and the server learns where to answer from the source of the first packet it sees.\nThe encryption is the point worth dwelling on, because it is what real tools do and it is why you will not catch this by looking inside the packet. The payload is sealed with AES-128-GCM under a pre-shared key before it is wrapped, so the bytes in the echo data are indistinguishable from the random padding a real ping carries. Strip the two crypto lines out and the tunnel still runs on nothing but the standard library — the only reason it needs pip install cryptography is the AES, and the AES is exactly the part that turns a readable tunnel into an unreadable one. Bring it up the same way as before, minus the masquerade. You do not need it to prove the point:\n# server: start it (creates tun0, then blocks), then configure from the same shell sudo python3 pingvpn.py --server \u0026amp; sudo ip addr add 10.9.0.1/24 dev tun0 sudo ip addr add 2001:db8:9::1/64 dev tun0 sudo ip link set tun0 mtu 1400 up sudo sysctl -w net.ipv4.icmp_echo_ignore_all=1 sudo sysctl -w net.ipv6.icmp.echo_ignore_all=1 # client sudo python3 pingvpn.py --client \u0026lt;server-public-ip\u0026gt; \u0026amp; sudo ip addr add 10.9.0.2/24 dev tun0 sudo ip addr add 2001:db8:9::2/64 dev tun0 sudo ip link set tun0 mtu 1400 up sudo sysctl -w net.ipv4.icmp_echo_ignore_all=1 sudo sysctl -w net.ipv6.icmp.echo_ignore_all=1 sudo ip route add \u0026lt;server-public-ip\u0026gt; via \u0026lt;gateway\u0026gt; dev \u0026lt;iface\u0026gt; sudo ip route replace default dev tun0 The reason to write it out is not the tool. It is that the tool is disposable. A short file, no dependencies until you add the encryption, and every copy someone types looks a little different on the wire. So a signature that catches this one catches nothing next week. You cannot block your way out of this by naming the software, because there is no software to name.\nIt does not even need a shell. The logic is bytes in, bytes out, so it ports to anything that can open a socket. Even WebAssembly: the browser sandbox normally refuses a page a raw socket outright, but where that barrier is lifted (a browser handed the permission by the user, on a machine where the user runs with privilege enough for a raw socket to be opened) the same short program runs inside a tab. Which is the uncomfortable half of it. You cannot trust your users here. The host driving the tunnel is on the inside, held by somebody you decided was safe because they sit behind the firewall, and the firewall is the thing being tunnelled through. Perimeter trust assumes the threat is outside the wall. This one starts inside it, every time.\nMind the MTU, and IPv6\u0026rsquo;s 1280 Floor The tunnel is not free on the wire. Every packet you carry gains an outer IP header and an ICMP header before it leaves, so the inner interface has to sit below what the path can carry. On IPv4 that costs the attacker almost nothing. Set the inner MTU low (icmptunnel\u0026rsquo;s 1472 is 1500 less 20 for the outer IP header and 8 for the ICMP header, and the Python above drops to 1400 for slack), and where the number is still wrong, IPv4 fragments the oversized packet and reassembles it at the far end rather than dropping it. Between the low floor you can pick and fragmentation papering over the rest, the tunnel runs across almost any path. That flexibility is exactly what makes IPv4 the comfortable place to do this.\nIPv4's fragmentation and MTU slack let the tunnel run anywhere; IPv6 has neither, so it fails closed Why it runs anywhere on IPv4 and fails closed on IPv6 IPv4 IP hdr 20 B ICMP 8 B inner packet tun MTU 1472 B Too big? IPv4 fragments and reassembles \u0026#8212; the tunnel runs across almost any path. IPv6 IPv6 hdr 40 B ICMPv6 8 B inner packet cannot drop below 1280 B 1280 B hard floor (RFC 8200) Too big? IPv6 routers drop it, no fragmentation \u0026#8212; the tunnel fails closed. IPv4's fallbacks \u0026#8212; a floor you can pick, fragmentation for the rest \u0026#8212; are what make the tunnel reliable. IPv6 has neither, so the attack that runs almost anywhere on IPv4 is fragile on IPv6. The safer family here. Packet Too Big \u0026#8212; an error you keep, not an echo \u0026#8212; is what lets a sender find the size that fits. Not to scale. The tunnel loses an outer IP and ICMP header off every packet. IPv4 lets you shave the inner MTU and fragments whatever is still too big, so it runs almost anywhere; IPv6 sets a hard 1280-byte floor and its routers do not fragment, so an oversized packet is dropped and the tunnel fails closed — which makes IPv6 the safer family here. IPv6 is less forgiving, and for once that is on your side. The floor is hard: RFC 820011 §5 requires that \u0026ldquo;every link in the Internet have an MTU of 1280 octets or greater\u0026rdquo;, and IPv6 routers do not fragment in transit. That removes both of IPv4\u0026rsquo;s fallbacks at once, and it bites the tunnel in both directions. Run it over ICMPv6 and every link on the path has to carry 1280, so a path that drops below that takes the transport down, and there is no shaving under the floor the way you can on IPv4. Carry IPv6 inside the tunnel and you meet the same wall from the other side: the inner interface cannot go below 1280 either, while the outer ICMPv6 wrapper, 40 bytes of header and 8 of ICMPv6, is already spending the budget, so the path has to spare 1280 plus the overhead. Miss it in either place and the packet is dropped, not cut down to fit: the tunnel comes up, small things work, anything full-size hangs. So IPv6 is the safer family here, not the riskier one. The attack that runs across almost anything on IPv4 is fragile on IPv6, and fails closed. Safer is not the same as safe, mind: the echo channel is open on IPv6 too, so the rule still drops it on both families — the attacker just cannot lean on IPv6 the way they lean on IPv4.\nThe recovery from that is the ICMP error you were careful to keep. Packet Too Big is what lets a sender find the working size, and it is an error, not an echo, so the rule this post argues for leaves it alone. Drop the tunnel and keep the diagnostics — that is the whole design, and the MTU is one more place it earns its keep.\nThe Rule, Narrowed to Echo The last post gave the full transit ruleset: permit the errors under a rate limit, drop the echo. I am not going to reprint it. The change this post argues for is one pair of lines, and it is the pair that does the work:\n# permit the ICMP errors — these are load-bearing, keep them ip protocol icmp icmp type { destination-unreachable, time-exceeded, parameter-problem } \\ limit rate 100/second accept ip6 nexthdr ipv6-icmp icmpv6 type { destination-unreachable, packet-too-big, \\ time-exceeded, parameter-problem } limit rate 100/second accept # drop the tunnel — echo, both directions, both families ip protocol icmp icmp type { echo-request, echo-reply } drop ip6 nexthdr ipv6-icmp icmpv6 type { echo-request, echo-reply } drop The rule is not Linux\u0026rsquo;s, though. It is the same intent on any firewall worth the name: permit the errors, cap their rate, drop echo both ways on both families. So here it is in the dialects you are more likely to be holding.\nBSD pf (pfSense, OPNsense, OpenBSD, FreeBSD):\n# permit the errors; block echo both directions, both families pass in proto icmp icmp-type { unreach, timex, paramprob } pass in proto icmp6 icmp6-type { unreach, toobig, timex, paramprob, \\ routersol, routeradv, neighbrsol, neighbradv } block in proto icmp icmp-type { echoreq, echorep } block in proto icmp6 icmp6-type { echoreq, echorep } Keep the neighbour-discovery types on the icmp6 line; those are the ones that take the segment down if you lose them.\nCisco IOS (extended ACLs, applied at the edge):\nip access-list extended ICMP-EDGE permit icmp any any unreachable permit icmp any any time-exceeded permit icmp any any parameter-problem deny icmp any any echo deny icmp any any echo-reply ! ipv6 access-list ICMP6-EDGE permit icmp any any packet-too-big permit icmp any any unreachable permit icmp any any time-exceeded permit icmp any any parameter-problem permit icmp any any nd-ns permit icmp any any nd-na deny icmp any any echo-request deny icmp any any echo-reply On IPv4 the echo type is echo; on IPv6 it is echo-request. Rate-limiting lives in CoPP, not the ACL.\nJuniper Junos (firewall filter; the inet6 filter mirrors this and keeps neighbour discovery):\nfirewall family inet filter icmp-edge { term errors { from { protocol icmp; icmp-type [ unreachable time-exceeded parameter-problem ]; } then { policer icmp-cap; accept; } } term drop-echo { from { protocol icmp; icmp-type [ echo-request echo-reply ]; } then discard; } } MikroTik RouterOS — drop the two echo types, then accept the rest of ICMP, which keeps the errors and, on v6, neighbour discovery:\n/ip firewall filter add chain=forward protocol=icmp icmp-options=8:0 action=drop comment=\u0026#34;echo request\u0026#34; add chain=forward protocol=icmp icmp-options=0:0 action=drop comment=\u0026#34;echo reply\u0026#34; add chain=forward protocol=icmp action=accept comment=\u0026#34;keep the errors (add limit= to cap)\u0026#34; /ipv6 firewall filter add chain=forward protocol=icmpv6 icmp-options=128:0 action=drop comment=\u0026#34;echo request\u0026#34; add chain=forward protocol=icmpv6 icmp-options=129:0 action=drop comment=\u0026#34;echo reply\u0026#34; add chain=forward protocol=icmpv6 action=accept comment=\u0026#34;keep errors and ND\u0026#34; Different syntax, one rule. Permit the messages that carry the truth, drop the one that carries your bytes.\nTwo cautions, both of which will bite you if you skim.\nDrop echo at the border, not on the wire between a host and its own router. On IPv6, neighbour discovery is ICMP — types 133 to 137, nd-router-solicit through nd-redirect — and it is how the segment does the job ARP does in IPv4. Carry an echo-drop onto a link-local or host chain without keeping those and the segment stops working within minutes, and it will not look like a firewall fault. Filter echo where traffic leaves your network, and leave the internal chains alone.\nDrop it in both directions and it stops being useful either way. Blocking only the request stops your hosts being pinged but still lets a reply out, and a tunnel can be built on replies alone with a little more effort. Drop request and reply, on IPv4 and IPv6, and the channel is shut both ways.\nWhat Dropping Ping Actually Costs You Be honest about the loss, because a control you have oversold is a control somebody quietly reverses the first time it is inconvenient.\nYou lose ping across the boundary. That is the cost, stated in full. It is a real cost. ping is the reflex, it is in everyone\u0026rsquo;s fingers, and the day after you ship this someone will say the internet is down because their ping to the gateway times out. It is not down. You told the gateway to stop answering a request that was only ever a convenience.\nReachability testing does not need echo. A TCP connection to a port you know is open tells you the host is up and the path works, and it tells you more than a ping does because it proves a service answered, not just a kernel:\n# \u0026#34;is it up and reachable?\u0026#34; without sending a single echo nc -zv \u0026lt;host\u0026gt; 443 # did the TCP handshake complete? traceroute -T -p 443 \u0026lt;host\u0026gt; # walk the path on TCP, read the errors back Both of those ride on the ICMP errors you kept and the TCP a real service already speaks. Neither sends an echo. So the honest trade is this: you give up the least informative tool in the box, the one that answers \u0026ldquo;is a kernel willing to reply\u0026rdquo; and nothing more, and in exchange you close the only part of ICMP that an attacker can turn into a route off your network. The diagnostics you actually reach for on a bad day are all on the other side of the line, untouched.\nThe Fisher-Price OS (Windows) Is Not a Way Out of This If you run a Fisher-Price OS (Windows) shop you might be reading this as somebody else\u0026rsquo;s problem. It is not. The liability is in the protocol, not the operating system. The standard obliges every host that answers a ping to hand your bytes back, and it does not stop to ask what the host is running first.\nA box on that OS makes a perfectly good tunnel endpoint. Hans ships a Windows client, so an inside machine can be the client end and pass its traffic out inside echo like any other. What it cannot easily be is the server end, because that wants a tun device and a raw ICMP socket held open, and on this OS only members of the Administrators group can create sockets of type SOCK_RAW12 — Microsoft\u0026rsquo;s own words. So the server sits on a real operating system, as mine did, and the box behind the firewall quietly pinging its way out is the one you were told was safe because it was on the inside.\nThat is the point for anyone defending it. The border rule does not care what the inside runs. Drop echo where the traffic leaves the network and every host behind it is covered — the ones you administer, and the one you were assured was hiding behind NAT. I do not run the thing, and I am not going to pretend the tunnel respects it. The fix is the same rule, in the same place, whatever the endpoint is booted into.\nIf the box itself is what you are hardening (the box, not the network), PowerShell will at least stop it answering or emitting a ping on its own account:\nforeach ($t in 8,0) { New-NetFirewallRule -DisplayName \u0026#34;Drop ICMPv4 Echo $t\u0026#34; -Protocol ICMPv4 -IcmpType $t -Direction Inbound -Action Block } foreach ($t in 128,129) { New-NetFirewallRule -DisplayName \u0026#34;Drop ICMPv6 Echo $t\u0026#34; -Protocol ICMPv6 -IcmpType $t -Direction Inbound -Action Block } Add -Direction Outbound twins to stop it reaching out as a tunnel client rather than sitting there as the target, and leave the ICMPv6 error and neighbour-discovery types alone, for the reason they are left alone everywhere else on this page. But do not mistake it for the control. A host firewall is not a transit firewall. It hardens the box and changes nothing about the border, and the border is where this actually gets shut.\nRaise This With Whoever Runs Your Border, Today You have probably read this far thinking of a network that is not yours to change. Most people are. So the useful thing to do is not to reach for the firewall, it is to put the question to whoever holds it — your own network team, your MSP, or the vendor whose box sits at the edge — and put it today, because this has been an open door since 1996 and one more week of it is a choice.\nAsk plainly, and ask expecting a blank look, because I would wager the whole of the UK GDP that nobody there has given echo a second thought. It is the one packet everybody allows and nobody owns, and an unowned rule is exactly the one that has been sat wrong for years with no one to notice.\nAsk them four things, in these words, so there is no room to nod along and change nothing:\nDo we drop ICMP echo request and echo reply at the border, on IPv4 and IPv6 both? Not \u0026ldquo;do we allow ICMP\u0026rdquo; — the specific message, both directions, both families. If the answer is one family only, it is no answer, because the tunnel just moves to the other. Do we still permit the ICMP errors? Destination Unreachable, Time Exceeded, Parameter Problem, and Packet Too Big on IPv6, under a rate limit rather than a block. If they cannot say, the odds are even they have either left everything open or blocked the lot, and both are wrong. Have we kept IPv6 neighbour discovery on the internal chains? Types 133 to 137. This is the one that turns \u0026ldquo;we hardened ICMP\u0026rdquo; into a segment that dies quietly a week later, and it is the mistake a rushed change makes. Show me the rule. Not a policy statement, the actual line on the actual box, and the date it went on. A control nobody can point to is a control that is not there. Will the box have arrived from the vendor with this already shut? It will not. The default nearly everywhere is to pass echo and record it as a packet count, which is the whole reason the tunnel works at all. It is not exploiting a bug, it is using the configuration that ships as standard. Nobody closes a door they were never told was open.\nThe reason to insist on the wording is that the lazy fix to this post is \u0026ldquo;right, we\u0026rsquo;ll block ICMP\u0026rdquo;, and that fix is worse than the hole. It breaks path MTU discovery and traceroute, it will have you chasing hung transfers with nothing in the logs, and on IPv6 it takes parts of the internet off the table entirely. If the person you ask reaches for that, stop them. The instruction is narrow on purpose: drop echo, keep the errors, keep neighbour discovery. As such, anyone who runs firewalls for a living should be able to make that change in an afternoon and tell you it is done.\nAnd if they are an MSP you pay to run this: a supplier who cannot tell you your own echo-and-error posture off the top of their head, or who answers \u0026ldquo;we block ICMP\u0026rdquo; as though that were the safe option, has just told you something about the rest of the estate. The next time a fault sits between two networks and both say they are clean, remember which of them could not describe what their own firewall does to a ping.\nThe Diagnostic That Was Never a Diagnostic Ping has had a good run. Mike Muuss wrote it in 1983 to check whether a host was answering, named it after sonar, and it did that one job so well that it became the first thing everyone reaches for and the last thing anyone questions. Forty years on it is muscle memory, and muscle memory is exactly how a liability survives. Nobody re-examines the thing they have always done.\nBut look at what it actually is, stripped of the habit. A message that carries no fact about the network, that every host is obliged to answer with your own bytes handed back, that runs over a protocol the controls do not inspect and the logs do not read. Everything that makes it feel harmless — it is only a diagnostic, it is only a keepalive, everyone allows it — is the same thing that makes it the cleanest way off a filtered network that exists. The industry blocks the ICMP errors, which carry the truth and break the network when they are gone, and it waves through echo, which carries whatever you load into it. It has the protocol exactly upside down.\nThe fix is not clever. Keep the messages that tell you the truth, and drop the one that just hands your bytes back. It costs you a command you had no business trusting for a diagnosis anyway, and it shuts a door that has been standing open since 1996, documented in a hacker zine, packaged for two decades, and left wide because closing it would mean someone could not ping. Know what a thing costs, not what it is priced at. Ping is priced at nowt. It costs you the one channel you cannot see.\nLoki, Phrack 49 — a command channel carried inside ICMP echo payloads, 1996.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPtunnel — carries a TCP session inside ICMP echo.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 792 — ICMP: \u0026ldquo;the data received in the echo message must be returned in the echo reply message\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 4443 — ICMPv6: echo data \u0026ldquo;MUST be returned entirely and unmodified in the ICMPv6 Echo Reply message\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 4890 — the ICMPv6 messages a firewall must not drop.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 1812 — router requirements: Time Exceeded is a MUST, named for traceroute.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nHans — IP over ICMP echo tunnel, by Friedrich Schöller.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nicmptunnel — tunnels IP traffic through ICMP echo and reply packets, by Dhaval Kapil.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nLinux kernel — IP sysctl — icmp_echo_ignore_all: \u0026ldquo;the kernel will ignore all ICMP ECHO requests sent to it\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nLinux commit e6f86b0f — \u0026ldquo;ipv6: Add icmp_echo_ignore_all support for ICMPv6\u0026rdquo; — the IPv6 equivalent, by Virgile Jarry.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRFC 8200 §5 — IPv6 requires every link to have an MTU of at least 1280 octets.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMicrosoft — TCP/IP raw sockets — \u0026ldquo;only members of the Administrators group can create sockets of type SOCK_RAW\u0026rdquo;.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blogs.damiendye.uk/en/networking/ping-the-diagnostic-tool-that-opens-a-whole-lot-more/","summary":"Ping, not the rest of ICMP, is the liability: echo is a channel every host must answer with your own bytes, so a network that \u0026lsquo;only allows ping\u0026rsquo; already has a full VPN out. This walks the threat first — what it costs your egress, and how a visitor on your WiFi or an unlocked ethernet port can open one — then three working tunnels built on ping alone (Hans, icmptunnel, and a short Python one with AES-128), the MTU and IPv6 catches, and the rule that shuts it: drop echo, keep the errors, in nftables, pf, Cisco, Junos, MikroTik and Windows.","title":"Ping: The Diagnostic Tool That Opens a Whole Lot More"},{"content":" Is Your MSP Lying To You — 3 parts\nIs Your MSP Lying To You To Sell You Premium Products?you are here What Your MSP Built You, And Who Else Can Reach It When It Breaks, Who Actually Carries It? Short answer: some of them are.\nThe longer answer is worse, and it is the one worth your time. Most of them never have to lie, because the arrangement does it for them. They are paid by the vendors whose products they recommend, at rates that move with which product you take and how much of it gets consumed, and nobody is obliged to mention a word of it. Put a business in that position for thirty years and dishonesty becomes unnecessary. The shortlist writes itself.\nFrom where you sit, on the receiving end of the invoice, a lie and a rigged process cost exactly the same.\nThis part is the selling: who pays the person advising you, what never reaches the shortlist, why the better answer is the one they will not offer, and what you are quietly renting. Part two is what gets built once the paperwork is signed and who else can reach it. Part three is who carries it when the thing falls over, and what leaving costs. Some of it I have watched happen. The rest is on the public record with a regulator\u0026rsquo;s name on it, and every claim here carries a link.\nNone of it needs you to be technical. You only have to ask.\nThe Easy Fix Is The Tell Take a common one. A home worker cannot get their laptop to hold a VPN tunnel back to the office. The provider looks at it and comes back with something for the customer to change at the home end. Not at their end. At the home end.\nThe tunnel is L2TP over IPsec onto a Cisco Meraki MX, and what is actually wrong is NAT traversal. The client is still negotiating on UDP 500 rather than moving to UDP 4500, because NAT-T was never configured at the office end. RFC 3947 is blunt about it: once a NAT is detected, the initiator \u0026ldquo;MUST set both UDP source and destination ports to 4500\u0026rdquo;. It never does. The tunnel dies in the NAT. Every time.\nThe answer is sat on their own firewall, switched off. The same MX runs AnyConnect, which \u0026ldquo;will attempt to connect using both TLS, and DTLS (Datagram TLS) over TCP and UDP 443 respectively\u0026rdquo;. Ordinary TLS on an ordinary port. A NAT understands that perfectly — there is nothing to traverse and nothing to configure — and it would work the afternoon somebody enabled it.\nNobody takes a packet capture. Nobody checks which port the client is actually talking on. Those are the first two things you do, and they would end it in a minute.\nNow a second one from the same provider, and there is no network anywhere in it. A label printer, and a shipping label that has to come out in the right format. That is the whole of the request. It is also the only thing a label printer does.\nThe answer that came back was that it is not possible.\nIt was an option in the print driver. A box, in a settings dialogue, on software they administer, and the whole job was ticking it and running a test print. Nobody had to buy anything. Nobody had to design anything. They would not tick it.\nNotice that \u0026ldquo;not possible\u0026rdquo; is an answer with no measurement anywhere in it, and it is the one answer that closes a ticket without anybody having to do owt. It also gets harder to walk back the longer it stands, because going to look now means admitting there was something to look at.\nThat is the shape to watch for, and it is the same shape twice. The fault is free to fix, the fix is sat inside kit the provider already owns and is already being paid to run, and what comes back instead is either an instruction to change something at the customer\u0026rsquo;s end or a flat statement that the thing cannot be done. When the cheap answer gets waved away without a measurement, you are not being given a diagnosis. You are being managed.\nAn engineer who has found the fault tells you what the fault is. Somebody who has not found it tells you it cannot be done, or tells you what to buy.\nWho Pays Your Adviser Start with the money, because everything else follows it.\nWhen your MSP recommends a hyperscaler, they are not neutral. Microsoft\u0026rsquo;s own billing documentation describes partner earned credit — a credit applied against the charges on your Azure consumption, earned by the partner who holds admin rights to manage it. Your bill goes up, they get a slice. Sat alongside it are reseller margins, volume tiers, certification rebates, market development funds and quarterly targets, across every vendor in the stack, not just that one.\nNone of this is secret, and none of it is against any rule. It is published, it is normal, and it is how the channel has worked for thirty years.\nTwo people pay your adviser, and only one of them is you You Ask which product to buy Your provider Writes the shortlist you will choose from The vendor Sets what a recommendation earns a fee one option rebate Rebate, margin, deal registration, sales target None of it has to be mentioned to you A financial adviser may only be paid by the client (FCA Handbook, COBS 6.1A). An IT adviser may be paid by both, and nothing obliges them to say which pays more. Two parties pay for the same recommendation. Only one of them is sat in the meeting, and only one of the payments has to be mentioned. Here is the bit that ought to bother you. Nobody has to tell you about any of it — not the rate, not the targets, not the credit. The person recommending the product is paid by the company that makes the product, at a rate that depends on which product they talk you into and on how much of it you then go on to consume, and there is no obligation anywhere to put a word of that in the proposal you are reading.\nAnother industry looked at exactly this arrangement and banned it. Under the FCA\u0026rsquo;s rules a financial adviser must \u0026ldquo;only be remunerated for the personal recommendation \u0026hellip; by adviser charges\u0026rdquo; and must \u0026ldquo;not solicit or accept \u0026hellip; any other commissions, remuneration or benefit\u0026rdquo;. You pay your adviser. The product provider does not. That rule exists because the regulator worked out something obvious. You cannot tell advice from selling when the seller is paid by the manufacturer.\nIT never had that reckoning. The CMA has been through the cloud market and looked hard at committed spend, egress and switching, but nobody has yet looked at the layer in the middle. The outfit sat between you and the vendor, holding both a duty to you and a target from them.\nWorth calling the thing by its name, as well.\nAn incentive is a payment from one party to shape a decision that a different party is relying on. That is the mechanism, whatever the programme is called on the vendor\u0026rsquo;s website. The vendor pays, the customer relies, and the recommendation moves.\nIt is coercive at the MSP\u0026rsquo;s end too, which is the half nobody looks at. Miss the tier and you do not simply forgo a bonus. Your buy price goes up on everything you sell for the next twelve months. So the pressure is not \u0026ldquo;sell this and get a treat\u0026rdquo;. It is \u0026ldquo;sell this or your whole business gets dearer\u0026rdquo;. Nobody in that position is choosing freely, and it was never designed to feel like a choice.\nEnglish law already knows the shape of this. The Bribery Act 2010 needs no public official anywhere near it — section 3 covers \u0026ldquo;any activity connected with a business\u0026rdquo;, and it bites where the person performing that activity is expected to do so \u0026ldquo;in good faith\u0026rdquo;, or \u0026ldquo;impartially\u0026rdquo;, or is \u0026ldquo;in a position of trust\u0026rdquo;. Read those three conditions. Then read the proposal on your desk.\nI am not accusing anybody\u0026rsquo;s account manager of an offence. What is going on is duller than that and harder to fix. The arrangement sits a hair on the right side of the line, and the only thing holding it there is that nobody has ever established that an MSP owes you impartiality in the first place. Financial advice established it, and the commission stopped. Nobody has established it here, so it has not.\nCall an incentive a mild form of corruption and people bristle. Put the payment, the target and the shortlist on the same page, then ask what else to call it.\nThe Small Supplier Never Gets Named The list of suppliers you were shown is not the list of suppliers that exist. It is the list your MSP already has an account with.\nTo get onto that list a vendor needs a partner programme. Tiers, accreditation, rebates, deal registration, a distributor willing to carry the line. That is a machine. Running it costs money that has nothing whatsoever to do with how good the product is. There are plenty of small outfits in this country building better kit and better software than the badge on your proposal, and they will never appear on it, because they have twelve staff and no channel team.\nWatch what the machine does to the recommendation.\nDeal registration ties your MSP to a vendor before anybody has spoken to you about requirements. They log the opportunity, they get a better buy price and protection from another partner quoting you the same thing, and from that moment there is a reason to steer the design towards that vendor which has nowt to do with your business.\nTiers do the rest. Gold, platinum, whatever it is called this year, the status runs on annual volume, and it sets their discount on everything else they sell all year. Your project can be the thing that gets them over the line. You will not be told that.\nThen the distributor decides what is left. An MSP buys through a distie, the distie carries the lines it holds agreements on, and a vendor that is not on the price list may as well not exist.\nWhat that costs you is not abstract. A smaller supplier will usually put you on the phone to the people who wrote the software rather than a first-line script, will change something because you asked them to, and is still answerable to you next year because you are a real part of their revenue instead of a rounding error. None of that fits in a comparison matrix, and none of it pays anybody a rebate.\nSo ask what it would take to get one of those smaller outfits onto the shortlist. The answer tells you who the shortlist was drawn up for.\nAnd The Mandate Comes From One Country Look at the badges on the proposal. The hypervisor, the cloud, the network kit, the firewall, the office suite, the CRM, the backup target, the monitoring. Nearly all of them American.\nNone of that is a verdict on quality. It is what happens when the route to market is a channel programme, because the companies big enough to run one of those at global scale sit in one country. So the targets your MSP is carrying, the rebates that shape your shortlist and the tier that sets their margin are all written in the United States.\nWhich makes it worth knowing how that country currently treats the rules on winning business abroad.\nOn 10 February 2025 the President signed an executive order titled \u0026ldquo;Pausing Foreign Corrupt Practices Act Enforcement to Further American Economic and National Security\u0026rdquo;. It instructed the Attorney General, for 180 days, to \u0026ldquo;cease initiation of any new FCPA investigations or enforcement actions\u0026rdquo;, on the stated reasoning that enforcement against American businesses \u0026ldquo;for routine business practices in other nations\u0026rdquo; harms American competitiveness.\nThe FCPA prosecutes bribing foreign officials. That is the practice being described.\nWhat followed is on the record. New Department of Justice guidelines on 9 June 2025 narrowed enforcement to cases touching cartels, direct harm to American companies or national security. Across 2025 the Securities and Exchange Commission brought no civil FCPA actions at all and disbanded its FCPA unit, while the Department of Justice closed roughly half of its active investigations. It did not turn up to the OECD Working Group on Bribery\u0026rsquo;s March 2025 meeting either.\nSame statute, different direction. It took $772 million off a French engineering company in 2014, and it is the reason the French state commissioned a report on whether American extraterritorial law works as a commercial weapon. I went through that in how far one country\u0026rsquo;s law reaches. So the law reaches across borders for foreign firms and is stood down for domestic ones, by the government of the country your entire product list comes from.\nDo not expect the UK to cover the gap either. The law here is not the weak part — the Bribery Act 2010 goes further than the American statute in places, and section 7 makes it an offence for a commercial organisation to fail to prevent bribery by anyone acting on its behalf. Strict, broad, and on the books for fifteen years.\nThe weak part is that enforcement at this size only ever works as a joint operation. Airbus is the largest bribery settlement this country has been part of, and Airbus\u0026rsquo;s own announcement sets out the shape of it: EUR 3,598 million in penalties on 31 January 2020, going EUR 2,083 million to the French Parquet National Financier, EUR 984 million to the Serious Fraud Office, EUR 526 million to the Department of Justice and EUR 9 million to the State Department, with the SFO and the PNF running it as a joint investigation team. The money crosses several countries and so does the evidence, and no single agency can compel the lot on its own. It needs everybody to turn up.\nNote the date on that. January 2020 sits inside the first Trump administration, and the Department of Justice of the day was happy enough to take its EUR 526 million share of it. The pause came five years later, from the same president in his second term. This is not a difference between administrations. It is a decision taken in 2025.\nOne of them has stopped turning up, and it runs deeper than an empty chair. A prosecutor that is not opening cases is not producing anything to share. No subpoenas, no document productions, no cooperating defendants, no witnesses put under pressure — the evidence that made Airbus possible existed because somebody went out and got it. Close half a docket and bring no new actions, and there is nothing sat in the pot for anybody else to draw on. This is not a country declining to hand things over. It is a country that no longer has anything to hand over.\nIt shuts the other door as well. Airbus\u0026rsquo;s own announcement puts the outcome down to \u0026ldquo;reporting, cooperation and new compliance standards\u0026rdquo; at the company. Firms come forward because of what happens to them if they do not. Remove what happens to them, and the self-reports that start most of these cases stop arriving. In London as much as in Washington.\nThe Bribery Act does not get any weaker when that happens. It just loses the supply of evidence that made it usable on precisely the cases it was written for.\nEighteen months of that, and it has not moved a single shortlist in this industry. Nobody has it in a risk register, nobody asks it on a procurement form, and nobody selling you a five-year subscription has mentioned it.\nThe Training Budget Went First There is a second reason the product keeps winning, and it is less cynical than the first. A lot of the people selling to you could not do the other thing.\nEmployer training investment in this country has been falling for twenty years. The Learning and Work Institute\u0026rsquo;s read of the 2024 Employer Skills Survey puts it at 36% less per employee in real terms than in 2005 — GBP 1,700 against GBP 2,634. Since the 2022 survey alone it is down another 13%. The apprenticeship levy arrived in 2017 to reverse exactly this, and spend per employee including the levy has fallen 23% since.\nThen look at which sectors cut hardest between 2022 and 2024. Public administration at 50%, financial services at 47%, and information and communications at 30%. That last one is ours. Nearly a third gone in two years, from the industry that changes fastest.\nMeanwhile the thing being defended grew. I counted the published vulnerabilities either side of a decade straight out of the National Vulnerability Database: 6,595 CVEs published in 2015, and 49,972 in 2025. Seven and a half times as many in ten years, against a training budget that went down by a third. Those two lines are heading in opposite directions and they have been for years.\nThe controls that would show up the difference are not in place either. The government\u0026rsquo;s own Cyber Security Breaches Survey 2025/2026 puts two-factor authentication at 47% of businesses — the same control the five-eyes agencies named for MSP accounts in 2022, and the same one the ICO fined Advanced over. The same survey has formal cyber security policies down from 59% to 52% in a year, and business continuity plans covering cyber security down from 53% to 44%. Not holding steady. Falling.\nAnd then put a cloud tenancy underneath the lot. A small business does not administer its own. The MSP holds the global admin, builds the identity model, sets the conditional access, decides which storage is public and which is not, and owns the console. That is precisely what it was hired for. So when a tenancy ends up misconfigured, the hands on it were the contractor\u0026rsquo;s. Not a customer clicking the wrong box.\nThe blast radius changed as well. A server misconfigured in 2005 reached about as far as the wire it was plugged into. A tenancy misconfigured today is on the internet the second it is saved, it is the same tenancy for every system that business runs, and the credential that administers it sits with somebody they have never met, at a company they are not allowed to audit.\nThe security agencies of five countries wrote an advisory about this in 2022, and the very first action on their list concerns the accounts a provider uses to get into your systems. They put it first because that is the way in.\nThen there is what this industry calls training. A vendor certification is product training. It teaches you where the buttons are in one company\u0026rsquo;s console and what that company has decided to call its features, it is written and priced by the company whose products it covers, and holding enough of them is a condition of the partner tier that sets the margin. It is a sales channel wearing a mortarboard.\nKnowing a console is not knowing how the thing works. An engineer with five certifications may never have read an RFC, never taken a packet capture, and never once worked out why something failed from first principles. Put that person in front of a tunnel that will not come up and the honest answer is not available to them. Blaming the IPv6 on the customer\u0026rsquo;s home network is.\nSo the two halves meet. They are paid to sell a product, and more and more the product is the only answer they have got. A customer paying for expertise ends up with neither.\nNobody Wrote Down What You Needed Ask to see your requirements. Not the proposal, not the quote, not the architecture diagram with your logo on it. The requirements.\nA proper capture is boring and it is not short. What does the service have to do, and for whom. How many people, from where, on what. What is the busiest hour and what is the growth over three years. How long can it be down before it costs real money, and how much data can you afford to lose. What must never leave the country, and under whose law. What has to survive a fire in one building. What are you contractually on the hook for to your own customers. What is the budget, capital and revenue, split out. What do you already own that still has life in it. Who keeps it running afterwards, and what do they already know how to run.\nThat is a morning\u0026rsquo;s work with the right people in the room — and it should end in a document you sign off before anybody draws a single box.\nIf nobody asked you most of that, you have not been given a design. You have been given the thing they already sell, with your company name on the cover.\nWatch for the tell in the other direction as well. If the sizing exercise happened in the first meeting, before anybody looked at what your load actually is, then the numbers came from a template rather than from your estate, and real sizing needs data. Somebody looking at what your kit is actually doing now, for long enough to see a month-end and a quarter close.\nOne Option Is Not A Choice A design is a set of choices with the reasoning attached. Which means options, and costs on all of them, not just the one they want you to take.\nYou should get the do-nothing, priced, including what it costs you when it breaks. You should get the cheapest thing that meets the requirements. You should get the recommendation, and you should get the one that is over-specified for you, so you can see where the line is. Each one with what it costs to buy, what it costs to run for five years, what it does not do — and what you would have to do next if you outgrew it.\nAnd crucially, you should get the list of what was ruled out and why. That is the part that shows somebody actually thought.\nIf you received exactly one answer, and that answer happens to be the vendor they are certified in and the licensing model that pays monthly, you did not get a design. You got a quote wearing a design\u0026rsquo;s clothes. Nothing more.\nWhat the shortlist filters out before you ever see it Four answers to the same requirement Open source, on a support contract Margin is your provider's own labour A smaller British supplier No partner programme to be on What you already own, configured Nothing to invoice at renewal The vendor on the partner programme Rebate, deal registration, target Does it pay? Is it on the programme? What reaches your desk One option, quoted No comparison, no costed alternative The other three were never priced, so you never learned what they cost and their absence looks like there was nothing to say The filter runs before you see anything. Three answers to the same requirement never get priced, so their absence reads as though there was nothing to compare. There is a simple question that flushes this out, and I would ask it in the room. What else did you consider, and what did it cost? Somebody who did the work has the numbers to hand and rather enjoys being asked. Somebody who did not will tell you the alternatives are not supported, not enterprise-grade, or not something they would put their name to. None of those is a number.\nOpen Source Never Makes The List It is not that they hate it. There is no margin on it, no rebate against it, no certification to sell and no quarterly target that it moves.\nThe objection is always the same, and it is the one claim in this post that is flatly untrue — \u0026ldquo;it is not supported\u0026rdquo;. All of it is supported, commercially, with a contract and an SLA and somebody to ring — Proxmox sell subscriptions per socket, and Red Hat, SUSE and Canonical sell support for the stack. You are choosing who supports it rather than choosing whether it is supported at all. What you stop paying for is the right to use software you already have.\nThen there is the kit already on the floor. It gets missed entirely. At my last employer I built a private cloud out of decommissioned HPC nodes — Proxmox, Ceph and a full chain of Ansible on top — and ended up with 7 servers, 480 cores, 15 TiB of RAM and 1.8 PiB of storage, on no budget and with three people. An MSP quoting that same requirement would have costed new hardware and a subscription — there is nothing in it for them in kit you have already bought and paid for.\nThe other half of this is the kit already sat on your floor, and it is the half that never gets costed at all. A server does not stop working on the day its support contract runs out. \u0026ldquo;End of support\u0026rdquo; is a date the vendor chose, not a measurement anybody took of the hardware, and a box with five good years left in it is worth more to you than to anybody selling its replacement. Open source is what lets you keep using it, because the licence does not care how old the CPU is or whether the badge on the front is still in warranty.\nThat is the bit that pays nobody. There is no rebate on hardware you already own, no tier credit for a machine that stays where it is, and no renewal on a licence nobody had to buy. As such it does not get proposed, and the phrase reached for instead is \u0026ldquo;end of life\u0026rdquo;, which sounds like engineering and is a sales date.\nAsk for the open source option to be costed properly, support included, alongside the others — and ask for the version that reuses what you have got, priced against the version that does not. Not to be talked out of the commercial one. To see the gap, so the decision is yours.\nWhat The Alternative Actually Looks Like There is a supported open source answer to nearly everything on a proposal, and it is worth starting with the part that costs you most, because it is never the servers.\nThe per-seat software. This is where the recurring money is, and where an alternative is never mentioned. Office suite: LibreOffice, ONLYOFFICE, or Collabora Online, which sells supported deployments and gives you the browser-based editing people think only comes from one place. Mail, calendars and shared contacts: grommunio speaks MAPI, so Outlook connects to it like it would to Exchange, and it is sold with support; SOGo and mailcow cover the same ground differently. Files, sharing and the bits people actually use a cloud drive for: Nextcloud, with an enterprise contract behind it. Chat and meetings, which is the Teams half: Mattermost and Rocket.Chat are the closest like-for-like, both self-hostable with support to buy, and Zulip is Apache-licensed and threads properly. Matrix with Element, Nextcloud Talk and Jitsi cover the rest. And the SharePoint half — intranet, document libraries, team sites — is XWiki, BookStack, OpenProject, Seafile and Nextcloud between them.\nThe business applications. ERP: Odoo Community, ERPNext, Dolibarr. CRM: SuiteCRM, which the company behind it sells support for, or EspoCRM. Accounting: GnuCash, or the ledger built into Dolibarr and ERPNext. Document management: Paperless-ngx. Reporting: Metabase. Service desk and asset tracking, which your MSP is charging you for as a product: GLPI and Zammad.\nIdentity and secrets, which everything else hangs off. Keycloak, FreeIPA, or Samba as a domain controller. For passwords, Bitwarden can be self-hosted, Vaultwarden is a lightweight AGPL server that speaks to the same clients, and Passbolt and KeePassXC cover the same job differently. This is a per-user monthly line on most proposals.\nDevice management, which is the Intune line on your bill. Fleet does inventory, policy and enrolment across the platforms an estate actually contains, built on osquery. MicroMDM handles Apple enrolment, Headwind handles Android, and Ansible, Puppet or Salt do the configuration underneath any of it. For the remote-access half — the thing a later section of this post is entirely about — MeshCentral is Apache-licensed and RustDesk is AGPL, and both run on a server you own. Which means the answer to \u0026ldquo;which remote tool is on my machines, and who can log into the console\u0026rdquo; can be one you host and patch yourself, rather than one you find out about afterwards.\nAnd the updating itself is no longer the dark art it gets sold as. Even on the Fisher-Price OS (Windows), application updates now run through winget, which is MIT-licensed, against a community manifest repository under the same licence; Chocolatey and Scoop have been doing the same job for longer, and every Unix has had it since the nineties. So when patch management turns up on a proposal as a premium managed line, look at what is actually being sold. The mechanism is free and shipped by the platform vendor, and setting it up is an afternoon. After that it is a scheduled task. A cron entry, or whatever the console calls one, running on a machine that was already there. You are paying a monthly fee for a job that runs itself.\nThe answer to that is supposed to be that you are paying for somebody to watch the result and act when it fails, and that would be worth the money. So look at the evidence for the watching. At Capita the alert was raised in ten minutes and acted on the best part of three days later, against a one-hour target, by a team the regulator found understaffed. Out on the internet there are firewalls still carrying vulnerabilities that went into the exploited catalogue years after a fix shipped. The watching is the part that would justify the invoice, and it is the part with the least evidence of happening.\nThe phone system, which is one of the oldest per-extension monthly lines there is. Asterisk has been doing this for twenty-five years, FreePBX puts a web interface on top of it, FreeSWITCH is the other engine, and Kamailio and OpenSIPS handle SIP routing at carrier scale. If you want it packaged rather than assembled, Wazo, FusionPBX and Issabel all ship it built.\nIt is also worth remembering what happened to the proprietary one most providers propose. In March 2023 CISA published an alert stating that \u0026ldquo;3CXDesktopApp — a voice and video conferencing app — was trojanized, potentially leading to multi-staged attacks against users employing the vulnerable app\u0026rdquo;. Nobody had to find a vulnerability and exploit it. It arrived as the vendor\u0026rsquo;s own signed application through the vendor\u0026rsquo;s own update channel, onto every desk a partner had rolled it out to.\nAdobe, which is a per-seat subscription like everything else. Most businesses are not paying for a creative suite at all. They are paying for Acrobat and for signatures. Stirling PDF is a self-hosted toolkit doing the merging, splitting, redaction, OCR and form filling that Acrobat Pro gets bought for, and Okular or LibreOffice Draw cover the rest. For signing, Documenso and DocuSeal are both AGPL and both self-hostable — and a per-envelope charge for signing a document is about as clean an example as this post has of metering something your own server would do for nothing. Where there really is a design team: GIMP and Krita for images, Inkscape for vector, Scribus for layout, darktable and RawTherapee for photography, Kdenlive for video, Blender for 3D and compositing, Audacity and Ardour for audio.\nVideo, which is worth a paragraph of its own. Cameras get sold as a managed product with a licence per channel and a recorder you are not allowed inside. Frigate is MIT-licensed and does object detection locally on hardware you own; ZoneMinder has been GPL for twenty years. For conferencing, Jitsi and BigBlueButton. And remember which product it was that Cisco paid $8.6 million and another $6 million to settle over, a couple of sections down from here. Video surveillance software, sold to government bodies. The premium camera stack does not come with safety attached.\nStorage and filesystems, and note that nobody ever offers you a choice here at all. OpenZFS for checksummed storage with snapshots and send/receive, Btrfs, XFS, CephFS. The filesystem under your data is an engineering decision with real consequences for how much of it you get back after a bad day, and it usually arrives as whatever the appliance shipped with.\nThe infrastructure, last, because it is the cheapest part of the bill. Hypervisor and cluster: Proxmox VE. Storage: Ceph, which will happily run on the disks you already own instead of a new array. Firewall and routing: nftables, OPNsense. VPN: WireGuard or strongSwan. For the mesh overlay everybody now sells, it is worth getting the picture right. Tailscale\u0026rsquo;s clients are open source and its own hosted coordination server is not — the company states that it \u0026ldquo;remains proprietary as part of our managed service\u0026rdquo;. But there is an open coordination server, and it is a serious one: Headscale is BSD-licensed, \u0026ldquo;an open source, self-hosted implementation of the Tailscale control server\u0026rdquo;, and Tailscale employs its head maintainer while saying it \u0026ldquo;does not set Headscale\u0026rsquo;s product direction\u0026rdquo;. So the whole thing can be run on your own kit. NetBird and Nebula are the other routes. Load balancing: HAProxy and keepalived. Monitoring: Prometheus, Grafana, Zabbix. Backup: Proxmox Backup Server, Bareos, restic. Config management: Ansible. And for the integration work that gets quoted either as bespoke development or as a per-run cloud automation subscription, Node-RED is Apache-licensed, runs on a box you own, and does not bill you per execution.\nThree places where I will not pretend the swap is clean. First, Teams and SharePoint. The destination is not the problem, whatever you get told — the products above are mature and businesses run on them. The migration is the problem: years of accumulated sites, permission inheritance nobody ever documented, Power Automate flows somebody built and then left, and co-authoring behaviour staff expect without being able to name it. That is real work, and it should be costed as real work rather than waved away in either direction. Though notice what you stop carrying as well. Those flows and those sites are business processes running in somebody else\u0026rsquo;s datacentre, on a platform you cannot restart. When they stop, you do not fix them. You wait, and you tell your own customers you are waiting. And UK accounting has a hard edge: VAT has to be filed through software HMRC recognises, and HMRC publishes the list. The open source routes to Making Tax Digital exist — GnuCash has a community bridge, ERPNext has a UK VAT module — but they are community-maintained rather than a product with a support contract behind them. That is a real limitation and it belongs in the comparison. Third, a working design agency lives or dies on file interchange, and clients who send you a .psd and expect one back are a real problem rather than a matter of principle. For everybody else, who needs to fill in a PDF and get it signed, there is no swap to make at all. It is just stopping paying.\nWhy The Better Business Never Gets Sold So why is none of it offered? Because it would be better business for them, not worse, which makes the refusal more interesting rather than less.\nThere are two ways to make money out of a customer. Resell a licence, and the margin is set by somebody else, capped by them, repriced by them at renewal, and paid whether you did any work that month or not. Deploy and run an open stack, and the margin is your own labour at your own rate, with nobody taking a cut on the way through and no vendor able to change the number next April. The second is worth more per customer, and it builds something. An engineer who can run Ceph, or stand up a mail server that Outlook talks to, is worth more next year than they are this year. Somebody who has only ever driven a console is worth exactly the same next year, and only for as long as that console exists.\nIt is also harder, and that is the whole of it. It has to be earned again every month. It needs engineers rather than administrators, and engineers are expensive, take years to grow, and can walk out and set up on their own. A licence never leaves.\nThat is a commitment, and commitment is the thing being avoided. Reselling asks nothing of anybody. Pick a vendor this year, pick a different one next year, and when it falls over it was never your design in the first place. Running the stack yourself means choosing it, learning it properly, and standing behind it in front of a customer at two in the morning. One of those needs you to know something. The other needs a portal login.\nIt also breaks the arithmetic the model runs on. A managed service is priced on leverage — as many customers as possible per engineer, junior staff working down runbooks, escalation only when the runbook runs out. That works because the job has been reduced to steps. Put in a stack that somebody has to actually understand and the ratio collapses, the wage bill climbs, and the business finds itself depending on people who could leave and take customers with them.\nSo the question underneath the shortlist was never which product is better. It is whether a firm is willing to be the sort that employs people who know things.\nWhich is where the training figure from earlier comes back round. A sector that has cut its spend on skills by 30% in two years cannot sell skill, so it sells licences. Selling licences means it never has to acquire the skill, so the training stays cut, and next year there is even less it can offer. That is a loop, and it only turns one way.\nThe rest is incentive, and we have been through it already. The vendor pays a rebate on the licence and nothing at all on the labour. The salesperson is compensated on product. And when a hyperscaler has an outage it is the hyperscaler\u0026rsquo;s fault, while a cluster you built yourself is yours — so reselling buys somebody a place to put the blame. That is worth real money to a business that would rather not be accountable.\nNone of that makes it the right answer for you. It just explains why the comparison never gets written.\nThe IPv4 Tax This one is the cleanest example in the whole post, because you can put a price on both sides.\nThey will build you an IPv4-only service. Then they will sell you the public addresses you need to reach it, per address, per month, forever. Ask for another and there is a form, a justification and a line on the renewal.\nMeanwhile IPv6 costs nothing extra. The allocation comes with the registry membership you are already paying for, and RIPE\u0026rsquo;s 2026 charging scheme is a flat EUR 1,800 per LIR account whatever you hold. RIPE-690 says an end site gets a /48 or a /56. There is no per-address meter on the v6 side. There is nothing to meter.\nThe numbers on the other side are public too. AWS charges $0.005 per hour for every public IPv4 address, attached or not, which is about $43.80 a year each. On the transfer market the average price in the first half of 2026 was $20.04 per address, with the going lease rate about $0.59 per address per month. So an address is worth roughly twenty dollars to buy outright, and rents wholesale for about seven a year — and it is being resold to you as a scarce resource. A monthly line on your bill and a form to fill in when you want another.\nThe scarcity is real. The reason you are still paying for it is not. Dual-stacking a service is an afternoon\u0026rsquo;s work, and it turns a recurring charge into a one-off change request, which is precisely why it never gets proposed.\nIf you want the full picture of who in this country has actually bothered, I counted the lot: We Never Ran Out of Addresses. We Ran Out of Effort.\nCloud By Default, When Everything You Have Is On Site Think about where the work actually happens. A manufacturing site. A car mechanic. A catering business.\nEvery user is in the building. Every bit of data is made in the building — the machines, the tills, the job cards, the stock, the CAD, the CNC programs, the orders coming over the counter. Everything that consumes that data is in the building as well. No second site, no field force, no customers hitting a web front end, and no month where the load doubles.\nPut the application in somebody else\u0026rsquo;s datacentre and every one of those bytes now leaves the premises and comes straight back. Same users. Same data. Same processing. Plus a WAN in the middle of it, a monthly bill, and a dependency on a line you do not own.\nNothing cloud is genuinely good at applies here. Elastic scale is for load that moves, and a shop floor running two shifts does not move. Global reach is for users who are somewhere else, and yours are stood at the machine. Somebody else\u0026rsquo;s building is a real answer for disaster recovery, but disaster recovery is a backup target, not the place you run the line from.\nWhat it does buy you is a new way to stop. On prem a dead broadband line is an inconvenience, and the work carries on until somebody fixes it. In the cloud it is a stoppage — the line goes, the garage cannot pull up a job, the kitchen cannot take an order, production stands there, and you are waiting on somebody else\u0026rsquo;s engineer against an SLA you never negotiated.\nWhere the work actually goes, once the application leaves the building On site Your building The people Stood at the machine The data Made here, all of it The application Also here Nothing leaves the premises. A dead broadband line is an inconvenience, and the work carries on until somebody fixes it. Cloud by default, and the route it actually takes Your building The same people, the same data The exchange where your line lands Your ISP's core and whoever they buy transit from A peering point or the provider's edge network The datacentre The application lives here now Three buildings you do not own, cannot ring, and never chose and every byte makes the return trip as well, for every save and every lookup A dead line is now a stoppage. The garage cannot pull up a job, the kitchen cannot take an order, production stands there, and you are waiting on somebody else\u0026#8217;s engineer against an SLA you never negotiated. Same users, same data, same processing. The difference is four other people\u0026rsquo;s networks in the middle, none of which you can ring when it stops. And your broadband is only half of it. The other half is theirs, and it goes down too. AWS\u0026rsquo;s own post-event summary for October 2025 describes a disruption in its primary Northern Virginia region running from 11:48 PM on 19 October to 2:20 PM on 20 October — the better part of fifteen hours — where customers and other AWS services \u0026ldquo;were unable to establish new connections\u0026rdquo;. The cause, in their words, was \u0026ldquo;a latent defect within the service\u0026rsquo;s automated DNS management system\u0026rdquo;. Nine days later Azure Front Door took Microsoft 365, Outlook and the Azure portal down with it for most of a working day. Microsoft keeps a running history of these, and it is not a short document.\nDowntime is not the whole of it either. The other thing sold on the front of the proposal was capacity on demand, and that has a documented failure mode of its own. Microsoft publishes a page titled \u0026ldquo;Troubleshooting Azure VM allocation failures\u0026rdquo;. AWS publishes \u0026ldquo;How do I troubleshoot InsufficientInstanceCapacity errors when I start or launch an EC2 instance?\u0026rdquo;. Read those titles again. Both vendors maintain standing documentation for the case where you ask for a machine and there is not one — no fault, no outage, just nothing spare in that region today. On top of that sit quotas, set per subscription, which is a second way to be told no.\nSo the elastic scale on the proposal carries a caveat the proposal does not. Elastic within whatever is spare, in that region, on the day you ask. And providers keep signing customers into constrained regions regardless, because a signature counts this quarter and a capacity check does not. You find out at deployment, which is the worst moment available — after the migration is committed, after the on-premises kit has gone, and after the fallback was torn down to pay for the move.\nRead what all of that means for the garage and the kitchen. On your own kit, an outage is somebody you can ring, or a machine somebody can walk over to and restart. In somebody else\u0026rsquo;s cloud there is no lever at all. You cannot escalate it, you cannot prioritise your own recovery, and neither can the provider you are paying. They are refreshing the same status page as everybody else. What you bought as resilience turns out to be a dependency you share with several million other customers, and on the day it fails your position is the same as theirs.\nSo why is it always the recommendation? Because a capital purchase pays your MSP once. A subscription pays them every month, at a percentage, with a credit from the vendor on top. Migrating you also moves the hardware off their plate, which is the part of the service they find hardest to staff. Three reasons pointing the same way, none of them yours.\nThen there is where your data ends up. Under 18 U.S.C. § 2713, added by the CLOUD Act, a US provider must produce data in its \u0026ldquo;possession, custody, or control\u0026rdquo; when properly served, \u0026ldquo;regardless of whether such communication, record, or other information is located within or outside of the United States\u0026rdquo;. So \u0026ldquo;it is in the UK region\u0026rdquo; is a true answer to a question you did not ask. The region tells you where the disk is. The statute follows who owns the company.\nNobody put that to you as a decision. It arrived as an assumption, inside a proposal, written by somebody who gets paid more when you say yes.\nI have written about that dependency at length in what renting your technology costs.\nAll Of That Happened Before Anything Was Built Notice when it takes place, the whole of it. Before a machine is racked, before anybody has logged into anything, in meetings and on spreadsheets you were mostly not in. The rebate, the requirements nobody captured, the single option, the missing open source line, the addresses you will now rent for the life of the contract, the datacentre nobody in the building needed — none of it is technical, and all of it is settled in the first fortnight.\nWhich is why it is worth your attention even if you have never opened a switch. It is also why it is so hard to unpick afterwards. Everything downstream inherits that fortnight. The estate gets built the way the shortlist said it would. The tooling turns up because it is what the provider already owns and already knows. The keys end up wherever their process puts them. And the contract signed at the end of it decides, years ahead of time, who is carrying the loss on the morning something stops.\nSo part two is what actually got built: the box they will not be talked out of, the perimeter kit with the worst record on the exploited list, the basics that were the thing you bought, and the tooling that reaches every machine you own from a console you have never seen. Two of the failures in it have a regulator\u0026rsquo;s finding attached. None of them happened to a customer. They happened to a provider, and the customers were downstream.\nIs Your MSP Lying To You — 3 parts\nIs Your MSP Lying To You To Sell You Premium Products?you are here What Your MSP Built You, And Who Else Can Reach It When It Breaks, Who Actually Carries It? Sources Retrieved 28 August 2026.\nHow the money works.\nMicrosoft Partner Center billing documentation — partner earned credit applied against charges on customer Azure consumption. FCA Handbook, COBS 6.1A — adviser charging: the rule that a firm must only be paid for a recommendation by the client, and must not accept commission from the product provider. CMA cloud services market investigation — the UK competition regulator\u0026rsquo;s work on the cloud market. Addresses.\nRIPE NCC charging scheme 2026 — EUR 1,800 per LIR account, flat. RIPE-690 — /48 or /56 to an end site, October 2017. AWS public IPv4 charge — $0.005 per address per hour, from February 2024. IPv4 transfer market, first half of 2026 — average $20.04 per address, lease rate about $0.59 per address per month, summarising CircleID\u0026rsquo;s analysis of publicly priced transactions. The VPN.\nCisco Meraki\u0026rsquo;s AnyConnect troubleshooting guide — TLS and DTLS on 443, with no NAT problem to solve. RFC 3947 and RFC 3948 — NAT traversal for IPsec, and the move to UDP 4500. Support you can actually buy.\nProxmox VE subscriptions — one example of commercial support for open source infrastructure. Training and skills.\nLearning and Work Institute, 24 December 2025 — employer training investment down 36% per employee in real terms since 2005, and down 30% in information and communications between 2022 and 2024, on the 2024 Employer Skills Survey. National Vulnerability Database — CVE publication counts by year, summed per quarter: 6,595 in 2015 against 49,972 in 2025. Cyber Security Breaches Survey 2025/2026 — two-factor authentication at 47% of businesses, formal policies down to 52%, continuity plans covering cyber down to 44%. CISA alert, 30 March 2023 — the trojanised 3CX desktop application. Law and policy.\nExecutive order, 10 February 2025 — pausing Foreign Corrupt Practices Act enforcement, in the administration\u0026rsquo;s own words. Just Security, on the year that followed — the June 2025 enforcement guidelines, the SEC\u0026rsquo;s disbanded FCPA unit, and the closed investigations. Bribery Act 2010 and section 7 — the UK statute, and the corporate failure-to-prevent offence. Airbus, 31 January 2020 — the company\u0026rsquo;s own account of the tripartite settlement and how the penalties split between the PNF, the SFO, the DoJ and the DoS. Jurisdiction.\n18 U.S.C. § 2713 — the CLOUD Act provision: disclosure regardless of where the data is stored. When the platform stops.\nAWS post-event summary, October 2025 — the DynamoDB and DNS disruption in us-east-1, in Amazon\u0026rsquo;s own words. Azure status history — Microsoft\u0026rsquo;s own running record of its incidents. ","permalink":"https://blogs.damiendye.uk/en/random/is-your-msp-lying-to-you-part1/","summary":"Part 1 of 3. Some are lying. Most never have to, because they are paid by the vendor whose product they are recommending and nobody has to tell you. The tells that say you are being sold to rather than engineered for, and what never makes it onto the shortlist.","title":"Is Your MSP Lying To You To Sell You Premium Products?"},{"content":" Is Your MSP Lying To You — 3 parts\nIs Your MSP Lying To You To Sell You Premium Products? What Your MSP Built You, And Who Else Can Reach Ityou are here When It Breaks, Who Actually Carries It? Part one was the sale. This is the estate.\nThe paperwork is signed, the invoice runs monthly, and there is kit. Some of it yours, some of it theirs, most of it picked before anybody asked what the business actually does between eight and six. What follows is the kit itself: what gets picked, what got left switched on inside it, and who else can reach the machines you paid for from a console you have never logged into.\nTwo of the failures here carry a regulator\u0026rsquo;s finding with a number attached. None of them happened to a customer — they happened to a provider, and the customers were downstream. Every claim has a link on it.\nThe Box They Will Not Be Talked Out Of Now the awkward one. I want to do this with evidence rather than assertion, because it is the section where people reach for the word conspiracy — and reaching for that word is how the subject gets dropped without anybody having to look at it.\nCisco is the default recommendation across most of this industry. So look at what the vendor itself has published about undocumented access into its own kit.\nIn March 2018 they published an advisory for CVE-2018-0141: Prime Collaboration Provisioning could be logged into over SSH because of \u0026ldquo;a hard-coded account password on the system\u0026rdquo;. Eight months later, CVE-2018-15439 on the Small Business switches, where \u0026ldquo;the affected software enables a privileged user account without notifying administrators of the system\u0026rdquo;, with no fixed software available when it was published and a workaround offered instead. In October 2023, CVE-2023-20101: Emergency Responder shipped with a root account holding \u0026ldquo;default, static credentials that cannot be changed or deleted\u0026rdquo;, credentials \u0026ldquo;typically reserved for use during development\u0026rdquo;.\nAn account nobody was told about, that you cannot remove, with root on a box in your estate. Call it what you like. It is what the vendor\u0026rsquo;s own advisory describes.\nThen there is the Snowden material, which is now more than a decade old and has never been retracted. In December 2013 Der Spiegel published the ANT catalogue, reporting that an NSA division \u0026ldquo;has burrowed its way into nearly all the security architecture made by the major players in the industry \u0026ndash; including American global market leader Cisco and its Chinese competitor Huawei\u0026rdquo;. The same reporting described how kit gets there: shipments diverted to secret workshops in a process the NSA calls interdiction, where \u0026ldquo;at these so-called \u0026rsquo;load stations,\u0026rsquo; agents carefully open the package in order to load malware onto the electronics, or even install hardware components that can provide backdoor access\u0026rdquo;. An internal NSA newsletter from 2010, published in 2014 with the photographs, set out the process in the agency\u0026rsquo;s own words: devices \u0026ldquo;being delivered to our targets throughout the world are intercepted\u0026rdquo;, then \u0026ldquo;re-packaged and placed back into transit to the original destination\u0026rdquo;.\nWhat the reporting does not show is complicity. Der Spiegel said plainly that nothing in the documents suggested the manufacturers knew or helped, and Cisco denied any involvement at the time, and in May 2014 its general counsel put it on the record: \u0026ldquo;as a matter of policy and practice, Cisco does not work with any government, including the United States Government, to weaken our products\u0026rdquo;. They complained to the President about it. On the face of it, they were a victim here too.\nTake that at face value. It changes nothing about the box on your rack, because interdiction never needed the vendor\u0026rsquo;s help. And it is worth knowing what weight the denial carries, because Cisco\u0026rsquo;s word has been tested in court. In 2019 they settled False Claims Act proceedings for $8.6 million federally, plus $6 million across a group of states, over video surveillance software sold to government bodies with \u0026ldquo;flaws that would permit unauthorized access to the system, with the potential to control and otherwise manipulate security cameras and the recorded footage\u0026rdquo;. The New York Attorney General\u0026rsquo;s account is that Cisco knew about the flaws in 2009 and did not fix them until 2013, after the investigation had started. Settlements are not admissions of liability and I will not pretend otherwise. They are still four years of selling a product to police forces and public bodies with a known way in, and not saying so.\nSo the record is: undocumented root accounts by their own admission, a decade-old catalogue naming them, a shipping chain demonstrated to be compromisable, and a period where they knew about a hole and kept selling. Any single one of those you could shrug off. Together they are a pattern. A statement from an interested party does not settle a pattern.\nThe answer to that is controls, not a ban. The risk belongs in the design document with everything else, and the alternatives get priced instead of dismissed. Ask which competing kit was evaluated and what it cost. Ask what the firmware verification and update policy is, and who checks it. Ask how the kit is received and inspected, and by whom. Ask what happens on a serial number that does not match the purchase order.\nAnd notice which vendors get this scrutiny in your sector and which do not. The same industry that put Huawei on a risk register over a capability nobody has produced in public will write Cisco into the low-level design without a line of justification, on the strength of a partner badge. I have written about what happens to that argument when you apply the standard evenly. A partner badge is a commercial relationship with targets attached. As such, it is not a security finding.\nThe Firewall Is The Way In Widen that from one vendor to the category, because the firewall is the product an MSP sells hardest. It is the line item that justifies the security part of the contract.\nCISA keeps a catalogue of vulnerabilities known to be exploited in the wild — not theoretically dangerous, actually used against somebody. I pulled the catalogue and counted it by vendor. Version 2026.08.27 holds 1,685 entries. Cisco accounts for 96 of them, second only to Microsoft\u0026rsquo;s 386, and 42 of those sit on the firewall and edge lines. Fortinet has 29. Palo Alto Networks has 15. For scale, Ivanti has 35 and SonicWall 17.\nLook at what the entries actually are, because the pattern never varies. It is the management interface or the VPN, which is to say the part deliberately exposed to the internet.\nThe only part of this timeline that is yours to close One vulnerability, from published to patched on your box Published a CVE number exists Patch available the vendor's part is done Confirmed exploited added to CISA's catalogue Your box patched by whoever you pay Your exposure Known to be used against somebody, still sat on your perimeter Cisco 96 entries, Fortinet 29, Palo Alto 15. Those are the vendors' numbers and they are public. The length of that shaded span is yours, and nobody is asking for it. Every provider has the number. It is the same measurement the ICO took at Capita. The vendor\u0026rsquo;s total is public and is not the number that decides whether you get hurt. The shaded span is, and it belongs to whoever you pay. Palo Alto\u0026rsquo;s CVE-2024-3400 was a command injection in the GlobalProtect feature of PAN-OS, scoring 10.0, exploited as a zero-day. Later the same year CVE-2024-0012 let \u0026ldquo;an unauthenticated attacker with network access to the management web interface\u0026rdquo; straight past authentication at 9.8, and it was chained with a command injection in the same interface.\nFortinet\u0026rsquo;s CVE-2022-40684 was an \u0026ldquo;authentication bypass using an alternate path or channel\u0026rdquo; at 9.8. Its CVE-2018-13379, an SSL VPN path traversal, went into the catalogue in November 2021. Years after the patch existed, because boxes were still sat on the internet unpatched and being walked into.\nCisco\u0026rsquo;s CVE-2023-20198 in the IOS XE web UI scored 10.0. CVE-2025-20333, in the VPN web server of its Secure Firewall ASA and FTD software, scored 9.9 last September.\nAnd then there is CVE-2026-20316, added to the catalogue on 29 July 2026. Cisco Secure Firewall Management Center, \u0026ldquo;use of hard-coded password\u0026rdquo;, letting \u0026ldquo;an unauthenticated, remote attacker to log in to an affected device\u0026rdquo;. A static credential, in the box that manages the firewalls, confirmed as being exploited, a month before this post. The same failure as 2018 and 2023 in the section above, on the machine that administers your perimeter.\nNow the useful part, because \u0026ldquo;some vendors are worse than others\u0026rdquo; is not really the lesson. Every one of these has a public record and none of it appears in a proposal. More to the point, the vendor\u0026rsquo;s total is not the number that decides whether you get hurt. Your exposure is the gap between a vulnerability going into that catalogue and your box being patched. That gap is not the vendor\u0026rsquo;s to close. It belongs to whoever you pay to run the thing.\nWhich is the same measurement the ICO took at Capita. Alert at ten minutes, action at fifty-eight hours, target of one. Nobody is asking their provider for that number on firewall patching, and it is a number every provider has.\nWhen The Basics Are The Thing You Bought Part one was about the sale. This is what came after it, in two cases where a regulator did the investigating and published what it found.\nIn March 2025 the Information Commissioner fined Advanced Computer Software Group £3.07 million over a ransomware incident in August 2022. Advanced \u0026ldquo;provides IT and software services to organisations, including the NHS and other healthcare providers\u0026rdquo;. The attackers got in \u0026ldquo;via a customer account that did not have multi-factor authentication\u0026rdquo;. NHS 111 was disrupted, healthcare staff could not reach patient records, and the personal information of 79,404 people was taken. The ICO\u0026rsquo;s finding on the cause is worth reading slowly: \u0026ldquo;while Advanced had installed multi-factor authentication across many of its systems, the lack of complete coverage meant hackers could gain access\u0026rdquo;. It was also the first penalty the ICO has issued against a data processor rather than the organisation whose data it was.\nIn October 2025 the same regulator fined Capita £14 million over the 2023 attack that took the personal information of 6.6 million people. A malicious file landed on an employee device on 22 March 2023. A high priority alert fired within ten minutes. The device was not quarantined for 58 hours, against a target response time of one hour, and the ICO found that the Security Operations Centre \u0026ldquo;was understaffed, and in at least six months before the incident fell well below the target response times for responding to security alerts\u0026rdquo;, alongside inadequate penetration testing and risk assessment.\nRead what those two findings have in common. Not a clever adversary. Not a zero-day. Not owt nobody could have seen. Multi-factor authentication that was bought but not finished, and an alert queue nobody had staffed. Both are line items somebody signed off and reported green.\nAnd none of this crept up on the industry. On 11 May 2022, three months before the Advanced incident, the cyber security agencies of the UK, Australia, Canada, New Zealand and the United States put out a joint advisory on threats to managed service providers and their customers, because they were \u0026ldquo;aware of recent reports that observe an increase in malicious cyber activity targeting managed service providers\u0026rdquo;. The first tactical action on the list is to enforce multi-factor authentication on the MSP accounts that reach into the customer\u0026rsquo;s environment.\nFive countries\u0026rsquo; security agencies wrote it down, in public, and named the control. Three years later a regulator is still fining people for not finishing it.\nNeither of those is a small business story. Here is one. On Friday 24 November 2023 the legal sector IT provider CTS went down in a cyber incident. The Law Society\u0026rsquo;s own paper reported that around 80 firms were unable to complete transactions, with systems offline and contracts stuck. These are conveyancing practices, most of them small firms, and their customers were people in the middle of moving house. One of those put it plainly: \u0026ldquo;This leaves us with so much uncertainty — with movers needing to be booked, days needing to be taken off work at potentially short notice and lives put on hold.\u0026rdquo; CTS could say only that it was \u0026ldquo;unable to give a precise timeline for full restoration\u0026rdquo;.\nNobody in that chain picked CTS. The firm picked it, and every client downstream of the firm inherited the decision without ever being asked.\nThat is what you are buying when you buy a managed service.\nAnd How Would You Have Found Out? Notice something about every incident in the section above. None of them happened to the customer. They happened to the provider, and the customers were sat downstream of somebody else\u0026rsquo;s bad week.\nSo put the question directly. If you are breached, do we hear it from you? And how fast?\nThere is a legal floor, and it is lower than people assume. Where they process personal data on your behalf, Article 33 says \u0026ldquo;the processor shall notify the controller without undue delay after becoming aware of a personal data breach\u0026rdquo;. No fixed number of hours. Meanwhile you, as the controller, have 72 hours to tell the Information Commissioner once you become aware. Read those two together. Every hour they spend deciding whether to tell you is an hour off a clock that runs against you, not them.\nAnd that floor only covers personal data. An intrusion into their management console, with no evidence yet that anything of yours moved, may generate no duty to tell you at all — while the credentials that reach every machine you own sit in somebody else\u0026rsquo;s hands. That is the gap. It is not a small one.\nThink about who gets told before you. Their lawyers, because privilege. Their insurer, because the policy says so. Their regulator, on the statutory clock. You are further down that list than you would like to be, and the incentive at every step is to say less until more is known.\nWhich leaves the alternatives, and they are all worse. You find out because your systems stop, which is how the conveyancing firms found out. You find out because a journalist rings. Or you find out because somebody spots the provider\u0026rsquo;s name on a ransomware crew\u0026rsquo;s leak site, which is not a notification process, it is an accident of somebody else\u0026rsquo;s marketing.\nSo contract for it, and be specific, because vague clauses default to the statutory floor. Notification of any material incident on their network within a stated number of hours, in writing, whether or not your data is confirmed to be involved. Notification if any credential with access to your systems may have been exposed. Notification if they appear on a leak site. A named contact on your side, not a general mailbox.\nThen ask the question that tells you most, and ask it while they are still selling to you. Have you had an incident? What happened, when did customers find out, and how did they find out? A provider who has been through one and handled it properly will tell you about it. It is the best evidence they have. Somebody who says it has never happened is either very lucky, very small, or not counting owt.\nThe Tool That Reaches Every Customer At Once No MSP makes money touching machines one at a time. The economics need one console that reaches everything, and that is what remote monitoring and management software is for. RMM sits on every machine under contract, running as system, and the provider drives the lot from a single login.\nOne console reaches every machine, at every customer The tooling vendor Ships the agent, and every update to it a signed update is trusted on arrival Your provider's console One login, every customer they have Another customer Agent on every machine Your building Agent on every machine, running as system Another customer Agent on every machine The machine trusts the console. The console trusts the vendor. Your firewall was told to allow all of it. The reach is the product. One console, every customer, every machine — which is why an attacker who gets the console gets all of it at once. Turn that round and look at it from your side. There is software on your machines that you did not choose, probably cannot name, and have never seen the patching record for. As such, it can do anything, to any of them, at any time.\nThe record on those tools is not comforting. Start in 2021.\nKaseya, July 2021. CISA and the FBI described responding to ransomware \u0026ldquo;leveraging a vulnerability in the software of Kaseya VSA on-premises products — against managed service providers (MSPs) and their downstream customers\u0026rdquo;. Around fifty providers were hit, and up to 1,500 businesses sat underneath them, most of which had never heard of Kaseya and had no say in it being there.\nIn January 2023 CISA, the NSA and the MS-ISAC issued a joint advisory to \u0026ldquo;warn network defenders about malicious use of legitimate remote monitoring and management (RMM) software\u0026rdquo;, after a campaign the previous October in which attackers phished people into installing ScreenConnect and AnyDesk, then used them to run a refund scam against victims\u0026rsquo; bank accounts. Nobody had to breach an MSP for that. The tool works exactly as well for them as it does for the provider.\nIn February 2024 ConnectWise ScreenConnect turned out to carry CVE-2024-1709, which \u0026ldquo;may allow an attacker direct access to confidential information or critical systems\u0026rdquo;. It scored 10.0. Note the weakness class on it: authentication bypass using an alternate path or channel, word for word the same category as the FortiOS bypass further up. Different vendor, different product, same mistake. Full marks, in the software that reaches every customer a provider has.\nThen SimpleHelp. CISA\u0026rsquo;s advisory of June 2025 describes ransomware actors \u0026ldquo;leveraging unpatched instances of a vulnerability in SimpleHelp Remote Monitoring and Management (RMM) to compromise customers of a utility billing software provider\u0026rdquo;, reaching \u0026ldquo;downstream customers\u0026rdquo; for double extortion. The flaw underneath it, CVE-2024-57727, lets an unauthenticated attacker pull \u0026ldquo;server configuration files containing various secrets and hashed user passwords\u0026rdquo; straight off the host.\nNotice where the loss lands every time. Not on the vendor. Not really on the provider either. On the customers underneath, who never bought the tool, were never told which one it was, and had no say whatever in when it got patched.\nA caveat on those sources, because it matters. They are American advisories, and that is because CISA publishes incident detail while the NCSC generally does not name. That is a difference in disclosure, not in conduct. The tooling is the same tooling, out of the same catalogue, and a small IT support firm in this country is running ScreenConnect or SimpleHelp or one of their competitors on your machines this afternoon.\nSo ask which one it is. Ask what version, when it was last patched, who can log into the console, whether every account on it has multi-factor authentication, and what the plan is the next time one of these ships a ten.\nThen ask one more, because the answer tells you something the others do not. Does it work over IPv6?\nIt is a fair question in 2026 and it lands harder than it looks. If the management tool only speaks IPv4, then IPv4 is what your network has to keep — not because your business needs it, but because their tooling does. That is the addressing plan being set by a supplier\u0026rsquo;s software rather than by your requirements, and you are paying monthly for the addresses that make it possible. A remote worker on a mobile connection may well be on an IPv6-only network already, in which case the whole arrangement is leaning on translation somebody else maintains.\nAnd if the answer is that they have never checked, you have learned the thing you were actually asking about.\nA plainer one first, which gets forgotten in all the security talk. What load does the agent put on the machine, and where is the evidence?\nThe answer you will get is \u0026ldquo;negligible\u0026rdquo;. That is an adjective. What you want is a number, measured on hardware like yours rather than lifted off a datasheet — average and peak CPU, resident memory, disk activity during an inventory sweep or a patch scan, and daily network. Taken on the oldest machine in the estate rather than the newest.\nAnd measure the whole stack, not one agent on its own. Remote management, antivirus, endpoint detection, the backup client, monitoring, asset tracking. Every vendor says under one per cent, there are six of them, and the person who actually has to work on that laptop is the one who finds out what six of them add up to. Near zero is the right target. Evidence is the only way to establish it.\nAsk for the figures before rollout, and ask to take your own afterwards on a machine you pick. A provider confident in their tooling offers both without being pushed.\nOne more on the tool itself, and it is the one people find strangest until they think it through. Do you get told when the agent updates on your machines?\nThe agent is the highest-privilege software on your estate. It runs as system, on everything, and it changes version whenever the provider or the vendor decides it should. Software is being installed across your entire company by a third party, silently, and on most contracts nobody tells you it happened.\nNow put that next to what actually went wrong at Kaseya. A malicious update pushed down a trusted management channel to every machine at once. From where you sit, on the day, that is indistinguishable from a routine agent update — same channel, same privilege, same silence. The only difference is intent. Intent is not something you can observe.\nSo ask for version-change notifications, ask who approves an agent update and whether it is tested anywhere before it reaches you, and ask the one that matters most: what would tell you that a push had happened which should not have? If the honest answer is nothing, then the control you are relying on is the provider\u0026rsquo;s own vigilance, which is the thing this whole section is about.\nAnd The Keys That Open It Then there is what opens the console. That gets even less attention than the console does.\nAn MSP holds administrative credentials for every customer on its books. That is the job — it is the whole product. So the standard applied to its own credentials ought to be higher than the one it sets for yours, and in practice it is routinely lower. Shared accounts, because individual logins for twelve engineers across two hundred tenants is a faff. Passwords in a vault the whole service desk can read. SSH keys with no passphrase sat on laptops that go home on the train. Break-glass accounts nobody has touched since the person who made them left the company.\n\u0026ldquo;We use MFA\u0026rdquo; is where that conversation normally stops, and it should not, because the forms of it are not equal. CISA ranks them strongest to weakest: FIDO/WebAuthn and PKI-based at the top, which it calls \u0026ldquo;the gold standard\u0026rdquo;; app-based one-time passwords and push notifications below that, \u0026ldquo;vulnerable to push bombing attacks as well as user error\u0026rdquo;; and SMS at the bottom, which \u0026ldquo;should only be used as a last resort MFA option\u0026rdquo;. Only the top row is phishing-resistant. Everything under it can be relayed, fatigued or intercepted while the person approving it believes they are logging in as normal.\nThe attackers have read that document too. CISA\u0026rsquo;s advisory on Scattered Spider says the group \u0026ldquo;targets large companies and their contracted information technology (IT) help desks\u0026rdquo;. Not the customer. The help desk that can reset the customer\u0026rsquo;s credentials. That is a great deal easier, and it gets you everybody at once.\nSo the question is narrower than whether they use MFA. Is every account that can reach your systems on a hardware token — a token, not an app — and what happens when somebody rings the service desk at eleven at night saying they are an engineer who has locked themselves out?\nThe fix here is unusually simple, which is what makes the absence of it hard to forgive. Google told KrebsOnSecurity in 2018 that it \u0026ldquo;has not had any of its 85,000+ employees successfully phished on their work-related accounts since early 2017\u0026rdquo;, when it began requiring physical security keys instead of passwords and one-time codes. Not fewer incidents. None. The same article notes the basic key retailed at twenty dollars.\nNow scale that to a managed service provider. They do not need to issue keys to your staff. They need to issue them to their own engineers, and there are perhaps a dozen of those holding the credentials that reach every customer on the books. A dozen keys, bought once, against a blast radius that covers every business they touch. Twenty quid each. When somebody tells you hardware tokens are impractical, ask how many people would actually need one.\nAnd there is a related question you should ask about your own estate, because most customers have never thought to. Who holds the break-glass account for your systems? On a great many contracts the honest answer is that the provider does, and you do not have a set at all. You cannot get into your own infrastructure without ringing them. That is not a security control, it is a dependency. It bites on the worst days rather than the ordinary ones. The day you give notice. The day they are bought by somebody you did not choose. The day they are the ones who have been compromised and you need to act without them.\nYou should hold your own break-glass credentials, sealed, written into the contract, and tested on a date somebody can point at. If your provider is reluctant, the reluctance itself has told you something.\nThe same question runs all the way down to the desk, and this is the one that gets forgotten completely. Every laptop and desktop they built for you has a BIOS administrator password on it, set during the build by somebody at their end. You own the machine. They hold the key to it.\nThink about what that password actually governs. Boot order. Secure boot. Whether the thing will start off a USB stick at all. Which is to say whether you can reinstall a machine you own, recover one that will not boot, hand a batch to somebody else, or wipe them properly before they go out of the door. On a laptop it may also be what stands between a thief and the disk. None of that is exotic. It is the ordinary business of owning computers, and on a great many estates the owner cannot do any of it without ringing the supplier.\nAsk for the lot. Every laptop, every desktop, every server — the BIOS or firmware password, the BMC login on anything that has one, and the bootloader password on anything where GRUB or its equivalent has been locked down, which is the same lock one layer up and is what stands between you and a rescue boot on a server that will not start. If they are the same password on all of them, that is worth knowing too. And if the answer is no, ask why not, because the reasons on offer are thin: the honest one is that it makes their build easier, and the rest is dressed up as security. A password you are not allowed to have is not protecting you from anybody. It is protecting them from you leaving.\nThen there is the account you actually log in with to fix a machine. Local administrator on every Windows box, root on every Linux one, and something equivalent on every switch, firewall and hypervisor in the building. Two questions cover the lot, and they cover the firmware and bootloader passwords above as well. Are they different on every machine, and whose password management system holds them — yours, or theirs?\nDifferent matters more than people expect. One local administrator password across two hundred machines is not two hundred passwords, it is one, and it is sat on the least defended laptop in the company as well as on the finance server. That is not a theoretical weakness, it is the standard move — get onto anything, read the password, walk to everything. Microsoft ships the answer in the box. Windows LAPS sets a different password on every machine, rotates it on a schedule, and backs it up into Active Directory or your own Entra tenancy. Microsoft\u0026rsquo;s own page lists the benefit first as \u0026ldquo;protection against pass-the-hash and lateral-traversal attacks\u0026rdquo;, and the feature is free on every supported version of Windows, with nothing further to pay to store the passwords in your own directory. So if the answer is one password everywhere, it is not a licensing problem and it is not a tooling problem. Somebody never turned it on.\nWhose system holds the password to your machine The common arrangement Their vault, or the one bolted to their tooling Every administrator password in your company held by a company you have no account with Your machines Servers, laptops, switches, firewalls, hypervisors The log of who read one is theirs You cannot audit what you cannot log into Withdrawing it means asking nicely, at the worst moment The arrangement you want Your directory, or your vault A different password on every machine, rotated, escrowed where your own access control reaches it The same machines The provider is granted access, not given custody You keep the record of who read what You can see it today, without asking anybody You can withdraw it on a Tuesday afternoon One local administrator password across two hundred machines is not two hundred passwords. It is one, and it sits on the least defended laptop in the company as well as on the finance server. Microsoft ships the answer to that for nowt, and it puts the password in your directory rather than theirs. Same machines, same passwords, two different answers to who is holding them and who can see when one is read. Whose system is the one to press on, and notice where LAPS puts the password by default: in your directory, under your access control, with your record of who read it. That is the shape you want on everything. The credential to your machine lives in something you own, and your provider is granted access to it — access you can see, and withdraw on a Tuesday afternoon without asking permission.\nThe other shape is the common one. The passwords sit in their password manager, or in the privileged access vault bolted onto their remote management tool, and you have no login for it. Then every administrative password in your company is held by a company you have no account with, the log of who read one is theirs, and on the day the relationship sours you are asking nicely for the keys to machines you own.\nSo ask the follow-up, and ask it now rather than at the exit meeting. How do these get into our system? There are three honest answers. Move the escrow into our directory or our vault, and take delegated access to it. Or give us read access to yours today, with an export we can run ourselves, in a format our own vault will swallow. Or hand us a sealed copy on a schedule, dated, that we open and test. \u0026ldquo;It is all in our system and you can have it when you leave\u0026rdquo; is not on the list, because the day you leave is the day they have the least reason to be quick about anything.\nAnd ask what happens to those passwords afterwards. A credential their engineers have known for four years is not made safe by an email saying it is gone. Every one of them needs rotating on the way out, by you, on machines you now hold the firmware password for.\nAnd then the live version of all of it. Do you get told when a password to one of your systems is read?\nNot a log they keep and could show you if you asked. A message that arrives at your end: which credential, which engineer, which ticket, and when. The capability is not in doubt anywhere — every vault worth the name records a checkout, and Microsoft\u0026rsquo;s own escrow does it in your own tenancy, where recovering a password is written to the Entra audit log as \u0026ldquo;Recover device local administrator password\u0026rdquo; against the account that did it. So the only real question is whether the record points anywhere you can see.\nAsk, and if the answer is no, ask why not. There is an honest no in here somewhere — an alert on every checkout in a busy estate would fire forty times a day and you would stop reading it by Wednesday. That has an answer rather than being the end of the conversation. Alert on the ones that ought to be rare: the domain administrator account, the break-glass credentials, the firmware and bootloader passwords, the recovery keys. And alert on any read with no ticket number attached to it, because that is either sloppy record-keeping or precisely the thing you want to hear about, and neither of those is well served by finding out at the quarterly review.\nWhat They Can Do Once They Are In One more, and it is the one that most often gets a blank look. Do you get told when one of their engineers goes into your systems?\nNot a log they keep. A notification you receive — who went in, when, for how long, and against which ticket. Ask for it, and if the answer is no, ask why not.\nThere is a precedent, and it is sitting inside the products they resell you. Microsoft\u0026rsquo;s Customer Lockbox makes a Microsoft engineer request your explicit approval before they can reach your content in a support case. Google publishes Access Transparency logs of its own staff touching your data, and Access Approval makes them ask first. So the largest suppliers on earth, with far narrower access than your provider holds, have built approval workflows and access logs for their own employees, and hand you the record.\nYour MSP has domain administrator. Ask what they hand you.\nAnd there is a harder reason than accountability. A notification that arrives at your end is the only independent sign you will ever get that their credentials are being used by somebody who is not them. If an attacker walks in through a service desk, every log that would show it sits inside the organisation that has just been compromised. One that lands in your inbox sits outside it.\nOne level down from that again, at the machine itself. Can the remote tool connect to somebody\u0026rsquo;s desktop without that person agreeing to it?\nFor a server at three in the morning, unattended access is the entire point and nobody sensible objects. For a machine somebody is sat at, with their mail open and their work on the screen, it is a different act. Every serious remote management product can be set to prompt before connecting, to show a visible indicator while a session is live, and to let the person decline. Whether yours does any of that is a configuration setting, and the setting was chosen by the people it suits.\nSo ask three things, and keep them separate. Can your engineers reach a staff desktop with no prompt? Does the person sat at it see anything at all while somebody is connected? Can they refuse?\nIf the first is yes and the other two are no, that is not a technical constraint. It is a default nobody revisited, on a product bought by the party it favours. It is also worth putting in front of whoever carries data protection in your organisation, because somebody watching an employee\u0026rsquo;s screen without their knowledge is a decision that ought to have a name against it.\nThen the question that decides whether any of this is provable afterwards. Are their engineers recorded while they are working on your systems?\nSession recording is not exotic and it is not a big ask. It is a headline feature of every privileged access product on the market, which means your provider is quite possibly paying for it already and has never switched it on. A recorded session gives you a video, or a keystroke and command log, or both, tied to a named engineer and a ticket number. That is what accountability looks like when it is real rather than promised — not an assurance that engineers behave, but a record that would show it if one did not.\nSo ask whether it is on, and then ask how you get at them, because a recording you cannot obtain is not evidence, it is a rumour. There are three answers worth having, in descending order. The recordings land in storage you own, written as they are made. Or you have read access to their system today, with an export you can run yourself. Or there is a request route with a stated turnaround — hours, in writing, in the contract — that you have tested at least once on an ordinary session rather than for the first time during an argument.\nWhat you do not want is the common arrangement: recordings held only by the provider, retention set by the provider, deletion in the provider\u0026rsquo;s gift. That is a control that works perfectly until the day it is needed against the provider. So ask who can shorten the retention, who can delete one, and whether watching a recording is itself logged. And ask what happens to the lot on the day you leave.\nBe fair about the other side of it, because there is one. A recording of an engineer fixing a laptop is also a recording of whatever your staff had on that screen, and now and again of a credential typed in plain sight. The recordings are sensitive in their own right and want the same handling as the vault: encrypted, access controlled, logged when viewed, kept for a stated period and no longer. A provider who raises that with you before you raise it with them has thought about the problem. One who has never considered it has told you something as well.\nThe same tool almost certainly moves files, in both directions. Ask whether it does, and then ask what has been put around it.\nOutbound is the one nobody prices. A transfer over the management channel is encrypted, trusted, and sits outside every control you have already paid for — the data loss prevention, the egress monitoring, the policy about USB sticks. Anybody with console access can take a copy of anything on any machine, and on most deployments there is no record of it that you will ever be shown.\nInbound is how the incidents earlier in this post actually happened. Pushing a file to every endpoint at once is not a flaw in these products, it is the headline feature. Ransomware simply used it the way it was built to be used.\nSo ask whether file transfer is enabled at all, whether it can be turned off on the machines that never need it, whether every transfer is logged with the file, the direction, the machine and the engineer — and whether that log comes to you, or joins the others in their inbox.\nAnd last on this, because it is the one nobody thinks to ask at all. What is the agent actually collecting, and what happens to it?\nThese tools gather a great deal more than a patch level. Hardware and software inventory, event logs, performance telemetry, often session recordings and screenshots, sometimes a good deal more depending on what is switched on. That is a detailed picture of how your business works and what your staff do all day. It is leaving your premises continuously.\nSo ask what is collected, where it is stored and under whose jurisdiction, how long it is kept, and who can see it — because the answer is usually the provider and the tool\u0026rsquo;s vendor, which is a second company you never chose and have no contract with. Ask whether any of it is used for anything beyond supporting you: product analytics, benchmarking, model training. Ask what happens to all of it on the day the contract ends, and get that in writing rather than in a conversation.\nAnd note whose data it is. Records about your staff and your systems, held by somebody processing on your behalf, is a phrase with obligations attached, and they are yours rather than theirs.\nThe Chip That Answers When The Machine Is Off There is one more layer beneath all of that, and it is worth asking about by name. Are they using Intel vPro, or the Active Management Technology underneath it?\nIf you have not met it, the short version is that management firmware runs on a separate controller inside the chipset, with its own network stack. It answers while the machine is powered down, provided there is mains and a cable. It can power the box on, change BIOS settings, mount a remote image and reinstall, and on the right configuration take the screen and keyboard at hardware level — before the operating system has loaded, and whether or not it ever does.\nThere are real reasons to want that. A machine that will not boot, a BIOS setting on a device three hundred miles away, a reimage without sending anybody out. In a large spread-out estate it is useful and I would not pretend otherwise.\nBut look at where it sits. Underneath the operating system, which means underneath every control you have bought. Your endpoint protection cannot see it, because it is not running in the operating system. Your host firewall does not filter it, because the traffic never reaches the operating system\u0026rsquo;s network stack — it is handled on its own ports before anything else gets a look. Your logging does not cover it. Nothing you have installed can tell you a session took place.\nWhere your controls stop, and what carries on underneath them One machine, from the top down Your applications and your data The thing the business actually uses Your controls can see it Endpoint protection, logging, host firewall This is the part you are paying for Your controls are this The operating system Everything above runs inside it Your controls run here Below this line, nothing you installed is watching Firmware and BIOS Updated by the manufacturer, on their timetable Not in your patching Management engine, vPro or a BMC Own processor, own network stack, own ports Answers with the machine off management traffic arrives here Everything you bought runs in the operating system. The two layers underneath it do not, and nothing installed above the line can tell you a session took place. The security record is not reassuring either. CVE-2017-5689 scored 9.8, and the description is worth reading slowly: \u0026ldquo;an unprivileged network attacker could gain system privileges to provisioned Intel manageability SKUs\u0026rdquo;. CISA issued an alert and CERT/CC a vulnerability note. Note that word \u0026ldquo;provisioned\u0026rdquo; — dormant it is not a route in, and switched on it is one. And the firmware carrying it does not update through the normal patching you are paying for. It comes from the machine\u0026rsquo;s manufacturer, on their timetable.\nThen there is the part that ought to concern anybody responsible for staff rather than servers. Hardware-level screen control means somebody can watch a display while the operating system has no idea. Ask whether the user consent prompt is enforced and whether the visible session indicator is switched on, because both are configuration and both can be turned off by whoever provisioned it.\nSo the questions are short. Is it provisioned on our machines, and who did that, and when — was it part of a build nobody mentioned? What specific jobs need it, and how many machines actually need those jobs? Which network can reach the management ports, and is that segmented from everything else? Is consent enforced, is the indicator on, and where is the audit trail? And can it be unprovisioned on every machine that does not need it?\nIf the answer to the first is \u0026ldquo;we do not know\u0026rdquo;, that is worth knowing on its own, because it means the capability is sat there configured by somebody and watched by nobody.\nThe Risk Went In For Free Turn the whole of this part the other way up and it says one thing. Every capability in it arrived as a convenience and was priced as one. The agent that patches a thousand machines is the thing that can put a file on a thousand machines. The chip that saves a two-hundred-mile drive answers when the machine is off and tells your logging nothing. The convenience was quoted, itemised and signed off. The capability came along in the same box, unpriced. It appears on no document you have ever been shown.\nNone of that is an argument for going without any of it. Estates need patching, and somebody has to be able to reach a dead machine. It is an argument for knowing what is in the building, who can reach it, from where, and what it would take for the person holding that reach to be somebody other than the firm currently holding it. An invoice will never tell you, because an invoice is a list of what you are paying for. It is not a list of what you are exposed to. Nobody in this arrangement has ever been asked to produce the second one.\nPart three is what happens when one of them goes off. What the contract actually promises, who ends up carrying the loss, what the excuses sound like, and what it costs to leave when you have finally had enough.\nIs Your MSP Lying To You — 3 parts\nIs Your MSP Lying To You To Sell You Premium Products? What Your MSP Built You, And Who Else Can Reach Ityou are here When It Breaks, Who Actually Carries It? Sources Retrieved 28 August 2026.\nTraining and skills.\nNational Vulnerability Database — CVE publication counts by year, summed per quarter: 6,595 in 2015 against 49,972 in 2025. The kit.\nCVE-2018-0141 — hard-coded account password, Cisco Prime Collaboration Provisioning, March 2018. CVE-2018-15439 — Small Business Switches, a privileged account enabled without notifying administrators, November 2018. CVE-2023-20101 — Cisco Emergency Responder, static root credentials that cannot be changed or deleted, October 2023. Der Spiegel, 29 December 2013 — the ANT catalogue, naming Cisco and Huawei among others, from the Snowden documents. Der Spiegel, 29 December 2013 — inside TAO: interdiction, load stations, and what happens to a diverted shipment. Ars Technica, 14 May 2014 — the 2010 NSA internal newsletter and photographs of a Cisco router being implanted, published in Glenn Greenwald\u0026rsquo;s No Place to Hide. Cisco, 13 May 2014 — the company\u0026rsquo;s response, in its own words. New York Attorney General, 2019 — the multistate settlement over video surveillance software sold to government bodies, and the 2009 to 2013 timeline. The perimeter kit.\nCISA Known Exploited Vulnerabilities catalogue — version 2026.08.27, 1,685 entries; the per-vendor counts in this post are my own tally of that file. CVE-2024-3400, CVE-2024-0012 — PAN-OS, GlobalProtect and the management web interface. CVE-2022-40684, CVE-2018-13379 — FortiOS authentication bypass and SSL VPN path traversal. CVE-2023-20198, CVE-2025-20333, CVE-2026-20316 — Cisco IOS XE web UI, the ASA VPN web server, and a hard-coded password in Secure Firewall Management Center. CISA, Implementing Phishing-Resistant MFA — the ranking of MFA forms, strongest to weakest. KrebsOnSecurity, July 2018 — Google on 85,000+ employees and no successful phishing after mandating physical security keys. Joint advisory AA23-320A — Scattered Spider, and its targeting of contracted IT help desks. Windows LAPS overview — a different local administrator password per machine, rotated and backed up to your own Active Directory or Entra tenancy; free on every supported version of Windows. When it goes wrong.\nICO, March 2025 — Advanced Computer Software Group fined GBP 3.07 million, the regulator\u0026rsquo;s first penalty against a data processor. The enforcement page carries the detail. ICO, October 2025 — Capita fined GBP 14 million over the 2023 breach: the ten-minute alert, the 58-hour response against a one-hour target, and the understaffed Security Operations Centre. Law Society Gazette, 28 November 2023 — around 80 conveyancing firms unable to complete transactions after their IT provider went down. CISA and the FBI on the Kaseya VSA attack, July 2021 — ransomware against managed service providers and their downstream customers. Joint advisory AA23-025A, 25 January 2023 — CISA, the NSA and the MS-ISAC on malicious use of legitimate RMM software. CVE-2024-1709 — ConnectWise ScreenConnect authentication bypass, CVSS 10.0, February 2024. CISA advisory AA25-163A, June 2025 and CVE-2024-57727 — ransomware actors reaching downstream customers through unpatched SimpleHelp RMM. Joint advisory AA22-131A, 11 May 2022 — the UK, Australian, Canadian, New Zealand and US agencies on threats to managed service providers and their customers. ","permalink":"https://blogs.damiendye.uk/en/random/is-your-msp-lying-to-you-part2/","summary":"Part 2 of 3. What actually gets built once the paperwork is signed: cloud for a business with one building, the box they will not be talked out of, the basics that were the thing you bought, and the agent on every machine that answers to somebody else\u0026rsquo;s console.","title":"What Your MSP Built You, And Who Else Can Reach It"},{"content":" Is Your MSP Lying To You — 3 parts\nIs Your MSP Lying To You To Sell You Premium Products? What Your MSP Built You, And Who Else Can Reach It When It Breaks, Who Actually Carries It?you are here Part one was how the shortlist gets written. Part two was what got built off the back of it. This part is the morning it stops working, and everything that follows from there.\nWho is contractually on the hook, and for how much. What gets said in the room when the answer is nobody. What a provider worth keeping does instead. And what it takes to leave one that is not, which turns out to be the part nobody plans for until they need it that week.\nWhat Did The Premium Actually Buy The evidence is in part two and I am not walking you back through it. Take it as read that the expensive option did not buy competence, did not finish the basics, and did not get you a box that was any harder to break into.\nSo what is the money attached to?\nNot the silicon, which gets cheaper every year. Not the addresses. Not the engineering hours, in a sector that cut its training spend by nearly a third in two years. The premium is attached to the channel — a vendor big enough to run tiers, rebates, deal registration and a distributor network, which means a queue of people who all get paid before anything is built, and every one of them on your invoice.\nThat is the seller\u0026rsquo;s half, and on its own it does not explain much. Plenty of buyers are perfectly capable and sign anyway. So here is the other half, and it is the more uncomfortable one.\nThe premium buys cover. Not for the company. For the person who signs.\nPutting the well-known badge in is defensible in a way that choosing the open one never is. If the market-leading firewall gets breached, so did everybody else\u0026rsquo;s, and you were unlucky. If the thing you assembled yourself gets breached, you chose it, and you will be asked why. The outcome for the business is the same. The consequence for the individual is not, and every buyer above a certain grade understands that without anybody having to say it out loud.\nAs such, the ring closes. The seller is paid more for recommending the expensive option. The buyer is personally safer for accepting it. Neither of them is acting against their own interests at any point — and the only party carrying the cost of that arrangement is the business they both work for.\nThat is the answer to the question this series started with, and it is worse than lying. A liar knows what the truth is. This arrangement does not need anybody to know. It just needs the shortlist to keep coming out the same way, and it does.\nThe Reports You Never Get Every recurring line on a managed service invoice ought to produce a document. Most of them produce nothing, and almost nobody asks.\nTake infrastructure patching, the line that sits above the desktop one. The honest way to update a fleet is a playbook — Ansible or whatever else — kept somewhere you can see it, run on a schedule, leaving output behind. So ask to see the playbook, and ask to see the run log. On a great many contracts neither exists. The updates are manual when they happen, they happen when somebody remembers, and the line is billed every month either way. What you were sold as automation is a note in a calendar.\nWhat each line on the invoice ought to produce ON THE INVOICE THE DOCUMENT THAT PROVES IT WHAT USUALLY ARRIVES Patching The playbook, and its run log kept where you can read it, run on a schedule A compliance percentage Backup A restore test, dated what came back, onto what, how long, who checked it A report full of green ticks Monitoring The alerts, sent to you as well straight from the system, on a schedule, unedited A dashboard, curated monthly Documentation Records in your own systems your IPAM, your asset database, your wiki An export, on their timetable A backup nobody has restored is not a backup. It is a hypothesis about a backup. Ask for the most recent restore test. The answer, or the pause before it, is the whole of the review. Every recurring line ought to leave a document behind. Most of them leave a percentage, a green tick, or nothing at all. Then backups. This is the worst of them, because it is the line people believe in most.\nYou will get a backup report. Green ticks, jobs completed, bytes written, all of it arriving weekly and none of it read. That report is not the one that matters. A backup nobody has restored is not a backup. It is a hypothesis about a backup.\nThe document you want is a restore test. What was restored, onto what hardware, how long it took end to end, and who opened the data afterwards and confirmed it was the data. Ask for the most recent one.\nThen ask the same thing again about the offline copy, because it is a separate claim and it needs separate proof. Everybody says they hold one. Ask them to show it to you — the media, where it lives, the last write date, who has the credentials — and ask when a restore was last performed from that copy rather than from the live backup system. Those are different media in different formats, and testing one proves nothing whatever about the other. The offline copy is the one that survives a bad week, and it is the one least likely ever to have been read back. Ask how often they are done, and against which systems, and whether anybody from your side has ever seen the result. On a lot of contracts it has never been done at all, because doing it costs a day of somebody\u0026rsquo;s time and nobody bills for it.\nThere is a prior question to all of this, and it decides whether any of the rest means anything. Do these reports come to you, or only to them?\nMost of the time the tooling generates them, they land in the provider\u0026rsquo;s inbox, somebody reads them if there is time, and you get told when there is a problem. That is not accountability. It is self-assessment with an invoice attached, and it asks you to take their word for the exact thing you are paying them to do.\nSo ask for the reports to come to you as well. Directly from the system, on a schedule, in whatever form the tooling spits out — not a slide assembled by the person being measured, and not a green dashboard curated for the monthly meeting. Backup successes and, more importantly, failures. The patch run and what it skipped. Alert volumes and response times against target. The restore test when it happens, and the exceptions list as it stands.\nYou do not have to read every one. You have to be able to, and they have to know you can. That is the entire mechanism, and it is the difference between trusting somebody and being in a position to check.\nAnd if it turns out to be difficult, ask why. Every report on that list already exists and is already being sent somewhere. Adding a second recipient is a line in a distribution list, not a project. Reluctance here is not a technical problem, and it is worth more than the answer would have been.\nThere is a bigger version of the same question, and it is the one that decides whether you can ever check anything for yourself. Where does the documentation live?\nAsk whether they will record your estate in your systems — your IPAM, your asset database, your wiki — or in theirs. Most will say theirs, and it will be presented as efficiency. Their tooling, their templates, their process, nothing for you to run.\nLook at what that arrangement does. The addressing plan, the network diagram, the asset register, the runbooks and the licence keys are all records of your business, created with your money, describing kit you own. Held somewhere you cannot see them. You cannot audit what you cannot read, so you have no way of knowing whether the documentation matches the estate — and documentation drifts from reality constantly, which is normal and forgivable, but invisible drift is how you find out during an incident instead of during a review.\nThen there is what happens at the end, which the section on leaving comes back to. Documentation in your own systems is already yours, already current, already in a format you use. Documentation in theirs is an export, on their timetable, in whatever shape their tool offers, usually a PDF of something that was a database.\nSo ask the question plainly, and ask for read access today rather than in principle. A provider working in your systems has made a decision that costs them a little and gives you a lot, and they will tell you so. One that insists on keeping the picture in their own tooling has made the opposite decision, and it is worth asking which part of it they found attractive.\nThat gap is not academic. Every incident in part two lands on it. The Kaseya customers, the eighty conveyancing firms, the pension schemes behind Capita — for each of them the only question that mattered on the day was whether the restore worked. That is when a report nobody asked for becomes the only document in the building worth having.\nIf your provider cannot produce a restore test, you are not buying backup. You are buying copying, and you will find out which on the worst morning of the year.\nNobody Is On The Hook Buying somewhere to put the blame deserves more than the clause it got in part one, because accountability is the thing this whole arrangement is built to route around. Not morally. Contractually.\nStart with the service level agreement, and read what it actually promises. Almost always it is a response time. Four hours to acknowledge, eight to attend, that sort of thing. Responding is entirely within their control, so it can safely be promised. Fixing is not, so it is not. Look for a line committing them to an outcome — the service works, the data comes back, the site trades — and on most contracts there is not one anywhere.\nThen find the liability cap. It will be there, and it is usually the fees you have paid over the preceding twelve months, sometimes less. So the worst thing that can happen to them is refunding what you already gave them. The worst thing that can happen to you is the business. Those two numbers are not in the same universe, and the gap between them is the actual risk position you are in.\nWhat the agreement promises, and what each side stands to lose What the agreement actually commits them to Acknowledge the fault within four hours Promised Attend within eight hours Promised The service works again, the data comes back, the site trades Not promised And what each side loses on the worst day Them capped at about twelve months of the fees you already paid You the orders, the payroll, the customers, the business Responding is inside their control, so it gets promised. Fixing is not, so it does not. The two bars are the same contract, seen from each end. Resale moves the rest of it upstream. If the hyperscaler is down it is the hyperscaler\u0026rsquo;s fault. If the firewall ships a ten it is the vendor\u0026rsquo;s fault. If a trojanised installer arrives signed and shipped as an ordinary update, that is the vendor as well. Every one of those is true, and none of them is any use to you, because you never had a relationship with any of those companies. You had one with the firm that chose them on your behalf and took a rebate for it.\nAnd underneath all of that sits the quietest mechanism of the lot. You cannot fail to deliver a requirement that was never written down. The requirements document nobody wrote in part one is not only sloppiness — it is protection. No captured requirement, no measurable promise. No measurable promise, no breach. A design nobody signed off cannot be departed from.\nLook at what happened when accountability finally did arrive somewhere. Capita was fined £14 million, and it came from the Information Commissioner under data protection law, over an alert nobody actioned for fifty-eight hours. Not from a customer, and not under a service contract. The money went to the state. The 6.6 million people whose data went out of the door, and the 325 pension schemes carrying the consequences, were not the ones who brought it and were not the ones paid.\nSo the arrangement, end to end. The vendor sells a product under a licence disclaiming fitness for anything in particular. The provider sells hours against a capped liability and a response-time promise. The insurer prices whatever is left. And the loss settles on the business that cannot move, which is the only party in the chain that never got to disclaim anything.\nEverybody in the chain disclaims their share, and it all lands in one place Each layer passes the consequence down and keeps its margin The vendor Licence disclaims fitness for anything The distributor Moves the product, takes a cut The provider Hours sold against a capped liability The insurer Prices whatever is left over The business that has to open the doors on Monday Makes something, employs people, cannot move quickly and is the only party in the chain that can disclaim nothing Every layer above is legally in the clear. The loss still has to sit somewhere. Every layer is legally in the clear, and each one is paid before anything is built. The consequence keeps travelling until it reaches somebody who cannot pass it on. There is a simpler test in all of that, and it needs no lawyer.\nIf a firm believes in what it has recommended, it should be willing to stand behind it. That is what a recommendation is. You did not buy the box — anybody with a website can sell you a box. You bought somebody\u0026rsquo;s judgement that this was the right box for you, and they were paid for that judgement, and paid again by the vendor whose box it turned out to be.\nSo when it fails and the answer comes back that the vendor let everyone down, look at what has just happened. The part you actually purchased has been disclaimed. The judgement evaporates at the exact moment it is tested, and what is left is a firm that passed a product through and took a margin on the way past.\nNobody is asking an MSP to underwrite Microsoft. But there is a wide gap between underwriting a hyperscaler and standing behind your own advice, and every other trade lives somewhere in it. An electrician who fits a consumer unit does not get to blame the manufacturer for having chosen it. A structural engineer who specifies a beam owns the specification. They carry their judgement, because the judgement was the service.\nSo ask what your provider is prepared to carry. Not the vendor\u0026rsquo;s product — their own recommendation. The answer, or the length of the pause before it, tells you what they privately think of it.\nAnd Then It Is Your Fault The contract is the quiet half of this. The loud half turns up on the day something actually breaks, and it gets there before the fix does.\nWatch which way responsibility travels. Outward, in every direction that is available, and roughly in this order. The vendor shipped a bad patch. The circuit was the carrier\u0026rsquo;s. It is a known issue and everyone is seeing it. Nobody could have foreseen it, said about a flaw that had been on the exploited list long enough to have grown a beard. The previous provider left it like that, which is a fair answer for six months and is still being given in year three. Then the outward directions run out, and there is one left. You never approved the upgrade. That was out of scope, your staff clicked the link, and you never raised a ticket.\nHere is the uncomfortable bit. Some of those are true. A customer who has turned down the same replacement three years running does own that decision, and a provider who says so is being straight with you. So the question is not whether the excuse is accurate. It is when it was first said. A risk put to you in writing before the outage, with what it would cost to fix and what it would cost not to, is a provider managing your estate. The same sentence sent the week after is a defence being assembled. Same words, different date, opposite meaning.\n\u0026ldquo;You never raised a ticket\u0026rdquo; is the one worth stopping on, because it is not an excuse at all. It is the operating model said out loud. Nothing is anybody\u0026rsquo;s job until you notice it, which makes you the monitoring, and you are paying a monthly fee to a firm whose detection layer is a customer ringing up. Once you have heard that answer once, you know what the service is.\nThere is a more active version of this, and it works far better than it should. You call a review to ask why the last six months have gone the way they have, and the agenda that comes back is about something else entirely: an urgent security matter, a licensing deadline nobody had mentioned, an end-of-life notice on a box that has been end-of-life for two years and has suddenly become pressing this fortnight. It is presented with real concern and a quote attached, and it eats the hour. You leave having agreed to spend money, and the thing you called the meeting for never got said out loud at all.\nNotice what the manufactured crisis always has in common. It needs a purchase, it needs it this quarter, and it requires nobody in the room to admit anything. Real urgency looks different, because it comes with an identifier, a date it was published, the specific machines in your estate that have it, and what they have already done about it while waiting to tell you. The invented sort arrives as a vendor deadline and a rounded figure. Ask when it first appeared on their radar. If the answer is the same week you started asking awkward questions, that is not a coincidence and it was never meant to be one.\nSo keep your own file, and start it before you need it. Every time you are told a thing is fine, ask for that in an email. Every time you decline something, write down what you were shown and what it was quoted at. It takes a minute and it means that on the day it becomes your fault, you can ask for the date, and the answer either exists or it does not. A firm doing this job properly gets there first anyway. They open with we missed this, here is what we are changing, and they say it before you have finished asking. It costs them one sentence, which is precisely why so few of them will spend it.\nAnd when you get the invented crisis, or the blame lands on you for something nobody ever put to you, say it in the room. Not as a complaint afterwards and not as a note for the file. Ask them whether they consider that professional, whether it is the answer they would accept if it were given to them, and where the integrity is in it. Somebody will be uncomfortable, and that is the entire purpose of asking. A firm with any self-respect left takes it, says fair enough, and comes back different next month. You will know inside a minute which sort is sat across from you.\nThen start looking anyway. That week. One bad meeting does not decide it, but the deflection was never a bad day. It was a decision about you: that managing how you feel is cheaper than fixing what you bought, and that you will put up with it. Firms do not un-make that calculation because a customer pulled a face. Replacing a provider properly takes months you have not begun to spend, and the time to be doing it is while you are still calm enough to choose well, rather than in the fortnight after the outage that finally makes your mind up for you.\nAsk These In The Next Meeting None of this costs you owt, and you do not need to be an engineer to ask any of it.\nI have put the questions in a spreadsheet rather than down the page — 120 of them across fifteen areas, each with what a competent answer sounds like set against what a deflection sounds like, and a column to record which one you got.\n\u0026#8615; Supplier questions, the long version msp-supplier-questions.xlsx · 26 kB Treat it as a prompt sheet, not a script. Read down it, mark the dozen that actually bear on your contract, and ask those. Nobody sits through a hundred questions in an hour and you would learn less if they tried. Send it across before the meeting rather than producing it in one — a provider doing the job properly will be glad of the notice, and how they take the notice is itself an answer.\nOne question in there asks which of their recommendations earn them a rebate, a credit, a margin or a target from the vendor. That is the one that changes the temperature in a room, and it should not. A provider with nothing to hide answers it straight, because they were going to recommend the same thing anyway and would rather you knew. Watch what happens when you ask. The answer matters less than the reaction.\nA provider doing the job properly has all of this to hand. Already, today, without preparing. The requirements written down, the options costed, the open source line on the comparison, the address plan dual-stacked, the remote management tool patched and the risks in the register where they belong. Ask, and you find out inside one meeting which sort you are paying for.\nAnd if it is the asking itself that upsets them, you have found out everything you needed to.\nEvery Account Goes Quiet All of this assumes you can leave if you need to. That is worth testing before you need to, because leaving is where a managed service contract stops being about technology.\nAssume you will need it eventually. Every one of these arrangements drifts the same way, whoever you sign with. The attention you got while they were winning the work thins out once the direct debit is running, and you settle into their books as a monthly figure rather than an estate somebody is thinking about. Nobody sits down and decides to stop caring. The best engineer goes where the noise is, the reviews quietly stop happening, and your kit sits on whatever was the right answer the year you signed while the rest of the trade moves on without you.\nUnderstand what a quiet account actually is, because the phrase sounds like a compliment and is not one. A quiet account is one nobody has checked. The backups run green every night and nobody has restored one to a spare machine and watched it come up. The failover pair has never been failed over, because doing it means booking an outage and somebody would have to be there. The firmware is whatever shipped, the certificate has been renewed by whoever remembered, and the documentation describes a network that was decommissioned two moves ago. The licence count has not been looked at since the year you had eleven more staff. None of that generates a ticket, so none of it reaches anybody\u0026rsquo;s screen, and the account sails through every internal review because the only thing being measured is whether you rang.\nThen you find out. Not at a review, because there was not one. You find out on the morning the restore is needed, or when the auditor asks for the diagram, or when a thing that has run untouched for six years stops and nobody left in the building knows what it was doing. A quiet account pays the same as a busy one and costs far less to serve. That is the whole of the incentive, and it points away from anybody ever opening the lid.\nThe Advice That Costs Them Money Set the test the other way up, and ask what you are actually meant to be buying. It is not the absence of trouble. Anybody can be absent. What you are paying for is somebody who knows the business well enough to arrive with things you did not ask for: a way of doing the month end that stops falling over, a licence you are paying for and nobody has opened since the year before last, a cheaper way to hold the data, a plan for the box everyone has quietly agreed not to reboot. Some of that costs them revenue to tell you, which is exactly why it is the honest signal. Ask yourself when your provider last put a recommendation in front of you that made their own invoice smaller.\nThe commonest form that takes is not a discount. It is somebody going through what you already run. Most outfits are not short of software, they are short of anybody who has ever sat down properly with the software they have got: the licence tier that already includes the feature being quoted for, the module paid for in the original project and never rolled out, the workflow in the finance system that would end the re-keying if somebody gave it two days, the second subscription bought because nobody asked the first supplier whether their product did that as well, the reporting nobody ever built so the whole company exports to a spreadsheet and does the work there instead. A genuine provider starts in that pile. What you own, what it can be made to do, what it will never do, and how much of the problem goes away if the thing you are already paying for is configured the way it was meant to be, all of that before anybody opens a price list.\nWhich tells you how they intend to earn from you, because that work is the awkward sort. It means reading somebody else\u0026rsquo;s documentation, learning a system they do not resell and get nothing for knowing, and sitting with the people who use it all day to find out what actually happens at four o\u0026rsquo;clock on a Friday. No rebate arrives on the back of any of it. It is also the whole of what you thought you were buying. And sometimes the answer at the end really is a new system, in which case buy the new system. Spending is not the problem. The order is: what you run, what it will still do, where it genuinely falls short, and only then what to go out and get. A recommendation that skips the first three was never an assessment. It was a catalogue with your name typed at the top of it.\nAnd a provider who has never sat down and learned what the business does cannot do any of it. There is a difference between knowing you have ninety mailboxes and knowing which hour of which day it would genuinely hurt to lose the system that takes the orders. One is inventory, the other is understanding, and only the second produces advice worth the money. A firm that has proposed nothing in three years, that cannot say what you sell or when your busy season falls, that keeps the lights on and raises the invoice on the same day every month, is not a partner and has stopped pretending to be. It is a subscription with a phone number on it. Not one to keep.\nThe opposite failure turns up in a better suit. That is the provider who is never quiet at all, who has something for you every quarter, and whose answer to every question you have ever asked came back with a part number attached to it. It looks like attention. What separates advice from selling is not the volume of it, and it is not how good the slides are either — it is whether the recommendation is capable of coming back as leave it alone, you do not need anything this year, and that money is better spent on the thing in the corner nobody has budgeted for. If not one recommendation in the whole relationship has ever cost them a sale, you are not being advised. You are being worked through a list, and part one is about where those lists come from. A partner talks you out of spending sometimes. To everybody else you are an account that buys what it is shown, which is a cash cow with a service desk attached, and nobody involved has to be dishonest about it for that to be exactly what is happening.\nGetting Rid Of A Bad One Work through what actually has to happen. All of it. Administrative credentials handed over for every system. Documentation that describes how the estate is built, assuming it exists. Control of the domain names, and of whatever account the DNS lives in. Certificate private keys. Subscriptions moved out of the provider\u0026rsquo;s partner agreement and into a tenancy that is yours. Their agent removed from every machine, by them. And the incumbent cooperating with their replacement, in detail, for weeks, while being paid nothing to do it and having just been told they are finished.\nWhat has to move on the way out, and who is holding it Held by the firm you are leaving Every administrative credential Documentation, if any was written The domains, and the DNS account Certificate private keys Licences under their partner deal Their agent, on every machine you own And weeks of their engineers' time must move unpaid Yours, or your replacement's Before the notice period ends And on that day They have just been told they are finished The engineer who knew your estate is on a job that still pays Half of what you ask for turns out to be chargeable None of it is sabotage The worse they have been, the less of that list exists to hand over. A firm that never wrote your requirements down did not write the handover pack either, and the mess is the moat. So ask for the pack today, while everybody is still getting on, and check one page of it against a live machine. All of it sits with the firm that has just been told it is finished, and none of the moving is work anybody is paying them for. Now the part that ought to worry you most. The worse a provider is, the harder that exit gets, because the failings are the same failings. A firm that never wrote your requirements down did not write the handover pack either. A firm that kept the credentials in a shared vault has no clean way to give them to somebody else. A firm with no playbook and no restore test has nothing to hand over but access and good luck. The mess is the moat. Nobody sat down and designed it that way, and it works better than if they had.\nSo assume the documentation line does not hold. Documentation is the first thing to go unwritten and the last thing anybody checks, and what arrives on the way out will describe an estate that moved on without it. That is the good case. The bad one is a pack that looks complete and is wrong: a diagram with a switch on it that went to the tip two years ago, a runbook for a server that has been rebuilt since, a credentials sheet full of accounts nobody can log in as. Wrong documentation is worse than none, because you act on it. None at least sends you to go and look.\nPrice the exit as though nothing survives it. Somebody walks the racks, reads the configuration off the live kit, exports the firewall rules and the DHCP scopes and the zone files, and writes down what is actually running. That is weeks of work, and in an exit it is weeks you are spending under notice, while the only people who know the answers have already been told they are finished. Do it now instead. Ask for the pack today, then take one page of it and check it against a live machine. If it holds, you have learned something worth knowing. If it is not worth the bits used to store it, you have learned that as well, and you have learned it with a year to put it right rather than a fortnight.\nThat cooperation line is the one that will actually hurt, and it is the one nobody prices. Read it back as a request: a firm that has just lost the account is being asked to spend weeks explaining it to the people who took it off them. Nobody has to refuse. The account moves to the leavers\u0026rsquo; pile, the engineer who knew your estate goes on a job that still pays, and your questions land with whoever is left. Replies come back in days instead of hours. The handover call gets booked three weeks out. Half of what you ask for turns out to be chargeable, and the one person who built the thing left in March. None of it is sabotage. All of it costs you exactly what sabotage would.\nYour replacement carries it, and then you carry it again. Their first month goes on working out what they have inherited. That is discovery they have nothing to show you for, so they either price it honestly and look dear against the incumbent\u0026rsquo;s renewal, or they swallow it and start the job already behind. You pay either way. So buy the cooperation before you need it. Put it in the contract on the way in rather than on the way out: a defined handover period, a named person doing it, a day rate agreed while they still want your signature, and access that stays live until the incoming firm says it can go. Pay for an overlap and count it cheap. And never arrive at the switchover with the outgoing provider holding the only key to something you own.\nThe paperwork does its share too. Notice periods measured in months rather than weeks, auto-renewal dates that pass while you are still deciding, and a handover priced as professional services at a day rate somebody sets after you have already given notice. None of that is unusual, and none of it breaks any rule.\nSo test the exit while the relationship is fine and nobody is upset. Ask for the handover pack now, in writing, as a deliverable rather than a promise. Ask who the registrant of your domains actually is. Ask whether your licences sit in your own tenancy or under their partner agreement, and what moving them involves. A provider doing the job properly will answer all three off the top of their head, because a firm confident in its work has no reason to make leaving difficult.\nAnd if the answers are vague, remember what the next section says about who you complain to.\nThere Is No Regulator For The Selling Something should have been nagging by now. Every failure in this series that came with a formal finding attached is a security failure. The ICO on Advanced. The ICO on Capita. CISA on the tooling. On the selling — the shortlist, the rebate, the single option, the renewal — there is nothing. Not one enforcement action against a British MSP.\nThat is not because the selling is clean. It is because nobody has the job of looking at it.\nConsumer law stops at your front door. The Consumer Rights Act 2015 protects consumers, and a limited company buying a managed service is not one. The Digital Markets, Competition and Consumers Act 2024 handed the CMA direct enforcement powers in April 2025, pointed squarely at aggressive sales practices, misleading information and contract terms that are plainly unbalanced. They are consumer powers. The CMA\u0026rsquo;s own account of the first year is worth reading for the wording as much as the numbers: fourteen investigations, two settlements, £760,000 \u0026ldquo;refunded to consumers\u0026rdquo;, £4.7 million in fines, 157 advisory and warning letters. Consumers, every time. Not the thirty-person manufacturer who signed five years of managed service last spring.\nTelecoms is the one corner a British regulator has been anywhere near. It shows what enforcement looks like when somebody holds the brief. In July 2015 Ofcom fined a small business telecoms provider £200,000 for mis-selling landline services to a base of \u0026ldquo;around 100,000 small businesses\u0026rdquo;, and made it compensate the customers affected and rewrite its sales materials. Right behaviour, right size of customer, wrong industry, and eleven years ago.\nThere is no equivalent for IT services. Add it up from where you are sitting. No cooling-off period. No ombudsman to escalate to. No regulator with jurisdiction. No duty on anybody to tell you what the vendor pays them. No published findings to learn from, because there is nowhere for a finding to come from. The contract was drafted by the supplier, and the only remedy in it is to sue — which means costs, years, and a legal budget a small business has not got, as the supplier is well aware.\nFinancial advice had this exact problem and dealt with it. The fix was not complicated — say who pays the adviser, and stop the product provider being the answer.\nThere will not be an RDR for IT. Nobody is coming to make your MSP declare what the vendor pays them, and the Cyber Security and Resilience Bill currently going through Parliament, which would pull managed service providers inside the NIS Regulations, is about detection and reporting rather than about who is paying for the advice.\nSo when somebody tells you there is no evidence of a problem in how technology gets sold in this country, they are right. It means nothing at all. There is no evidence because there is no inspector, no complaint route that ends in a public document, and no register of what happened to anybody else who signed the same contract. In a market nobody supervises, absence of evidence is just absence of anybody looking.\nWhat It Says About The Trade Strip the invoices away and look at who is actually stood in this arrangement.\nAt one end, businesses that make and do things. A firm machining parts. A garage. A kitchen feeding people. Solicitors moving somebody into a house on a Friday. Every one of them produces something you can point at, and every one of them is carrying the risk in this post, because they are the only party in the chain who cannot disclaim anything.\nAt the other end, several layers that produce nothing at all. A vendor whose licence disclaims fitness for any particular purpose. A distributor moving a box between warehouses and taking a cut. A partner programme paying somebody to prefer one badge to another. A provider selling hours against a capped liability. Each takes its margin and passes the consequence down, and the consequence keeps travelling until it reaches the only person who has to open the doors on Monday.\nA trade is a body of people who know how to do something, who are paid for knowing it, and who stand behind what they say because their name is on it. Held against that, most of this industry is a distribution channel with certifications.\nLook at what has been given away to get there. Skill first, because you cannot cut a third out of what you spend on training and still sell expertise with a straight face. Then judgement, sold at the front of the engagement and disclaimed the moment it is tested. And last the plain willingness to say \u0026ldquo;that is not the right answer for you\u0026rdquo; when the right answer pays less — which is the only thing separating advice from selling, and costs nothing but nerve.\nThe way out of it is unglamorous and entirely available. Own the kit you can own. Hold your own keys. Keep the documentation somewhere you can reach without ringing anybody. Learn enough about your own systems to know when you are being told something daft — not to run them yourself, just to recognise an adjective arriving where a number was asked for. That is not nostalgia for everyone having a server in a cupboard. It is the only leverage on offer, and it is cheap.\nThere are people in this trade who never stopped doing the job properly, and they are not hard to spot once you know what to look for. They quote the boring option. They tell you what a thing costs rather than what it is priced at. They put the risk in writing before you ask, and they are relieved when somebody finally checks.\nThe rest have arranged matters so that nobody ever does. You are allowed to ask what it costs, who is paying whom, and what happens when it breaks. Nobody in this arrangement is going to volunteer it, and that is not the same thing as you not being owed it.\nIs Your MSP Lying To You — 3 parts\nIs Your MSP Lying To You To Sell You Premium Products? What Your MSP Built You, And Who Else Can Reach It When It Breaks, Who Actually Carries It?you are here Sources Retrieved 28 August 2026.\nLaw and policy.\nCyber Security and Resilience (Network and Information Systems) Bill 2024-26 — House of Commons Library briefing on bringing managed service providers into the NIS Regulations. Who is allowed to complain.\nConsumer Rights Act 2015 and the Digital Markets, Competition and Consumers Act 2024 — the protections, and who they are for. CMA, direct consumer enforcement one year on — April 2025 to April 2026: 14 investigations, GBP 760,000 refunded to consumers, GBP 4.7 million in fines. Ofcom fines Unicom, 31 July 2015 — GBP 200,000 for mis-selling to small businesses, the nearest thing to an enforcement precedent, in telecoms rather than IT. ","permalink":"https://blogs.damiendye.uk/en/random/is-your-msp-lying-to-you-part3/","summary":"Part 3 of 3. Response times instead of outcomes, a liability cap set at the fees you already paid, and a provider whose excuses eventually arrive at you. The questions that flush it out, what a genuine one does instead, and what it takes to get rid of a bad one.","title":"When It Breaks, Who Actually Carries It?"},{"content":"A connection that times out tells you almost nothing. The far end might have nothing listening, a firewall three hops away might be dropping your SYN in silence without generating so much as a log line, or the route might simply not exist — and from where you are sitting every one of those looks the same. You wait. Nothing happens.\nSo the ticket gets written as \u0026ldquo;port 445 is blocked somewhere\u0026rdquo;, and \u0026ldquo;somewhere\u0026rdquo; is what makes it bounce. Your provider checks their edge, finds it clean, and hands it back. You check your host firewall, find it clean, and hand it back. A week goes by. Nothing is fixed.\nThe word doing the damage is \u0026ldquo;somewhere\u0026rdquo;. It does not have to be there. Every router between you and the destination is obliged to tell you when it is the one, and that obligation has been written into the standards since 1995. You just have to ask in the right way.\nWhy This Needs Spelling Out Because too many people in this industry cannot do the basics, and it costs their employers real money every week.\nI do not mean juniors. I mean people with years behind them, certifications on the wall and senior in the job title, whose diagnosis of a timed-out port stops at \u0026ldquo;it\u0026rsquo;s blocked\u0026rdquo; and goes no further. They run ping. It fails, or it works, and either way they have learned nothing about the port they were asked about. Then the ticket goes to the provider, the provider bounces it, and a fortnight of somebody\u0026rsquo;s salary goes into a thread that could have been one command.\nNone of this is hard. Telling a drop from a reject, reading the TTL on a reply, knowing that a star in a traceroute means nothing on its own — that is an afternoon to learn and it lasts a career. The reason people do not know it is not that they are thick. It is that nobody teaches it. Vendor training teaches you a vendor\u0026rsquo;s console. Certifications teach you the exam. The fundamentals underneath get assumed on day one and never actually covered, so people arrive at senior roles having never been shown, and by then it is embarrassing to ask.\nSo rather than moan about it, here it is written down. This is the basics, spelled out, with the commands and a script you can run today.\nAnd a word on why I bothered: I ran this against my own line while writing it and found three faults I did not know I had. An SMB drop eleven hops out. Forged SMTP resets one hop away. A firewall rule that exists on IPv4 and not on IPv6. An afternoon, no root, on a line I look after myself and pay attention to. Have a think about what is sat unnoticed on the networks somebody is being paid to run.\nWhy the Fisher-Price OS (Windows) Is Not In Here Two reasons it is not here. One of them is technical and one of them is not, and I would rather give you both than pretend it is all engineering.\nThe technical one is that it cannot do this. tracert sends ICMP echo requests and nothing else — Microsoft\u0026rsquo;s own reference describes it as \u0026ldquo;sending Internet Control Message Protocol (ICMP) echo Request or ICMPv6 messages to the destination with incrementally increasing time to live (TTL) field values\u0026rdquo;, and there is no port parameter anywhere in its syntax. pathping is that same tool with statistics bolted on. Test-NetConnection will tell you a TCP port is open or shut and nothing whatsoever about how far away the answer came from. Not one of them can probe the port you actually care about at a chosen distance, which is the entire method in this post.\nNor can you script your way round it. Setting the TTL on a socket is easy enough in .NET, but reading the ICMP error back is the hard half, and there is no equivalent of the trick this post leans on — no way to have the kernel report the error on the ordinary socket that caused it. That leaves a raw socket, and Microsoft\u0026rsquo;s own documentation says \u0026ldquo;only members of the Administrators group can create sockets of type SOCK_RAW\u0026rdquo;. So the unprivileged route does not exist and the privileged one wants a local admin token. You are into third-party downloads before you have started.\nThe other reason is that I do not care, and I would rather say that than dress it up. The name is not a cheap shot either, it is a description. It is the OS you get handed when you have never had another, it hides the machine from you as a design goal rather than an accident, and the moment you want to ask the network a precise question it turns out the tool was never built — because the people it is built for were never expected to ask. Thirty years on and tracert still cannot aim a packet at a port.\nIt is not a proper operating system for real IT people, and if it is the only one you have ever used then you are not doing this job at the level this post is written for. I am aware that lands badly. I am not writing to be liked, I am writing down what I measure for people who measure things, and I do not much care that somebody who has only ever run it would rather I put it another way.\nThe tool list is the argument, not the opinion. Every Unix in the table further down will let you set a hop limit and choose a protocol, four of them from a single command, and Linux will do the whole measurement without so much as sudo. That is not because they are harder to use. It is because they were built by people who expected whoever was at the keyboard to want to know things. An operating system whose diagnostic ends at ping has told you plainly what it thinks of the person using it, and thirty years of people accepting that is how we ended up with an industry that cannot locate a dropped packet.\nSo I do not run it, I have not run it in anger for years, and I am not writing it a section of its own to look even-handed about a gap that is real. Everything I build on and everything worth measuring from is Unix, and that is where this post lives.\nIf the broken box happens to be running the Fisher-Price OS (Windows), that changes nothing about the method. Get a shell on anything else — a Linux VM, a Mac, a Raspberry Pi on the same switch — and walk the hop limit towards it. The measurement does not care in the slightest what the far end runs. It only cares what is in between.\nTTL Is a Hop Budget, and Every Router Owes You a Receipt The IPv4 header has an 8-bit Time to Live field. The name is a leftover: it was specified in seconds, and nothing has treated it as seconds for decades. RFC 1812 settled the argument in 1995 and made the hop-count reading normative:\nEach router (or other module) that handles a packet MUST decrement the TTL by at least one, even if the elapsed time was much less than a second. Since this is very often the case, the TTL is effectively a hop count limit on how far a datagram can propagate through the Internet.\nThen the part that matters here, from the same section:\nIf the TTL is reduced to zero (or less), the packet MUST be discarded, and if the destination is not a multicast address the router MUST send an ICMP Time Exceeded message, Code 0 (TTL Exceeded in Transit) message to the source.\nThat is a MUST. Not a nicety, not a suggestion. RFC 792 in 1981 only said a gateway \u0026ldquo;may also notify the source host\u0026rdquo;; thirteen years later the requirement was tightened, and RFC 1812 says why in as many words:\nICMP Time Exceeded messages are required because the traceroute diagnostic tool depends on them.\nIPv6 dropped the pretence and renamed the field. RFC 8200 calls it Hop Limit — \u0026ldquo;8-bit unsigned integer. Decremented by 1 by each node that forwards the packet\u0026rdquo; — and the expiry message became ICMPv6 type 3, code 0, \u0026ldquo;hop limit exceeded in transit\u0026rdquo;.\nRead that as an instrument rather than a rule and it says something useful. Every router on the path is a beacon you can address by distance. Set the hop limit to 4 and the fourth router identifies itself. You do not need to know the topology, you do not need access to anything, and you do not need the operator\u0026rsquo;s cooperation. You need one packet per hop.\nTraceroute has done exactly this since the late 1980s. What it does badly is the thing you actually care about, because by default it probes UDP ports up in the 33434 range, which is a port nobody filters and nobody serves, so it tells you about a path that nothing real ever uses. A clean traceroute to a host you cannot reach on TCP/445 proves only that UDP/33434 gets there. Which was never the question.\nSo probe the port you care about.\nRead What Came Back, Not Whether Something Came Back Before counting hops, look at what the far end does when you reach it with a normal hop limit. There are five distinct answers and people routinely collapse them into one.\nWhat comes back What it means Who sent it SYN-ACK, the connection opens the port is open the host, or something answering for it TCP RST actively refused the host with nothing listening, or a device configured to reject ICMP 3/13, communication administratively prohibited a device is refusing on policy and saying so that device — its source address is your answer ICMP 3/1, 3/2, 3/3 host, protocol or port unreachable the last router, or the host nothing at all somebody is dropping in silence unknown, so go and measure it The third row is the one worth changing your habits over. When a firewall is configured to reject rather than drop, it puts its own address in the source field of the ICMP and hands you the culprit for free. On Linux that is what nft ... reject with icmpx admin-prohibited produces, and what iptables -j REJECT --reject-with icmp-admin-prohibited has always produced. Most tools throw it away and print \u0026ldquo;filtered\u0026rdquo;. nmap --reason will show it. So will the script further down.\nRows two and five are the interesting failure. A silent drop is a policy decision to tell you nothing, and because it is the default on nearly every commercial firewall it is the one you will actually meet. Expect silence.\nRow two deserves suspicion too. A reset is not proof the host sent it, and I will come back to that with a live example, because I found one on my own line while writing this.\nWalk the Port You Care About, Twice The method is two runs and a diff. That is the whole of it.\nWalk the hop limit up from 1 to 20 using the exact protocol and port that fails, and write down which router answers at each hop. Do the same with something that works — ideally the same host and a port that opens. The hop where the answers stop on run one, and keep going on run two, is the device dropping you. Run two gives you its address. Why the answers stop rather than change: a router applies its inbound access list before it does anything else with the packet, so if the policy says drop then the packet is gone before the forwarding path ever looks at the hop limit, no Time Exceeded is generated, and the device never puts its name to what it did. As such, the silence starts at the offending hop, not after it.\nI wrote a small tool for the walking because none of the native ones do it portably. It sets IP_TTL (or IPV6_UNICAST_HOPS) on an ordinary socket, connects, and reads back the ICMP error. On Linux IP_RECVERR reports that error on the very socket that provoked it, which means the whole thing runs unprivileged — no root, no raw sockets, no capabilities. On macOS, the BSDs and Solaris the kernel will not hand you the ICMP that way, so it falls back to a raw socket and needs root.\n\u0026#8615; hopfind.py — the TTL walker hopfind.py · 11 kB It is 279 lines of standard library and nothing else, and the whole of it is printed at the end of this post if you would rather read it than download it.\npython3 hopfind.py example.net 445 # the port under suspicion python3 hopfind.py example.net 443 # the reference run python3 hopfind.py example.net 53 --proto udp python3 hopfind.py 2001:db8::1 443 -6 Use the same protocol for both runs. Comparing a TCP trail against an ICMP trail is comparing two paths, because load balancing hashes on the five-tuple and ICMP has no ports to hash. Same host, same protocol, different port is the honest comparison.\nEverything below is real output from my own line on 28 August 2026, run as an unprivileged user on Fedora. Every address in it has been rewritten into the documentation ranges — RFC 5737 for IPv4, RFC 3849 for IPv6, with the one interface identifier altered as well. So 198.51.100.x is my own router and my ISP, 203.0.113.x is transit and peering, 192.0.2.x is the far network, and 2001:db8::/32 is the whole IPv6 path. Structure is preserved throughout: same prefix boundaries, same host-part shapes, same number of distinct networks. The hop numbers, the timings, the ICMP types and which hop went quiet are exactly as measured.\nOne Host, Three Ports, Three Different Faults Same destination throughout: a host on the public internet, fourteen hops away, with 443 open. First the reference run.\n$ python3 hopfind.py 192.0.2.4 443 --max 15 walking to 192.0.2.4 TCP/443 hop limit 1-15 1 198.51.100.254 0.3 ms ICMP 11/0 time exceeded in-transit 2 198.51.100.133 28.8 ms ICMP 11/0 time exceeded in-transit 3 * 4 198.51.100.153 6.3 ms ICMP 11/0 time exceeded in-transit 5 * 6 203.0.113.240 16.5 ms ICMP 11/0 time exceeded in-transit 7 203.0.113.188 16.5 ms ICMP 11/0 time exceeded in-transit 8 203.0.113.185 16.5 ms ICMP 11/0 time exceeded in-transit 9 203.0.113.15 15.6 ms ICMP 11/0 time exceeded in-transit 10 203.0.113.125 16.1 ms ICMP 11/0 time exceeded in-transit 11 192.0.2.31 26.6 ms ICMP 11/0 time exceeded in-transit 12 * 13 * 14 192.0.2.4 19.5 ms connected Verdict: TCP/443 is open. It answered at hop 14. Note hops 3, 5, 12 and 13. Four routers on a path that plainly works said nothing, because plenty of kit is configured not to generate ICMP for itself, or rate-limits it hard. A star is not evidence of a firewall. Take that one away if nowt else. The signal is never the presence of stars in a single run; it is where two runs stop agreeing.\nNow the same host on 445, which times out from here.\n$ python3 hopfind.py 192.0.2.4 445 --max 15 walking to 192.0.2.4 TCP/445 hop limit 1-15 1 198.51.100.254 0.3 ms ICMP 11/0 time exceeded in-transit 2 198.51.100.133 8.0 ms ICMP 11/0 time exceeded in-transit 3 * 4 198.51.100.153 6.4 ms ICMP 11/0 time exceeded in-transit 5 203.0.113.76 5.6 ms ICMP 11/0 time exceeded in-transit 6 203.0.113.240 15.9 ms ICMP 11/0 time exceeded in-transit 7 203.0.113.188 15.8 ms ICMP 11/0 time exceeded in-transit 8 203.0.113.185 16.6 ms ICMP 11/0 time exceeded in-transit 9 203.0.113.15 15.9 ms ICMP 11/0 time exceeded in-transit 10 203.0.113.125 15.0 ms ICMP 11/0 time exceeded in-transit 11 * 12 * 13 * 14 * 15 * Verdict: answers stop after hop 10 (203.0.113.125). Whatever swallows TCP/445 is hop 11. Walk a port that works and read off the address at hop 11. Hop 11 answered the 443 run in 26.6 ms and said nothing at all to the 445 run. Same box, same path, same ten routers in front of it. Cross-reference the run that works and hop 11 has a name: 192.0.2.31. That is the device dropping SMB, three hops short of the destination and eight hops past my provider\u0026rsquo;s edge. Not mine, and not my ISP\u0026rsquo;s.\nHop 5 makes the point about stars from the other direction. It was a star on the 443 run and answered on the 445 run — the opposite way round from the fault. That could be ICMP rate limiting, or it could be the two runs taking different paths through a load balancer. I do not know which, and neither will you. Repeat both runs before you believe either.\nTwo hop-limit walks to the same host, and the hop where they stop agreeing Two walks to the same host, and the hop where they stop agreeing One probe per hop limit. A shaded cell means that router sent back ICMP Time Exceeded and named itself. router answered nothing came back silent from here on 1 2 3 4 5 6 7 8 9 10 11 12 13 14 hop limit set on the probe TCP/443 reference, it opens \u0026#8226; \u0026#8226; * \u0026#8226; * \u0026#8226; \u0026#8226; \u0026#8226; \u0026#8226; \u0026#8226; \u0026#8226; * * \u0026#10003; opens TCP/445 under test, it times out \u0026#8226; \u0026#8226; * \u0026#8226; \u0026#8226; \u0026#8226; \u0026#8226; \u0026#8226; \u0026#8226; \u0026#8226; * * * * hop 11 answered one walk and not the other Hops 3, 5, 12 and 13 said nothing on a path that plainly works, so a star on its own means nowt. Hop 11 replied to the reference walk in 26.6\u0026#160;ms and never replied to the test walk at all. That is the boundary, and the reference walk is what gives it an address:\u0026#160;192.0.2.31. The same two runs, side by side. The only cell that matters is hop 11, and what matters about it is the disagreement: it answered one run and not the other. Then port 25. This one stopped me.\n$ python3 hopfind.py 192.0.2.4 25 --max 15 walking to 192.0.2.4 TCP/25 hop limit 1-15 1 192.0.2.4 0.4 ms TCP reset Verdict: a reset came back to a probe with a hop limit of 1. Nothing more than one hop away can have sent it, so check the reply TTL before you believe the host did. A hop limit of 1 means the packet died at my own router. It went one hop. It cannot have travelled fourteen. Yet a TCP reset came back in 0.4 ms with the destination\u0026rsquo;s address on it, and my kernel dutifully reported connection refused. Without the hop limit set I would have read that as \u0026ldquo;the far end has no mail server\u0026rdquo; and closed the ticket.\nSomething one hop away is forging resets for outbound SMTP and signing them with the destination\u0026rsquo;s address. Blocking outbound 25 is an entirely ordinary thing for a consumer router or an ISP to do, and doing it with a reset rather than a drop is arguably the polite version, but the reset carries somebody else\u0026rsquo;s address and I had no idea mine did it. The hop limit is what caught it, and nothing else in the reply would have.\nOff By One: Where the Access List Sits in the Pipeline The verdict above says the dropper is hop 11 because the answers stopped after hop 10. Be careful with that arithmetic, because it depends on the order the offending device does two jobs.\nMost kit applies the inbound policy first and the hop-limit check second. Deny hits, the packet is discarded, and no Time Exceeded is ever generated, so the device never appears and the silence begins at its own hop number. That is the case above. It is also the common one.\nSome platforms handle hop-limit expiry in the fast path before policy is evaluated. There, the device answers the probe addressed to it and only drops probes aimed past it, so it appears normally and the silence begins one hop later.\nWhy the offending hop usually does not appear: policy is evaluated before the hop-limit check Inside the hop that is dropping you, the order of two checks decides what you see Your probe arrives with one hop left on its budget, and it matches a rule that says deny. Policy first \u0026#8212; nearly all firewalls, and every case measured in this post the router at hop N probe in inbound policy hop-limit check forward discarded before anything looks at the hop limit No ICMP is generated, so this router never puts its name to the drop. Your walk goes quiet\u0026#160;at\u0026#160;hop\u0026#160;N. Expiry first \u0026#8212; some platforms handle it in the fast path the router at hop N probe in hop-limit check inbound policy forward expired, so ICMP 11/0 goes back with this router's address It answers for itself and only swallows probes aimed past it. Your walk goes quiet from hop\u0026#160;N+1. Why the hop doing the dropping usually stays invisible. On nearly all firewalls the deny is evaluated first, so the probe is gone before the forwarding path notices the hop limit expired and no ICMP is ever generated. On kit that handles expiry in the fast path, the same device answers the probe addressed to it and only swallows the ones aimed further on. So the honest reading of \u0026ldquo;answers stop after hop N\u0026rdquo; is: the dropper is hop N+1 if it binned your probe before noticing the budget had run out, or hop N itself if it answered the probe addressed to it and swallowed everything aimed past. Two adjacent devices, and the run that works names them both. Quote the addresses, not the hop count — a hop count means nothing to the person reading your ticket, who is counting from somewhere else.\nThe Reply\u0026rsquo;s TTL Tells You Who Really Answered The SMTP reset above was caught by the hop limit going out. There is a second, independent check available in every reply that comes back, and it costs nothing.\nInitial TTL values are not standardised, but in practice there are three:\nStarts at Typical sender 64 Linux, macOS, the BSDs, illumos, most hosts 255 Cisco IOS, Junos, Solaris, most network kit\u0026rsquo;s own traffic 128 the Fisher-Price OS (Windows), which you still need for reading a reply off one Subtract the TTL you received from the next value up and you have the hop count back. A reply arriving with TTL 50 started at 64 and came 14 hops. One arriving with TTL 250 started at 255 and came 5. ping prints it without being asked:\nping -c1 192.0.2.4 # ttl=50 → 14 hops away tcpdump -n -v \u0026#39;icmp\u0026#39; # -v prints the ttl of every packet it shows Two things fall out of that. Both are free.\nA reply whose reverse hop count does not match the host\u0026rsquo;s other replies was not sent by the host. If an echo reply from a server comes back 14 hops out and the RST on port 25 comes back 1 hop out, a middlebox wrote the RST. Same trick as the section above, from the other end, and it works even when you cannot set the outbound TTL.\nA reply that started at 255 came from network kit, not from a server. Useful when you are trying to work out whether the thing rejecting you is the host or the router in front of it.\nTo see the field on TCP rather than ICMP you need a capture, and the filter is worth memorising:\n# every hop-limit expiry coming back to you, IPv4 and IPv6 tcpdump -n -v \u0026#39;icmp[icmptype] == 11 or icmp6[icmp6type] == 3\u0026#39; # who is resetting you, and from how far tcpdump -n -v \u0026#39;tcp[tcpflags] \u0026amp; tcp-rst != 0\u0026#39; tcpdump prints the expiry as ICMP time exceeded in-transit — the same phrase the standards use, and the same event that shows up as \u0026ldquo;TTL expired in transit\u0026rdquo; on platforms that word it that way.\nThe Same Trick With Native Tools, on Five Unixes If you would rather not run a script, the native tools will do most of it. They just disagree with each other more than you would expect. -P in particular means three different things depending on whose traceroute you are holding, and one of them will quietly ruin your test.\nLinux (traceroute 2.1.x) macOS / FreeBSD OpenBSD NetBSD Solaris 11 TCP probes -T -P tcp not usable no no ICMP probes -I -I -I -I -I UDP to a fixed port -U -p N -e -p N no no no destination port -p N (constant for TCP) -p N (increments without -e) -p N (increments) -p N (increments) -p N (increments) what -P means not used probe protocol numeric protocol, \u0026ldquo;will not work reliably for most protocols\u0026rdquo; set DF and probe path MTU pause between probes, in seconds needs privileges yes, for -T and -I yes yes yes yes Every one of those needs raw sockets, so every one needs privileges — though macOS and the BSDs usually ship traceroute setuid root, so you may not have to type sudo in front of it. Linux does not, and Fedora does not, which is half the reason the script above exists.\nThree traps in that table, and I have watched all three waste an afternoon.\nOn the BSDs and macOS, -p is a base port that increments with every probe. So traceroute -P tcp -p 445 host tests 445, then 446, then 447, and by hop 10 you are asking about a port nobody has ever heard of. You want -e as well, which the man page calls firewall evasion mode and which actually just means \u0026ldquo;keep the port still\u0026rdquo;:\nsudo traceroute -P tcp -e -p 445 example.net # macOS, FreeBSD On Linux, plain -p does the same thing for the default UDP method and you need -U -p for a constant UDP port. For TCP, -T -p is already constant — the man page is explicit that \u0026ldquo;for TCP and others specifies just the (constant) destination port to connect\u0026rdquo;.\nsudo traceroute -T -p 445 example.net # Linux sudo traceroute -U -p 53 example.net # Linux, UDP/53 specifically On Solaris and NetBSD there is no way to pin the port at all, and on Solaris -P is a pause in seconds, so a copied Linux command line will run without error and measure nothing you asked about. Solaris also has no TCP probe mode. This is the case where you want the script.\nRedox is the odd one out and worth a sentence because I expect the question. Its whole network toolkit is netutils — dns, ifconfig, nc, ping, telnetd, wget. No traceroute, no tcpdump, nowt to capture with. If a Redox box is one end of the problem, measure from the other end and point the walk at it.\nmtr deserves a mention too, because it does the repeat-and-average part that the tables above make you do by hand:\nsudo mtr -T -P 445 --report --report-cycles 20 example.net Run that against the failing port and again against a working one, side by side. Same method, prettier output.\nNone Of This Works If Somebody Blocks ICMP Every measurement in this post is made of ICMP errors travelling back to me. Block those and the entire diagnostic goes dark — and so does a good deal else.\nBlocking ICMP wholesale is still treated as a security posture in places. It has not been a defensible one since the 1990s. The attacks it is imagined to stop were ping of death and smurf, both of which were fixed in the stacks rather than at the border, and both of which were fixed before some of the engineers still repeating the advice were born. What blanket blocking stops now is diagnosis. Nothing else.\nRFC 1812 does not leave room for interpretation on the point: Time Exceeded is a MUST, and the standard states its reason is that traceroute depends on it. Drop it and you have broken a tool the internet\u0026rsquo;s own router requirements document names as the justification for the message existing.\nPath MTU discovery is the expensive one. It needs ICMP 3/4, fragmentation needed, to come back to the sender. Filter that and you get the fault every network engineer has chased at least once, the one where the handshake completes and small transfers work and anything carrying a full-size packet hangs forever. SSH connects and scp stalls. The page loads and the image never arrives. Nothing in the logs. Nothing to grep for.\nOn IPv6 it stops being a matter of taste. RFC 4890 §4.3.1 lists the messages a firewall must not drop:\nDestination Unreachable (Type 1) - All codes Packet Too Big (Type 2) Time Exceeded (Type 3) - Code 0 only Parameter Problem (Type 4) - Codes 1 and 2 only and on Packet Too Big it is blunt about the consequence: \u0026ldquo;Effectively, parts of the Internet will become inaccessible.\u0026rdquo; IPv6 routers do not fragment. If Packet Too Big cannot reach the sender, there is no recovery path.\nThe control is a rate limit, not a drop. Permit type 3 and type 11 inbound, count them, cap them at something like a hundred a second, log whatever exceeds the cap, and you have kept the diagnostics, kept path MTU discovery working, and kept every bit of the protection the blanket rule was imagined to be providing in the first place. Restricting how much of something you accept is a control. Refusing all of it and calling that hardening is just refusing to be measured.\nFor anyone running an MSP: if your customer\u0026rsquo;s line drops ICMP errors, you have removed their ability to prove which network a fault is in — and your own. The next time a fault sits between two providers who both say it is clean, that is the bill for the policy.\nExcept Echo. Drop That. Everything above is about ICMP errors. Echo is a different animal, and it is the one part of the protocol I would take off the wire at the border.\nLook at what the standard requires of it. In IPv4, RFC 792 says of an echo request that \u0026ldquo;the data received in the echo message must be returned in the echo reply message\u0026rdquo;. IPv6 is blunter still — RFC 4443 defines the field as \u0026ldquo;zero or more octets of arbitrary data\u0026rdquo; and then requires that it \u0026ldquo;MUST be returned entirely and unmodified in the ICMPv6 Echo Reply message\u0026rdquo;.\nRead that as an attacker rather than as an operator. The standard obliges every host on earth to accept a block of bytes you choose and hand it straight back to you. That is not a side effect. It is a mandated, bidirectional, arbitrary-payload channel, running over a protocol most firewalls pass without inspecting and most logging records as a packet count rather than as content.\nPeople have been building tunnels on it for thirty years. Loki did it in Phrack 49 in 1996. Ptunnel will carry a full TCP session inside ping and has been a package away for two decades. If your egress policy is \u0026ldquo;block everything, allow ICMP because the network team need it\u0026rdquo;, you do not have an egress policy. You have a VPN with extra steps, and the traffic leaves looking like somebody testing whether the internet is up.\nRFC 4890 disagrees with me, and it is worth saying so plainly rather than quoting only the half that suits. §4.3.1 puts Echo Request and Echo Response in the same must-not-drop list as the errors. Then read the justification it gives:\nFor Teredo tunneling [RFC4380] to IPv6 nodes on the site to be possible, it is essential that the connectivity checking messages are allowed through the firewall.\nThe stated reason to keep echo open is that somebody needs to build a tunnel through your firewall with it. That is my argument, written down by the people making the opposite case.\nSo the policy is narrow, not blanket:\n# transit rules. permit the errors, drop the ping ip protocol icmp icmp type { destination-unreachable, time-exceeded, parameter-problem } \\ limit rate 100/second accept ip protocol icmp icmp type { echo-request, echo-reply } drop ip6 nexthdr ipv6-icmp icmpv6 type { destination-unreachable, packet-too-big, \\ time-exceeded, parameter-problem } limit rate 100/second accept ip6 nexthdr ipv6-icmp icmpv6 type { echo-request, echo-reply } drop On IPv6, do not carry that pattern onto a link-local or host chain without keeping neighbour discovery. Types 133 to 137 — nd-router-solicit through nd-redirect — are how IPv6 does the job ARP does in IPv4. Drop those and the segment stops working within minutes, and it will not look like a firewall fault. Filter echo at the border, not on the wire between a host and its own router.\nWhat does that cost you? ping across the boundary, and nothing else. Everything in this post keeps working, because not one measurement here sends an echo request. hopfind.py walks TCP and UDP and reads the errors that come back; traceroute -T and -U do the same. Path MTU discovery needs Packet Too Big, which is an error. The reverse-TTL trick works on any reply, and a TCP handshake will give you one. The only thing you lose is the least informative tool in the box, and this whole post is an argument about why stopping at ping is the problem in the first place.\nMy own line already does exactly this, though I doubt it was deliberate. ping to my default gateway gets 100% loss, and ICMP Time Exceeded from that same gateway comes back in 0.3 ms — as every trace in this post shows. Echo shut, errors open. Whoever shipped that firmware got the right answer, and I only found out they had by going looking.\nThe Same Rule, Written Once Here is the fault I did not expect to find in my own house. A public DNS resolver, walked on TCP/443 over both families, minutes apart.\n$ python3 hopfind.py 192.0.2.53 443 --max 10 walking to 192.0.2.53 TCP/443 hop limit 1-10 1 * 2 * 3 * 4 * 5 * 6 * 7 * 8 * 9 * 10 * Verdict: nothing answered at all, not even the first hop. Nothing at all. Not one hop. My own router did not even report the expiry it must have generated, the same expiry it reported in 0.3 ms for every other walk in this post — so the drop happens at hop 1, before the hop-limit check ever runs, and hop 1 is mine.\nThe path is fine, which the same destination proves on UDP:\n$ python3 hopfind.py 192.0.2.53 53 --proto udp --max 10 1 198.51.100.254 0.3 ms ICMP 11/0 time exceeded in-transit 2 198.51.100.133 5.4 ms ICMP 11/0 time exceeded in-transit 3 * 4 198.51.100.167 14.5 ms ICMP 11/0 time exceeded in-transit 5 203.0.113.50 6.3 ms ICMP 11/0 time exceeded in-transit 6 203.0.113.174 5.7 ms ICMP 11/0 time exceeded in-transit 7 203.0.113.201 6.4 ms ICMP 11/0 time exceeded in-transit 8 * Seven hops of clean answers to the same address. So it is not routing and it is not the destination — something on my line drops TCP to that host and passes UDP to it.\nThen the same resolver, same port, over IPv6:\n$ python3 hopfind.py 2001:db8:53::53 443 -6 --max 10 walking to 2001:db8:53::53 TCP/443 hop limit 1-10 1 2001:db8:1:ee:beef:abcd:fec0:2a30 0.5 ms ICMP 3/0 hop limit exceeded in-transit 2 2001:db8:1::15c 24.6 ms ICMP 3/0 hop limit exceeded in-transit 3 * 4 2001:db8:2:200::50 5.5 ms ICMP 3/0 hop limit exceeded in-transit 5 2001:db8:2:2::4 5.7 ms ICMP 3/0 hop limit exceeded in-transit 6 2001:db8:2:991:: 16.6 ms ICMP 3/0 hop limit exceeded in-transit 7 2001:db8:53::53 5.7 ms connected Verdict: TCP/443 is open. It answered at hop 7. Straight through. Seven hops, no drama. Same service, same port, same intent, and the rule exists on one address family only.\nWhoever wrote that rule wrote it for IPv4 and never wrote the twin. I have no idea what it was meant to achieve — my guess is something about keeping DNS local — but whatever it was, it has been achieving it on half the traffic for as long as this line has had IPv6. If it was there for a reason, it does not work. If it was not, it should not be there.\nThat is the everyday version of the dual-stack problem, and it is far more common than the arguments about whether to deploy IPv6 at all. Two rulebooks. One maintained.\nThe Script, In Full No dependencies, no install, nothing but the standard library. Python 3.6 or later, and on Linux no privileges at all.\nThe two classes are the whole of the portability story. ErrorQueue is the Linux path: arm IP_RECVERR on the socket, and after the probe fails, read MSG_ERRQUEUE and pull the router\u0026rsquo;s address out of the sock_extended_err structure the kernel appends it to. RawIcmp is everywhere else: open a raw ICMP socket, read whatever arrives, take the source address off the packet. The first needs nothing, the second needs root, and the rest of the program does not care which one it was handed.\nOne detail is worth pointing out because it is the difference between a right answer and a plausible one. In probe() the error queue is drained before SO_ERROR is consulted. An ICMP error reaches a TCP socket as a plain errno — ICMP 3/3 port unreachable arrives as ECONNREFUSED, exactly like a real reset — so checking SO_ERROR first would have reported the forged SMTP reset further up this post as an honest refusal from the far end. Read the queue first and ee_origin tells you a router spoke.\n#!/usr/bin/env python3 \u0026#34;\u0026#34;\u0026#34;hopfind - work out how many hops away the thing blocking your port is. Walks the IPv4 TTL, or the IPv6 hop limit, up from 1 and records which router answers at each step - using the protocol and port you actually care about instead of traceroute\u0026#39;s default UDP high ports. Run it twice. Once against something that works, once against the port that does not. The hop where the answers stop is the device dropping you, and the run that works gives you its address. python3 hopfind.py example.net 445 # the port under suspicion python3 hopfind.py example.net 443 # the reference run python3 hopfind.py example.net 53 --proto udp python3 hopfind.py 2001:db8::1 443 -6 On Linux this needs no privileges at all: IP_RECVERR and IPV6_RECVERR hand the ICMP errors back on the ordinary socket that caused them. On macOS, the BSDs and Solaris the errors have to be read off a raw ICMP socket, which means root. Written for https://blogs.damiendye.uk/networking/how-far-away-is-the-firewall/ Public domain. Do what you like with it. \u0026#34;\u0026#34;\u0026#34; import argparse import errno import os import select import socket import struct import sys import time # Linux socket options. Absent from the socket module on some builds, so they # are spelled out rather than looked up. IP_RECVERR = 11 IPV6_RECVERR = 25 # ee_origin values from linux/errqueue.h. Anything else means the errno came # from the local stack rather than from a router. SO_EE_ORIGIN_ICMP = 2 SO_EE_ORIGIN_ICMP6 = 3 ICMP_V4 = { (11, 0): \u0026#34;time exceeded in-transit\u0026#34;, (11, 1): \u0026#34;fragment reassembly time exceeded\u0026#34;, (3, 0): \u0026#34;net unreachable\u0026#34;, (3, 1): \u0026#34;host unreachable\u0026#34;, (3, 2): \u0026#34;protocol unreachable\u0026#34;, (3, 3): \u0026#34;port unreachable\u0026#34;, (3, 4): \u0026#34;fragmentation needed\u0026#34;, (3, 9): \u0026#34;net administratively prohibited\u0026#34;, (3, 10): \u0026#34;host administratively prohibited\u0026#34;, (3, 13): \u0026#34;communication administratively prohibited\u0026#34;, (5, 0): \u0026#34;redirect\u0026#34;, } ICMP_V6 = { (3, 0): \u0026#34;hop limit exceeded in-transit\u0026#34;, (3, 1): \u0026#34;fragment reassembly time exceeded\u0026#34;, (1, 0): \u0026#34;no route to destination\u0026#34;, (1, 1): \u0026#34;communication administratively prohibited\u0026#34;, (1, 3): \u0026#34;address unreachable\u0026#34;, (1, 4): \u0026#34;port unreachable\u0026#34;, (2, 0): \u0026#34;packet too big\u0026#34;, } def describe(family, icmp_type, icmp_code): table = ICMP_V4 if family == socket.AF_INET else ICMP_V6 return table.get((icmp_type, icmp_code), \u0026#34;unrecognised\u0026#34;) def is_expiry(family, icmp_type): \u0026#34;\u0026#34;\u0026#34;Was this the router saying \u0026#39;your hop budget ran out here\u0026#39;?\u0026#34;\u0026#34;\u0026#34; return icmp_type == (11 if family == socket.AF_INET else 3) class ErrorQueue: \u0026#34;\u0026#34;\u0026#34;Linux. The kernel reports the ICMP error on the socket that provoked it.\u0026#34;\u0026#34;\u0026#34; def arm(self, sock, family): if family == socket.AF_INET: sock.setsockopt(socket.IPPROTO_IP, IP_RECVERR, 1) else: sock.setsockopt(socket.IPPROTO_IPV6, IPV6_RECVERR, 1) def extra_readers(self): return [] def collect(self, sock, family): try: _, ancillary, _, _ = sock.recvmsg(0, 1024, socket.MSG_ERRQUEUE) except OSError: return None wanted = (socket.IPPROTO_IP, IP_RECVERR) if family == socket.AF_INET \\ else (socket.IPPROTO_IPV6, IPV6_RECVERR) for level, kind, data in ancillary: if (level, kind) != wanted or len(data) \u0026lt; 16: continue # struct sock_extended_err, then the sockaddr of the router that # sent the error - SO_EE_OFFENDER in the kernel headers. _, origin, icmp_type, icmp_code = struct.unpack_from(\u0026#34;=IBBB\u0026#34;, data, 0) if origin not in (SO_EE_ORIGIN_ICMP, SO_EE_ORIGIN_ICMP6): return None addr = None if len(data) \u0026gt;= 24: offender_family, = struct.unpack_from(\u0026#34;=H\u0026#34;, data, 16) if offender_family == socket.AF_INET: addr = socket.inet_ntoa(data[20:24]) elif offender_family == socket.AF_INET6 and len(data) \u0026gt;= 40: addr = socket.inet_ntop(socket.AF_INET6, data[24:40]) return addr, icmp_type, icmp_code return None class RawIcmp: \u0026#34;\u0026#34;\u0026#34;macOS, the BSDs, illumos, Solaris. Read the ICMP off a raw socket, as root.\u0026#34;\u0026#34;\u0026#34; def __init__(self, family): proto = socket.IPPROTO_ICMP if family == socket.AF_INET else socket.IPPROTO_ICMPV6 self.sock = socket.socket(family, socket.SOCK_RAW, proto) self.sock.setblocking(False) def arm(self, sock, family): pass def extra_readers(self): return [self.sock] def collect(self, sock, family): try: packet, peer = self.sock.recvfrom(1500) except OSError: return None if family == socket.AF_INET: # BSD raw sockets hand back the IP header too. header_len = (packet[0] \u0026amp; 0x0F) * 4 packet = packet[header_len:] if len(packet) \u0026lt; 2: return None return peer[0], packet[0], packet[1] def probe(dest, port, proto, family, hop_limit, timeout, listener): \u0026#34;\u0026#34;\u0026#34;One probe at one hop limit. Returns (icmp, socket_state, note).\u0026#34;\u0026#34;\u0026#34; kind = socket.SOCK_STREAM if proto == \u0026#34;tcp\u0026#34; else socket.SOCK_DGRAM sock = socket.socket(family, kind) if family == socket.AF_INET: sock.setsockopt(socket.IPPROTO_IP, socket.IP_TTL, hop_limit) else: sock.setsockopt(socket.IPPROTO_IPV6, socket.IPV6_UNICAST_HOPS, hop_limit) listener.arm(sock, family) sock.setblocking(False) try: if kind == socket.SOCK_DGRAM: sock.connect((dest, port)) sock.send(b\u0026#34;\\x00\u0026#34; * 32) else: try: sock.connect((dest, port)) except BlockingIOError: pass except OSError as exc: sock.close() return None, None, \u0026#34;local error: %s\u0026#34; % exc.strerror readers = [sock] + listener.extra_readers() writers = [] if kind == socket.SOCK_DGRAM else [sock] deadline = time.time() + timeout icmp = state = None while time.time() \u0026lt; deadline: ready_r, ready_w, ready_x = select.select( readers, writers, [sock], max(0.01, deadline - time.time())) if not (ready_r or ready_w or ready_x): continue # Drain the error queue first, always. An ICMP error reaches a TCP # socket as a plain errno, so SO_ERROR on its own cannot tell you # whether a router spoke or the far end did. icmp = listener.collect(sock, family) if icmp: break if ready_w: err = sock.getsockopt(socket.SOL_SOCKET, socket.SO_ERROR) if err == 0: state = \u0026#34;connected\u0026#34; elif err == errno.ECONNREFUSED: state = \u0026#34;TCP reset\u0026#34; else: state = os.strerror(err) break sock.close() if icmp or state: return icmp, state, None return None, None, \u0026#34;no reply\u0026#34; def walk(dest, port, proto, family, first, last, timeout, listener): print(\u0026#34;walking to %s %s/%d hop limit %d-%d\u0026#34; % (dest, proto.upper(), port, first, last)) answered = [] for hop in range(first, last + 1): started = time.time() icmp, state, _ = probe(dest, port, proto, family, hop, timeout, listener) rtt = (time.time() - started) * 1000 if icmp: addr, icmp_type, icmp_code = icmp print(\u0026#34; %2d %-39s %7.1f ms ICMP %d/%d %s\u0026#34; % (hop, addr or \u0026#34;?\u0026#34;, rtt, icmp_type, icmp_code, describe(family, icmp_type, icmp_code))) if is_expiry(family, icmp_type): answered.append((hop, addr)) else: return answered, hop, \u0026#34;icmp-reject\u0026#34;, addr elif state: print(\u0026#34; %2d %-39s %7.1f ms %s\u0026#34; % (hop, dest, rtt, state)) return answered, hop, state, dest else: print(\u0026#34; %2d *\u0026#34; % hop) return answered, None, \u0026#34;silent\u0026#34;, None def main(): parser = argparse.ArgumentParser(description=__doc__.splitlines()[0]) parser.add_argument(\u0026#34;host\u0026#34;) parser.add_argument(\u0026#34;port\u0026#34;, nargs=\u0026#34;?\u0026#34;, type=int, default=443) parser.add_argument(\u0026#34;--proto\u0026#34;, choices=(\u0026#34;tcp\u0026#34;, \u0026#34;udp\u0026#34;), default=\u0026#34;tcp\u0026#34;) parser.add_argument(\u0026#34;--first\u0026#34;, type=int, default=1, help=\u0026#34;hop limit to start at\u0026#34;) parser.add_argument(\u0026#34;--max\u0026#34;, type=int, default=20, help=\u0026#34;hop limit to stop at\u0026#34;) parser.add_argument(\u0026#34;--wait\u0026#34;, type=float, default=2.0, help=\u0026#34;seconds to wait per hop\u0026#34;) parser.add_argument(\u0026#34;-6\u0026#34;, dest=\u0026#34;v6\u0026#34;, action=\u0026#34;store_true\u0026#34;, help=\u0026#34;force IPv6\u0026#34;) parser.add_argument(\u0026#34;-4\u0026#34;, dest=\u0026#34;v4\u0026#34;, action=\u0026#34;store_true\u0026#34;, help=\u0026#34;force IPv4\u0026#34;) args = parser.parse_args() family = socket.AF_INET6 if args.v6 else socket.AF_INET kind = socket.SOCK_STREAM if args.proto == \u0026#34;tcp\u0026#34; else socket.SOCK_DGRAM dest = socket.getaddrinfo(args.host, args.port, family, kind)[0][4][0] if sys.platform.startswith(\u0026#34;linux\u0026#34;): listener = ErrorQueue() else: try: listener = RawIcmp(family) except PermissionError: sys.exit(\u0026#34;%s cannot report ICMP errors on a normal socket, so this \u0026#34; \u0026#34;needs a raw one. Run it as root.\u0026#34; % sys.platform) answered, stop, why, who = walk(dest, args.port, args.proto, family, args.first, args.max, args.wait, listener) what = \u0026#34;%s/%d\u0026#34; % (args.proto.upper(), args.port) print() if why == \u0026#34;connected\u0026#34;: print(\u0026#34;Verdict: %s is open. It answered at hop %d.\u0026#34; % (what, stop)) elif why == \u0026#34;icmp-reject\u0026#34;: print(\u0026#34;Verdict: %s at hop %d is refusing %s on policy, and is honest \u0026#34; \u0026#34;enough to say so.\u0026#34; % (who, stop, what)) elif why == \u0026#34;TCP reset\u0026#34;: print(\u0026#34;Verdict: a reset came back to a probe with a hop limit of %d.\u0026#34; % stop) print(\u0026#34; Nothing more than %s away can have sent it, so check the reply\u0026#34; % (\u0026#34;one hop\u0026#34; if stop == 1 else \u0026#34;%d hops\u0026#34; % stop)) print(\u0026#34; TTL before you believe the host did.\u0026#34;) elif answered: last_hop, last_addr = answered[-1] print(\u0026#34;Verdict: answers stop after hop %d (%s).\u0026#34; % (last_hop, last_addr)) print(\u0026#34; Whatever swallows %s is hop %d.\u0026#34; % (what, last_hop + 1)) print(\u0026#34; Walk a port that works and read off the address at hop %d.\u0026#34; % (last_hop + 1)) else: print(\u0026#34;Verdict: nothing answered at all, not even the first hop. Either the\u0026#34;) print(\u0026#34; first hop is the one dropping you, or the ICMP errors are being\u0026#34;) print(\u0026#34; filtered on the way back. Walk a port that works to tell those\u0026#34;) print(\u0026#34; two apart.\u0026#34;) if __name__ == \u0026#34;__main__\u0026#34;: main() What This Cannot Tell You The method is cheap and it is honest about most things, but it is not a topology scanner. Be straight in the ticket about what you actually measured.\nDifferent five-tuples can take different paths. ECMP hashes the source and destination ports into the choice of next hop, so two runs on two different ports are not guaranteed to traverse the same routers at all, which is one of the two things that could explain hop 5 above answering one run and staying quiet on the other. Repeat both runs. A boundary that moves is unproven.\nMPLS hides hops. A label-switched core can present as one hop, or as none at all. Any count across somebody else\u0026rsquo;s backbone is a lower bound.\nICMP generation is rate limited nearly everywhere. Probe faster than the router will answer and you manufacture your own stars. hopfind.py sends one probe per hop and waits; that is deliberate.\nThe return path need not match the outbound one. The hop count out is not the hop count back, and reverse-TTL arithmetic measures the return leg only.\nAnycast means the host at hop N may not be the same box twice. Public resolvers and CDNs, in particular.\nA completed handshake does not mean the session survives. A stateful firewall can permit the SYN and drop what follows on inspection. If the connection opens and then dies, this is the wrong instrument — go and capture.\nYou have found the first device that drops, not the one anybody will admit to. In a CGN or a carrier network the address at hop N+1 may be one of several boxes behind one address. It is still the right thing to quote, because it is a fact about the path.\nWhat It Is Actually For Ending the bouncing. That is the whole return on the exercise.\n\u0026ldquo;Port 445 is blocked somewhere\u0026rdquo; is an invitation to hand the ticket back. This is not:\nTCP/445 to 192.0.2.4 is dropped silently at hop 11, address 192.0.2.31. Hop 11 answers ICMP Time Exceeded to TCP/443 on the same path in 26 ms and answers nothing at all on 445, so the drop is a policy decision on that device, not a routing fault. Ten hops in front of it are clean. Reproduced four times over twenty minutes, from an unprivileged shell, script attached.\nNobody hands that back. It names a device, states what it did, states what it did not do, and shows the working. Whether they choose to change it is still their call. But the week of ping-pong is over, and it took four commands.\nEverything in it came out of an 8-bit field that was specified as a timer in 1981, has never once been used as one, and quietly turns \u0026ldquo;somewhere\u0026rdquo; into an address.\nWorth learning to read.\n","permalink":"https://blogs.damiendye.uk/en/networking/how-far-away-is-the-firewall/","summary":"\u0026ldquo;Port 445 is blocked somewhere\u0026rdquo; is not a diagnosis, and it is why firewall tickets bounce between you and your provider for a week. Every router on the path owes you an ICMP Time Exceeded when your hop budget runs out, and that turns a timeout into a distance. I walked the TTL up on my own line and found three faults I did not know I had: an SMB drop eleven hops out, forged SMTP resets one hop away, and an IPv4 rule with no IPv6 twin.","title":"The Firewall Is Eleven Hops Away"},{"content":"There is a story the UK industry tells about IPv6, and it goes like this. The move is hard. The kit is old. The customers do not ask for it. There is no money in it. One day, when the business case turns, we will get to it.\nEvery part of that is a lie the industry tells itself so it does not have to do any work.\nIPv6 was specified in December 1995. I got onto it through the 6bone, the experimental testbed that carried it before the real internet would, and ran both the Linux stack and Microsoft Research\u0026rsquo;s stack on Windows XP to see how they differed. My access came from Hurricane Electric.\nThe 6bone was switched off on 6 June 2006, so I moved to 6to4 automatic tunnelling, and later to a Hurricane Electric tunnel — still free, and they will route you a /48 for the asking. Native IPv6 arrived at my house in 2017, when I changed ISP to Zen.\nSo for the best part of twenty years my IPv6 came from an American transit company giving it away, rather than from any of the British ISPs I was paying. Hurricane Electric handed out routed /48s to anyone who wanted one. My own provider would sell me a static IPv4 for a fiver a month.\nThe dates say the rest. The IETF killed the 6bone in 2006 and deprecated 6to4\u0026rsquo;s anycast relays in May 2015, calling the mechanism \u0026ldquo;unsuitable for widespread deployment and use in the Internet\u0026rdquo;. I outlived two official transition mechanisms waiting for a British ISP to hand me an address. On the day the 6bone shut down, thirty-seven of the forty UK providers in the chart below had not yet asked the registry for an allocation. Twenty-two of them — more than half — did not ask until 2015 or later, the year the IETF gave up on 6to4 as well.\nIt has been switched on by default in every operating system anyone runs since Windows Vista in 2007. It costs nothing extra from the registry. The biggest ISP that ever tried it in this country finished the job in three years with a team you could fit in a meeting room, and won an award for it.\nThirty years past the specification. Fourteen years past the day the internet permanently switched it on. And the answer in this country was to break the internet on purpose, wrap the broken bit in more machinery, and bill the customer for the inconvenience.\nThat is not a cost problem. It is a can\u0026rsquo;t-be-bothered problem, and it has been going on for twenty years.\nWhat I Measured, and How Everything below is either somebody else\u0026rsquo;s work, linked, or a number I produced myself. Where it is mine, the script that counted it is in the download below, run as published. The one exception is the address totals, and I explain how those are worked out in the caveats. Four public sources, all free: the registry delegation files, the RIPE routing table dump, the RIPE database, and the DNS. Where I picked a sample instead of measuring the lot, I say so.\nThe routing numbers come from two sets of public files.\nThe first is the delegation files, one per regional registry, which list every address block and AS number that registry has handed out, the country it is registered to, and an opaque identifier for the organisation holding it. RIPE\u0026rsquo;s covers Europe and the Middle East, and it is the one that matters for the UK — but a few dozen UK-registered AS numbers sit in the ARIN and APNIC files instead, and the comparison further down needs the lot. Mine were generated 26 and 27 August 2026.\nThe second is the RIPE Routing Information Service\u0026rsquo;s dump of the global routing table, which lists every prefix in BGP and the AS number announcing it. IPv4 and IPv6 come as separate files. Mine was generated at 18:06 UTC on 27 August 2026.\nPut them together and you can answer a question nobody in the UK industry wants asked out loud: of the networks this country registered, how many have actually turned IPv6 on?\n\u0026#8615; The scripts, ready to run ipv6-uk-2026-scripts.zip · 10 kB Take the AS numbers registered to GB out of the delegation files, take every AS number originating a prefix out of the RIS dumps, and comm the two lists against each other per address family. That gives the first table below. One trap worth naming: sort lexically, not with sort -n. comm compares strings, and numerically sorted input silently gives you the wrong answer rather than an error you would notice.\nSwap GB for any other country code and you get that country\u0026rsquo;s row in the comparison table further down. 05-country-row.sh does exactly that.\nThe organisation-level numbers use the eighth field, which is the registry\u0026rsquo;s opaque handle for the account holding each resource. These stay on the RIPE file alone — the handles are local to each registry, so concatenating five of them would count the same company twice rather than merging it. UK organisations are RIPE members, so RIPE is where they are. This is how many hold an AS number and no IPv6 at all:\n03-org-no-ipv6.sh counts them. And 04-silent-holders.py is the one that matters most. The organisations that hold IPv6, are live in BGP, and announce none of it.\nFour caveats before the numbers, because they matter and I would rather say them than have them thrown at me.\nThe organisation handles are per registry account, so a company running several accounts counts more than once.\nAnnouncing an IPv6 prefix in BGP is not the same as handing IPv6 to a customer. It is the floor, not the ceiling. A network that announces nothing has certainly not deployed it. A network that announces something might still be sitting on it.\nThe address totals collapse overlapping prefixes. A network announcing a /16 alongside four /17s out of it is announcing 65,536 addresses, not 196,608, and counting the prefixes naively inflates the big holders by two or three times. I use Python\u0026rsquo;s ipaddress.collapse_addresses before totalling.\nThe list of fifty websites later on is a sample I picked by hand, not a measurement of the whole country. A different fifty would give a different fraction. It illustrates a pattern rather than proving a proportion, and I name the ones I am talking about as I go.\nThe first two of those make the routing numbers kinder to the industry than the truth.\nThe Count UK AS numbers, and how many carry IPv6 UK networks in the global routing table, 27 August 2026 Source: RIPE NCC delegation file and RIPE RIS routing table dump registered to UK organisations 3,106 visible in the routing table 2,248 announcing IPv4 2,078 announcing IPv6 1,048 IPv4 and no IPv6 at all 1,200 57.7% of live UK networks Of those 1,200, a total of\u0026#160;463 hold IPv6 space the registry has already issued to them and have never announced a single prefix of it. Another 1,113 UK organisations holding an AS number have never asked for IPv6 at all, though it costs nothing on top of the membership they already pay. Every AS number registered to a UK organisation, measured against the global routing table on 27 August 2026. The gap on the right is the whole argument: 1,200 UK networks are live on the internet with no IPv6 at all, and 463 of them are holding registry-issued IPv6 space they have never announced. AS numbers registered to UK organisations 3,106 Visible in the global routing table 2,248 Announcing IPv4 2,078 Announcing IPv6 1,048 Announcing IPv4 and no IPv6 1,200 — 57.7% Nearly six in ten live UK networks do not carry IPv6 at all. Not partially. Not behind a flag. Not in a lab. Not one prefix.\nNow the part that ends the cost argument for good.\nOf the 2,363 UK organisations holding an AS number, 1,113 — 47.1% — hold no IPv6 allocation of any kind. They have never asked the registry for it.\nA RIPE NCC membership costs EUR 1,800 a year for 2026, flat, and that fee covers your allocations. An IPv6 /29 gives you 524,288 subnets the size of the entire IPv4 internet. It is free with a membership these organisations are already paying for, and it arrives in a couple of days.\nHalf of them never filled the form in.\nAnd of the ones who did, 463 UK organisations hold IPv6 space, announce IPv4 to the world every day, and announce no IPv6 whatsoever. That is 44.7% of the UK IPv6 holders that are live in BGP.\nRead that again, because it is the whole post in one sentence. They asked for the addresses. They were given the addresses. They put them in a spreadsheet. Then nobody could be bothered to type them into a router.\nYou cannot explain that with money. Nobody spent anything. There is no invoice, no procurement, no business case, no capital request. There is a free thing sitting in a registry account, and an engineering department that has not opened the ticket in fourteen years.\nWho Is On That List These are the largest UK networks announcing IPv4 and no IPv6, by the amount of address space they actually announce, on 27 August 2026. The names come from the RIPE database, which will tell you who holds any of them:\ncurl -s https://rest.db.ripe.net/ripe/aut-num/AS15914.json \\ | python3 -c \u0026#39;import sys,json; a=json.load(sys.stdin)[\u0026#34;objects\u0026#34;][\u0026#34;object\u0026#34;][0][\u0026#34;attributes\u0026#34;][\u0026#34;attribute\u0026#34;]; print(next(x[\u0026#34;value\u0026#34;] for x in a if x[\u0026#34;name\u0026#34;]==\u0026#34;org\u0026#34;))\u0026#39; The largest UK networks with no IPv6 The biggest UK networks announcing IPv4 and no IPv6, 27 August 2026 IPv4 addresses announced in BGP, overlapping prefixes collapsed. Names from the RIPE database. 0k 50k 100k 150k 200k Vodafone Limited AS25310 229,376 Nationwide Building Society AS8698 131,072 British Airways plc AS15914 131,072 Convergence Group (Metronet) AS42973 94,976 Rackspace Ltd AS24867 86,016 Lloyds Banking Group AS49758 81,920 QinetiQ Limited AS24775 69,632 Barclays Bank plc AS12701 68,608 MUFG Securities EMEA AS8651 65,792 University of Warwick AS201773 65,792 NatWest Markets plc AS21054 65,536 PricewaterhouseCoopers Services AS21296 65,536 London Borough of Hackney AS39400 65,536 Wireless Logic Limited AS51320 39,424 Between them the UK networks announcing no IPv6 sit on\u0026#160;4,232,232 IPv4 addresses. Barclays has held 141.228.0.0/16 since August 1990. The fourteen largest UK networks announcing IPv4 and no IPv6 on 27 August 2026, by the address space they actually announce. Overlapping prefixes collapsed before totalling. Look at that list and try to say the words \u0026ldquo;cost barrier\u0026rdquo; without laughing.\nFour of the clearing banks. A global accountancy firm whose entire product is telling other people how to run their affairs. A defence technology company. A hosting provider whose customers pay it to know this. An internet-of-things connectivity business, selling SIM cards, with no IPv6.\nBetween them, the UK networks announcing no IPv6 are sitting on 4,232,232 IPv4 addresses. The transfer market averaged $20.04 an address across the first half of 2026, so that is a holding worth somewhere north of eighty million dollars. Which is the actual reason none of them have moved: they are rich in addresses, so the shortage is somebody else\u0026rsquo;s problem, and the internet\u0026rsquo;s long-term plumbing is nobody\u0026rsquo;s job in particular.\nThat is not a strategy. That is being comfortable.\nThe delegation file carries the date each block was handed out, so you can see exactly how comfortable:\ngrep -E \u0026#39;\\|ipv4\\|(141\\.228|155\\.131|155\\.136|161\\.2)\\.0\\.0\\|\u0026#39; \\ delegated-ripencc-extended-latest | cut -d\u0026#39;|\u0026#39; -f4,5,6 Barclays has held 141.228.0.0/16 since 6 August 1990. Nationwide and NatWest took theirs in November 1991, four days apart. British Airways got 161.2.0.0/16 in April 1992. These are class B blocks from before the web existed, handed out when addresses were free and nobody counted.\nEveryone who came after them pays for that. AWS started charging $0.005 an hour for every public IPv4 address on 1 February 2024 — $43.80 a year, each — and said plainly why: the cost of acquiring one \u0026ldquo;has risen more than 300% over the past 5 years.\u0026rdquo; The scarcity is real and it has a price. It is just not being paid by the people holding four million addresses they got for nowt in 1991.\nHow We Look Against Countries Like Us Two numbers per country. The first is the share of its people who reach Google over native IPv6, which is Google\u0026rsquo;s measurement on 25 August 2026. The second is the share of its live networks that announce an IPv6 prefix, which is mine, from the same files as above. I have kept it to developed economies. Comparing ourselves with countries that got the internet late tells you nothing about us. Ordered by people.\nCountry Users on IPv6 Networks with IPv6 Live networks France 85.6% 48.5% 1,368 Germany 76.6% 63.9% 2,291 Belgium 72.8% 45.8% 273 United States 56.6% 25.9% 18,453 Japan 56.1% 56.4% 721 United Kingdom 53.7% 42.3% 2,078 Norway 52.6% 66.9% 278 Netherlands 51.9% 61.8% 1,023 Canada 43.6% 33.1% 1,578 Ireland 38.1% 41.5% 195 Australia 37.2% 26.3% 1,652 Sweden 36.1% 53.6% 642 South Korea 18.1% 5.3% 916 Italy 17.6% 35.8% 1,078 Spain 13.3% 26.7% 934 Sixth of fifteen. France has two thirds again as many of its people on IPv6 as we do, off the same European supply chain, under the same equipment vendors, with the same customers telling them nobody is asking for it. Germany is twenty-three points ahead of us on users and twenty-two points ahead on networks.\nThe two percentage columns do not agree with each other, and the disagreement is the story.\nUsers on IPv6 against networks carrying IPv6, fifteen developed economies People on IPv6, against networks carrying IPv6 Google's user measurement, 25 August 2026, against my count of live networks announcing an IPv6 prefix, 27 August 2026. share of people share of networks France 85.6% 48.5% Germany 76.6% 63.9% Belgium 72.8% 45.8% United States 56.6% 25.9% Japan 56.1% 56.4% United Kingdom 53.7% 42.3% Norway 52.6% 66.9% Netherlands 51.9% 61.8% Canada 43.6% 33.1% Ireland 38.1% 41.5% Australia 37.2% 26.3% Sweden 36.1% 53.6% South Korea 18.1% 5.3% Italy 17.6% 35.8% Spain 13.3% 26.7% 0% 20% 40% 60% 80% 100% A long blue bar over a short pink one means three or four carriers did the work and the rest of the country did not. France, Belgium, the United States and the United Kingdom all have that shape. Norway, Sweden and Japan do not. The same fifteen countries on both measures, ordered by the share of people using IPv6. Where the network bar is much shorter than the user bar, a few large carriers are carrying the country and nobody else has bothered. That is the shape of the United States, Belgium, France — and the United Kingdom. A country\u0026rsquo;s user percentage is set by three or four companies. The network percentage is set by everybody else. When the first is high and the second is low, it means the big access networks did the job and the rest of the country free-rode on them.\nThe United States is the clearest case: 56.6% of its people are on IPv6 and only 25.9% of its networks are. The cable and mobile carriers carry nearly everyone. The other eighteen thousand American networks did nothing.\nOurs is the same trick with smaller numbers — 53.7% of users against 42.3% of networks. That 53.7% is not a national achievement. It is Sky and BT, and a rounding error of everyone else.\nNorway and Sweden are the honest counter-shape: fewer users on IPv6 than us, more networks carrying it. More of their industry has actually done the work, and it is the consumer ISPs that are lagging rather than the trade.\nAnd one row deserves a closer look, because it is the one people reach for when they want to feel better about us.\nSouth Korea is the worst country on this list, by a distance. Of 916 live Korean networks, 61 announce IPv6. Sixty-one.\nSome of the fastest domestic broadband on earth, a chip industry that prints money, and 94.7% of its networks never turned it on. The three big carriers — KT, SK Broadband and LG U+ — all announce it, which is why 18.1% of Korean users have it. The other eight hundred and fifty networks did nowt.\nWhatever the excuse is there, it is not money, it is not capability, and it is not the state of the fibre.\nTwenty Years Of Bolting Things On Here is what the industry built instead of typing in the addresses.\nWhen the addresses started running short, the answer was carrier-grade NAT: put hundreds of customers behind one public IPv4 address and translate between them. Everything below exists to make that survivable, and every one of these documents is a piece of engineering work somebody chose to do rather than deploy IPv6.\nBolt-on What it is for RFC 6598 (2012) Burns an entire /10 — four million addresses — as \u0026ldquo;shared address space\u0026rdquo;, so the shortage workaround needs its own addresses RFC 6333 (2011) DS-Lite: tunnel IPv4 over the IPv6 network you built but did not give the customer RFC 6877 (2013) 464XLAT: translate IPv4 to IPv6 and back again on the same journey RFC 6888 (2013) The list of requirements a carrier-grade NAT must meet not to be dangerous RFC 7021 (2013) A full study of the applications carrier-grade NAT breaks RFC 7422 (2014) Deterministic address mapping, invented purely to stop the logging volume bankrupting the provider RFC 7597 / 7599 (2015) MAP-E and MAP-T: two more ways to carry IPv4 over IPv6 without admitting you have IPv6 Twenty years of workarounds against the one change they replace What we built instead, and what it was instead of Avoiding IPv6 Shared address space — a whole /10 burned RFC 6598, 2012 DS-Lite — tunnel IPv4 over an IPv6 core RFC 6333, 2011 464XLAT — translate out of IPv4 and back RFC 6877, 2013 Rules to make carrier-grade NAT survivable RFC 6888, 2013 A study of the applications it breaks RFC 7021, 2013 Deterministic mapping to cut the log volume RFC 7422, 2014 MAP-E and MAP-T — IPv4 over IPv6 again RFC 7597/9, 2015 Plus the NAT tier itself: session tables, port block allocation, application gateways, failover, capacity planning, a logging pipeline — and a retention law passed to paper over the attribution it destroyed. Doing IPv6 Dual-stack One more address family, on the routing protocol and firewall policy you already run. No new tier in the traffic path. No session state to size. No logs kept to satisfy a statute. Free from the registry. Both columns are engineering work, and the left one is larger. The difference is that the work on the left can be bought from a vendor, and the work on the right has to be understood by the people who own the network. Two decades of standards work, hardware and logging, all of it in service of not doing the thing on the right. Dual-stack is one address family added alongside the one you already run. Everything on the left exists to avoid it. Look at the shape of that. Every item on the list is harder than dual-stack. Tunnelling IPv4 inside IPv6 is strictly more work than routing IPv6, because you have to route the IPv6 anyway to carry the tunnel. Translating between families is more work than not translating. A carrier-grade NAT is a stateful box in the middle of your network, with capacity planning, failover, session tables, port block allocation, application-layer gateways for the protocols it breaks, and a logging pipeline sized for a legal obligation.\nDual-stack is an address family, a routing protocol you already run, and a firewall policy you already wrote.\nThe industry looked at those two options and picked the expensive one, twenty years running, because the expensive one could be bought and the cheap one had to be understood. Buying a box is a procurement exercise. Turning on IPv6 means somebody in the building has to know how the network works.\nWhat This Actually Breaks For anyone who thinks this is aesthetics, here is what a shared address costs your users, in the order they will phone you about it.\nNothing can reach in. No port forwarding, so no self-hosted anything, no games console acting as host, no site-to-site VPN without a relay, no security camera without a vendor cloud, no remote access to the thing at the other site. Every one of those gets replaced by a third-party rendezvous service, which is another company holding your data because your provider would not give you an address.\nYou inherit strangers\u0026rsquo; reputations. Share an address with a few hundred people and you share their behaviour. Rate limits, CAPTCHAs, Wikipedia blocks, streaming geo-errors and fraud scoring all land on you for something someone else did.\nPorts run out. A carrier-grade NAT has 65,535 ports per public address per protocol, and one modern browser session eats dozens. Oversubscribe and the failure is not a clean error. It is a slow, intermittent, unreproducible fault that looks like everything except what it is, and it burns days of support time per incident.\nEvery workaround has to be maintained forever, by people who could have spent that time on the fix.\nAnd then there is the one that stopped being an inconvenience and became everybody\u0026rsquo;s problem. Nobody can tell who did what.\nThe Bolt-On That Reached Parliament Once hundreds of customers share one address, an address no longer identifies anybody. So the police cannot resolve an IP address to a person, and the answer to that was not IPv6. It was legislation.\nSection 21 of the Counter-Terrorism and Security Act 2015 amended the data retention regime specifically so the Secretary of State could compel providers to retain the extra data needed \u0026ldquo;to link the unique attributes of a public Internet Protocol (IP) address to the person (or device) using it at any given time.\u0026rdquo; The explanatory notes are blunt about why it was needed: providers \u0026ldquo;may share IP addresses between multiple users, and the providers generally have no business purpose for keeping a log of who used each address at a specific point in time.\u0026rdquo;\nRead that as an engineer rather than a lawyer. The industry broke attribution to save itself some work, and Parliament passed a law obliging it to build a logging system to paper over the breakage.\nTwo years later Europol put it plainly. In October 2017 it published a call for the industry to stop using carrier-grade NAT, with figures: 90% of mobile internet access providers and 50% of fixed-line providers had adopted a technology that stopped them identifying their own subscribers. Europol\u0026rsquo;s then executive director said CGN \u0026ldquo;has created a serious online capability gap in law enforcement efforts to investigate and attribute crime\u0026rdquo;, and noted that it \u0026ldquo;forces judiciary and law enforcement authorities to investigate many more individuals than would normally be necessary.\u0026rdquo;\nEuropol also said the quiet part. Carrier-grade NAT \u0026ldquo;was supposed to be a temporary solution until the transition to IPv6 was completed\u0026rdquo;. Instead the industry kept increasing its use of it while the replacement sat there, finished, free and ignored.\nSo the cost of not deploying IPv6 includes: a piece of primary legislation, a nationwide retention obligation, innocent people pulled into investigations because they shared an address with somebody who was not innocent, and an ongoing capability gap that the police describe as a public safety problem.\nNobody put that on the business case. It never appears in the \u0026ldquo;IPv6 has no ROI\u0026rdquo; slide, because it is not paid for by the people who caused it.\nIt Does Not Just Hide Criminals. It Helps Them. The attribution argument is the one law enforcement makes, and it is about catching people after the event. There is a second argument that gets made far less often and is worse: address sharing actively degrades the defences that stop attacks happening at all.\nThis is not my analysis. The IETF published the catalogue in RFC 6269, Issues with IP Address Sharing, in June 2011. That was before the UK deployed most of the carrier-grade NAT it now runs. Its plain words: address sharing \u0026ldquo;creates a vector for attack amplification in numerous ways.\u0026rdquo;\nHere is what it warned would break, and did.\nRate limiting and lockouts stop working. The standard defence against password guessing and credential stuffing is to count failures per address and put the offender in a penalty box. Share that address between hundreds of people and the counter is measuring a crowd. RFC 6269 is blunt about the outcome: \u0026ldquo;In the presence of widespread large-scale address sharing, penalty box solutions to service abuse simply will not work.\u0026rdquo; One user\u0026rsquo;s failed logins lock out everybody else, so operators raise the thresholds, and raising the thresholds is what the attacker wanted.\nBlocklisting becomes collateral damage. Block the spammer and you block the road they live on. So the sensible operator stops blocking, and the abuse continues from an address nobody dares touch.\nInfected machines stay infected. Abuse feeds and malware notifications arrive as an address and a timestamp. Behind a CGN with no port logging, the provider cannot tell which of its customers is running the bot, so the customer never gets told and the infection stays up. Worse, RFC 6269 notes the reverse problem: \u0026ldquo;someone else\u0026rsquo;s worm can interfere with the ability to access the service for other subscribers sharing the same IP address.\u0026rdquo;\nAddress-based access control fails. Every allow-list built on source address is now admitting a crowd rather than a customer.\nAnd one defence is measurably weakened rather than merely blunted. Blind TCP attacks depend on guessing the five-tuple, and the industry\u0026rsquo;s mitigation is randomising the source port (RFC 6056). A carrier-grade NAT hands each subscriber a slice of the port range instead of the whole of it. In RFC 6269\u0026rsquo;s words, \u0026ldquo;with shared IPv4 addresses, the port selection space is reduced.\u0026rdquo; The workaround for the address shortage takes entropy directly out of an anti-attack mechanism.\nThen there is the part that ought to worry anybody, regardless of what they think about policing. If the server did not log source ports and the NAT did not log destinations, RFC 6269 spells out what a provider must do when a lawful request arrives: it \u0026ldquo;would need to disclose the identity of all subscribers who had active sessions on the NAT during the time period in question. This may be a large number of subscribers.\u0026rdquo;\nThe alternative to identifying one guilty subscriber is handing over the identities of several hundred innocent ones. That is the actual privacy outcome of address sharing, and it is the opposite of the one its defenders claim for it.\nThree Things Have To Line Up, And Nobody Is Required To Provide Any Of Them People assume the logs exist somewhere and it is a matter of asking. They mostly do not, and the reason is arithmetic rather than ill will.\nTo turn a shared address back into a household, three separate things must all have gone right:\nThe far-end server logged the source port. RFC 6302 asked internet-facing servers to log source port and timestamp alongside the address, in 2011. It is a recommendation. Nobody enforces it, and a great many servers still log the address alone. At which point the trail is dead before it reaches the British end. The provider kept the mapping. Every session, for months. The clocks agreed. RFC 6269 warns that on a busy CGN \u0026ldquo;even very small amounts of clock skew between a third party\u0026rsquo;s server and the CGN operator will result in ambiguity about which customer was using a specific port at a given time.\u0026rdquo; Miss any one and you have nothing. And the middle one is where it falls apart, because the standards documents contain the sums.\nRFC 7422 put real numbers on it. Operators reported roughly 33,000 connections per household per day. At about 150 bytes per log entry that is 5 MB per subscriber per day, 150 MB a month. For a provider with a million subscribers: 150 terabytes of logs a month, 1.8 petabytes a year — to be kept for the six to twelve months the law expects, and searched on demand.\nAnd it is never one log. NAT444, the case RFC 7422 sizes its entries for, puts one translation in the customer\u0026rsquo;s router and another at the carrier, and every gate a packet crosses has to write down what it did. Reconstructing a single session means correlating separate tables, kept by separate parties, against the clock problem above. The evidence arrives in pieces from different systems, or it does not arrive.\nAnd the money is only half of it. Capturing session records at that rate, shipping them somewhere, indexing them so a lawful request comes back in hours rather than weeks, and holding the lot for a year is a data engineering project. There is no dashboard for it and no box to buy. It has to be built by somebody who understands what they are building, and that is not point and click, which in this industry is close enough to saying it does not get built.\nNobody was ever going to pay for it either. And the IETF knew it, which is why RFC 6888 tells operators the opposite of what public safety needs: \u0026ldquo;A CGN\u0026rsquo;s port allocation scheme SHOULD minimize log volume\u0026rdquo;, justified because \u0026ldquo;huge log volumes can be problematic to CGN operators.\u0026rdquo; RFC 7422 exists for no other purpose than cutting that bill down.\nSo the design advice to the industry is log less, the economics say 1.8 petabytes a year is unaffordable, and the legal expectation is a complete record. Those cannot all be true at once, and the one that gives is the record.\nThat is why Europol found the majority of access providers cannot identify a subscriber when served with a legal order. Not because they are obstructive. Because the thing being asked for was never economically possible to keep, and nobody was ever made to.\nTwo things follow from that, and they are mine rather than anybody\u0026rsquo;s quote.\nFirst: a large share of British internet connections are unattributable by construction. Mobile is the clearest case, and independent measurement puts it higher than Europol did, at 95%. So the anonymity that used to require Tor, or a VPN somebody had to buy, is now the factory setting on a British mobile connection — issued free with the SIM, to everybody, including the small number of people the whole apparatus is meant to find.\nSecond, and worse: breaking the cheap targeted method is what produces demand for the expensive untargeted one. When you can serve a warrant on one address and get one household, you need nothing else. When that stops working, the state does not shrug. It reaches for something broader. That is what the 2015 Act was: a retention duty across the whole subscriber base, to answer questions about a handful of people.\nNone of that exists on the other side. Nothing is translated, so there is no per-connection record to keep at all. The address in the far-end server\u0026rsquo;s log is already the subscriber\u0026rsquo;s prefix: one record, written once when the line was provisioned, in one system. Even a provider rotating prefixes daily writes a few hundred a year per customer, against the twelve million that 33,000 connections a day comes to. It is not that IPv6 logs less. There is nothing to log.\nThe people who wrote the workaround knew it. In the middle of a specification written for no reason but to make CGN affordable to log, they stopped to record that \u0026ldquo;native IPv6 will offer subscribers a better experience than CGN\u0026rdquo;.\nAn industry declined to spend a fortnight per network on a free protocol, and the country got a data retention regime instead.\nWould Fixing This Protect Children Better Than The Online Safety Act? I want to be careful here, because it is easy to make this argument badly and the badly-made version deserves the kicking it would get.\nStart with how a child abuse investigation actually runs. A platform detects the material and reports it. The report carries an address and a timestamp. Police serve the access provider to turn that into a subscriber, and the subscriber is an address in the real world with a door on it. That is the whole chain, and every step after it depends on the one before.\nNow put a carrier-grade NAT in the middle. The report still arrives. The address still resolves. To several hundred households, and Europol found investigations \u0026ldquo;dropped or delayed\u0026rdquo; as a result. Their case examples include a prosecutor unable to identify the members of a forum supporting ISIS, so the prosecution did not happen, and HMRC tracing bulk tax fraud to mobile addresses and finding the leads \u0026ldquo;frustrated from the outset\u0026rdquo;.\nNCMEC\u0026rsquo;s CyberTipline took 21.3 million reports in 2025 and referred more than 18.8 million to law enforcement, including over 53,000 involving a child in immediate danger. NCMEC also records that more than 10% of industry reports arrived with information too poor to work out which jurisdiction to send them to. That figure is not CGNAT\u0026rsquo;s doing and I am not claiming it is — but it tells you where in this pipeline cases die. They die on metadata.\nSo the honest version of the comparison is this. Carrier-grade NAT breaks the Online Safety Act\u0026rsquo;s own last mile. Parliament has imposed duties to detect and report, and left the access network unable to resolve what gets reported. You may pass as many reporting duties as you like. If the final step returns a crowd, the report is paper.\nAnd the cost of the two things is not remotely comparable. The Act is the largest piece of internet regulation this country has attempted — thousands of services in scope, a regulator writing codes for years, age assurance running at millions of checks a day, and a circumvention problem large enough that Parliament has debated VPN use in the Lords. IPv6 costs nothing from the registry and takes a competent team a couple of weeks. One of those has been required of the entire industry. The other has never been asked of anyone.\nThree things need saying plainly, because the argument is worthless without them.\nOne. It is not a substitute, and I am not proposing it as one. IPv6 does nothing to stop a twelve-year-old finding pornography. It does nothing about recommender systems, autoplay or live-streaming. It does nothing about material hosted in another country, which is most of it. Those are the problems the Act was written for and no protocol change touches them.\nTwo. IPv6 is not an identity layer, and anyone selling it as one is overselling it. Privacy extensions rotate a device\u0026rsquo;s address by design, so the address of the machine is not the stable thing. What is stable is the prefix delegated to the line — the /56 that Sky has been handing every subscriber since 2016. That resolves to a subscriber, which is exactly the resolution a lawful request needs, and no more than that. It is a restoration of what a single IPv4 address per line used to give, not a new surveillance capability.\nThree. The property that frustrates the police also frustrates everyone else who tracks you, and some people value that. Being one of five hundred behind a shared address is genuine crowd cover against commercial profiling. I do not think it is worth what it costs — it is cover bought by making abuse unattributable and rate limiting useless, and it is cover the platforms mostly see through anyway with cookies and fingerprinting. But it is a real argument and it deserves stating rather than ignoring.\nSo no, this is not IPv6 instead of the Online Safety Act. It is that Britain wrote the most expensive online safety law in its history on top of plumbing it knew was broken, when the fix was free, well documented, and available for the whole time the Bill was being drafted.\nWhat To Actually Ask For Not a ban. A ban is the wrong instrument and it would backfire.\nThere is no IPv4 left to hand out — RIPE has been empty since November 2019, and a new provider gets a single /24 from a waiting list. Outlaw address sharing tomorrow and the small operator cannot connect customers at all, while the outfits sitting on class B blocks from 1990 carry on untouched. It would entrench exactly the people this post is about.\nThe better instrument already exists and somebody has already run the experiment.\nIn 2012 Belgium\u0026rsquo;s federal police, its telecoms regulator, its Council of Prosecutors-General and its ISP association signed a two-page voluntary code of conduct. Maximum 16 subscribers behind one IPv4 address. Limit the use of CGN. Start adopting IPv6.\nBy 2017 most Belgian operators were inside the limit, one had gone down to 8, and Belgian police were seeing an average of four users per mobile address. Europol\u0026rsquo;s own summary of why is the part worth reading twice: the biggest providers \u0026ldquo;are quickly moving towards IPv6 because no financial interest to invest in CGN anymore.\u0026rdquo; Cap the oversubscription and the economics of the workaround collapse, because a NAT that can only stack sixteen people is not cheaper than the protocol that needs none.\nThat year Belgium had the highest IPv6 adoption in the world at 49%, when Britain and France were on 14% and Spain and Italy were under 1%. Belgium is still third in the country table above, on 72.8%.\nThat is the ask. Not \u0026ldquo;you may not share addresses\u0026rdquo; — you may not sell a connection that is both unattributable and has no IPv6. Shared addressing alongside working IPv6 is fine. That is how every mobile network on earth operates. Shared addressing with no IPv6 is selling a broken service and charging the public for the consequences.\nSky Proved It Could Be Done, Eleven Years Ago If it were genuinely hard, nobody in the UK would have managed it.\nSky started an internal IPv6 project in early 2013 and finished in 2016, with roughly 90% of its fixed-line base — about five million users — picking up IPv6 and using it. Their engineer wrote the whole thing up on RIPE Labs: 6PE across the MPLS core, dual-stacked peering and transit, RADIUS attributes to enable it per subscriber, firmware work across seven CPE models including five legacy ones, and capacity upgrades on RADIUS and DNS.\nThree years, one ISP, and the same Openreach copper everybody else was selling over. ISPreview reported the finish in September 2016, with Sky expecting 95% of its base by the end of that year, and Sky took the Jim Bound IPv6 Award for it.\nTheir advice was: \u0026ldquo;Do not underestimate the work required to enable IPv6, and do not leave it to the last minute to begin the journey.\u0026rdquo;\nEleven years later, most of the industry is still at the last minute, and treating it as a place to live.\nEverybody Has Had The Addresses For Years Sky was the first of the big ISPs. It was nowhere near the first in the country, and not one provider on this list can say it was waiting on the registry.\nRIPE stamps the allocation date into the name of the block, so you can check any of them yourself:\nwhois -h whois.ripe.net 2a01:4b00::/32 | grep -E \u0026#39;netname|^org:\u0026#39; # netname: UK-BCUBE-20110225 -\u0026gt; Hyperoptic, allocated 25 February 2011 Every date below came from that lookup against the provider\u0026rsquo;s own allocation, cross-checked against the delegation file. The big ISPs are in bold, the rest are the full-fibre builders. Whether customers actually get IPv6 is from ISPreview\u0026rsquo;s survey as updated in March 2025, and from the mobile tracker for the phone networks.\nHow long each UK provider has held IPv6, and whether customers get it How long each UK provider has held IPv6 — and whether customers actually get it Allocation dates from the RIPE delegation file and its netname stamps. Whether it reaches customers: ISPreview, March 2025, and the community mobile tracker. customers get it partly — one network only still not shipping it (17 of 40) 2002 2002 2006 2006 2010 2010 2014 2014 2018 2018 2022 2022 2026 2026 TalkTalk (as Opal Telecom) Andrews \u0026amp; Arnold Vodafone UK EE (as T-Mobile) Sky Gigaclear O2 (Telefónica UK) BT Virgin Media (as NTL) KCOM Hyperoptic (as Bcube) Zen Internet B4RN Trooli (as Call Flow) Exascale Three UK Community Fibre Ogi (as NetSupport) WightFibre Truespeed Airband G.Network Wessex Internet Quickline FibreNest Wildanet Fibrus (as B4B Networks) Zzoomm GoFibre (as Borderlink) Toob Netomnia / YouFibre Pine Media BeFibre Squirrel Internet brsk Lit Fibre (as Broadreach) iDNET * Grain Hey! Broadband Octaplus Every one of the forty has held IPv6 space for at least three years, and most for over a decade. 17 of them still do not give it to a customer. Forty UK providers, ordered by the date the registry gave them IPv6. Each bar runs from that allocation to today. The blue bars ship it to customers; the pink ones never have. Forty providers. Every one has held IPv6 address space for at least three years, most of them for over a decade — and seventeen of them still do not give it to a customer.\nHyperoptic has held 2a01:4b00::/32 since February 2011. Fifteen years of building fibre into blocks of flats, selling gigabit connections, and putting people behind carrier-grade NAT with an unused IPv6 allocation on the books. Trooli has had theirs thirteen years. Truespeed and Airband ten.\nAndrews \u0026amp; Arnold is the one to hold the rest against. A small ISP in Bracknell with a fraction of the customers and engineers of anybody else on that list, giving IPv6 to every line since 2002, and one of the outfits behind 6UK. TalkTalk took its allocation three months earlier and still does not ship it to consumers.\nVirgin Media took its block three weeks before Zen took theirs. Zen shipped it, and I have been on the end of one of their /48s ever since. Virgin is still saying \u0026ldquo;when we are ready\u0026rdquo;.\nAnd look at the bottom of the table. Squirrel, brsk, Lit Fibre and Octaplus all got their allocations in the last six years and all ship IPv6, while providers holding space since 2011 do not. Starting late is not the obstacle. Starting at all is.\nTwo names from that survey are not in the chart, and the reason is the same for both. Cuckoo is a retail brand buying wholesale access over Openreach, CityFibre and others, and Freedom Fibre is a wholesale network whose customers come through retail partners. Neither holds its own address space, so IPv6 is somebody else\u0026rsquo;s decision to make for them. iDNET is marked with an asterisk because theirs is a provider-independent assignment rather than an allocation of their own.\nThe One That Settles It If you want the argument reduced to a single company, it is Plusnet.\nBT bought Plusnet in January 2007. Plusnet sits under British Telecommunications\u0026rsquo; own RIPE account, so it has had access to BT\u0026rsquo;s IPv6 allocation since June 2010. BT ships IPv6. EE, the other sister company, ships IPv6. ISPreview noted that BT and Plusnet even use nearly identical customer routers, calling the gap \u0026ldquo;somewhat of a peculiarity\u0026rdquo;.\nPlusnet trialled IPv6 in 2011 and publicly urged the rest of the industry to get on with it. In 2019 it said it would launch in spring 2020. In 2021 it expected \u0026ldquo;to make good progress over the coming year\u0026rdquo;. In November 2023 it ran a three-month trial across two sites in Chesterfield and Sheffield, with about twenty staff and friendly customers on it.\nIn April 2026 its own customer forum was still asking where IPv6 had got to.\nOne parent company. One address allocation. Near-identical hardware. Engineers who work for the same group and can walk down the corridor to the people who already did it. Three brands, and one of them cannot manage in fifteen years what the other two finished.\nWhatever is stopping this, it is not the technology, the money, the kit, or the address space. It is somebody deciding it is not their problem this quarter, fifteen years running.\nVirgin Media, Sixteen Years Of \u0026ldquo;When We Are Ready\u0026rdquo; The other end of the scale deserves naming, because the timeline is a matter of public record and it is remarkable.\nMarch 2010: a customer asks on Virgin Media\u0026rsquo;s own forum when IPv6 is coming. The answer is \u0026ldquo;when we are ready\u0026rdquo;.\nNovember 2016: Virgin tells ISPreview it plans to adopt IPv6 by mid-2017. It does not.\nJune 2018: a consumer trial is reported. December 2018: a third presentation to the UK IPv6 Council, hinting at 2019.\n2021: a statement that they are \u0026ldquo;continuing to plan our IPV6 deployment having tested several solutions and intend to introduce IPV6 for our customers in future.\u0026rdquo;\nFebruary 2024: Virgin Media locks the fourteen-year-old forum thread.\nAugust 2026: still nothing.\nSixteen years. In that time the company was bought, merged with O2, rebuilt its core twice and replaced its entire router estate. At no point did anyone add an address family. Locking the thread is the most honest thing in that list. It is the moment they stopped pretending and started managing the complaint instead of the problem.\nThe Altnets Had No Excuse At All The full-fibre builders were the chance to start clean. New networks, new kit, no legacy, engineers hired this decade. Look at where they sit in the allocation table and most of them took the address space and stopped.\nSo a company that raised institutional money to build a brand-new fibre network dug up the roads, blew fibre to a hundred thousand homes, bought new routers, wrote a new provisioning stack. And put its customers behind a shared address on a network with no IPv6, in 2026, with the address space already sat in its own registry account.\nAnd then several of them charge £5 a month for a static public IPv4.\nThey took away the working thing, declined to ship the free replacement, and turned the resulting breakage into a line item on your bill. There is a word for a business model that manufactures a fault and then sells the fix, and it is not \u0026ldquo;innovation\u0026rdquo;.\nThe Websites Give The Game Away The routing table shows what networks do. The DNS shows what everybody else does. So on 27 August 2026 I put fifty of the UK\u0026rsquo;s best-known sites in a file — central government, the banks, the big retailers, the telcos, transport and a few universities — and asked each one whether it answers on IPv6:\nwhile read -r d; do n=$(dig +short AAAA \u0026#34;$d\u0026#34; | grep -c \u0026#39;:\u0026#39;) printf \u0026#39;%-46s %s\\n\u0026#39; \u0026#34;$d\u0026#34; \u0026#34;$([ \u0026#34;$n\u0026#34; -gt 0 ] \u0026amp;\u0026amp; echo AAAA || echo none)\u0026#34; done \u0026lt; sites.txt | sort -k2 Then I checked every answer against a second resolver, because one recursive server having a bad day is not a finding:\ndig @1.1.1.1 +short AAAA www.tesco.com | grep -c \u0026#39;:\u0026#39; Seventeen out of fifty had an AAAA record. Thirty-three did not, and both resolvers agreed on every one.\nThe ones with no IPv6 include www.bbc.co.uk, www.nhs.uk, www.hmrc.gov.uk, www.hsbc.co.uk, www.barclays.co.uk, www.lloydsbank.com, www.santander.co.uk, www.tesco.com, www.johnlewis.com, www.marksandspencer.com, www.britishairways.com, tfl.gov.uk, monzo.com — a bank founded in 2015, with no legacy anything — and, my favourite, www.sky.com.\nSky. The company that put five million customers on IPv6 and won a prize for it. Its own website does not answer on IPv6.\nNow the bit that proves the thesis beyond argument.\nFifteen of the seventeen that do have IPv6 got it from a supplier, not from themselves. I resolved each one and looked up who owns the address that answered:\nwhois -h whois.radb.net -- \u0026#34;$(dig +short AAAA www.sainsburys.co.uk | grep \u0026#39;:\u0026#39; | head -1)\u0026#34; | grep -i descr www.gov.uk and www.cam.ac.uk answer from Fastly. ico.org.uk, www.parliament.uk, www.ofcom.org.uk, www.asda.com, www.autotrader.co.uk, www.nationalrail.co.uk and www.jisc.ac.uk answer from Cloudflare, which turns IPv6 on for everyone by default. natwest.com and nationwide.co.uk answer from Azure Front Door. www.legalandgeneral.com and www.screwfix.com answer from CloudFront. www.sainsburys.co.uk and www.next.co.uk answer from Akamai.\nTwo did it themselves: Imperial College London, answering from its own address space, and the UK IPv6 Council\u0026rsquo;s own website. A university and the people whose entire purpose is IPv6. That is the list.\nAnd www.tesco.com, www.sky.com and www.nhs.uk also sit on Akamai — the same CDN, the same product — and have no IPv6 at all.\nAkamai has been explicit about this since June 2022: \u0026ldquo;Akamai has enabled IPv4+IPv6 dual-stack as the default for our CDN delivery products for many years, meaning that customers have needed to opt-out for content to be IPv4-only.\u0026rdquo; They added that they made it easy to switch, including through the API.\nSame vendor. Same platform. Default on. One organisation left it alone and one went in and turned it off, or kept an ancient config nobody has read since. Sainsbury\u0026rsquo;s has IPv6 and Tesco does not, and the difference between them is a single configuration flag and somebody\u0026rsquo;s attention.\nAfter that there is no cost argument and no complexity argument left standing. There is only whether anyone was paying attention.\nAcross the whole sample the rule holds: where IPv6 arrives as a supplier\u0026rsquo;s default, Britain has it. Where a British organisation would have had to decide something, it does not. Two sites out of fifty, and one of those was the IPv6 Council.\nWhich Brings Us To The Managed Service Industry The consumer ISPs get the blame for CGNAT, and they have earned it. But the layer doing the most damage is the one that sells expertise: the managed service providers, the integrators, the outsourced network teams, the consultancies who write the low-level design.\nGo back to the list of the biggest UK networks with no IPv6 — the banks, British Airways, PwC, QinetiQ. Those are not scrappy startups. They are outfits that pay a great deal of money for somebody else to run their network, or employ a large team to run it themselves. Every one of those AS numbers has a design document behind it, a change process, an architecture review board, and a supplier with \u0026ldquo;network\u0026rdquo; in its name. Not one of them produced an IPv6 plan.\nThe pattern is the same everywhere you look at it:\nThe template is IPv4. The build standard, the firewall ruleset, the monitoring checks, the IPAM, the runbook, the DR plan, the customer handover pack — all IPv4, written once, cloned for a decade. Adding an address family means editing all of it, and nobody is paid to edit it.\nNobody asked for it. This is the sentence that ends every IPv6 conversation in this industry, and it is a confession. Nobody asked for TLS 1.3 either. Nobody asked you to stop using SMBv1. Customers buy the outcome and pay you to know what the outcome requires. \u0026ldquo;The customer did not ask\u0026rdquo; means \u0026ldquo;I do not want to learn it and they cannot tell\u0026rdquo;.\nRFC 1918 feels infinite. Ten-dot is 16.7 million addresses, so an internal network never feels short, so there is never a forcing event. Then the merger arrives, both estates are on 10.0.0.0/8, and the answer is another decade of overlapping-subnet NAT and a document explaining which fake address means which real one — more machinery, again, to avoid the address family that would have made it a non-problem.\nIPv6 exposes competence. This is the real one. Dual-stack does not let you hide.\nYou have to know what your firewall policy actually is, because you have to write it twice. You have to know what your DNS looks like. You have to understand neighbour discovery, prefix delegation, and what your CPE does with a /56.\nAn engineer who has been getting by on NAT as an accidental security control finds out, in front of people, that it never was one. There are twenty-year careers in this industry built on that single confusion.\nSo it does not get proposed. Not because it costs money — it does not — but because proposing it means owning it, and owning it means learning it.\nThat is what I mean by bone idle. Not lazy in the sense of not working hard. This industry works extremely hard. It works hard at carrier-grade NAT, and at explaining to a customer why their CCTV will not connect from outside any more. It will do any amount of work, as long as the work is the kind you can buy rather than the kind you have to understand.\nAnd that is about to stop being a matter of taste. The Cyber Security and Resilience Bill now going through Parliament would amend the NIS Regulations to pull in, among others, \u0026ldquo;managed service providers (organisations that provide third-party IT services to other businesses)\u0026rdquo;. It is not law yet. When it is, detection, logging and incident reporting stop being product lines this layer sells and become duties it has to discharge.\nSit that next to the chain further up. Resolving any abuse report starts with a far-end server having logged a source port, and the far-end server is very often one of these outfits\u0026rsquo; boxes. The people who will shortly have to prove they can detect and report an incident are the same people who cannot presently be got to enable an address family, or to read a packet capture when a tunnel will not come up.\nThe Three Excuses You will hear the same three every time, and none of them lasts a minute.\n\u0026ldquo;Dual-stack is two of everything.\u0026rdquo; I ran it in production at Nominet, on F5 load balancers in front of the .uk registry, so I know what the objection is worth. So is the NAT tier you bought instead, and that one sits in the traffic path with a session table, a capacity model, a failover story and a logging obligation attached to it. You were never choosing between complexity and simplicity. You picked the complexity that came with an invoice.\n\u0026ldquo;The kit does not support it.\u0026rdquo; In 2006, fair enough. In 2026 it means your kit is out of support, which is a worse thing to admit than the one you were trying to avoid saying.\n\u0026ldquo;There is no revenue in it.\u0026rdquo; There is no revenue in backups either.\nThe Things People Post The excuses above are what you hear in a meeting. Underneath them sits a layer of technical claims that get repeated in forum threads, comment sections and LinkedIn replies every single time IPv6 comes up, and most of them have been wrong for over a decade.\nSome of this is honest confusion and some of it is a person who has decided not to learn something reaching for a reason. Either way it is worth going through, because these claims are doing real work. They are what an engineer repeats to a manager who cannot check them.\n\u0026ldquo;NAT is my firewall. IPv6 puts every device straight on the internet.\u0026rdquo;\nThis is the big one and it is the wrong way round. The protection people credit to NAT comes from there being no mapping until something inside asks for one — which is a stateful firewall, and it is the firewall doing the work, not the translation. The IETF said so in RFC 4864 in 2007: that role, \u0026ldquo;often marketed as a firewall, is really an arbitrary artifact\u0026rdquo;, where a real firewall gives you \u0026ldquo;explicit and more comprehensive management controls.\u0026rdquo;\nEvery consumer IPv6 router ships with default-deny inbound. You get the same posture, from a policy somebody wrote down, rather than from a side effect of running out of addresses. And you can then permit exactly the one thing you meant to permit, instead of the port-forwarding séance.\nIf your entire security model is \u0026ldquo;attackers cannot find my devices\u0026rdquo;, you did not have a security model. You had NAT.\n\u0026ldquo;IPv6 is slower.\u0026rdquo;\nThis one deserves an honest answer rather than a dismissal, because the truth is mixed and the people making it are not simply wrong.\nContent providers who have optimised for it measure gains: Facebook reported page loads around 15% faster over IPv6, Akamai around 5% on mobile. APNIC\u0026rsquo;s broader measurement of the whole internet is less flattering and has IPv6 round-trip times running marginally higher on average — of the order of a millisecond or so, and improving over time.\nSo: broadly a wash, better where somebody has done the work, occasionally a hair worse where nobody has.\nIt is worth knowing where the cost actually sits, because the header is the thing people picture and the header is not the problem.\nThe IPv4 and IPv6 headers, and what each costs a router The two headers, and what each one costs a router Fields drawn to scale across 32 bits. Sources: RFC 791 (IPv4) and RFC 8200 (IPv6). work a router must do per hop wider lookup key IPv4 20 bytes, up to 60 with options 031 Version IHL Type of Service Total Length Identification Flags Fragment Offset Time to Live Protocol Header Checksum Source Address Destination Address Options — variable length, 0 to 40 more bytes IPv6 40 bytes, fixed. Always. 031 Version Traffic Class Flow Label Payload Length Next Header Hop Limit Source Address (128 bits) Destination Address (128 bits) IPv4 makes a router\u0026#160;recompute the header checksum on every hop, read a length field before it knows where the payload starts, and carry a fragmentation path. IPv6 drops all three — and asks it to match on\u0026#160;a key four times as wide. The two headers side by side, fields drawn to scale across 32 bits. Shaded orange is work a router has to do on every hop; shaded blue is the wider lookup key. Start with what IPv6 took off the router. IPv4 carries a header checksum. The Time to Live changes every hop, so the checksum has to change with it, and RFC 6583 lists \u0026ldquo;verifying and updating the checksum\u0026rdquo; as a step in the forwarding process itself. IPv6 has none. That step just goes.\nThen the length. An IPv4 header is variable, which is what IHL is for: a router reads a length before it knows where the payload starts. An IPv6 header is 40 bytes. Always. Every field sits at a fixed offset and nothing has to be worked out first.\nThen fragmentation. IPv4 routers can fragment in flight, which is why Identification, Flags and Fragment Offset are sat in the header at all. RFC 8200 is flat about it: \u0026ldquo;fragmentation in IPv6 is performed only by source nodes, not by routers along a packet\u0026rsquo;s delivery path\u0026rdquo;. So that path goes as well.\nOn header handling alone IPv6 is the cheaper protocol to forward. It was built to be.\nThe one place it does cost more is per route, and that turns out not to matter. The lookup key went from 32 bits to 128, so an IPv6 forwarding entry is wider and on a lot of kit takes two hardware slots where an IPv4 route takes one. Everybody stops the argument there. It is worth going one step further, because the full table is four times smaller.\nOn the RIS dump generated at 02:03 UTC on 28 August 2026 there were 1,229,166 IPv4 prefixes in the global routing table and 300,470 IPv6 ones. Four times as many IPv4 routes, each a quarter of the width. So the raw key storage comes out a dead heat: 4.92 MB against 4.81 MB. Now apply the two-slots-per-IPv6-route rule people worry about. A full IPv6 table still needs about half the hardware entries of a full IPv4 one.\nAnd the reason the IPv4 table is that big is the shortage itself. 767,543 of those 1.2 million routes are /24s — 62% of the entire IPv4 internet sitting at the longest prefix anybody will accept, because blocks got chopped up, sold off and announced in pieces by whoever bought them. Every one of those is a router somewhere holding an entry it would not need if the space had not run out.\nThe IPv4 routing table is four times the size of the IPv6 one The IPv4 routing table is four times the size — and most of it is the shortage Distinct prefixes in the global routing table, RIPE RIS dump generated 02:03 UTC, 28 August 2026. IPv432-bit key 767,543 of them are /24s 1,229,166 IPv6128-bit key 300,470 Four times as many IPv4 routes, each a quarter of the width — so raw key storage is a dead heat, 4.92 MB against 4.81 MB. Count an IPv6 route as two hardware entries, the way people worry about, and a full IPv6 table still needs about half. 62% of the IPv4 table is /24s\u0026#160;— blocks chopped up, sold off and announced in pieces because the space ran out. Every one is an entry a router would not be holding if it had not. Distinct prefixes in the global routing table on 28 August 2026, counted from the RIPE RIS dump. The solid part of the IPv4 bar is the /24s — deaggregation the address shortage forced. So the memory argument runs the other way round from how it gets told in meetings. Carrying IPv6 is cheaper on your FIB than carrying IPv4, and it gets cheaper every year the transfer market slices another /16 into sixteen /24s.\nAnd the CPU spikes people actually hit are neither of those. A modern router forwards both families in silicon at line rate. What hurts is anything that shoves a packet off that path into the control plane, which RFC 6583 calls \u0026ldquo;a \u0026lsquo;slower\u0026rsquo; software process running on a general purpose processor\u0026rdquo;. That processor was sized for routing protocols. Never for traffic.\nTwo things shove packets at it. The first is extension headers. They are a chain rather than a fixed block, so a box wanting the layer-4 ports for an ACL or an ECMP hash has to walk a variable-length list to find them, and a Hop-by-Hop Options header \u0026ldquo;may be examined or processed by any node along a packet\u0026rsquo;s delivery path\u0026rdquo;. On plenty of kit that means punted.\nThe second is Neighbour Discovery. A /64 covers trillions of addresses that will never be assigned, so scanning one sets a router resolving addresses that do not exist. RFC 6583 exists for that, and calls it a denial of service.\nBoth have known answers. Filter Hop-by-Hop at the edge, rate-limit ND, cap the neighbour cache. Neither is a reason the protocol is slow. They are reasons an unconfigured router is slow, and that points where the rest of this post points.\nNow weigh a millisecond against the alternative you actually deployed — a stateful translation box in the path of every connection, holding a session table, that breaks some of them outright. Nobody stalled for twenty years over a millisecond.\n\u0026ldquo;There is plenty of IPv4 about, you can just buy it.\u0026rdquo;\nYou can. That is what a shortage looks like. Blocks handed out for nothing in 1990 now change hands at around $20 an address, and AWS bills $43.80 a year for every one you use.\nA market in a thing does not prove there is plenty of it. It proves somebody worked out how to charge you for the shortage.\nNobody Was Ever Going To Make Them The UK has no policy on this at all, and it never has.\nThere was an attempt. 6UK was set up in 2010 with £20,000 of seed money from the Department for Business, Innovation and Skills, backed by Vint Cerf, with LINX, AAISP, Timico and Easynet behind it. In December 2012 its volunteer directors resigned at the AGM, nobody stood for the board, and it was wound up. Its parting verdict: free-market incentives are insufficient, \u0026ldquo;one factor appears to dominate IPv6 adoption rates, namely government support,\u0026rdquo; and \u0026ldquo;countries with hands-off governments fall behind.\u0026rdquo;\nFourteen years on, that is exactly what happened. The UK IPv6 Council is still going, but a forum is not a lever.\nCompare the United States, where OMB memorandum M-21-07 required 80% of federal IP-enabled assets to be IPv6-only by the end of the 2025 financial year. Agencies missed it. They still had a number to miss, a date to miss it by, and somebody who has to stand up and explain the miss. Here there is nothing to miss, so nobody has ever had to explain anything.\nThe UK government does not require IPv6 in its own procurement in any meaningful way. Ofcom does not measure it. No regulator asks about it. And so, predictably, www.nhs.uk and www.hmrc.gov.uk do not have it, while www.gov.uk does. And www.gov.uk only has it because it is served through Fastly, which turned IPv6 on years ago on somebody else\u0026rsquo;s behalf.\nI have written before about what happens when nobody holds a contract over an industry — the mechanisms that work turn out to be the ones somebody with resources chooses to operate, and if nobody does, nothing happens for a decade. IPv6 in the UK is that pattern again, without even a members\u0026rsquo; vote at the end of it.\nAnd It Is Not That Nobody Noticed The absence of a requirement would be disappointing if this had gone unspotted. It did not.\nThe state worked out what carrier-grade NAT does and legislated about it. Parliament looked straight at the problem, understood it well enough to write law about it, and wrote the law that accommodates the breakage. \u0026ldquo;Require them to deploy the protocol that removes it\u0026rdquo; was either never raised or was raised and dropped.\nSo we are a country whose declared position is that IP attribution matters enough for counter-terrorism legislation, and which asks not one provider to do the free thing that restores it. Fraud, account takeover, harassment, threats to kill, child abuse referrals and terrorism all arrive at the access network asking the same question, and for a lot of British connections the honest answer is \u0026ldquo;one of these several hundred households\u0026rdquo;.\nThe National Cyber Security Centre is part of GCHQ and publishes guidance on a great many things. It does not require IPv6 of anybody. Nobody in Britain does.\nwww.ncsc.gov.uk and www.gchq.gov.uk do both answer on IPv6, mind. So does the Internet Watch Foundation. All three because they sit behind Cloudflare, which turned it on for everyone by default. www.police.uk has none.\nAnd What That Means For The Seventeen Seventeen of the forty providers in this post sell connections with no IPv6, and about half the full-fibre builders put customers behind carrier-grade NAT.\nI am not accusing any of them of a crime, and nobody in those buildings is hoping for one.\nBut a connection behind carrier-grade NAT with no IPv6 cannot be resolved to a subscriber. That is not contested. The IETF wrote it down in 2011, Europol in 2016 and 2017, Parliament in 2015 — all of it published before most of this equipment was bought. The alternative was free, and available the whole time.\nThat is what \u0026ldquo;we will get to IPv6 eventually\u0026rdquo; means, once you follow it down.\nWhat To Do About It, Concretely Short, because none of it is hard. That is the point of the whole post.\nIf you buy connectivity: put IPv6 in the tender as a pass/fail requirement, not a nice-to-have. Ask for native dual-stack and a delegated prefix, in writing, and ask what size.\nIf the answer is a single /64, keep asking. RFC 6177 killed that one in 2011: handing a home site one /64 \u0026ldquo;precludes the expectation that even home sites will grow to support multiple subnets\u0026rdquo;, and it is \u0026ldquo;strongly intended that even home sites be given multiple subnets worth of space, by default\u0026rdquo;.\nWhat it did not do is name a size. It withdrew the old blanket /48, said the choice \u0026ldquo;is an issue for the operational community\u0026rdquo;, and dropped one worked example on the way past: a home default \u0026ldquo;of less than /48, such as a /56\u0026rdquo;.\nThe operators answered it themselves. RIPE-690 is their own practice document and it is blunt. A /48 each if you want a simple plan. A /48 for business and a /56 for residential if you want a pragmatic one. Anything longer than a /56 is \u0026ldquo;strongly discouraged\u0026rdquo;, and a /64 does not conform to IPv6 standards and will break customer LANs.\nSo the floor is a /56, and it is the IETF\u0026rsquo;s own number rather than anybody\u0026rsquo;s preference. Sky has handed every subscriber one since 2016. Zen hands out a /48, which is 65,536. A business should not accept less than a /48.\nIf a provider tells you a /48 to a house is extravagant, a British ISP has been doing it for years while they were still working out their position.\nIf you run an AS number: you probably already hold a /29 you have never announced. Check.\nAS=AS20712 # your AS number ORG=$(whois -h whois.ripe.net \u0026#34;$AS\u0026#34; | awk \u0026#39;/^org:/{print $2; exit}\u0026#39;) # what IPv6 the registry has already given you whois -h whois.ripe.net -- \u0026#34;-i org $ORG\u0026#34; | grep -i \u0026#39;^inet6num\u0026#39; # what you are actually announcing of it whois -h whois.ripe.net -- \u0026#34;-i origin $AS\u0026#34; | grep -i \u0026#39;^route6\u0026#39; If the first command prints a prefix and the second prints nothing, you are one of the 463.\nAnnounce it, dual-stack your border and one internal VLAN, and put a AAAA on one public service. That is a fortnight\u0026rsquo;s work for one engineer and it turns your organisation from a statistic in the table above into one that has started.\nIf you run a website: check for an AAAA record. If you are behind a CDN, it is probably a toggle you can turn on this afternoon at no cost. If it is off, somebody turned it off.\nIf you sell managed services: write IPv6 into the build standard and the low-level design template, once, and every customer after that gets it by default. Nobody has to ask for it, because nobody asks for TLS either.\nIf you are an engineer who has never done it: build a lab tonight. About 90% of my own traffic runs over native IPv6 and it is the least eventful thing about my network. Get a tunnel or a VPS with a /64, put addresses on things, break it, fix it. It takes an evening to stop being frightening and it is the single cheapest thing you can do to your career this year.\nThe Questions I Cannot Settle Everything above I can show you. This part is the bit I keep turning over, and I do not have a clean answer to any of it.\nWhy can some and not others? This is the one that matters, and the data makes it stranger rather than clearer.\nSky and Virgin Media sold broadband to the same country, over the same regulator, at the same time. One finished in 2016. The other has been saying \u0026ldquo;when we are ready\u0026rdquo; since 2010. Sainsbury\u0026rsquo;s and Tesco sit on the same CDN, on a product where dual-stack is the default, and one has IPv6 and one does not. Norway and Britain buy from the same vendors and 66.9% of Norwegian networks carry IPv6 against 42.3% of ours.\nEvery external factor you might blame is held constant in those pairs. Same country, same suppliers, same kit, same customers, same decade, same regulator, same money. And the outcomes are opposite.\nSo the cause is not in the circumstances. It is inside the building. Somewhere in Sky there was a person who made this their business and kept making it their business for three years. In the other places there was not, or there was and nobody above them cared. That is the whole variable, and it is not a technical one.\nWhich is an uncomfortable answer, because you cannot procure it, and you cannot put it in a strategy document.\nIs it training? Partly, and less than you would think.\nLook again at the 463 outfits holding IPv6 they have never announced. Somebody in each of those buildings knew enough to know they needed it, knew who to ask, filled the form in, and got it issued. The knowledge was there and the follow-through was not.\nTraining gets an engineer to the point of being able. It does not get them to the point of being made to. Nobody has ever had a bad appraisal for not deploying IPv6. Nobody has ever lost a contract over it. Until one of those is true, the training goes on the pile with everything else somebody learned on a course and never used.\nHow long until we are all on it? I can put a number on this one, and it is worse than I expected.\nGoogle has measured the share of its own visitors arriving over IPv6 since 2008. Taking mid-August each year, so it is like for like:\nWorld IPv6 adoption is decelerating short of halfway World IPv6 adoption is slowing down, not speeding up Share of Google's own visitors arriving over native IPv6, mid-August each year. Source: Google IPv6 statistics. 0% 10% 20% 30% 40% 50% 2017 18.1% 2018 2019 2020 2021 2022 2023 2024 2025 2026 48.1% Points added that year +3.6 +5.0 +4.4 +3.3 +4.3 +3.3 +2.3 +2.6 +1.1 This year added\u0026#160;1.1 points, the smallest gain in a decade, against 6.6 points in the year to August 2017. Google\u0026rsquo;s measurement of its own visitors arriving over native IPv6, taken mid-August each year so it is like for like. The line is the level. The bars are what each year added. We are not accelerating towards the finish. We are decelerating short of halfway. This year added 1.1 points, the smallest gain in a decade, against 6.6 points in the year to August 2017.\nStraight-line it from the last three years and the world reaches 100% in 2049. Straight-line it from this year alone and it is 2073. Neither is a forecast. A curve that is flattening does not reach the top by drift at all. It stalls somewhere in the sixties and the remainder never moves, because the networks that have not done it by then are the ones nothing was ever going to move.\nI would love to be wrong about that. The number has got smaller every year I have looked at it.\nDo we need a law? This is the one I have gone back and forth on most, and I have landed on \u0026ldquo;not the law people reach for\u0026rdquo;.\nAgainst a mandate: the Americans passed the strongest one anybody has, and missed it. A deadline is not a deployment.\nFor one: 6UK\u0026rsquo;s parting verdict in 2012 was that free-market incentives are insufficient and that the countries which fall behind are the ones with hands-off governments. Fourteen years of British data agrees with them.\nAnd here is the part that settles it. This country has already legislated about this problem — it just legislated in the wrong direction. We were willing to legislate to accommodate address sharing. We have never been willing to legislate to remove it.\nWhat I would ask for is not a ban and not a target, but the Belgian instrument described earlier: a hard limit on how many subscribers may share one address. It needs no new addresses, it does not shut anybody out of the market, and it works on the economics rather than on anybody\u0026rsquo;s good intentions.\nA mandate on its own produces one useful thing, and it is not deployment. It is a named person who has to explain the miss. We have never had one of those.\nHow did we move to IPv4 so quickly, then? Because somebody could switch the old one off.\nThe comparison is exact, and almost nobody makes it. The ARPANET ran the Network Control Program, which addressed hosts in 8 bits — 6 for the node and 2 for the host, so 64 nodes of 4 machines, 256 hosts in total. By the late 1970s that was obviously not enough, and the answer was a new protocol with a bigger address. Same problem we have now, forty-odd years earlier.\nJon Postel published the transition plan in November 1981. In March 1982 the US Department of Defense declared TCP/IP its official standard. Both protocols ran side by side, and on 1 January 1983 NCP was turned off. Hosts that had not converted lost access to the network. Vint Cerf remembers \u0026ldquo;I survived the TCP/IP switchover\u0026rdquo; pins being worn afterwards by the people who got through it.\nFourteen months from plan to flag day.\nNow count what made that possible. Roughly a couple of hundred hosts, not four billion. One network, not every network. One funder who owned every machine on it and paid the wages of everyone touching them. A single organisation able to set a date, and — this is the bit that matters — able to make the old protocol stop working on that date.\nNone of those things exist now, and that is the whole answer. IPv4 did not win because the migration was easy. It won because there was somebody in a position to end the argument.\nNobody is in that position today. There is no authority that can switch IPv4 off, and there never will be. Which means this transition cannot be finished the way the last one was — it can only be finished by several thousand outfits each deciding, on their own, to bother.\nOn this year\u0026rsquo;s numbers, that lands in 2073. Ninety years after the flag day.\nWe Used To Make Things Properly A standard is something you keep when nobody is checking. That is the whole of it. There is no inspector for this, no certificate, no auditor who turns up and asks to see your routing table, and twenty years have now shown exactly what this country does with an obligation nobody enforces.\nWe drop it, and then we buy something to cover the gap.\nNone of that is a technical failure and I will not pretend it is. Britain can do this work. The skills are here, the kit is here, the address space is issued and sat waiting in accounts we are already paying for. What has gone is the instinct to do a job properly because it is the job — without being paid extra for it, and without somebody stood over you making you.\nAsk what actually stops it and you land on the money, but not in the way people mean.\nA carrier-grade NAT has a purchase order. It has a vendor, a quote, a discount, a support contract and a renewal date. It goes in the capital plan, it depreciates over five years, and somebody\u0026rsquo;s name sits on the business case. Delivering it is a visible thing a manager can point at in an appraisal.\nIPv6 has none of that. No invoice, no supplier, no renewal, nothing to put in a budget and nothing anybody can be seen to have bought. It is just work, done properly, by people who know what they are doing, for no return this quarter. Nobody in this industry has ever been promoted for a thing that never appeared in a budget.\nAs such the cheaper option loses, every year, for twenty years. Not because anyone weighed it up and chose wrong, but because a management culture has grown up here that can only see the parts of engineering that arrive with a price on them. Cheap was never the obstacle. Unbillable was. That is what it looks like when a business stops caring about standards and starts caring only about what it can put on an invoice, and it is a choice being made by people paid well enough to know better.\nThen there is what they sold us instead of an address, which ought to make people angrier than it does.\nThe internet was built so that any machine could reach any other machine directly. Not a detail of the design. The design. It is why anybody with a connection and an idea could put something up that the whole world could reach. Put a customer behind a shared address and that is gone. You can ask, but you can never answer. You are a consumer of other people\u0026rsquo;s services, permanently, and never a provider of your own.\nThat is not an unfortunate side effect of a shortage. It is a re-architecture, and it suits everybody selling it. The camera that now needs the manufacturer\u0026rsquo;s cloud. The remote access that now needs somebody\u0026rsquo;s relay. The thing people used to run at home that is now a monthly subscription. Every one of those is somebody who owned a thing being converted into somebody who rents it, and a connection quietly downgraded from a place on the internet to a window onto somebody else\u0026rsquo;s.\nWe gave the middle of the internet away to a handful of companies on another continent, then stood about looking surprised that it ended up centralised. You cannot be self-reliant on a connection that will not let you host owt.\nSelf-reliance is the part I keep coming back to, because this country has stopped expecting it of itself. The instinct now is to wait. For a vendor, a regulator, a grant, a mandate, a customer who rings up and asks. None of those are coming. There is no market signal on its way, no policy in drafting, no deadline anybody will have to explain missing.\nWhich leaves it where it has been the whole time. A free allocation, sat in a registry account with your company\u0026rsquo;s name on it, and a fortnight between you and having done the job right.\nNobody is coming to make you. That is exactly why it counts.\nYou can check whether it is your building. The script is in the download at the top of the post.\nSources Everything below was retrieved on 27 August 2026.\nThe data I measured from. Every number of mine comes from these. They are free, they are public, and you can repeat the whole thing in an afternoon.\nRegistry delegation files, one per regional registry: RIPE NCC, APNIC, ARIN, LACNIC, AFRINIC. RIPE\u0026rsquo;s was generated 26 August 2026, the rest 27 August 2026. RIPE RIS routing table dumps, riswhoisdump.IPv4.gz and riswhoisdump.IPv6.gz, generated 18:06 UTC on 27 August 2026. RIPE database REST interface, for the organisation behind each AS number. The public DNS, for the AAAA sweep, cross-checked against a second resolver. Other people\u0026rsquo;s measurements.\nGoogle IPv6 statistics — per-country native IPv6 among Google\u0026rsquo;s own visitors, figures as of 25 August 2026. APNIC on IPv6 performance and on IPv6 security misconceptions — the measured picture rather than the forum one. IPv4 transfer market prices, first half of 2026, summarising CircleID\u0026rsquo;s analysis of publicly priced transactions. Standards. The bolt-ons, in the order they were published.\nRFC 3056 — 6to4 automatic tunnelling, February 2001. RFC 3701 — the 6bone phaseout plan, setting its shutdown for 6 June 2006. RFC 7526 — deprecating the 6to4 anycast relays and moving them to Historic, May 2015. RFC 4864 — what NAT does and does not give you, and why the firewall people think they have is \u0026ldquo;an arbitrary artifact\u0026rdquo;, 2007. RFC 4941, RFC 7217 and RFC 8981 — temporary and opaque addresses, which is why a device\u0026rsquo;s IPv6 address is not a stable identifier. RFC 801 — Jon Postel\u0026rsquo;s NCP/TCP transition plan, November 1981, setting the 1 January 1983 flag day. RFC 1883 — the original IPv6 specification, December 1995. RFC 6333 — DS-Lite, 2011. RFC 6056 — source port randomisation, the defence CGN weakens, 2011. RFC 6269 — Issues with IP Address Sharing, June 2011. The IETF\u0026rsquo;s own catalogue of what CGN breaks, including abuse logging, penalty boxes, blacklisting, port randomisation and traceability. RFC 791 and RFC 8200 — the two header formats, and why IPv6 dropped the checksum, the variable length and in-flight fragmentation. RFC 6583 — Neighbour Discovery cache exhaustion on a /64, and the forwarding-plane versus control-plane split that decides what costs a router CPU. RFC 6177 — how much address space an end site should get, 2011. Obsoletes RFC 3177\u0026rsquo;s blanket /48, rules out the single /64, and hands the actual number to the operational community. RIPE-690 — the European operators\u0026rsquo; own answer to that question, October 2017: /48 or /56 to an end user, never a /64. RFC 6302 — log the source port, timestamp and protocol, 2011. RFC 6598 — shared address space, 2012. RFC 6877 — 464XLAT, 2013. RFC 6888 — carrier-grade NAT requirements, 2013. RFC 7021 — the impact of carrier-grade NAT on applications, 2013. RFC 7422 — deterministic mapping to cut CGN logging, 2014. RFC 7597 and RFC 7599 — MAP-E and MAP-T, 2015. World IPv6 Launch, 6 June 2012. Hurricane Electric\u0026rsquo;s free tunnel broker — where a lot of us got IPv6 while our own ISPs had none. Law and policy.\nCyber Security and Resilience (Network and Information Systems) Bill 2024-26 — House of Commons Library briefing on the bill that would bring managed service providers inside the NIS Regulations. Counter-Terrorism and Security Act 2015, section 21 — retention of relevant internet data, and its explanatory notes. Europol, October 2017 — law enforcement calling for the end of carrier-grade NAT, with the 90% mobile and 50% fixed figures. OMB memorandum M-21-07 — the US federal IPv6-only requirement, November 2020. Online Safety Act 2023. Europol EC3, Carrier Grade NAT and crime attribution online — Gregory Mounier\u0026rsquo;s presentation to RIPE 74, with the August 2016 survey of EU law enforcement, the case examples, and the Belgian code of conduct and its results. NCMEC CyberTipline data — 2025 report volumes and referrals. A Multi-perspective Analysis of Carrier-Grade NAT Deployment, ACM IMC 2016 — the independent measurement of CGN use by mobile and fixed providers. RIPE NCC charging scheme 2026 — EUR 1,800 per LIR account, flat. Vendors, in their own words.\nAkamai, June 2022 — dual-stack is the default and customers have to opt out of it. AWS, 2023 — the public IPv4 charge, and why. Reporting and the record.\nSky\u0026rsquo;s own write-up on RIPE Labs — how five million users were moved. ISPreview, September 2016 — Sky completing the rollout. CircleID, September 2016 — the Jim Bound IPv6 Award. ISPreview, December 2012 — 6UK winding itself up. ISPreview altnet IPv6 and CGNAT survey — April 2024, updated to March 2025. ISPreview on Plusnet\u0026rsquo;s IPv6 trial, November 2023 — the 2011 trial, the missed 2020 launch, and the note that BT and Plusnet ship near-identical routers. A community-maintained tracker of UK mobile networks and IPv6 — user-reported rather than official, and the source for which mobile networks hand out IPv6 today. havevirginmediaenabledipv6yet.co.uk — the Virgin Media timeline, 2010 to now. Internet Society, September 2016 — Sky at 90% of its base, each subscriber getting a /56. Internet Society, September 2016 — Ron Broersma\u0026rsquo;s account of the 1983 NCP to TCP/IP migration, including the 256-host limit and what happened to anyone who missed the deadline. The Register, January 2013 — thirty years on from the flag day, and the pins people wore afterwards. UK IPv6 Council. ","permalink":"https://blogs.damiendye.uk/en/networking/we-never-ran-out-of-addresses/","summary":"IPv6 has been finished, free and switched on by default in every operating system for the best part of twenty years. The UK\u0026rsquo;s answer was carrier-grade NAT, a law about logging, and a £5 a month charge to give you back the address you used to have. I counted every UK network in the global routing table to find out who has actually turned IPv6 on — 1,200 of them have not, and 463 of those are sitting on address space they asked for and never used.","title":"We Never Ran Out of Addresses. We Ran Out of Effort."},{"content":"Disclosure up front: I work for croit, which sells Ceph and appears in the tables below. I have tried to be as critical about us as about everyone else. Every number comes from the public git history, Ceph\u0026rsquo;s own in-tree records, or a named public statement — all reproducible, all listed in the references at the end.\nCeph holds up a lot of kit that nobody thinks about. Proxmox clusters, OpenStack, Kubernetes, national research labs, banks, telcos, particle accelerators. Twenty years on, it is still the first answer when somebody wants block, file and object storage out of one cluster built from ordinary hardware.\nSo who writes it, who maintains it, and who runs it?\nNot \u0026ldquo;who is on the mailing list\u0026rdquo; or \u0026ldquo;who spoke at Cephalocon\u0026rdquo;. I cloned ceph/ceph, took every non-merge commit authored in the ten years to 2026-08-27, and mapped the authors onto companies using Ceph\u0026rsquo;s own .organizationmap plus a documented set of corrections. Then I took Ceph\u0026rsquo;s in-tree governance file and traced all 38 Steering Committee members to the employer in their own published address. Then I went looking for who says publicly that they run it.\nThat comes to 69,613 commits from 1,718 people across 478 organisations. It is a better picture than I expected going in, and not the one I set out to write.\nThe Decade Rank Organisation Commits Share 1 Red Hat 42,354 60.8% 2 SUSE 6,191 8.9% 3 IBM 4,308 6.2% 4 Intel 2,066 3.0% 5 ZTE 1,176 1.7% 6 QiAnXin 900 1.3% 7 Ceph Foundation 691 1.0% 8 Mirantis 566 0.8% 9 IONOS 498 0.7% 10 China Mobile 397 0.6% 11 Proxmox 254 0.4% 12 croit 242 0.3% 13 Cloudbase Solutions 239 0.3% 14 Bloomberg 231 0.3% 15 XSKY 225 0.3% — Clyso 181, Huawei 159, Inspur 155, Cafe Bazaar 130, CERN 121, SK Telecom 120, UMCloud 106, Deutsche Telekom 105, EasyStack 103, ISCAS 99, Tencent 83, and 460 more organisations — No employer visible in the address 6,295 9.0% Rows 1 and 3 are the same team. On 4 October 2022 Red Hat and IBM announced that Red Hat\u0026rsquo;s entire Ceph team was moving to IBM, with IBM taking over Red Hat\u0026rsquo;s Foundation sponsorship and helping fund the upstream test lab. The people did not change; their email addresses are still migrating, one engineer at a time, four years later.\nAdd them: 46,662 commits. 67.0% of the decade.\nSit with that before anything else, because it is the deal that makes Ceph exist. One company has paid dozens of engineers to build and look after distributed storage it then gives away. Ten years of it, through two takeovers and a change of parent. Nobody made them do it.\nThe Last Three Years Rank Organisation Commits Share 1 Red Hat 7,273 45.0% 2 IBM 3,818 23.6% 3 QiAnXin 597 3.7% 4 Ceph Foundation 526 3.3% 5 IONOS 498 3.1% 6 Intel 279 1.7% 7 Proxmox 249 1.5% 8 Bloomberg 220 1.4% 9 croit 157 1.0% 10 Clyso 153 0.9% 11 Cafe Bazaar 118 0.7% 12 ISCAS (Institute of Software, Chinese Academy of Sciences) 99 0.6% — No employer visible in the address 2,008 12.4% 16,160 commits. Red Hat plus IBM: 11,091, or 68.6%.\nOne company's share of Ceph, over a decade and over three years The share does not move. What is inside it does. Ten years to 2026-08-27 69,613 commits Red Hat\u0026#160;\u0026#160;60.8% IBM 6.2% SUSE 8.9% everyone else\u0026#160;\u0026#160;24.1% Red Hat + IBM: 46,662 commits, 67.0% Last three years 16,160 commits Red Hat\u0026#160;\u0026#160;45.0% IBM\u0026#160;\u0026#160;23.6% SUSE: 13 everyone else\u0026#160;\u0026#160;31.4% Red Hat + IBM: 11,091 commits, 68.6% SUSE was the second-largest contributor of the decade. In the last three years it wrote 13 commits, and in 2025 and 2026 none at all. The one-company share barely moves between the decade and the recent window — 67.0% against 68.6%. What changes is everything around it. The Resilience Nobody Talks About Here is the bit I did not expect, and it is the best thing in the whole exercise.\nCeph has already survived the two events this data would tell you to fear, and it did not flinch.\nIts creator, Sage Weil, wrote 7,517 commits over the decade — more than any organisation except Red Hat, SUSE and IBM. In 2017 alone he wrote 2,184 commits, 21.5% of the entire project by himself. He stepped back in October 2021 after 17 years; his last commit is dated 2022-01-20.\nIts second-largest contributor of the decade, SUSE, wrote 6,191 commits and peaked at 1,839 in one year. Then it cancelled SUSE Enterprise Storage for Rancher\u0026rsquo;s Longhorn and wound down: 463 in 2021, 110 in 2022, 14 in 2023, 10 in 2024, nothing since.\nBoth inside the same five years. If a project this concentrated were brittle, that is when it would have snapped.\nIt did not. The commit rate this year is 14.6 a day against 14.8 last year. Flat. Releases kept coming. The governance was rebuilt into an Executive Council and a 38-person Steering Committee, and it held.\nYear Total commits Red Hat + IBM Share Everyone else 2016 10,286 5,834 56.7% 4,452 2017 10,158 6,942 68.3% 3,216 2018 8,099 5,265 65.0% 2,834 2019 9,000 5,934 65.9% 3,066 2020 8,361 5,230 62.5% 3,131 2021 7,381 4,517 61.1% 2,864 2022 4,731 2,937 62.0% 1,794 2023 4,634 3,315 71.5% 1,319 2024 5,410 3,571 66.0% 1,839 2025 5,389 3,730 69.2% 1,659 2026 3,495 2,329 66.6% 1,166 (2026 runs to 27 August, about eight months.)\nBetween 57 and 72 per cent from one vendor, every year for a decade. Volume is down on the 2016 peak, which is what happens when a project stops rebuilding its foundations and starts looking after them. The last three years are flat, or a shade up.\nCeph commit volume by year — the share holds, the total halves Commits per year, and who wrote them Ceph main branch, non-merge commits, by author date 0 3k 6k 9k 12k 2016 57% 2017 68% 2018 65% 2019 66% 2020 63% 2021 61% 2022 62% 2023 72% 2024 66% 2025 69% 2026* 67% Red Hat + IBM — one team since October 2022 everyone else SUSE peaks: 1,839 Sage Weil leaves SUSE: 110 in 2022, 14 in 2023, 0 by 2026 * 2026 runs to 27 August; plotted at its current daily rate of 14.6 commits, against 14.8 for 2025. Commit volume by year, split between the Red Hat/IBM team and everyone else. Marked: SUSE\u0026rsquo;s peak and exit, and Sage Weil\u0026rsquo;s departure. Ceph absorbed both without a change in cadence. Ceph Outlived the Companies That Built It This is the decade\u0026rsquo;s real story, and you cannot see it in a three-year window at all.\nLook at rows 2, 5, 8, 10, 13 and 15. SUSE, ZTE, Mirantis, China Mobile, Cloudbase Solutions, XSKY — plus Inspur, EasyStack, UMCloud, Kylin, UnitedStack, Xtao, Istuary and more further down. Between them, well over 10,000 commits. Nearly all of it has stopped.\nSUSE: 6,191 commits, peaked 2020, now zero. ZTE: 1,176 commits, 1,071 in 2016 alone, last commit 2020. Mirantis: 566 commits, gone. XSKY, EasyStack, Inspur, UMCloud, Kylin, UnitedStack: the OpenStack-era storage cohort, all wound down. Every one of those was a company betting a product on Ceph. The products got cancelled or pivoted. And Ceph is still here, shipping at the same rate, with their code still in the tree and somebody else looking after it.\nThe best illustration is the dashboard. Over the decade, src/pybind/mgr/dashboard looks like this:\nsrc/pybind/mgr/dashboard, ten years Commits Share SUSE 1,408 36.6% Red Hat 1,177 30.6% IBM 734 19.1% No employer visible 478 12.4% The Ceph dashboard was mostly SUSE\u0026rsquo;s work. SUSE then left the project entirely. In the last three years the same directory is 57.5% IBM and 30.3% Red Hat, and the dashboard is still shipping and still gaining features.\nThat is upstream-first development doing the exact job it is for. A vendor put a lot in, the vendor left, and the users kept the software. Had SUSE built that dashboard as a closed layer bolted on top, the way plenty of storage vendors would have, it would have died with the product line. It went upstream instead, so it lived.\nThe Next Decade Is Being Built by Several Companies at Once Crimson is the ground-up rewrite of the OSD — the daemon that owns your disks — on the Seastar framework, aimed at the thread-per-core model modern NVMe demands. It is the biggest bet on Ceph\u0026rsquo;s next ten years. Over the decade:\nsrc/crimson, ten years Commits Share Red Hat 3,038 52.1% Intel 1,269 21.8% QiAnXin 824 14.1% No employer visible 450 7.7% And over the last three years, Red Hat and IBM together are a minority of it at 45.4%, with QiAnXin on 28.4% and Intel on 13.4%.\nThe most important thing being built in Ceph right now really is a multi-company job. Not one vendor\u0026rsquo;s roadmap with a few contributors bolted on — three outfits doing heavy engineering on the same subsystem, in public, for years.\nQiAnXin deserves a note, because it is not a name most storage people know. It is a Chinese enterprise cybersecurity company, founded as a Qihoo 360 subsidiary in 2014 and spun out around 2016; Qihoo 360 sold its remaining 22.6% stake to firms affiliated with China Electronics Corporation in April 2019, and CEC held 38.3% by the 2020 IPO filing. The corporate history is legible in the git log: their lead contributor, Xuehan Xu, has commits under @360.cn in 2017 and 2018 and @qianxin.com from 2021. A security company with no storage product to sell has put 900 commits into Ceph\u0026rsquo;s future OSD. That is an open project working the way it says on the tin.\nWho Maintains Ceph, and Who Pays Them Writing code is one thing; holding the keys is another. Ceph keeps its governance in the repository, in doc/governance.rst, and the Ceph Steering Committee is listed there by name and email. That makes the maintainer-to-employer question answerable from a primary source instead of guesswork.\nThirty-eight seats, traced to the employer in each member\u0026rsquo;s own listed address:\nEmployer Seats Red Hat 16 IBM 9 Clyso 3 Personal address (Anthony D\u0026rsquo;Atri, Myoungwon Oh) 2 croit — Igor Fedotov 1 Bloomberg — Joseph Mundackal 1 Intel — Yingxin Cheng 1 Ceph Foundation — Zac Dover 1 XSKY — Haomai Wang 1 ZTE — Xie Xingguo 1 Ubiquiti — Yehuda Sadeh 1 11:11 Systems — David Orman 1 Red Hat and IBM hold 25 of 38 seats — 65.8%. Against 67.0% of the decade\u0026rsquo;s commits and 68.6% of the last three years\u0026rsquo;. The governance body mirrors the code almost exactly, which is the healthy way round: the people doing the work have the say, and as such nobody holds a veto they have not earned.\nCeph Steering Committee seats by employer — 25 of 38 are one company Who maintains Ceph, by who pays them 38 seats on the Ceph Steering Committee, from the address each member lists in doc/governance.rst Red Hat IBM Clyso personal address one seat each 16 9 3 2 8 croit, Bloomberg, Intel, Ceph Foundation, XSKY, ZTE, Ubiquiti, 11:11 Systems Red Hat + IBM: 25 of 38 seats, 65.8% Against 67.0% of the decade's commits and 68.6% of the last three years'. The committee mirrors the code. XSKY and ZTE still hold seats. Neither company has committed a line to Ceph since 2020. Five of the 38 have no commit since 2021, or none at all. Steering is not the same job as writing code \u0026#8212; but two of those seats belong to companies that have left the project entirely. The 38 Ceph Steering Committee seats by the employer in each member\u0026rsquo;s listed address, from doc/governance.rst. The committee\u0026rsquo;s composition tracks the commit distribution closely. Two details are worth pulling out, and both say something good.\nXSKY and ZTE still hold seats. Neither company has committed a line since 2020 — Haomai Wang\u0026rsquo;s last commit is 2020-03-18, Xie Xingguo\u0026rsquo;s 2020-07-24 after 750 commits in the decade. Ceph has not turfed them out. A project that keeps a seat warm for the people who built large parts of BlueStore and the OSD, years after their employer walked off, is not one that treats contributors as disposable.\nYehuda Sadeh sits on the committee with an @ui.com address. He wrote RADOS Gateway — the S3 and Swift front door onto RADOS, the Reliable Autonomic Distributed Object Store that everything else in Ceph sits on — starting at DreamHost in 2008, through Inktank, Red Hat and IBM: 972 commits in the decade alone and thousands before it. He was still writing cephx crypto code in July 2025. Then in June 2026 he made exactly one commit, doc: governance/csc: update email address, changing his own entry to Ubiquiti. He changed employer and kept his seat. Your standing travels with you here, and that is one of the better things about working in the open.\nThe Component Leads The component leads own each subsystem day to day:\nComponent What it is Lead Employer Cephadm Cluster deployment and management Adam King Red Hat CephFS The POSIX file system Venky Shankar Red Hat Crimson The next-generation OSD Matan Breizman Red Hat Dashboard The web management interface Afreen Misbah IBM RADOS The object store everything else sits on Radosław Zarzyński Red Hat RBD RADOS Block Device — virtual disks Ilya Dryomov Red Hat RGW RADOS Gateway — the S3 and Swift layer Adam Emerson, Eric Ivancich Red Hat NVMe-oF NVMe over Fabrics gateway Aviv Caro IBM Seastore Crimson\u0026rsquo;s storage backend Yingxin Cheng Intel Ten named leads, nine at Red Hat or IBM, and Seastore — Crimson\u0026rsquo;s storage backend — led from Intel. Above them sits the three-person Executive Council created when Sage Weil left: Dan van der Ster (Clyso), Neha Ojha (Red Hat), Patrick Donnelly (IBM). Clyso holds a third of the top governance body on 0.3% of the decade\u0026rsquo;s commits — the council was built around judgement and standing, not headcount.\nThe Maintainers Moved, and the Code Stayed The strongest argument for Ceph being a real commons rather than one company\u0026rsquo;s product is what happens when its maintainers change jobs. Every one of these is traceable through addresses in the repository:\nMaintainer (current employer) Career, per the git log Last commit Igor Fedotov (croit) — BlueStore Mirantis (2015–17) → SUSE (2017–24) → croit (2021–) 2026-08-24 Kefu Chai (Proxmox) — core, build Red Hat (2015–22, 4,770 commits) → XSKY → Proxmox (2025–) 2026-08-18 Radosław Zarzyński (Red Hat) — RADOS lead Mirantis (2015–17) → Red Hat (2017–) 2026-06-16 Dan van der Ster (Clyso) — Executive Council CERN (2013–22) → Clyso (2023–) 2026-03-18 Zac Dover (Ceph Foundation) — documentation independent → Clyso → Ceph Foundation 2026-07-23 Yehuda Sadeh (Ubiquiti) — RGW author DreamHost → Inktank → Red Hat → IBM → Ubiquiti 2026-06-08 Mark Nelson (Clyso) — performance DreamHost → Inktank → Red Hat → Clyso 2024-04-16 Xuehan Xu (QiAnXin) — Crimson Qihoo 360 (2017–18) → QiAnXin (2021–) 2026 Three employers each, four in some cases, and the work carried on through every move. Igor Fedotov has now outlived two of his employers\u0026rsquo; Ceph strategies and still maintains the engine that puts your bytes on disk. So the honest answer to \u0026ldquo;what if a vendor leaves\u0026rdquo; is this: the engineers keep going.\nWho Actually Runs Ceph Contributions are only half of it. Here is who says out loud that they run Ceph, with the numbers they published themselves.\nOrganisation Publicly stated deployment CERN, the European Organization for Nuclear Research 19 production clusters, ~73 PB raw, plus 5 more in a new datacentre — the storage backbone under CERN\u0026rsquo;s IT cloud Bloomberg Object stores from hundreds of TB to over 8 PB; added 6 PB of raw capacity live, a 50% increase to an online cluster, in under an hour Wikimedia Foundation Five production Ceph clusters — block for Cloud VPS, S3 via multisite RGW, and CephFS for Airflow, Dumps and ML-Lab DigitalOcean Ceph powers its Block Storage service via RBD, with \u0026ldquo;hundreds of enterprise-class SSDs\u0026rdquo; per region and 3× replication across servers and racks OVHcloud \u0026ldquo;Persistent storage for virtual machines is ensured by Ceph RADOS Block Device\u0026rdquo; in its On-Prem Cloud Platform Proxmox Ships Ceph as the built-in hyperconverged storage option in Proxmox VE CERN is worth a closer look, because it is the most detailed public account anybody has published. From a September 2024 CERN IT talk by Enrico Bocchi:\nCERN Ceph, by application Raw size Clusters Blocks — OpenStack Cinder/Glance, HDD 3× replica 25.1 PB 5 Blocks — Flash, EC 4+2 976 TB 2 File system — OpenStack Manila, K8s/OKD, HPC, HDD 3× replica 13.4 PB 5 File system — Flash, 3× replica 1.7 PB 4 Objects — S3, Swift, Backups, HDD EC 4+2 28.2 PB 2 Objects — multi-site, EC 4+2 3.6 PB 1 EC 4+2 is erasure coding, four data chunks to two parity. HPC is high-performance computing, K8s is Kubernetes and OKD is its upstream distribution. The figures are raw capacity, before replication and coding overhead.\nNineteen production clusters, run on the stated principle \u0026ldquo;don\u0026rsquo;t put all your eggs in the same basket\u0026rdquo;, with five more going into a new datacentre. The service history is a quiet advert for the software: 300 TB proof of concept in 2013, 3 PB in production for RBD by the December, 3 PB to 6 PB expanded with no downtime in 2016, S3 and CephFS in production in 2018, an entire CephFS cluster physically relocated with no downtime in 2022, kernel RBD in production in 2023.\nWhat it carries at CERN is the telling part: GitLab, OpenStack, OpenShift, Kubernetes, Harbor, Jenkins, Grafana, Kafka, OpenSearch, InfluxDB, HTCondor, Slurm, Jupyter, Spark, Zenodo, Indico, and the virtualisation of NFS, AFS and CVMFS. Ceph is not a side experiment there. It is the floor the rest of the building stands on.\nAnd the community-wide figure, from the Linux Foundation\u0026rsquo;s own Squid release announcement: 1 exabyte of data across more than 3,000 Ceph clusters.\nThat Exabyte Is a Floor, Not a Total This is the part worth stopping on, because it changes how you read every adoption number about Ceph.\nCeph\u0026rsquo;s telemetry is opt-in. You only get counted if somebody runs ceph telemetry on --license sharing-1-0. Everybody who never ran it is invisible, and in practice that is most people, because it is not the default and nothing nags you about it.\nNow add Proxmox. Proxmox VE ships Ceph as its hyperconverged storage option: three nodes, a few clicks in the web UI, pveceph under the hood, and you have a Ceph cluster. A very large number of people are running Ceph in production without ever thinking of themselves as Ceph users at all. They are Proxmox users. They never joined a mailing list, they will never write a testimonial, they have not turned telemetry on, and they appear in none of the tables above.\nEvery small hosting company, managed service provider, university department, homelab that quietly went into production and three-node office cluster in that bracket is a real Ceph deployment that no figure in this post counts. Same goes for anybody getting Ceph through Rook on Kubernetes, or inside a vendor appliance that never says what is under the lid.\nSo 1 EB across 3,000 clusters is the number from the clusters that put their hand up. The real installed base is a good deal bigger and nobody knows by how much. That is an odd spot for infrastructure software to be in — normally the vendor knows, because you had to buy a licence — and it is a direct result of the thing being free.\n478 Organisations Have Put Code In The commit log doubles as a roster of who runs Ceph at scale, because a company that sends patches is nearly always a company running the thing. Over the decade 478 distinct organisational email domains turn up, 140 of them with five commits or more.\nThe names, grouped, from the git log alone:\nChip, disk and hardware makers: Intel, Samsung, Seagate, SanDisk, Western Digital, Quantum, Mellanox, Lenovo, Fujitsu, Hitachi, Nokia, Arm, Linaro, HiSilicon, Synology, 45Drives Clouds and hosts: DigitalOcean, OVH, IONOS, Akamai, Linode, Hetzner, Binero, City Network, iland, 11:11 Systems, Vexxhost, StackHPC, Canonical, Deutsche Telekom, China Telecom, China Unicom, China Mobile, Chunghwa Telecom Internet and enterprise users: Bloomberg, eBay, GoDaddy, Flipkart, Wikimedia, Naver, LINE, Kakao, SK Telecom, Alibaba, Tencent, Baidu, ByteDance, Kuaishou, UnionPay, SenseTime, Sangfor, Micro Focus, MITRE, Igalia, Walmart Labs Storage vendors and integrators: SUSE, Mirantis, XSKY, EasyStack, Inspur, UMCloud, Kylin, UnitedStack, H3C, Xtao, Eisoo, Cloudin, Istuary, ProphetStor, SoftIron, Bigtera, Cloudbase Solutions, Digiware, Bisect, 42on, croit, Clyso, Proxmox, DreamHost Research and education: CERN, the Institute of Software at the Chinese Academy of Sciences (ISCAS), Pennsylvania State University, Boston University, the University of Michigan, Carnegie Mellon University, plus the Foundation\u0026rsquo;s Associate members — FAS Research Computing at Harvard, the Greek Research and Technology Network (GRNET), Monash University, the South African Radio Astronomy Observatory (SARAO), the Science and Technology Facilities Council (STFC), SWITCH, SLAC at Stanford, and the Center for Research in Open Source Software (CROSS) at UC Santa Cruz Not all of those are current, and that is the point of looking at a decade. It shows the full span of who has leaned on this software hard enough to send patches back, and how wide that spread has been.\nThe Players Who Say They Back Ceph, Against What They Ship The Foundation\u0026rsquo;s tiered membership is where companies declare support. The tiers do not track engineering, and the clearest illustration comes from the three Diamond members quoted in the Linux Foundation\u0026rsquo;s own Ceph Squid release announcement.\nDiamond member What they said publicly Commits, last 3 years IBM — Vincent Hsu, IBM Fellow, CTO \u0026amp; VP of IBM Storage \u0026ldquo;reinforce our trust in Ceph and our commitment to open source\u0026rdquo; 11,091 (with Red Hat) Bloomberg — Matthew Leonard, Head of Storage Engineering \u0026ldquo;Our Diamond Membership is a symbol of our commitment to the future of Ceph and its growing community\u0026rdquo; 220 45Drives — Doug Milburn, Co-founder and President \u0026ldquo;our unwavering commitment to open-source excellence\u0026rdquo; 0 IBM\u0026rsquo;s statement is backed by the largest engineering commitment in the project\u0026rsquo;s history, and then some. Bloomberg\u0026rsquo;s is backed by 220 commits, a Steering Committee seat, and an 8 PB production estate they talk about openly — a serious contribution by any measure. 45Drives builds and sells Ceph hardware appliances and funds the shared infrastructure; that is a genuine contribution too, and it is not code.\nThe full picture across the tiers:\nMember Tier Commits, last 3 years IBM Diamond 11,091 (with Red Hat) Bloomberg Diamond 220 CLYSO Diamond 153 45Drives Diamond 0 Western Digital Platinum 0 42on Gold 1 croit Silver 157 DigitalOcean Silver 13 Canonical Silver 9 OVHcloud, Sony, OSNexus, CloudFerro Silver 0 each And in the other direction — four of the top seven contributors are not members at all:\nContributor Commits Member? QiAnXin 597 No IONOS 498 No Intel 279 Not any more Proxmox 249 No Foundation tier against commits — the money and the code are unrelated What they pay, against what they wrote Ceph Foundation tier vs commits to main, 2023-08-27 to 2026-08-27 200 400 600 commits DIAMOND IBM, with Red Hat 11,091 Bloomberg 220 CLYSO 153 45Drives nothing at all PLATINUM Western Digital nothing at all GOLD 42on 1 SILVER croit 157 DigitalOcean 13 Canonical 9 six other Silver members nothing at all, all six NOT MEMBERS QiAnXin 597 IONOS 498 Proxmox 249 Cafe Bazaar 118 The six Silver members with nothing: OVHcloud, Sony Interactive Entertainment, OSNexus, CloudFerro, Intelligent Systems, LongVan. Two of the four top-tier members wrote nothing. The third- and fifth-largest contributors to Ceph are not members at all. Foundation tier against commits over the last three years. Sponsorship and engineering are different contributions — the tiers measure the first, not the second. I do not read that as hypocrisy and I would rather nobody else did either. Foundation money pays for the upstream test lab, the continuous integration that gates every pull request, Cephalocon and the community staff — things Ceph could not do without, and things a hardware vendor shipping Ceph appliances is right to fund. Western Digital, DigitalOcean and OVHcloud all sell products that lean on Ceph, and they pay into the commons that keeps it going. That is a fair trade.\nThe practical takeaway is narrow: read the members page as a list of who funds the shared infrastructure, and the commit log for who writes the code. They are different questions with different answers, and both answers are useful.\nThe Founders, Eight Years On The Foundation launched on 12 November 2018. The roster is on the record twice — the Linux Foundation announcement and the wire copy — and they agree exactly: thirteen Premier members, ten General, eight Associate.\nAgainst the members page today: four of thirteen Premier members are still listed under their own name (Canonical, DigitalOcean, OVHcloud, Western Digital), five if you count IBM as Red Hat\u0026rsquo;s seat. Two of ten General members remain — croit and Intelligent Systems.\nAnd all eight Associate members are still there. Their full names, as the founding announcement gives them:\nBoston University Information Services and Technology CERN — the European Organization for Nuclear Research FAS Research Computing, Harvard University The Greek Research and Technology Network (GRNET) Monash University, Melbourne The South African Radio Astronomy Observatory (SARAO) The Science and Technology Facilities Council (STFC) at UK Research and Innovation (UKRI) The Center for Research in Open Source Software (CROSS) at the University of California, Santa Cruz — where Ceph was written in the first place, as Sage Weil\u0026rsquo;s PhD work with Scott Brandt, Ethan Miller, Darrell Long and Carlos Maltzahn; the original 2006 paper is still hosted on ceph.io Eight for eight, over eight years.\nThe paying members churned. The universities and research labs, who join at no cost, have stuck it out eight years without one of them leaving. Those are the people running Ceph at scale for science, and not one has walked.\nThe Quincy documentation still carries the member list as it stood around 2022, which gives the halfway point: twelve of the twenty-four commercial members in that snapshot have since gone, exactly half. Ceph\u0026rsquo;s cadence across that period did not change.\nOn croit, since it is my employer and one of the survivors. croit GmbH joined as a founding General member on day one and is still a member eight years later — a longer run than Intel, SUSE, ZTE, Arm or Samsung managed. It is also 0.3% of the decade\u0026rsquo;s commits. Sticking around is not the same as building, and I am not about to dress up longevity as contribution for the company that pays me.\nThe Manual Is One Person, and the Foundation Pays for Him Row 7 of the decade table is \u0026ldquo;Ceph Foundation\u0026rdquo;, 691 commits. That is very nearly one man.\nZac Dover has 1,060 commits over the decade, almost entirely documentation. Over the last three years he is 28.4% of everything in doc/ — the single largest contributor, ahead of both Red Hat and IBM. His address history runs @gmail.com, then @clyso.com, then @proton.me, mapped in Ceph\u0026rsquo;s own records to the Ceph Foundation.\ndoc/, last three years Commits Share Ceph Foundation 499 28.4% No employer visible 383 21.8% Red Hat 379 21.5% IBM 333 18.9% doc/ is the one directory where the biggest single contributor is neither Red Hat nor IBM, and it shows. The Ceph manual is better than most infrastructure software this size manages. Paying for a technical writer who answers to no vendor is the smartest thing the Foundation does with the money.\nIndividuals Can Still Move the Needle The top individuals of the decade, grouped by author name:\nPerson Decade commits Employer(s) Sage Weil 7,517 Red Hat — creator, left 2022 Kefu Chai 5,538 Red Hat → Proxmox Casey Bodley 2,472 Red Hat Patrick Donnelly 2,330 Red Hat → IBM Jason Dillaman 1,629 Red Hat — left 2021 Radosław Zarzyński 1,607 Mirantis → Red Hat Samuel Just 1,415 DreamHost → Inktank → Red Hat John Mulligan 1,300 Red Hat Yingxin Cheng 1,202 Intel Zac Dover 1,060 Ceph Foundation Yehuda Sadeh 972 Red Hat → IBM → Ubiquiti Alfredo Deza 930 Red Hat — left 2019 Twenty-one people wrote half the decade\u0026rsquo;s commits; ninety-one wrote 80%. That is normal for a big C++ codebase, and it is also why individuals count for so much here.\nThe clearest proof the door is open: fifth place in the last three years is one engineer at IONOS. Max Kellermann has 498 commits since 2024 — more than Intel, Proxmox, Bloomberg, croit or Clyso managed as companies — across src/mds, src/common, src/mon, src/tools, src/librbd, src/mgr and src/rgw. He is 19.3% of all CephFS metadata work in the window, second only to Red Hat.\nNobody appointed him. He turned up and started fixing things, and three years on he is one of the busiest contributors to a project run by a Fortune 50 company. You can still walk into Ceph and matter.\nProxmox is the other side of the same coin: 249 commits in the last three years, 241 of them from Kefu Chai since 30 September 2025. Proxmox went from nothing to a top-ten Ceph contributor in eleven months by hiring one good engineer.\nRGW: Where the Work Actually Is If you want to know where Ceph\u0026rsquo;s engineering goes, the answer is object storage, and the reason is simple: RGW has the furthest to go before it matches the thing it competes with.\nsrc/rgw is the largest functional subsystem in Ceph over the decade — 6,958 commits, ahead of Crimson\u0026rsquo;s 5,827, the OSD\u0026rsquo;s 4,499, the dashboard\u0026rsquo;s 3,848, BlueStore\u0026rsquo;s 2,560 and CephFS\u0026rsquo;s 2,546. It is four times the size of the block layer\u0026rsquo;s effort. In the last three years it took 1,705 commits, second only to Crimson, and Crimson is a greenfield rewrite. On shipping code, RGW is the biggest ongoing feature programme in the project.\nAmazon\u0026rsquo;s S3 is a moving target with a huge API surface, and every year it grows features that customers then expect from anything calling itself S3-compatible. So RGW chases it. Count the last three years of RGW commit subjects by feature area and the shape of that chase is plain:\nRGW work in the last 3 years Commits mentioning it Multisite replication 94 Accounts 94 IAM — identity and access management 72 Policy 71 Bucket notifications 69 STS (temporary credentials) 67 Topics 50 Roles 47 Restore 41 POSIX / filesystem gateway 40 Server-side encryption (SSE) 38 Multipart upload 32 Lifecycle 16 KMS — key management service 12 S3 Select 11 Checksums 11 Cloud transition 10 Versioning, CORS, object lock, bucket logging, tagging 28 combined (Keyword counts over 1,705 commit subjects, so a commit can appear in more than one row — the point is the distribution, not a precise total.)\nThat is not maintenance. That is identity accounts, roles and policies, session tokens, bucket notifications and topics, SSE-KMS, object lock, lifecycle rules, S3 Select, checksums, cloud tiering and multisite replication — the AWS feature list, being built out. Read the recent subjects and you find SigV4 signature-verification work, x-amz-content-sha256 handling, presigned URLs. Fiddly compatibility detail, the sort that only matters because somebody\u0026rsquo;s client library expects Amazon\u0026rsquo;s exact behaviour and will fall over without it.\nIt is also why RGW has the most mixed contributor list of the big subsystems. Over the decade Red Hat is 64.3% of it, but Bloomberg (8.6% in the recent window) and Cafe Bazaar (6.2%) are in there too — companies running large object stores in production, fixing the things that bite them.\nIf you are weighing Ceph up for S3 work, this is the number that should settle you. The gap to Amazon is why RGW gets more attention than anything else in the tree, and the largest single engineering effort in the project is pointed at closing it.\nRBD: Stable Code, Not a Decline src/librbd is Ceph\u0026rsquo;s block device — what Proxmox uses, what OpenStack Cinder uses, what most Kubernetes Container Storage Interface drivers use. If you run Ceph, you probably run RBD. Its commit graph looks like this:\nRBD commits by year — a component settling into maintenance Commits to\u0026#160;src/librbd, by year Ceph's block device — what Proxmox, OpenStack Cinder and most Kubernetes CSI drivers use 0 100 200 300 400 431 2015 477 2016 258 2017 276 2018 209 2019 455 2020 165 2021 85 2022 49 2023 74 2024 45 2025 19 2026 Feature work substantially complete Jason Dillaman wrote 281 of 2020's 455 commits — persistent write-back cache and crypto — then RBD settled into maintenance. 2026 runs to 27 August. This is not decline — it is a mature component being maintained rather than rebuilt. Commits to src/librbd by year. The heavy feature work finished around 2020 — Jason Dillaman wrote 281 of that year\u0026rsquo;s 455 commits, on the persistent write-back cache and crypto — and the component has since settled into maintenance. 159 commits in the last three years, about one a week, against 1,705 for RGW and 1,944 for Crimson.\nThat is what stable code looks like, and it is a feature. Block storage over RADOS is a solved problem. RBD has had snapshots, clones, layering, mirroring, encryption, live migration and a persistent cache for years, and there is nothing like the S3 API racing off ahead of it, because what a hypervisor wants from a block device has barely changed in a decade. The feature work is done. What is left is upkeep: bug fixes, keeping step with the kernel, the odd performance win.\nSet it against RGW on purpose. RGW takes ten times the commits because it has ten times as far left to go. RBD has not, so it does not. A component that has stopped changing shape is not a component going to seed — and if librbd suddenly took 400 commits a year I would want to know what had gone wrong, because those are my virtual machine disks it is holding.\nThe one thing worth knowing is that it puts the expertise in very few heads. RBD is more or less Ilya Dryomov, who looks after both ends — upstream Ceph and the Linux kernel rbd driver. That is about the best setup going, and it is still one person deep. Which tells you who to ask, not whether to deploy.\nWhere the Work Sits The same mapping across the tree, decade and recent window side by side:\nArea Decade leader Share Last 3 years leader IBM group, last 3y src/mds — CephFS metadata Red Hat 78.5% Red Hat 56.4% 73.9% src/osd — current OSD Red Hat 69.7% Red Hat 56.9% 84.0% src/cephadm — deployment Red Hat 66.8% Red Hat 73.7% 92.0% src/rgw — S3 Red Hat 64.3% Red Hat 56.1% 66.7% src/librbd — block Red Hat 61.1% Red Hat 76.7% 83.0% src/crimson — next OSD Red Hat 52.1% Red Hat 42.0% 45.4% doc/ — the manual Red Hat 45.6% Ceph Foundation 28.4% 40.5% src/os/bluestore — engine Red Hat 36.9% IBM 46.5% 54.2% src/pybind/mgr/dashboard SUSE 36.6% IBM 57.5% 87.8% Two things to read off this. Ownership: the closer to the parts a vendor sells — deployment tooling, the GUI — the more it is one company; the further out — the next-generation OSD, the storage engine, the manual — the more crowded, and that is where the room is if you want to contribute somewhere not already owned.\nVolume: by total commits over the decade, the ranking is RGW 6,958, Crimson 5,827, the OSD 4,499, the dashboard 3,848, the monitors 2,896, BlueStore 2,560, CephFS 2,546, RBD 1,750, cephadm 1,685. Effort tracks distance-to-done, not deployment share. RGW is first because S3 parity is a long way off; RBD is near the bottom because block storage is finished.\nBlueStore deserves its own line. It is the engine that writes your bytes to disk, and over the last three years the second-largest contributor after IBM is croit at 18.1% — Igor Fedotov, its principal maintainer, at a company of a few dozen people. My employer, so weigh me accordingly. The point stands whoever signs his cheque: a small company can employ the maintainer of one of the most safety-critical components in the stack, and the project is better for it.\nWho Holds the Merge Button I counted merges too — 7,893 in the last three years, attributed to whoever pressed the button:\nOrganisation Merges Share Red Hat 3,892 49.3% IBM 1,597 20.2% Ceph Foundation 524 6.6% Proxmox 202 2.6% Intel 136 1.7% croit 76 1.0% 69.5% of merges against 68.6% of commits. The gate and the work are the same shape. This is not one company writing the code and another controlling what lands, which is the failure mode actually worth worrying about in corporate open source. Ceph does not have it.\nWhere the Numbers Come From All of it is reproducible. Ceph maintains its own contributor-to-organisation mapping in the repository — .organizationmap, alongside .mailmap, .peoplemap and .githubmap — and documents the command to use it.\ngit clone --filter=blob:none --no-checkout https://github.com/ceph/ceph.git cd ceph git show HEAD:.organizationmap \u0026gt; /tmp/orgmap # Ceph\u0026#39;s own documented method, over the last ten years git log --no-merges --since=2016-08-27 --until=2026-08-27 --pretty=\u0026#39;%aN \u0026lt;%aE\u0026gt;\u0026#39; \\ | git -c mailmap.file=/tmp/orgmap check-mailmap --stdin \\ | sort | uniq -c | sort -rn | head -30 # the maintainer-to-employer mapping, straight from the repo git show HEAD:doc/governance.rst | sed -n \u0026#39;/^.. _csc:/,/^\\.\\. _ctl:/p\u0026#39; \\ | grep -oE \u0026#39;\\* [^\u0026lt;]+\u0026lt;[^\u0026gt;]+\u0026gt;\u0026#39; Run the first and you get a smaller IBM than mine, because the official map is out of date. It does not know aainscow@uk.ibm.com, bill_scales@uk.ibm.com, ylifshit@ibm.com, rkachach@ibm.com, leonid.usov@ibm.com or the li-*.ibm.com machine-generated hostnames. Using the project\u0026rsquo;s own tooling understates the concentration.\nMy corrections on top of the map:\nAny address ending ibm.com — including uk.ibm.com, il.ibm.com, in.ibm.com, de.ibm.com and the li-*.ibm.com forms — is IBM. redhat.com and inktank.com are Red Hat, shown separately from IBM but the same team since October 2022. Seven personal addresses are attributed to employers where the repository itself proves it: sage@newdream.net (Red Hat), idryomov@gmail.com (listed as idryomov@redhat.com in Ceph\u0026rsquo;s own doc/governance.rst), max.kellermann@gmail.com (IONOS), xxhdx1985126@gmail.com (QiAnXin), yuvalif@yahoo.com (IBM), yingxincheng@gmail.com (Intel), shraddha.agrawal000@gmail.com (IBM). Everything else keeps its domain. Personal addresses stay \u0026ldquo;no employer visible\u0026rdquo; rather than being guessed at. Caveats I cannot fix. Commits are a rough unit — a careful 900-line refactor counts once, forty typo fixes count forty times, and nothing here is weighted. Email domains are imperfect: 9.0% of the decade shows no employer, and some of those people are certainly paid to write Ceph. Review is invisible in git — Ceph reviews in GitHub pull requests, not Reviewed-by: trailers, of which I found twelve in three years; the most important gate in the project leaves no trace in a clone. Figures are main only, so backports to stable branches are uncounted, which understates the maintenance work. And every adoption figure here is a floor, for the telemetry reason set out above.\nWhere a claim rests on a date I have used the author date; where it rests on somebody\u0026rsquo;s employer, an address they published themselves.\nWhat I Take From This Ceph is a corporate-funded project with a real community round the edges, and it has been that way its whole commercial life. That is not a dig. Somebody has to pay engineers to look after distributed storage at this scale, and for ten years somebody has.\nI will not pretend Ceph is typical, because I checked. LWN\u0026rsquo;s statistics for Linux 6.15 record 2,068 developers from at least 195 employers with the largest single company, Intel, on 12.0% of changesets. Ceph is 1,718 people over a decade with one company on 67.0%. The kernel spreads its corporate dependence across dozens of firms. Ceph packs it into one. That is a real difference, and \u0026ldquo;everyone does this\u0026rdquo; would be a lazy way to wave it off.\nBut here is what ten years of data says about whether that matters, and it is a better answer than I went looking for:\nCeph is tough in the way that counts. It lost the man who wrote it, who was doing a fifth of the work. It lost SUSE, its second-largest contributor and the author of the dashboard. It lost ZTE, Mirantis, XSKY, EasyStack, Inspur and half the Foundation\u0026rsquo;s founding members. The commit rate today is within two per cent of last year\u0026rsquo;s. Twenty years of storage engineering sits in that tree under LGPL-2.1 or LGPL-3, and nobody can close it, relicense it or take it back.\nThe maintainers are portable. Fedotov has kept BlueStore going through three employers. Kefu Chai went from Red Hat to Proxmox and carried on. Sadeh wrote RGW at DreamHost and changed his committee address to Ubiquiti this June. When a company walks, its people often stay.\nThe door is properly open. One engineer at IONOS became the fifth-largest contributor in three years. Proxmox got into the top ten on one hire. A security firm with no storage product is building a fifth of the next-generation OSD. 478 organisations have sent patches. If you want in, there is nothing stopping you but the work.\nThe userbase is far bigger than anybody can measure. An exabyte across 3,000 clusters is what put its hand up through opt-in telemetry, and CERN on its own accounts for nineteen production clusters. Every Proxmox hyperconverged cluster, every Rook deployment, every vendor appliance with Ceph under the lid is real production use that no published figure counts. Software this widely and this quietly deployed does not just disappear.\nEffort goes where the gap is, not where the users are. RGW is the biggest programme in the project — 6,958 commits over the decade — because catching Amazon on S3 is a long chase against a moving target. RBD sits near the bottom because block storage is done. A low commit count on a mature component is a finished job, not a warning, and reading those two numbers the wrong way round is the easiest mistake going with data like this. I made it myself on the first pass.\nJudge suppliers on commits, not on tiers. The history is public and four lines of shell will show you who actually looks after the thing you are about to depend on. Closed storage does not offer you that at any price.\nSo if you are weighing Ceph up: the concentration is worth knowing when you are planning five years out, and it is no reason to hold back. A mature block layer. The project\u0026rsquo;s biggest engineering effort aimed squarely at S3 parity. A future OSD being built by three companies at once. A manual better than most. Governance that came through losing its founder. A decade of review done in the open, 478 organisations\u0026rsquo; worth of work in the tree, and a licence whose worst case is a fork rather than a dead end. On the evidence of 69,613 commits, this project is in good health.\nAnd count it yourself if you doubt me. The commands are up there, the data is public, and nowt in this post needs taking on trust — mine or anybody else\u0026rsquo;s.\nReferences Ceph project sources\n\u0026ldquo;Ceph: A Scalable, High-Performance Distributed File System\u0026rdquo; — Weil, Brandt, Miller, Long and Maltzahn, OSDI \u0026lsquo;06, November 2006. The paper Ceph started as ceph/ceph on GitHub — the repository every commit figure comes from; .organizationmap, .mailmap, .peoplemap, .githubmap, doc/governance.rst and COPYING are all in-tree Ceph governance — Executive Council and Ceph Steering Committee membership, with addresses Ceph component team — component leads Ceph Foundation members — current membership by tier Ceph Foundation documentation — tier structure and no-cost Associate membership Ceph Foundation members, Quincy documentation — membership as it stood around 2022 Ceph telemetry module — confirms telemetry is opt-in \u0026ldquo;Red Hat\u0026rsquo;s Ceph team is moving to IBM\u0026rdquo;, 4 October 2022 Ceph Community Newsletter, November 2021 — Sage Weil stepping back Ceph Foundation and the Linux Foundation\nIntroducing Ceph Squid — the Diamond member statements quoted above, and the 1 exabyte / 3,000+ cluster figures The Linux Foundation Launches Ceph Foundation, 12 November 2018 — founding member roster The same release via PRNewswire — used to confirm the roster independently Deployments\n\u0026ldquo;Ceph: Infrastructure Storage at CERN\u0026rdquo; — Enrico Bocchi, CERN IT Storage, 27 September 2024. Every CERN figure above comes from this deck Why We Chose Ceph to Build Block Storage — DigitalOcean \u0026ldquo;We Added 6 Petabytes Of Ceph Storage and No Clients Noticed\u0026rdquo; — Matthew Leonard and Joseph Mundackal, Bloomberg, Cephalocon 2020 Ceph on Wikitech — Wikimedia Foundation\u0026rsquo;s five production clusters Ceph RBD block storage — OVHcloud\u0026rsquo;s own documentation Deploy Hyper-Converged Ceph Cluster — Proxmox VE shipping Ceph as its hyperconverged storage Companies\n\u0026ldquo;SUSE says tschüss to Ceph-based enterprise storage product\u0026rdquo;, The Register, 25 March 2021 — SUSE Enterprise Storage cancelled for Longhorn Qi An Xin files for $634m IPO, Global Venturing, 13 May 2020 — QiAnXin\u0026rsquo;s origins in Qihoo 360, the CEC stake and shareholding Comparison\nDevelopment statistics for the 6.15 kernel, LWN.net — kernel developer and employer counts used for the concentration comparison ","permalink":"https://blogs.damiendye.uk/en/ceph/who-actually-writes-and-uses-ceph/","summary":"I counted every commit to Ceph\u0026rsquo;s main branch for the last ten years — 69,613 of them from 1,718 people across 478 organisations — traced all 38 Steering Committee members to their employers, and gathered the publicly stated deployments. The result is a project that absorbed the loss of its founder and its second-largest contributor without dropping a beat, and is still shipping at the same rate today.","title":"Who Actually Writes \u0026 Uses Ceph"},{"content":"The Host Exists This Time. It Is Just on the Wrong Hypervisor Last time the problem was that the machine named in the inventory did not exist yet — no IP, no SSH, no Python, nothing to connect to. Every task had to be delegated away from it.\nA VMware migration inverts that and changes nothing. The machine exists, it is running, people are using it. You still never connect to it. It is a name and a bag of variables describing something to rebuild somewhere else. Every task still runs on the control node, and now there are two APIs on the other end instead of one.\nA note on what this is. This post is the design and the playbook, not a war story. I have not yet run it against a production estate end to end. Everything I say about module behaviour below was read out of the shipped code and checked, and I have said plainly where a claim comes from the source rather than from a run. When I have done a real migration with it, the numbers and the surprises will get their own post.\nVersions this was checked against:\n$ ansible --version | head -1 ansible [core 2.20.7] $ ansible-galaxy collection list | grep -E \u0026#39;vmware|proxmox\u0026#39; community.proxmox 1.6.0 community.vmware 6.2.1 vmware.vmware 2.9.0 The fragments below are cut from a single playbook — migrate.yml, a vSphere dynamic inventory in inventory/vmware.vms.yml, and one group_vars/all.yml. I have genericised the datacenter, node and storage names for readability.\nThe Shape of the Job Five steps, and only the last one costs anything.\nguests up, nothing at risk the outage discover read the distributed portgroups and their VLAN tags sdn one VLAN zone, one VNet per VLAN on Proxmox build diskless shells, right CPU, RAM, firmware and MACs relocate storage vMotion the VMDKs onto NFS, while they run cutover power off, import the disks, set boot order and start the guest Everything reversible is done before anything is switched off The slow part is the relocate, and it costs no downtime at all. By the time the window opens the disks are already on storage Proxmox mounts, so the cutover is a local import rather than a copy. A shell with no disk is cheap to delete, so a mistake before the cutover costs nothing but time. Abandoning the migration halfway leaves every guest still running on VMware, untouched. Everything reversible happens first. The slow stage is free, and the expensive stage is short, because by then the disks are already where they need to be. The ordering is the whole design. Discovery changes nothing. Building the network changes only Proxmox. Building the shells changes only Proxmox, and a shell with no disk is cheap to delete. Moving the disks is slow but live. Only the last play powers anything off.\nWalk away halfway through and every guest is still running on VMware, untouched.\nTwo Collections, and One of Them Is Being Retired You need both, and not for the reason you would guess.\ncollections: - name: community.vmware version: \u0026#34;\u0026gt;=6.2.1\u0026#34; - name: vmware.vmware version: \u0026#34;\u0026gt;=2.5.0\u0026#34; - name: community.proxmox version: \u0026#34;\u0026gt;=1.6.0\u0026#34; community.vmware is the old, broad collection and it is being taken apart. Its MANIFEST.json declares {\u0026quot;vmware.vmware\u0026quot;: \u0026quot;\u0026gt;=2.5.0\u0026quot;} as a hard dependency, so installing the first pulls in the second whether you asked for it or not. Modules are moving across one at a time, and the ones you reach for in a migration are at different stages of that move:\nvmware_dvs_portgroup_info — still only in community.vmware, and it is what reads your VLANs. vmware_vmotion — still only in community.vmware. vmware_guest_powerstate — deprecated, removed in community.vmware 7.0.0. Use vmware.vmware.vm_powerstate. vmware_vm_inventory — deprecated, removed in 7.0.0. Use vmware.vmware.vms. Ansible tells you about the module deprecations on the first run, which is decent of it:\n[DEPRECATION WARNING]: community.vmware.vmware_guest_powerstate has been deprecated. Use vmware.vmware.vm_powerstate instead. This feature will be removed from collection \u0026#39;community.vmware\u0026#39; version 7.0.0. It does not warn you about the inventory plugin, because inventory plugins are parsed before that machinery is running. You have to go and read the plugin.\nThere is a third trap in the split. vmware.vmware.vm_portgroup_info looks like exactly what a network migration wants — per-VM, per-NIC, gives you the portgroup and the VLAN. But it is built on ModuleRestBase and imports com.vmware.vapi, which means it needs the vSphere Automation SDK on the control node, not just pyVmomi. Its documented return is also stale: the RETURN block promises name and vlan_id, while the code actually builds portgroup_name and a vlan_info dict for the distributed case. I went a different way, below, and needed neither.\nThe Inventory Is the Discovery There is no \u0026ldquo;go and find the VMs\u0026rdquo; play in this playbook, because by the time the first task runs the inventory has already done it — in one property-collector query rather than a per-VM loop.\n# inventory/vmware.vms.yml plugin: vmware.vmware.vms hostname: \u0026#34;{{ lookup(\u0026#39;ansible.builtin.env\u0026#39;, \u0026#39;VMWARE_HOST\u0026#39;) }}\u0026#34; username: \u0026#34;{{ lookup(\u0026#39;ansible.builtin.env\u0026#39;, \u0026#39;VMWARE_USER\u0026#39;) }}\u0026#34; password: \u0026#34;{{ lookup(\u0026#39;ansible.builtin.env\u0026#39;, \u0026#39;VMWARE_PASSWORD\u0026#39;) }}\u0026#34; validate_certs: false search_paths: - /Datacenter-1 properties: - name - config.name - config.uuid - config.guestId - config.firmware - config.template - config.hardware.numCPU - config.hardware.numCoresPerSocket - config.hardware.memoryMB - config.hardware.device - summary.runtime.powerState gather_compute_objects: true hostnames: [\u0026#39;name\u0026#39;] filter_expressions: - \u0026#39;config.template\u0026#39; The filename matters. The plugin\u0026rsquo;s verify_file only claims files ending vms.yml, vms.yaml, vmware_vms.yml or vmware_vms.yaml. Call it vcenter.yml and it is silently not your inventory.\nsearch_paths filters before the query, not after. On a large estate that is the difference between seconds and minutes — unlike filter_expressions, which the docs are explicit about: it runs after collection and \u0026ldquo;does not affect the speed of the inventory plugin\u0026rdquo;.\nfilter_expressions drops a host when the expression is true. config.template therefore removes templates, which reads backwards the first time.\nAnd the important line is config.hardware.device, which is in no default property list anywhere. It is the whole hardware inventory of the VM, and it carries three things this migration cannot proceed without: the MAC of every NIC, the dvportgroup key each NIC is attached to, and the datastore path of every disk. Without it you are back to a vmware_guest_info loop, one round trip per VM.\nThe devices come back as JSON with their vSphere type preserved in _vimtype. That is worth knowing because it is how you tell a NIC from a disk. I checked the encoder rather than guessing:\n{ \u0026#34;_vimtype\u0026#34;: \u0026#34;vim.vm.device.VirtualVmxnet3\u0026#34;, \u0026#34;macAddress\u0026#34;: \u0026#34;00:50:56:87:a5:9a\u0026#34;, \u0026#34;backing\u0026#34;: { \u0026#34;_vimtype\u0026#34;: \u0026#34;...DistributedVirtualPortBackingInfo\u0026#34;, \u0026#34;port\u0026#34;: { \u0026#34;_vimtype\u0026#34;: \u0026#34;vim.dvs.PortConnection\u0026#34;, \u0026#34;portgroupKey\u0026#34;: \u0026#34;dvportgroup-1014\u0026#34; } } } So a compose block can pull the awkward paths up into flat hostvars:\ncompose: vm_moid: moid vm_firmware: config.firmware vm_memory_mb: config.hardware.memoryMB vm_num_cpu: config.hardware.numCPU # A virtual NIC is any device with a MAC. Filtering on _vimtype does not # work cleanly here, because VMXNET3, E1000 and SR-IOV cards are all # different types with no shared substring. vm_nics: \u0026gt;- config.hardware.device | selectattr(\u0026#39;macAddress\u0026#39;, \u0026#39;defined\u0026#39;) | selectattr(\u0026#39;macAddress\u0026#39;, \u0026#39;ne\u0026#39;, None) | list # Disks are one exact type, so this one can match on it. vm_disks: \u0026gt;- config.hardware.device | selectattr(\u0026#39;_vimtype\u0026#39;, \u0026#39;eq\u0026#39;, \u0026#39;vim.vm.device.VirtualDisk\u0026#39;) | list That asymmetry is real and it catches people. There is no VirtualEthernetCard type to match. That is the abstract base class, and what vCenter actually hands you is VirtualVmxnet3, VirtualE1000, VirtualE1000e, VirtualPCNet32 or VirtualSriovEthernetCard. There is no substring common to all of them. Having a MAC, though, is a thing only a NIC does.\nEvery Info Module Hides the Field You Need This is the through-line of the whole job, and once you have seen it three times you start checking every default before you write the task.\nvmware_dvs_portgroup_info has six show_* options. Five default to true. The sixth is show_vlan_info, and it defaults to false.\nshow_mac_learning=dict(type=\u0026#39;bool\u0026#39;, default=True), show_network_policy=dict(type=\u0026#39;bool\u0026#39;, default=True), show_teaming_policy=dict(type=\u0026#39;bool\u0026#39;, default=True), show_uplinks=dict(type=\u0026#39;bool\u0026#39;, default=True), show_port_policy=dict(type=\u0026#39;bool\u0026#39;, default=True), show_vlan_info=dict(type=\u0026#39;bool\u0026#39;, default=False), Leave it alone and you get MAC learning policy, teaming policy, uplink ordering and port policy for every portgroup in the estate. Everything except the VLAN tag, which is the only field a network migration is actually asking about. So the task is inside out from what you would write by instinct: turn the one thing on, turn the other five off.\n- name: Read the distributed portgroups community.vmware.vmware_dvs_portgroup_info: datacenter: \u0026#34;{{ vcenter_datacenter }}\u0026#34; show_vlan_info: true show_network_policy: false show_teaming_policy: false show_port_policy: false show_mac_learning: false show_uplinks: false register: dvs_pgs It is not a one-off. vmware.vmware.vms has gather_compute_objects, which populates cluster and esxi_host — default false. community.vmware.vmware_vm_info has show_allocated, which is the block holding CPU and memory — default false. In all three cases the expensive-to-collect field is the one the migration needs, and the default protects a read-only reporting use case that is not the one you are in.\nvlan_id Is Three Different Types Then you get the VLAN tags and find they are not one shape. Straight from get_vlan_info:\nif isinstance(vlan_obj, vim...TrunkVlanSpec): ... return dict(trunk=True, pvlan=False, vlan_id=vlan_id_list) elif isinstance(vlan_obj, vim...PvlanSpec): return dict(trunk=False, pvlan=True, vlan_id=str(vlan_obj.pvlanId)) else: return dict(trunk=False, pvlan=False, vlan_id=str(vlan_obj.vlanId)) An access portgroup gives you the string \u0026quot;100\u0026quot;. A PVLAN gives you a string. A trunk gives you a list of strings, each either \u0026quot;20\u0026quot; or \u0026quot;20-30\u0026quot;. And every distributed switch has at least one trunk on it whether you made one or not, because the uplink portgroup is a trunk carrying \u0026quot;0-4094\u0026quot;.\nSo | int is not available to you until you have thrown the other two shapes away:\naccess_pgs: \u0026gt;- {{ dvs_pgs.dvs_portgroup_info | dict2items | map(attribute=\u0026#39;value\u0026#39;) | flatten | rejectattr(\u0026#39;vlan_info.trunk\u0026#39;) | rejectattr(\u0026#39;vlan_info.pvlan\u0026#39;) | rejectattr(\u0026#39;vlan_info.vlan_id\u0026#39;, \u0026#39;in\u0026#39;, [\u0026#39;0\u0026#39;, 0]) | list }} Three rejects, in that order. Trunks go, PVLANs go, and then untagged portgroups go, which also disposes of the uplink groups and anything on VLAN 0.\nI am not translating trunks or PVLANs automatically and I would push back on anyone who did. A VMware trunk landing on Proxmox needs either a Q-in-Q zone or a VLAN-aware VNet, and which one is right depends on what the guest expects to see. That is a decision, not a mapping. The playbook prints them and moves on:\nTASK [Report what was found] ok: [localhost] =\u0026gt; { \u0026#34;msg\u0026#34;: \u0026#34;3 access portgroups -\u0026gt; [100, 200]. Not translated: [\u0026#39;dvs_001-uplink\u0026#39;] (trunks), [\u0026#39;isolated\u0026#39;] (PVLANs).\u0026#34; } Mirroring the VLANs into SDN One VLAN zone bound to a bridge, then one VNet per VLAN with the tag on it.\n- name: Create the VLAN zone community.proxmox.proxmox_zone: zone: \u0026#34;{{ sdn_zone }}\u0026#34; type: vlan bridge: \u0026#34;{{ sdn_bridge }}\u0026#34; mtu: \u0026#34;{{ sdn_mtu }}\u0026#34; state: present - name: Create one VNet per VMware VLAN community.proxmox.proxmox_vnet: vnet: \u0026#34;{{ sdn_vnet_prefix }}{{ item.vlan_info.vlan_id | int }}\u0026#34; zone: \u0026#34;{{ sdn_zone }}\u0026#34; tag: \u0026#34;{{ item.vlan_info.vlan_id | int }}\u0026#34; alias: \u0026#34;{{ item.portgroup_name }}\u0026#34; state: present loop: \u0026#34;{{ access_pgs | unique(attribute=\u0026#39;vlan_info.vlan_id\u0026#39;) }}\u0026#34; throttle: 1 VNet names are short and constrained, and VMware portgroup names are not. Production-Web-Tier-VLAN100 is a perfectly ordinary portgroup name and an impossible VNet name. So the name is generated — v100, from the tag — and the human-readable original goes in alias, where it stays visible in the UI and in pvesh output. Deriving the name from the VLAN rather than from the portgroup also means the mapping is reversible by inspection six months later.\nTwo portgroups on the same VLAN collapse into one VNet. That is correct — they were the same broadcast domain in VMware too — but you should see it happen, which is what unique(attribute='vlan_info.vlan_id') is doing. Two portgroups called prod-web and prod-web-b, both on VLAN 100, produce one v100.\nthrottle: 1 is not caution, it is the module. Every SDN write in community.proxmox takes a global cluster lock, applies the pending config and releases it — get_global_sdn_lock(), then apply_sdn_changes_and_release_lock(). Run them in parallel and they queue on the lock anyway; the throttle just stops you pretending otherwise. Worth knowing too that rollback on failure is version-dependent — the module checks is_lock_and_rollback_supported and, on older PVE, tells you it could not roll back rather than doing it.\nOne cosmetic thing that will make you doubt yourself. At 1.6.0 proxmox_vnet emits its entire params dict as an Ansible warning on every single create:\nself.module.warn(f\u0026#34;{vnet_params}\u0026#34;) self.proxmox_api.cluster().sdn().vnets().post(**vnet_params) That is a debug line somebody left in. It is noise, not a fault.\nBuild the Shells, With No Disks Now the VMs, and this is where the design earns itself. Every VM gets built in Proxmox with the right CPU count, the right memory, the right firmware and the right NICs on the right VLANs. No disks at all.\nA diskless shell is fast to create, free to delete, and boots to a PXE prompt if somebody starts it by accident. You can build four hundred of them in an afternoon, look at the result, decide it is wrong, delete the lot and do it again. Nothing has been copied, nothing has been powered off, and nobody has noticed.\nThe derived values are declarations, not tasks. Ansible evaluates them lazily against whichever host is in scope, so every VM gets its own without a single set_fact:\n# group_vars/all.yml pve_vmid: \u0026#34;{{ vmid_base | int + (vm_moid | regex_replace(\u0026#39;^vm-\u0026#39;, \u0026#39;\u0026#39;) | int) }}\u0026#34; pve_bios: \u0026#34;{{ \u0026#39;ovmf\u0026#39; if vm_firmware == \u0026#39;efi\u0026#39; else \u0026#39;seabios\u0026#39; }}\u0026#34; pve_cores: \u0026#34;{{ vm_cores_per_socket | int }}\u0026#34; pve_sockets: \u0026#34;{{ ((vm_num_cpu | int) / (vm_cores_per_socket | int)) | round(0, \u0026#39;ceil\u0026#39;) | int }}\u0026#34; The VMID comes from the vCenter MoID. vm-42 becomes 20042. That matters more than it looks: the cutover play has to find the VM the build play created, and a re-run must land on the same one rather than quietly building a second. Letting the API allocate the next free ID — which is what happens if you omit vmid, and which I wrote about last time — makes that impossible.\nMemory needs no conversion. VMware reports config.hardware.memoryMB and Proxmox wants MB. Sockets do: VMware gives you total vCPUs and cores-per-socket, Proxmox wants sockets and cores.\n- name: Create the VM shell delegate_to: localhost community.proxmox.proxmox_kvm: node: \u0026#34;{{ proxmox_node }}\u0026#34; vmid: \u0026#34;{{ pve_vmid }}\u0026#34; name: \u0026#34;{{ inventory_hostname }}\u0026#34; cores: \u0026#34;{{ pve_cores }}\u0026#34; sockets: \u0026#34;{{ pve_sockets }}\u0026#34; memory: \u0026#34;{{ vm_memory_mb }}\u0026#34; ostype: \u0026#34;{{ pve_ostype }}\u0026#34; bios: \u0026#34;{{ pve_bios }}\u0026#34; scsihw: \u0026#34;{{ default_scsihw }}\u0026#34; efidisk0: \u0026#34;{{ {\u0026#39;storage\u0026#39;: pve_target_storage, \u0026#39;efitype\u0026#39;: \u0026#39;4m\u0026#39;, \u0026#39;pre_enrolled_keys\u0026#39;: false} if pve_bios == \u0026#39;ovmf\u0026#39; else omit }}\u0026#34; agent: \u0026#34;enabled=1\u0026#34; onboot: false state: present onboot: false on purpose. Nothing should start by itself in the middle of a migration, least of all a machine whose disks are still being written to by another hypervisor.\nFirmware is not optional to get right. A UEFI guest imported onto a SeaBIOS VM will import perfectly and then refuse to boot, and you will spend an hour on it. config.firmware is efi or bios and maps straight onto ovmf and seabios. A UEFI guest also needs an EFI vars disk, which has to be created with the VM — see below for why.\nproxmox_kvm Will Not Fix a NIC, and Will Not Tell You Last time I wrote that proxmox_kvm declines to converge rather than updating. Here is the sharper version of that, which bit me while writing this and is worth being exact about.\nupdate defaults to false, so re-running against a VM that already exists does nothing. Fine, and documented. But set update: true and the module still refuses to touch some parameters:\n# If update, don\u0026#39;t update disk (virtio, efidisk0, tpmstate0, ide, sata, scsi) # and network interface, unless update_unsafe=True if update_unsafe is False: ... if \u0026#34;efidisk0\u0026#34; in kwargs: del kwargs[\u0026#34;efidisk0\u0026#34;] It deletes them from the request and carries on. So you correct a NIC in your inventory mapping, re-run with update: true, watch Ansible report changed, and the NIC is exactly as wrong as it was. The changed is true — something else in the payload was updated — but not the thing you were fixing.\nupdate_unsafe: true lifts the restriction, and the name is honest. The same guard covers disks, so on a VM that has disks, an unsafe update is a good way to acquire a second copy of one. That is not a switch to reach for during a migration.\nThe way out is to not use net at all. NICs go on with proxmox_nic, which is a module whose entire job is one interface and which converges properly:\n- name: Attach each NIC to its VNet delegate_to: localhost community.proxmox.proxmox_nic: vmid: \u0026#34;{{ pve_vmid }}\u0026#34; interface: \u0026#34;net{{ idx }}\u0026#34; bridge: \u0026#34;{{ sdn_vnet_prefix }}{{ pg_vlan[item.backing.port.portgroupKey] }}\u0026#34; mac: \u0026#34;{{ item.macAddress }}\u0026#34; model: \u0026#34;{{ default_net_model }}\u0026#34; state: present loop: \u0026#34;{{ vm_nics }}\u0026#34; loop_control: index_var: idx That is the same split I ended up at last time: proxmox_kvm to define the machine, proxmox_disk and proxmox_nic for the things that change afterwards. The parameter is mac, not mac_addr.\nefidisk0 cannot be moved out the same way — proxmox_disk has no efitype or pre_enrolled_keys — so it has to go on at create time and be right first time.\nCarry the MAC across. VMware hands out MACs from 00:50:56:... and Proxmox will take them without complaint. Keeping them means DHCP reservations still match, MAC-locked licences still validate, and any firewall rule written against a MAC still fires. Changing them means a day of small mysteries. proxmox_nic also accepts model: vmxnet3 if you need the guest to see the same NIC it saw before, but on KVM, virtio is the better card, and a Windows guest is going to want new drivers either way.\nRefuse Rather Than Guess A NIC on a standard portgroup has no backing.port at all. Its backing is a NetworkBackingInfo with a deviceName. It will not be in the map, and the right thing to do is stop:\n- name: Every NIC must sit on a distributed portgroup with a VNet ansible.builtin.assert: that: - vm_nics | rejectattr(\u0026#39;backing.port.portgroupKey\u0026#39;, \u0026#39;defined\u0026#39;) | list | length == 0 - vm_nics | map(attribute=\u0026#39;backing.port.portgroupKey\u0026#39;) | reject(\u0026#39;in\u0026#39;, pg_vlan.keys() | list) | list | length == 0 fail_msg: \u0026gt;- {{ inventory_hostname }} has NICs that do not map to a Proxmox VNet. Attaching it to the wrong network is worse than not building it. Two conditions rather than one, because the first has to run before the second: map(attribute=...) over a NIC with no port would explode on the undefined lookup. Reject the shapeless ones first, then check the rest against the map.\nOne Export, Mounted Twice Here is the part that makes the whole thing cheap.\nPut an NFS export where both hypervisors can mount it. vCenter sees a datastore called nfs-migration; the Proxmox nodes mount the same export and see /mnt/pve/nfs-migration. Now storage-vMotion the VMDKs onto it.\nStorage vMotion is live. The guest keeps serving traffic the entire time. Nothing is cut over, no window is needed, and it can be abandoned halfway with no consequence beyond wasted I/O. It is the slowest stage by a wide margin and it costs nothing.\n- name: Relocate to the NFS datastore delegate_to: localhost throttle: 2 community.vmware.vmware_vmotion: moid: \u0026#34;{{ vm_moid }}\u0026#34; destination_datastore: \u0026#34;{{ nfs_datastore_vmware }}\u0026#34; timeout: \u0026#34;{{ vmotion_timeout }}\u0026#34; timeout defaults to 3600 — one hour. A 2 TB VMDK will not make it, and the failure mode is nasty in a quiet way: the Ansible task fails while the vMotion carries on running in vCenter. You now have a playbook that says it failed and an estate that is still busy. Set it to something that reflects your actual storage.\nthrottle: 2, because the bottleneck is not the control node. Storage vMotion is bounded by the array and the network. Six at once does not give you six times the throughput; it gives you six slow migrations and an angry storage team.\nThe module is idempotent in the way you want — it sets storage_vmotion_needed = False if the VM is already on the target datastore — so re-running to pick up stragglers is safe.\nBy the time this finishes, the bytes are sitting on storage that Proxmox already mounts. As such, nothing else needs to copy them. Ever.\nThe Cutover This is the only play that costs downtime, and the order inside it is not negotiable.\nFirst, a problem that is easy to miss: the inventory is now stale. It was gathered before the vMotion, so vm_disks still holds the old datastore paths. Import from those and you are pointing Proxmox at a path it cannot see.\n- name: Re-read the inventory now the disks have moved ansible.builtin.meta: refresh_inventory Which is also why caching is switched off in the inventory config. A warm cache would hand refresh_inventory back exactly the stale data it was called to replace. That is a real trade — vCenter is not fast — but a wrong path here is a failed cutover in a window, and the round trip is cheap by comparison.\nThen power off. Importing a VMDK that an ESXi host still has open gives you a crash-consistent copy at best:\n- name: Shut the guest down in VMware delegate_to: localhost vmware.vmware.vm_powerstate: moid: \u0026#34;{{ vm_moid }}\u0026#34; state: \u0026#34;{{ \u0026#39;shutdown-guest\u0026#39; if vm_power_state == \u0026#39;poweredOn\u0026#39; else \u0026#39;powered-off\u0026#39; }}\u0026#34; timeout: 600 force: true shutdown-guest is a graceful shutdown through VMware Tools; force: true hard-stops anything that will not go within the timeout. On the new module the parameter is timeout, not state_change_timeout as it was on the deprecated one.\nThen the import, which is the pivot:\n- name: Import each VMDK onto its VM delegate_to: localhost throttle: 2 community.proxmox.proxmox_disk: vmid: \u0026#34;{{ pve_vmid }}\u0026#34; disk: \u0026#34;scsi{{ idx }}\u0026#34; storage: \u0026#34;{{ pve_target_storage }}\u0026#34; import_from: \u0026gt;- {{ item.backing.fileName | regex_replace(\u0026#39;^\\[[^\\]]+\\]\\s*\u0026#39;, \u0026#39;/mnt/pve/\u0026#39; ~ nfs_storage_pve ~ \u0026#39;/\u0026#39;) }} format: \u0026#34;{{ pve_target_format }}\u0026#34; timeout: \u0026#34;{{ import_timeout }}\u0026#34; create: regular state: present loop: \u0026#34;{{ vm_disks }}\u0026#34; loop_control: index_var: idx The regex_replace is doing the translation between the two worlds. vCenter names a disk [nfs-migration] app01/app01.vmdk; Proxmox reaches the same file at /mnt/pve/nfs-migration/app01/app01.vmdk. Same export, same bytes, no second copy. You keep the descriptor .vmdk and ignore the -flat.vmdk beside it — qemu-img reads the descriptor and follows it to the extent.\nThree things about import_from that are all in the module and all worth knowing before the window opens.\nIt only fires on create. In the update branch:\n# \u0026#39;import_from\u0026#39; fails on disk updates playbook_config = self.get_create_attributes() playbook_config.pop(\u0026#34;import_from\u0026#34;, None) If scsi0 already exists on that VM, the parameter is dropped and you get an ordinary update. So a re-run after a bad import does not re-import. It silently does nothing at all and reports success. If an import goes wrong, delete the disk before trying again.\ntimeout defaults to 600 seconds. Ten minutes, to import and convert a virtual machine\u0026rsquo;s disk. The module\u0026rsquo;s own documentation says to raise it; take the advice.\nAnd an absolute path needs root. The documentation is blunt about it:\n\u0026lt;STORAGE\u0026gt;:\u0026lt;VMID\u0026gt;/\u0026lt;FULL_NAME\u0026gt; or \u0026lt;ABSOLUTE_PATH\u0026gt;/\u0026lt;FULL_NAME\u0026gt;. \u0026lt;STORAGE\u0026gt;:import/\u0026lt;FULL_NAME\u0026gt; for PVE 9.x and later, to use storage\u0026rsquo;s import directory. Attention! Only root can use absolute paths.\nWhich lands awkwardly against the advice I gave last time, and still stand by: use a scoped API token, not root. That advice holds for every other stage here: discovery, SDN, building shells, setting boot order all work fine with a token. This one task does not, and no amount of privilege on the role will change it, because the restriction is on the user being root rather than on a permission.\nThere are three honest ways out, and no clever fourth:\nPVE 9.x: use \u0026lt;storage\u0026gt;:import/\u0026lt;file\u0026gt; and stay on the token. PVE 8.x: do this one task as root@pam, and only this one. PVE 8.x, no root over the API: run qm importdisk over SSH instead. The playbook takes the first two via a flag, because pretending otherwise would just move the problem to whoever runs it.\nFinally the boot order, which is an ordinary update and therefore untouched by the update_unsafe restriction:\n- name: Boot from the first imported disk delegate_to: localhost community.proxmox.proxmox_kvm: node: \u0026#34;{{ proxmox_node }}\u0026#34; vmid: \u0026#34;{{ pve_vmid }}\u0026#34; boot: \u0026#34;order=scsi0\u0026#34; update: true Nowt starts the guest. That is deliberate. Start it by hand, watch it come up, and only then think about deleting anything in VMware.\nRunning It The whole thing is one playbook, tagged by stage, because these are not steps you want to run together:\nansible-playbook migrate.yml --tags discover # look, change nothing ansible-playbook migrate.yml --tags sdn # build the VLANs ansible-playbook migrate.yml --tags build # build the diskless shells ansible-playbook migrate.yml --tags relocate # storage vMotion, live ansible-playbook migrate.yml --tags cutover # power off and import --limit is your friend throughout. Do one VM first. Do one cluster. The playbook has no opinion about how much you bite off, and the inventory gives you groups for free — power_poweredOn, cluster_\u0026lt;name\u0026gt;, plus vmware_windows and vmware_linux from the groups block.\nCheck what you are pointing at before you point at it:\nansible-inventory --graph ansible-inventory --host some-vm What I Would Still Watch For Things I expect to find when this meets a real estate, written down now so I cannot claim afterwards that I saw them coming:\nWindows guests will not boot cleanly off a VirtIO SCSI controller without the driver being present first. virtio-scsi-single is the right controller and the wrong one to hand a Windows VM that has never seen it. That is a whole problem of its own and it is not solved by anything above. VMware Tools should come off before the move, not after. Snapshots. A VM with a snapshot chain has more than one .vmdk per disk and importing the base gets you the state before the snapshot. Consolidate first. config.hardware.device ordering is what decides which disk becomes scsi0. It has matched the guest\u0026rsquo;s own ordering everywhere I have looked, but I would check it on a multi-disk database server before trusting it in a window. Independent and RDM disks will not storage-vMotion like ordinary ones. None of that changes the shape. Build the shells first, move the disks while everything is still running, and keep the outage to the one play that needs it.\n","permalink":"https://blogs.damiendye.uk/en/ansible/vmware-to-proxmox-ansible/","summary":"Reading a vSphere estate with the dynamic inventory, mirroring its VLANs into Proxmox SDN, and rebuilding every VM as a diskless shell before a single disk moves. Why every VMware info module hides the one field the migration needs, why vlan_id is three different types, and why one NFS export mounted twice turns the cutover into a local read.","title":"VMware to Proxmox with Ansible — Build the Shells Before You Move a Byte"},{"content":"I was a DNS registry system admin at Nominet, the .uk registry, from 2017 to 2019. What follows about ICANN is all public record and I have linked the lot. Where I am talking from the job instead, I say so.\nAsk most engineers who runs DNS and you get one of two answers. Either a shrug, or something about thirteen root servers. Both are wrong, and the second is wrong in a more interesting way, because it points at the machines instead of at the file.\nControl of DNS is not distributed across thirteen servers. It sits in one text file, and in the handful of outfits that decide what goes into it, edit it, sign it, and publish it. Everything else in the system — every resolver, every registrar, every zone you have ever run — is downstream of that file and takes its authority from it.\nThis post is about who holds that file, what the 2016 handover actually transferred, and what the people holding it have done with it. The machinery underneath — what a registry actually is, how names get into it, and who can take one away — is the follow-up.\nBefore ICANN, It Was a Phone Call None of the current arrangement is inevitable, and the history says what ICANN was actually built to fix.\nTo begin with there was no DNS at all. From 1972 there was one text file, HOSTS.TXT, holding every machine name on the ARPANET and the address it lived at. It was kept at Stanford Research Institute by Elizabeth Feinler and her team, and if you wanted your machine in it you rang the Network Information Center during office hours and asked. Everybody else fetched the file now and then and hoped it was current.\nThat does not scale, and by the early 1980s it plainly was not. DNS was built to replace it — a tree, delegated downwards, so that no single office had to hold the whole list.\nSomebody still had to hold the top of the tree. For years that somebody was one man. Jon Postel, at the University of Southern California, ran the name and number assignments on US government research money. IANA — the Internet Assigned Numbers Authority — was not an institution at that point. It was Postel and a handful of colleagues, and it worked because the people running the network trusted him.\nMoney arrived in 1993, when the National Science Foundation contracted InterNIC — Network Solutions among them — to handle registration. On 14 September 1995 free registration ended. Network Solutions charged $50 a year on a two-year minimum, and 30% of it went to a government fund that a court later ruled an illegal tax. One company, one price, nowhere else to go, and by 1997 an antitrust suit.\nThen in January 1998 Postel did the thing that tells you what the root\u0026rsquo;s authority is actually made of.\nHe emailed eight of the twelve root server operators, on nothing but his own standing, and asked them to point their servers at IANA\u0026rsquo;s machine instead of Network Solutions\u0026rsquo;. All eight did it. For about a week the authoritative root of the internet was wherever Jon Postel had asked people to look.\nHe called it a test. Plenty of people read it as a demonstration — that the root belonged to the engineers who built it rather than to a government contractor. The response settles which reading Washington took. Ira Magaziner, the presidential adviser on the matter, told Postel he would never work on the internet again. The test was reversed.\nICANN was incorporated in California that September. Postel died the following month.\nSo the arrangement this post is about was built to fix two real problems: a namespace being sold by an unaccountable monopolist, and a root whose authority rested on one man being widely trusted. Both were real problems. ICANN was the answer to them.\nTwo Codes That Outlived the Rule Two loose ends from that era before moving on, because between them they say more about how this system really works than anything in ICANN\u0026rsquo;s bylaws does.\nThe UK took the wrong code and kept it.\nRFC 920, in October 1984, said country top-level domains would be taken from the two-letter codes in ISO 3166. The United Kingdom\u0026rsquo;s ISO 3166 code is GB. By the rule as written, the UK\u0026rsquo;s domain should be .gb.\nIt is not, because the UK got there first. JANET, the academic network, had already settled on uk as its top-level identifier a few months before the ISO-derived list was drawn up, and .uk was registered on 24 July 1985. .gb was assigned as well, on the understanding that .uk would migrate across to it in time.\nThe migration never happened. Nobody made it happen. .gb then sat in the root for four decades having picked up, in its whole life, one second-level domain — hmg.gb, for Her Majesty\u0026rsquo;s Government — which was barely used. ISO eventually bent around the fact on the ground and exceptionally reserved UK at the United Kingdom\u0026rsquo;s request.\nNobody was robbed, incidentally. UK was not another country\u0026rsquo;s code — it is exceptionally reserved in ISO 3166 for the United Kingdom, at the United Kingdom\u0026rsquo;s own request, and no other state ever had a claim on it. That is exactly why nobody forced the issue: there was no injured party to complain.\nWhich is what makes it worth telling. The ISO 3166 rule is rigid for anybody trying to get in: no entry on the list, no country-code domain. That is why territories lobby to be added to ISO 3166 in the first place, and why places without recognition have no ccTLD at all. For an incumbent already in the root, the same rule turned out to be a suggestion.\nThere is even a fair argument the UK ended up with the better name. GB is Great Britain, which leaves out Northern Ireland. UK does not. The non-compliant code describes the state more accurately than the compliant one would have.\nAnd the Soviet Union is still in the root. This one is stranger.\n.su was delegated to the Soviet Union on 19 September 1990. The Soviet Union ceased to exist fifteen months later.\nIt is still there. Thirty-five years after the state it belongs to stopped existing, .su is live and taking registrations — something over 111,500 names as of May 2025, administered from Moscow.\nThe rule says ccTLDs come from ISO 3166. ISO 3166 does not list the Soviet Union. And it is not as though the rule was never applied: .dd for East Germany and .yu for Yugoslavia both went when those states did. .su was the one that did not, and nobody has ever been able to explain the difference in terms of the rule.\nWhat it became is predictable enough. When .ru tightened its registration checks in late 2011, the trade moved next door. Malicious sites in .su doubled in 2011 and doubled again in 2012, which is the same story as the cheap new gTLDs and the free ccTLDs later in this post: abuse is a fluid, and it flows to wherever the checks are weakest.\nTwo country codes, then, that outlived the rule that produced them. .uk because nobody made an incumbent move, .su because nobody made a delegation go when its country did. In both cases the rulebook is clear and in both cases it went unenforced, because enforcing it would have meant taking something away from somebody who already had it.\nThat is the whole character of authority in DNS, visible before ICANN existed, and as such none of what follows should surprise you.\nWhat that answer turned out to be is the rest of this post.\nThe Eighteen Years to the Handover ICANN was incorporated in California on 30 September 1998 and immediately signed a Memorandum of Understanding with the US Department of Commerce. It did not begin independent and grow American. It was American from the first day, by construction, and the arrangement was renewed in one form or another for eighteen years.\nStart with the thing it got right, because there is one and it matters.\nIn 1999 ICANN broke the registration monopoly. Network Solutions had been the only place to buy a .com; the Shared Registration System let other registrars sell names in the same zone, and the price came down and kept coming down. That is a real achievement, it is the reason a domain costs what it does today, and nothing later in this post cancels it.\nThen the pattern that runs through everything else starts.\n2000. ICANN held a global election in which internet users chose five board members directly. It was never repeated. The at-large structure that replaced it advises and does not vote. The first and last time the public got a binding say in ICANN, ICANN discontinued it.\n2005. At the World Summit on the Information Society in Tunis, a large part of the world objected to the US holding the root. What came out was the Internet Governance Forum — an annual conference with no authority over anything. The root arrangement did not change.\n2005. The .xxx affair, which ran for six years and is the sharpest single proof in this whole post that jurisdiction is not abstract. Socially conservative lobby groups in the United States pressed the Department of Commerce. The NTIA — the National Telecommunications and Information Administration, the arm of Commerce that held the agreement with ICANN — drafted letters to ICANN and, in its own words, marshalled its resources at ICANN. The board — which had been heading for approval — rejected the application, nine to five, then again eight to four, with the chairman and the CEO both reversing position. Viviane Reding, then the European commissioner responsible, called it the first clear case of political interference in ICANN by the US government. .xxx was finally approved in 2011, at which point the Department of Commerce announced it was disappointed.\nAn American domestic lobbying campaign, routed through an American federal agency, changed what top-level domains exist on the internet. No treaty, no court, no vote outside the United States.\n2009. The Joint Project Agreement with Commerce was replaced by the Affirmation of Commitments, widely written up as ICANN becoming independent. The IANA functions contract stayed exactly where it was.\nThen the thing that actually moved it, and it was nothing ICANN did.\n2013. Edward Snowden. On 7 October, months into the disclosures, the leaders of ICANN, the Internet Engineering Task Force, the Internet Architecture Board (IAB), the World Wide Web Consortium, the Internet Society and all five regional internet registries published the Montevideo Statement, calling for the globalisation of ICANN and the IANA functions and citing, in terms, the damage that pervasive surveillance had done to global trust. The internet\u0026rsquo;s own technical leadership — ICANN\u0026rsquo;s chief executive among the signatories — said out loud that American stewardship had become a liability.\n14 March 2014. NTIA announced its intent to transition its stewardship of the IANA functions.\n1 October 2016. The contract expired.\nSo the handover was not earned and it was not granted on the merits. It was conceded, eighteen years in, because a US intelligence scandal made the existing arrangement politically indefensible and the technical community said so in public.\nWhich is worth holding onto when you read what the transition actually did.\nWhat the 2016 Transition Changed, and What It Did Not You will hear that the Americans handed over the internet in 2016. Anyone arguing that ICANN is an instrument of US control gets told this, and if they have got their facts wrong they lose the argument there. So get them right.\nOn 1 October 2016 the IANA functions contract between NTIA and ICANN expired and was not renewed. That was real. The US government no longer holds a contract giving it approval over root zone changes, and new bylaws created an Empowered Community with the theoretical ability to reject budgets and remove board members.\nHere is what did not change, and it was not an oversight. It was written down as an objective.\nThe transition proposal stated that the legal jurisdiction in which ICANN resides was to remain unchanged. The new bylaws require ICANN to stay headquartered in California. The entire accountability structure built during the transition is built on California law. It works by making ICANN a Californian non-profit that Californian courts can be asked to hold to its own articles.\nSo after the great handover: ICANN is a California corporation, subject to US federal and Californian law, whose accountability mechanisms are enforceable in American courts and nowhere else, setting policy for a root zone edited and signed by an American company under an agreement with the American Department of Commerce.\nThe contract went. The jurisdiction was deliberately kept. And jurisdiction is the part that has teeth, because it does not require anyone to intervene. It applies automatically, all the time, by default.\nThe clearest demonstration is sanctions. OFAC — the Office of Foreign Assets Control, part of the US Treasury — runs American economic and trade sanctions, and decides who US persons and companies are allowed to do business with. ICANN is a California corporation, so OFAC binds it, and it constrains who ICANN may contract with and accredit.\nBe precise about the scope: that reaches generic top-level domain (gTLD) registries and registrars, because those hold ICANN contracts. It does not reach country-code (ccTLD) operations, which sit outside ICANN\u0026rsquo;s contractual structure entirely. But the effect leaks well past the legal boundary, because registrars outside the United States have applied OFAC restrictions to their own customers under the mistaken assumption that holding an ICANN contract requires it, or simply by copying American registrant agreements. American foreign policy propagates down the registrar chain by imitation as much as by law.\nThere is no version of this in which the answer to \u0026ldquo;who controls DNS\u0026rdquo; does not begin with the United States.\nWho that leaves able to do anything about an ICANN decision is the other half of the question, and it is better asked once there is a record to test it against. This post comes back to it at the end.\nThe Root Is a Text File So much for who is in charge. Here is the thing they are in charge of, and you can just fetch it. ICANN publishes the root zone by zone transfer — AXFR, the DNS mechanism for copying a whole zone rather than one record — to anyone who asks, no credentials:\ndig . AXFR @xfr.dns.icann.org Right now that is 1,578,790 bytes across 24,886 lines. It contains 1,439 delegations — every top-level domain that exists — of which 1,350 carry a DS record — the delegation signer, the fingerprint that ties a child zone\u0026rsquo;s signing key into its parent — and are therefore part of the signed chain.\nOne and a half megabytes. The entire namespace of the internet, small enough to email.\nThat file is the whole of the root\u0026rsquo;s authority. A resolver starting cold knows nothing except the addresses in its root hints, and the moment it gets an answer it is following delegations out of that file and nowhere else. Change a delegation in it and you have changed where an entire country\u0026rsquo;s traffic goes. There is no second copy with a different opinion, no consensus protocol, no vote at resolution time. There is the file.\nSo the question \u0026ldquo;who controls DNS\u0026rdquo; reduces to a much narrower one: who can change that file, and who signs it afterwards.\nThree Organisations Touch It The answer is a chain of three, and it is worth being clear about which does what, because the distinctions are where all the arguments live.\nPTI — Public Technical Identifiers, an ICANN affiliate — performs the IANA functions. It receives root zone change requests from TLD operators, checks them, and authorises them. This is the clerical layer, and deliberately so: the whole design intent is that IANA is a careful clerk with no discretion.\nVerisign is the Root Zone Maintainer. It takes the authorised change, edits the zone file, signs it with the root zone signing key, and publishes it for distribution. Verisign is an American public company, and it does this under a Cooperative Agreement with the US Department of Commerce.\nThe root server operators then serve it. They are the least powerful part of the chain and the only part anyone has heard of.\nNote where the discretion actually sits. Not with the operators. Not really with the clerk. It sits with whoever sets the policy that the clerk applies, which is ICANN, and with the company that holds the pen and the signing key, which is Verisign, under an agreement with the American government.\nTen of the Thirteen The root server letters are worth listing in full, because people cite the number thirteen as though it implies dispersal:\nLetter Operator Country A Verisign US B USC Information Sciences Institute US C Cogent Communications US D University of Maryland US E NASA Ames Research Center US F Internet Systems Consortium US G US Department of Defense (DISA) US H US Army Research Laboratory US I Netnod Sweden J Verisign US K RIPE NCC Netherlands L ICANN US M WIDE Project Japan Thirteen letters, twelve organisations, because Verisign holds both A and J. Ten of the thirteen are operated from the United States. Two of them are the American military.\nThat is not a conspiracy, it is fossilised history. These are the institutions that were on the network in the 1980s and never left. But an accident of history that leaves the US military running two of the internet\u0026rsquo;s root servers is still the US military running two of the internet\u0026rsquo;s root servers, and it is a strange thing to describe as a global system.\nThe operators also have no meaningful contract binding them. ICANN does not employ them and cannot in any straightforward way remove them. They serve the root because they always have. The system\u0026rsquo;s stability at this layer rests on goodwill and nowt sturdier, which works right up until the day it doesn\u0026rsquo;t.\nICANN Does Not Run the Root Zone This is the most important fact in the post and it is almost never said out loud, so it gets its own heading.\nICANN does not operate the root zone.\nIt decides what should go in it. It does not edit the file, it does not sign the file, and it does not serve the file. Verisign edits and signs. Twelve organisations serve. ICANN\u0026rsquo;s job is to say what the answer ought to be, and then get somebody else to make it so.\nThat split is the only thing keeping ICANN in check.\nPicture it without the split. One body sets the policy, holds the pen, owns the signing key and runs the servers. Between deciding a thing and that thing being true everywhere on earth, there is no other party, no second pair of hands, and nobody in a position to say no. Whatever you think of ICANN\u0026rsquo;s record below, that arrangement would be worse.\nAs it is, there are three brakes. Not one of them is in a bylaw.\nThe file is public. Anyone can pull the root zone over AXFR — the command is at the top of this post — and diff it against yesterday\u0026rsquo;s. You cannot change a delegation quietly. Somebody else has to make the change. The maintainer does the edit and the signing. That is one more organisation that has to agree to do it, and one more that could decline. The operators serve by consent. As above, ICANN has no meaningful contract with the root server operators. They distribute the zone because they always have. Nothing obliges them to distribute anything — and as Postel showed in 1998, consent is movable by somebody they trust asking them nicely. The last one is the real backstop, and it has been used one layer down within living memory. When Verisign wildcarded .com in 2003, the Internet Systems Consortium (ISC) shipped delegation-only in BIND and operators simply stopped honouring the answers. Nobody had to win an argument at a policy forum. The technical community\u0026rsquo;s ability to refuse is written down nowhere and everybody involved knows it is there.\nNow the uncomfortable part, because this is thinner than it sounds.\nNobody designed this check. It is not a separation of powers, it is an accident of how the work got divided up in the 1990s, and the 2016 transition neither strengthened it nor wrote it down. There is no rule saying the body that sets policy may not one day also hold the pen.\nAnd the party doing the checking is a commercial company that ICANN is doing business with. Verisign holds the root zone maintainer role, and the .com contract, and — as the rest of this post lays out — a $20 million agreement with ICANN signed in the same negotiation as a .com price rise. A check that depends on one party being willing to refuse the other stops working once the two of them are signing things together.\nSo the separation is the best thing about the current arrangement. It is also unwritten, unplanned, and held together by habit.\nSelling the Namespace Before any of the governance argument, something simpler shows what the people holding a piece of the namespace do with it when nothing stops them. It has happened repeatedly, at every layer, and the first time it happened at the top it lasted nineteen days.\nOn 15 September 2003 Verisign added a wildcard A record to the .com and .net zones:\n*.com. IN A 64.94.110.11 That address reverses to sitefinder.verisign.com. From that moment, every name in .com and .net existed. Every typo, every unregistered domain, every malformed string, every expired name — all of them resolved, to a Verisign search page carrying Verisign\u0026rsquo;s advertising.\nWhy This Is Not an Advertising Story The complaints at the time were mostly about the ads, and they missed the point. NXDOMAIN is not a user-experience feature. It is a load-bearing protocol signal, and an enormous amount of software above DNS is built on being able to ask \u0026ldquo;does this name exist?\u0026rdquo; and get a truthful answer.\nDelete the negative answer and things break in ways that have nothing to do with browsers.\nMail was the worst of it, and it is the part people still get wrong. Verisign did not publish a wildcard MX record. It did not need to. RFC 5321 §5.1 says that when an MX lookup returns nothing, the sender falls back to the domain\u0026rsquo;s address record and treats it as an implicit MX at preference 0. Verisign had just given every non-existent domain in .com an address record. So every MTA on the internet — every mail transfer agent, every machine that relays mail — following the standard correctly, now had a mail exchanger for soemcompany.com — and it was Verisign\u0026rsquo;s box.\nConnect to port 25 and it answered:\n220 snubby2-wceast Snubby Mail Rejector Daemon v1.3 ready Verisign\u0026rsquo;s stated intent was reasonable enough: reject the mail immediately so it did not sit in queues worldwide. The implementation was not. Snubby only gave up after the sending MTA had transmitted the message body, and returned a code that most MTAs read as a transient failure — so instead of an instant bounce, mail to mistyped addresses was retried for days before dying. Verisign later swapped it for a Postfix-based responder after operators complained on the NANOG list.\nAnti-spam broke at the same time, and more quietly. Checking whether a sender\u0026rsquo;s domain actually exists was, and still is, one of the cheapest and most effective filtering heuristics available. Overnight, every domain in the two largest TLDs existed. The check returned true for everything and stopped discriminating.\nAnd everything else that speaks DNS but not HTTP — mail relays, FTP clients, networked printers, monitoring systems — stopped getting \u0026ldquo;no such host\u0026rdquo; and started getting a web server, which mostly manifested as timeouts and hangs rather than clean failures. A dead name now looked like a broken service.\nOne company added one record to one zone file and changed the failure semantics of the internet.\nWhat Stopped It Not governance. Engineering, and then a threat.\nISC shipped a delegation-only feature in BIND within days, letting operators discard synthesised answers from TLD zones — the technical community routing around the registry rather than appealing to anyone. Plenty of ISPs deployed it.\nICANN asked Verisign to suspend the service. On 21 September Verisign refused. ICANN then demanded it on 3 October, with the contractual consequences made explicit, and Verisign pulled the records on 4 October 2003. The IAB published its architectural objection to registry wildcards, and ICANN\u0026rsquo;s own Security and Stability Advisory Committee reported on 9 July 2004 that the service should never have been deployed without review and that registries should phase wildcards out.\nThen Verisign sued ICANN, on 27 February 2004, arguing that ICANN had exceeded its authority by stopping it. The case was mostly dismissed that August, and the remainder settled on 1 March 2006 — a settlement that gave Verisign a new .com registry agreement.\nRead that sequence once more. The registry monetised the namespace it was contracted to operate, refused to stop, was forced to stop, sued the body that forced it, and came out of the settlement holding a renewed contract for the most valuable TLD in existence. It still holds it. It is the same company that today edits and signs the root zone.\nThen Everyone Else Did It Anyway Stopping the registry did not stop the idea, it just moved it one hop down. If the authoritative server will not lie about non-existence, the resolver will.\nFrom August 2006 Earthlink began redirecting NXDOMAIN responses to Barefruit, serving search pages and ads. Paxfire sold the same thing, and additionally redirected certain typed keywords to paying advertisers. Comcast\u0026rsquo;s \u0026ldquo;Domain Helper\u0026rdquo; did it at scale. In the UK, BT and Virgin Media both ran it. The economics are irresistible from an ISP\u0026rsquo;s side: mistyped domains are free inventory generated by your own customers\u0026rsquo; fingers.\nThe failure modes were worse than Verisign\u0026rsquo;s, because a resolver sees every query, not just one TLD. Barefruit\u0026rsquo;s implementation would hijack NXDOMAIN for private address space, breaking split-horizon lookups and VPN behaviour on corporate networks. Dan Kaminsky demonstrated cross-site scripting (XSS) against the redirect pages themselves, because now every non-existent hostname in the world resolved to attacker-reachable HTML served in a context the browser associated with somebody else\u0026rsquo;s domain. Monetising the error case had turned a failed lookup into an XSS surface.\nThe Protocol Fix Two things closed it off, and both are worth noting because they are the shape of every real fix in DNS: make the lie detectable, then make it contractual.\nDNSSEC provides authenticated denial of existence. NSEC and NSEC3 records let a signed zone prove that a name does not exist, and a validating resolver will reject a synthesised answer in its place. Non-existence stopped being the one answer nobody could verify. It is not airtight — a resolver that strips signatures on the way past can still rewrite the answer, which is exactly why validating on the client rather than trusting the resolver matters, and it is the argument this site has already made at length.\nAnd ICANN, to its credit, did learn this one. Specification 6 of the new gTLD registry agreement flatly prohibits wildcards, synthesised records and redirection for unregistered names, and requires authoritative servers to return Name Error, RCODE 3. Every one of the 1,200 strings from the 2012 round is contractually barred from doing what Verisign did to .com.\nThat is a real improvement, and it is worth being exact about what produced it: not the governance process, but nineteen days of visible breakage in 2003 that were embarrassing enough to be written into a contract a decade later.\nAnd None of It Touched the Country Codes Specification 6 binds gTLDs. It binds them because they sign a registry agreement with ICANN, and that agreement is the lever.\nA ccTLD signs nothing of the kind. No registry agreement, no Specification 6, no compliance function, no fee. Nothing in ICANN\u0026rsquo;s rulebook governs how a country-code registry operates its zone — which is why both of the things below were possible, and why nobody was in a position to stop them.\nThat is not the same as saying ICANN is absent, and I want to be precise about where it sits, because I spent two years on the receiving end of it.\nWhat ICANN holds over a ccTLD is the delegation itself. Every NS record, every piece of glue, every DS record and every contact change for .uk lives in the root zone, and the only route into the root zone is an IANA change request — verified against the registered administrative and technical contacts, and processed on IANA\u0026rsquo;s schedule rather than yours. Under RFC 1591 IANA also decides, in the last resort, who holds the delegation at all. Redelegations are rare. They are not hypothetical.\nSo a country-code registry is sovereign over how it runs, and completely dependent on a third party for anything that has to be visible in the root. The moments you most need a change to land — a nameserver moving, a key rollover whose DS has to be published before the old one goes — are exactly the moments you are waiting on somebody else\u0026rsquo;s queue. That is a live operational dependency rather than a governance abstraction, and it is a subject for the follow-up.\nNow notice what that combination produces. ICANN\u0026rsquo;s grip on a ccTLD is tight precisely where it inconveniences a registry that is behaving, and absent precisely where it might have restrained one that is not. It can hold up your DS record. It could not stop Cameroon pointing an entire top-level domain at an advertising page.\nSo the practice never stopped. It just moved somewhere the contract did not reach.\nCameroon wildcarded an entire top-level domain to farm typos.\nIn August 2006 the .cm registry pointed every unregistered name in the zone at a parking page of paid search links. There is nothing subtle about the play: .cm is .com with the o missed, so the target market was the fat-finger rate of the largest TLD in existence, and the operator was a government agency — ANTIC, under Cameroon\u0026rsquo;s Ministry of Posts and Telecommunications.\nIt paid well. NameJet reported over $500,000 of .cm sales on the first day and more than $2 million in the first week; hotels.cm went for $81,100 in 2009. It also did exactly what you would expect to the safety of the zone, because inbound typo traffic is the ideal delivery channel for a hostile download. In December 2009 McAfee rated .cm the riskiest TLD in the world, with 36.7% of its sites assessed as posing a risk.\nVerisign was forced to unwind the same trick in nineteen days. Cameroon ran it for years. The difference is not that one was worse. The difference is that one had signed a contract.\nTokelau became the largest country-code domain on earth by giving names away.\nTokelau is a New Zealand territory in the South Pacific with a population of about 1,500 people. Its ccTLD, .tk, was operated by Freenom, which gave registrations away for nothing. By 2016 it was the most-registered country-code domain in the world at 31,311,498 names — a figure that comes, as it happens, from a world map published by Nominet.\nFree was not free. Freenom\u0026rsquo;s terms required a free domain to carry regular traffic, and provided that if the redirect stopped working — or if the name started drawing visitors worth having — the registry could take it back and serve its own advertising on it. That is the whole business, and it is more elegant than Verisign\u0026rsquo;s. Do not try to guess which names are valuable. Give away the entire namespace at zero marginal cost, let the world discover the valuable ones for you, then repossess those and monetise the traffic. Roughly one sixth of Tokelau\u0026rsquo;s annual income came from it.\nThe externality landed on everybody else. Free registration with no verification is the ideal input to bulk abuse — the same economics as the cheap new gTLDs, taken all the way to zero. By the time Meta filed suit, Freenom\u0026rsquo;s five free ccTLDs — .tk, .ml, .ga, .cf, .gq — were the source of more than half of all new phishing domains coming out of country-code TLDs.\nWhat stopped it is the part that matters here.\nNot ICANN, which had no contract and no standing. Not Tokelau, which was collecting a sixth of its national income. Not New Zealand. Meta\u0026rsquo;s lawyers, in the Northern District of California, in March 2023, on cybersquatting and trademark claims.\nFreenom halted new registrations within days. Phishing originating from those extensions fell from over 60% to under 15%. Freenom settled in February 2024 and left the domain business, and by that March around 12.6 million domains — 99% of its portfolio — had stopped resolving.\nOne corporation\u0026rsquo;s legal department, in one American court, removed twelve and a half million names from the internet. No governance body in the history of DNS has ever exercised that much authority over the namespace, and it did not do it through governance.\nWhich is the third time in this post that the answer to \u0026ldquo;what actually enforces anything here\u0026rdquo; has turned out to be a court in California — and the second time that the enforcement was a private party acting in its own commercial interest, which happened on that occasion to coincide with everyone else\u0026rsquo;s.\nAnd .uk is a ccTLD too. Same absence of any contract governing how the zone is run, same absence of Specification 6, same freedom to wildcard it or give the namespace away. It did neither. It ran an abuse process instead.\nWhich is where the tidy division stops being tidy, and it is worth spoiling deliberately.\nNominet does not only run .uk. It runs generic top-level domains too — its own, and several dozen more on behalf of other operators — and for those it signs the ICANN registry agreement like anybody else, Specification 6 and continuous monitoring included. On its own platform the un-contracted zone was outnumbered by roughly thirty to one.\nSo ICANN\u0026rsquo;s requirements reached .uk anyway. Not by authority, which it did not have, but because nobody sanely runs two operational regimes side by side to preserve an exemption for one zone. You build the strict thing once and run everything on it.\nIt is the same shape as the OFAC problem earlier: ICANN\u0026rsquo;s formal reach stops at the contract, and its actual reach carries on past it, propagated by operators for whom complying everywhere is cheaper than maintaining the distinction. The set of registries effectively governed by ICANN is materially larger than the set that has signed anything.\nThat estate, what it was like to run, and what happened to Nominet afterwards is its own post.\nWhich leaves the question this post keeps arriving at from different directions: when the contracts do not reach, what actually stops a registry doing owt it likes?\nA follow-up will answer it from inside Nominet — the publication pipeline, EPP (the protocol registrars use to create and change names in a registry) and the economics underneath it, signing at registry scale, and who can really take a name away. This post is about the layer above it, and the layer above it does not come out well.\nThe Money The next charge is simpler and needs less interpretation.\nThe Product They Invented In 2012 ICANN opened applications for new generic top-level domains. Anyone could apply to run a new string to the right of the dot, for a non-refundable-in-large-part evaluation fee of $185,000.\nIt received 1,930 applications. That is over $350 million in evaluation fees, collected before a single string was delegated. Applicants who withdrew early got some of it back on a sliding scale; the great majority of it stayed.\nICANN described the fee as cost recovery.\nThen, where two applicants wanted the same string and would not settle privately, ICANN auctioned it between them and kept the proceeds, which came to another $240,590,128. So the programme charged you to apply, and charged you again to win.\nAsk the question that should have been asked in 2008: what problem was this solving?\nThe stated case was competition, choice and innovation. Fourteen years on, the results are measurable. Around 1,200 strings were delegated. As of August 2026 there are 1,112 new gTLDs holding about 48.7 million domains between them — against .com alone at more than ten times that. The incumbent monopoly was not disturbed in the slightest. It got a price rise instead.\nThe choice argument fails on its own evidence. 34% of the 2012 applications were for .brand strings — a company applying for its own trademark, largely so that nobody else could have it. Those are not new choices for anybody. Many were never used at all. McDonald\u0026rsquo;s never launched .mcdonalds. Intel took delivery of .intel in July 2016 and terminated it in November 2020, Symantec gave up .symantec two months earlier, and SC Johnson applied for eight strings — .scjohnson, .raid, .glade, .off, .duck among them — then terminated the lot in January 2022. Six years after the application window, more than one in ten new gTLDs had still not launched, 144 had not reached a sunrise period, and L\u0026rsquo;Oréal was sitting on strings it had never announced any plan for.\nThat is a lot of dead namespace. Here is why it does not bother ICANN.\nUnder the base registry agreement a gTLD operator pays ICANN a fixed fee of $25,000 a year, plus $0.25 per registration — but only once the TLD passes 50,000 transactions in a quarter. Below that threshold there is no transaction fee at all.\nRead what that means. ICANN\u0026rsquo;s income from a top-level domain with zero names in it is exactly the same as from one with forty thousand: $25,000 a year, every year, for a delegation nobody uses. A dead string is not a failure on ICANN\u0026rsquo;s books. It is an annuity with no support burden.\nThere was no financial reason for ICANN to care whether any of this worked, and it is hard to find evidence that it did.\nWhat the Internet Got Instead The programme did produce one measurable effect, and it is not the one in the prospectus.\nThe new strings that did sell, sold on price. Registries with no brand and no natural demand competed the only way available to them, at a dollar or less a name, in bulk, with minimal checks. That is a product, and it found its market.\nInterisle\u0026rsquo;s Cybercrime Supply Chain 2025 study found that new gTLDs carried 47% of reported cybercrime domains while making up 12% of the domain market — roughly a sixfold over-representation. The same study recorded 19.5 million unique domains used in attacks, up 126% year on year, with 7.3 million of them registered in bulk. The common factor it identifies in the most-abused domains is that they are cheap.\nICANN did not create phishing. But it manufactured 1,200 new places to do it from, priced the entry so that the only viable strategy for most of them was volume at near-zero cost, and took a fixed fee from each regardless of what came out.\nThe programme\u0026rsquo;s own showpiece failure makes the point better than any statistic. .sucks was delegated to Vox Populi, which charged trademark owners $2,499 a name during sunrise — a price set exactly because brands would have to pay it to stop somebody else. ICANN\u0026rsquo;s response was to report the registry to the US Federal Trade Commission for predatory pricing. The FTC found no rules had been broken, and observed that ICANN had already ignored several concerns the FTC had raised about the new gTLD programme.\nThat is the whole thing in one episode. ICANN designs the programme, ignores the regulator\u0026rsquo;s warnings about it, delegates the string, takes the fee, and then complains to the regulator about the predictable result.\nApplications for the next round opened in 2026. The fee is $227,000.\nThe Strings Too Dangerous to Delegate One more thing the programme produced, and this one is for anybody who has ever built an internal network.\nOrganisations have always invented top-level domains for internal use, on the assumption that a name which does not exist publicly never will. .corp. .home. .mail. .local. Pick something, put it in your Active Directory, nobody outside can see it.\nThe 2012 round proposed to delegate some of those for real, at which point every one of those private assumptions becomes a live security problem: internal names start resolving to somebody else\u0026rsquo;s servers, queries that used to fail start leaking your internal structure to a registry, and certificates issued for internal names become certificates for names a stranger now controls.\nNobody had checked. It only surfaced because researchers measured what was actually being asked of the root, and found that .home and .corp were among the most queried strings in existence — heavily used names that had never been delegated to anyone. ICANN\u0026rsquo;s own Security and Stability Advisory Committee raised it in 2013, after the applications were in.\n.corp, .home and .mail have never been delegated. They are still deferred, indefinitely, because delegating them would break too much. Three strings that were applied for and paid for turned out to be too dangerous to exist.\nThat is a programme that expanded the root without first establishing what the expansion would collide with, and found out afterwards from other people\u0026rsquo;s measurements. If you want the practical end of this, it is the reason making up an internal TLD is a bad idea — the namespace you invented is only private until somebody sells it.\nThe Auction Money Between June 2014 and July 2016 those contention auctions collected that $240,590,128, roughly $233 million after auction costs.\nThis is money obtained by selling pieces of a namespace ICANN does not own and holds in trust. There is a defensible answer to what should happen to it, and the community set up a cross-community working group to find one.\nWhile that working group was still sitting, ICANN\u0026rsquo;s board took $36 million of the proceeds and put it into ICANN\u0026rsquo;s own reserve fund, which was running $68 million short of its target. Not proposed — approved. And when the community objected, the position put to it was that the alternative was ICANN raising fees.\nThe working group carried on regardless. The board did not adopt its recommendations until June 2022 — six years after the last auction, during which ICANN held a quarter of a billion dollars of other people\u0026rsquo;s money and helped itself to $36 million of it while the people deciding what it was for were still in the room.\nThe .org Price Caps In March 2019 ICANN proposed renewing the .org registry agreement with the price caps removed. The caps were the thing that stopped the operator of .org charging whatever it liked to the charities, NGOs and non-profits that had been told for twenty years that .org was where they belonged.\nPublic comment ran. 3,252 comments opposed removal. Six supported it. The opposition included NPR, the YMCA, C-SPAN, the National Geographic Society, AARP and the National Trust for Historic Preservation — not the usual domain-industry commenters, but precisely the constituency .org exists for.\nOn 1 July 2019 ICANN signed the agreement. No public announcement. Comparison of the signed text against the proposed text showed no changes made in response to the comment period. Not \u0026ldquo;some concerns addressed\u0026rdquo; — the same document.\nIf a public comment process can run 542 to 1 against and change not one word, it is not a consultation. It is a formality that produces a paper trail.\nThen the Sale In November 2019, four months later, the Internet Society announced it was selling Public Interest Registry — the non-profit operator of .org — to Ethos Capital, a private equity firm, for $1.135 billion.\nThe price cap removal is what made PIR worth $1.135 billion. A registry that cannot raise prices is an annuity. A registry that can is a growth asset. ICANN had converted the second into the first four months earlier, against unanimous objection, and the market had immediately priced it.\nEthos Capital had been established in May 2019. The domain ethoscapital.com was registered on 8 May 2019 by Fadi Chehadé, ICANN\u0026rsquo;s former CEO — the week of the deadline for ICANN staff to publish their report on removing the price caps. His name appeared nowhere on Ethos Capital\u0026rsquo;s website when the deal was announced. His involvement became public because of WHOIS data — the public register of who owns a domain, which the next section is about — and the irony writes itself. Ethos then confirmed he had advised on the transaction, and in July 2020 he became its co-CEO.\nICANN did eventually block the sale, in April 2020. It is fair to record that. It is also fair to record what preceded it: months of ICANN insisting the matter was mostly outside its remit, sustained public campaigning, letters from US senators, and finally a letter from the Attorney General of California warning ICANN off the deal and citing the lack of transparency around Ethos Capital.\nICANN was not the safeguard here. ICANN removed the caps that created the opportunity, and was itself stopped, at the last minute, by a state law officer — which is one more demonstration that the real accountability mechanism in this system is Californian jurisdiction rather than anything in the bylaws.\nThe Verisign Arrangement This is the same counterparty as the wildcard, fifteen years on. In October 2018 NTIA and Verisign signed Amendment 35 to the Cooperative Agreement, lifting the .com price freeze and permitting increases of 7% a year in four years of every six.\nThat was the US government\u0026rsquo;s decision, not ICANN\u0026rsquo;s. But the increases still needed the .com registry agreement amended, and that is ICANN\u0026rsquo;s. In March 2020 ICANN agreed Amendment 3 — and alongside it a binding Letter of Intent under which Verisign pays ICANN $20 million over five years from 1 January 2021, for security and stability work.\nBoth things were negotiated before public comment opened. The comment period ran, was overwhelmingly hostile, and changed nothing — the same pattern as .org, in the same window, with the same result.\nTake the structure on its own terms. The body that decides whether a monopolist may raise prices negotiated, at the same time and with the same counterparty, a payment to itself. Public comment came afterwards and was decorative. Whatever the money is spent on, an arrangement in which the regulator is paid by the regulated in the same transaction as the price rise is one that no competent regulator would enter, and the word commenters reached for at the time — kickback — is the obvious one.\nWholesale .com has gone from $7.85 to $10.26 on the back of it, on a name with no technical need for a price rise and no competitor a registrant can move to.\nEight Countries Against One Company If you want a single episode that shows who ICANN actually answers to, it is .amazon, and it ran for seven years.\nAmazon the company applied for .amazon in the 2012 round. The Amazon Cooperation Treaty Organization objected — Bolivia, Brazil, Colombia, Ecuador, Guyana, Peru, Suriname and Venezuela, eight sovereign states whose territory the name describes and in which some 30 million people live. Their position was that a shared geographic and cultural name should not become one company\u0026rsquo;s private property.\nThey used the channel ICANN provides for governments. The Governmental Advisory Committee (GAC) issued consensus advice against the application, and in May 2014 ICANN\u0026rsquo;s board accepted it. The states had won, through the mechanism designed for exactly this.\nAmazon filed an Independent Review Process (IRP) claim.\nIn 2017 the IRP panel ruled for Amazon. It found the board had acted inconsistently with ICANN\u0026rsquo;s own bylaws, held that the board cannot treat GAC consensus advice as conclusive, directed it to re-evaluate the applications on the merits, and ordered ICANN to reimburse Amazon $163,045.51 in costs.\nIn May 2019 ICANN concluded there was no public policy reason for the applications not to proceed. Amazon got .amazon.\nRead the structure rather than the outcome, because the outcome is arguable and the structure is not.\nGovernments get the GAC, and the GAC advises. A corporation gets the Independent Review Process, and the IRP produces a binding declaration, a direction to the board, and a costs award. When those two channels met head-on, the panel\u0026rsquo;s finding was explicitly that the governmental one is not conclusive.\nSo ICANN does have a working accountability mechanism. It worked. It was used successfully by one of the largest companies on earth to overturn the collective objection of eight countries, and ICANN paid its legal costs for the privilege.\nThat is the same fact as the jurisdiction problem earlier in this post, wearing different clothes. The mechanisms are real, and they are shaped so that the parties who can afford to operate them are the ones who get results out of them. Eight governments could not make the advisory channel stick. One company made the legal channel work in three years.\nThe GDPR Fight: What ICANN Does When a Law Applies to It The strongest evidence about an institution is not its mission statement. It is what it does the first time a rule it did not write is enforced against it.\nFor ICANN that moment was GDPR, and the record is unambiguous.\nICANN did not lack warning. It had, by The Register\u0026rsquo;s count, more than a decade of letters telling it that WHOIS — publishing the name, postal address, email and phone number of every domain registrant, to anyone, with no access control — was incompatible with European data protection law. GDPR itself was adopted in 2016 with a two-year runway specifically so that organisations could prepare. ICANN arrived at May 2018 with no compliant model.\nWhat it did instead, in April 2018, was go to Brussels and ask the Article 29 Working Party for a one-year moratorium on enforcement, plus permission to keep publishing registrant email addresses in the meantime.\nSit with what that request actually is. Not an extension to file paperwork. A request that European regulators agree not to enforce a fundamental-rights regulation against one organisation and its global contracted parties, for a year, because that organisation had not got round to complying. There is no mechanism in GDPR to grant this. Data protection is a fundamental right under the Charter; no supervisory authority and not the European Data Protection Board has the power to suspend it for a single data controller. ICANN was not asking for a concession that was being withheld. It was asking for something that does not exist, having apparently not established whether it existed.\nWP29 refused both asks. ICANN\u0026rsquo;s own summary of the meeting conceded that registrant, administrative and technical contact email addresses must be anonymised, and simply omitted any mention of the moratorium it had requested.\nThen it sued.\nOn 25 May 2018, the day GDPR took effect, ICANN filed against EPAG — Tucows\u0026rsquo; German registrar — in Bonn. EPAG had decided to stop collecting Admin-C and Tech-C contact details, on the grounds that collecting personal data it had no use for was exactly what GDPR prohibits. Tucows\u0026rsquo; position was that in the overwhelming majority of registrations the registrant, admin and tech contacts are the same person anyway, so the collection was not merely unlawful, it was pointless.\nICANN\u0026rsquo;s legal theory was that GDPR\u0026rsquo;s \u0026ldquo;necessary for the performance of a contract\u0026rdquo; basis covered the collection, because ICANN\u0026rsquo;s own contract with the registrar required it. That is a remarkable argument: that an organisation can manufacture a lawful basis for processing other people\u0026rsquo;s personal data by writing a requirement to collect it into a contract with a third party. If it worked, Article 6(1)(b) would be a formality anyone could satisfy by drafting.\nIt did not work. ICANN lost in Bonn. It appealed. It lost again. In August 2018 the Cologne appellate court rejected it a third time, found the earlier rulings convincing, held there was no imminent emergency justifying an injunction, and — this is the part worth reading twice — refused ICANN\u0026rsquo;s request to refer the question to the Court of Justice of the European Union on the grounds that ICANN\u0026rsquo;s legal interpretation was not material to the decision. The court did not consider the argument close enough to be worth asking Luxembourg about.\nICANN spent members\u0026rsquo; money trying to push the case to that court regardless. In 2019 it gave up on WHOIS entirely.\nThe technical outcome was correct — WHOIS as it existed should not have existed. But look at how it was reached. Ten years of warnings ignored. A request for an exemption from a fundamental right. Litigation against its own contracted party, filed on the day the law took effect, to establish that the law did not apply the way the regulators said it applied. Three defeats. Then capitulation.\nThat is not an organisation that misread a statute. It is an organisation that did not accept, until courts told it three times, that the law was addressed to it at all.\nWhat It Cost Everyone Else Most of this post is about what ICANN and the registries did. This is about who paid for it, because it was very rarely them.\nEvery .com registrant on earth pays the increase. Wholesale .com went from $7.85 to $10.26. Verisign\u0026rsquo;s own reporting puts the .com base at 163.6 million names as of 31 March 2026. Multiply the two and that rise is worth on the order of $394 million a year, taken from every registrant of a .com anywhere in the world, for a name that needed no technical change to justify it. A business in Lagos or Manila pays the same increase as one in Palo Alto, decided by an American agency and an American non-profit that took $20 million from the beneficiary in the same negotiation.\nThe .org decision lands on charities worldwide. .org was sold to the non-profit sector for twenty years as the part of the namespace that belonged to them. Lifting the caps was a decision that whoever runs it may charge them what the traffic will bear. NPR and the YMCA can absorb that. A small NGO working on a grant cannot, and it was not asked either — 3,252 against, six in favour, signed without a word changed.\nThe abuse burden is carried by everyone who runs a mail server. New gTLDs are 12% of the market and 47% of reported cybercrime domains. Every one of those names arrives in somebody else\u0026rsquo;s inbox, somebody else\u0026rsquo;s abuse queue, somebody else\u0026rsquo;s fraud losses. ICANN collected $185,000 an application and collects $25,000 a year per string regardless. The cost of what the cheap strings then produced falls on every mail operator, every bank, every security team and every person taken in by a phishing page. That is an externality in the textbook sense, and the programme was designed without a line about who would bear it.\nSite Finder broke failure semantics for the entire planet at once. This is worth stating plainly because it is easy to read as an American story. There is one root and one .com. When Verisign changed what a non-existent name does, it changed it for every network on earth simultaneously — including every network with no relationship to Verisign, no say in the decision, and no way out except patching their own resolvers, which is what a great many of them ended up doing.\nWHOIS went dark on the people using it against abuse. Both harms here are real and both were avoidable. Publishing every registrant\u0026rsquo;s name, address, email and phone number to anyone who asked was a genuine, decades-long harm to people all over the world, and GDPR was right about it. But ICANN had more than ten years of warning and no plan, so May 2018 was not a managed move to tiered access. It was an abrupt shutdown. Anti-abuse researchers, security teams and law enforcement, including well outside Europe, lost a working tool overnight because ICANN spent the runway litigating instead of building the replacement. The exposure before and the vacuum after both belong to the same failure to prepare.\nAnd 12.6 million names stopped resolving. The Freenom collapse was a good outcome for the phishing figures. It was not a good outcome for everyone using a free .tk or .ml because they could not spare $10 a year, and there were a great many of those, disproportionately in places where $10 is not nothing. When the only free namespace on the internet is also the most abused, the people who lose it when it goes are not the criminals. They moved to the next cheap thing the week after.\nThe shape is consistent. The decisions are made in California, the revenue is collected in California, and the costs are distributed globally to people with no vote, no contract and no court they can reach.\nAnd Only in American Courts Now the record is in, come back to the jurisdiction point from the start of this post, because it does more work than the incorporation does.\nEverybody in the world is subject to what ICANN decides. The people who can do anything about it are the ones who can litigate in California.\nConsider who that shuts out. A ccTLD operator in Cameroon. A registrant in Tehran whose domain got dropped because a registrar over-applied OFAC that never bound it. A registrar in Bonn — which is exactly why the fight between ICANN and EPAG had to be run as a German case on German law, and not as an ICANN accountability matter at all. Any small registry that cannot fund American counsel to argue a point of Californian non-profit law against an organisation with a nine-figure budget.\nThen consider who it lets in. Every intervention in this post that actually changed something:\nSite Finder ended when ICANN threatened Verisign\u0026rsquo;s contract — and Verisign\u0026rsquo;s reply was to sue in a US court, and to walk out of the settlement holding a renewed .com. The .org sale was stopped after the Attorney General of California sent a letter. Freenom stopped because Meta sued in the Northern District of California, and 12.6 million names went dark behind it. .amazon went to Amazon because Amazon took ICANN through its own Independent Review Process and was awarded costs. Four interventions that worked. Four American actors — one state law officer and three corporations. Not one of them a route available to anybody outside the United States, and in three of the four the thing that moved was a private company\u0026rsquo;s commercial interest, which on those occasions happened to point the same way as everybody else\u0026rsquo;s.\nThe community did raise this. A jurisdiction subgroup of the accountability working group spent Work Stream 2 on it and produced recommendations that left things roughly where they found them, which is unsurprising given the transition proposal had already ruled the one change that mattered out of scope before anybody sat down.\nSo global multi-stakeholder governance resolves, in the only place it can be tested, to this: you may sue in California, if you can afford to.\nWhat the Record Actually Shows Put them together, because separately each has an excuse and together they do not.\nAn organisation that had to be told three times by German courts that European law applied to it, having first asked those courts\u0026rsquo; regulators to simply not enforce it.\nAn organisation that invented a product nobody had asked for, took more than $350 million in fees to consider applications for it, another $240 million auctioning the contested ones, and now collects $25,000 a year from top-level domains with nothing in them — while the strings that did sell became the cheapest place on the internet to buy a phishing domain.\nAn organisation whose public comment process has run 542 to 1 against and altered not one word of the document it was consulting on.\nAn organisation that lifted the price caps on the non-profit namespace four months before its former CEO\u0026rsquo;s private equity vehicle bid $1.135 billion for it, and that had to be stopped by a state Attorney General rather than by any mechanism of its own.\nAn organisation that took $20 million from the monopoly registrar in the same negotiation that let the monopoly registrar raise prices, and opened public comment afterwards.\nAn organisation that helped itself to $36 million of money it was holding in trust, while the group deciding what that money was for was still sitting.\nAnd an organisation that settled a lawsuit from the registry which had hijacked the two largest zones on the internet by handing that registry a renewed contract for .com.\nAn organisation whose one working accountability mechanism was used by a trillion-dollar company to overturn the unanimous objection of eight countries, with costs awarded against ICANN.\nAn organisation that came within a committee report of delegating .corp and .home to strangers, and found out what that would break from somebody else\u0026rsquo;s measurements.\nThe consistent thread is not incompetence. Incompetence is random. This is directional: every one of these went the way of the incumbents, the money, and ICANN\u0026rsquo;s own institutional interest, and the participation mechanisms — comment periods, working groups, empowered communities — produced documentation rather than outcomes.\nAnd this is the body that sets policy for the file. Not a standards body, not a court, not anything you elected. A California non-profit with a governance record like that, sitting on top of 1.5 megabytes of text that every network on earth resolves against.\nThe one thing standing between that record and the file is that ICANN does not hold the pen. It has to ask. Everything above is what an organisation does when it still has to ask — so the question worth carrying away is not whether ICANN behaves well. It plainly does not. It is what keeps the asking in place, given that nobody wrote it down and the party being asked is already on the payroll.\n","permalink":"https://blogs.damiendye.uk/en/dns/who-actually-controls-dns/","summary":"The root of the internet is a 1.5 MB text file that one American company edits and signs. Who really controls DNS, what the 2016 IANA transition did and did not change, and the documented record of how ICANN has used that control.","title":"Who Actually Controls DNS"},{"content":"I was a DNS registry system admin at Nominet, the .uk registry, from 2017 to 2019 — inside the building while the pressure that produced the 2021 revolt was building up. The vote itself came after I left, and that part is public record, linked as usual. Where I am talking about what it looked like from inside, I say so.\nThe companion post to this one is about ICANN, and it keeps arriving at the same question from different directions: when nobody holds a contract over a registry, what actually makes it behave?\n.uk is a good place to answer that, because nobody does. There is no ICANN registry agreement over it, no Specification 6, no compliance function, no external monitoring, and no regulator in the ordinary sense. Ofcom does not run it. The government does not run it.\nIts members do. And in March 2021 they used that.\nWhat Nominet Actually Is Nominet is a company limited by guarantee. It has no shareholders. It has members — registrars and other interested parties who pay a subscription and get a vote — and it was set up to run .uk for public benefit rather than for profit.\nThat structure is the entire story. A company with no owners to enrich still takes in money, and a registry running a national namespace with no competitor takes in a lot of it. What that surplus is for is a question the constitution answers vaguely and the board answers in practice.\nThe members are the only check. There is nowt else. Which is fine while the answers line up, and becomes the whole game when they stop.\nThe Estate, From Inside Here is the part people outside a registry rarely picture, and it matters for everything after.\nNominet was never just .uk. While I was there the platform carried:\n.uk, the country-code domain (ccTLD), under no ICANN contract at all. .cymru and .wales, generic top-level domains (gTLDs) Nominet holds in its own right — and those are under ICANN registry agreements. Other people\u0026rsquo;s gTLDs. In April 2016 Minds + Machines handed Nominet the back end for up to 28 of its strings — .london, .work, .law, .fashion, .cooking among them — which put Nominet in the top tier of registry operators by number of TLDs run. Add dot-brands like .bbc and .bentley. ICANN\u0026rsquo;s emergency role. Nominet is one of ICANN\u0026rsquo;s Emergency Back-end Registry Operators, the outfits ICANN hands a gTLD to when it takes it off whoever was running it. It joined in 2014, and in December 2017 ICANN used it — Nominet became emergency interim operator of .wed after its operator\u0026rsquo;s registration data service failed. The resolver end. Nominet built and ran Protective DNS for the National Cyber Security Centre (NCSC), the recursive resolver UK public sector bodies query, which declines to resolve names known to be malicious. Worth one clarification, because \u0026ldquo;Nominet runs .uk\u0026rdquo; is looser than it should be.\nNominet runs the .uk registry and most of what sits under it — .co.uk, .org.uk, .me.uk and second-level .uk itself. It has never run all of it. .ac.uk belongs to Jisc, successor to the academic network that named uk in the first place, which has administered and registered names under it since 1996. Through the years this post covers, .gov.uk was Jisc\u0026rsquo;s too — Nominet only took it over in 2024, and even now the approvals sit with the Central Digital and Data Office rather than the registry.\nIt goes further than that, and in a direction you would not guess. Jisc also provides the administration for .gov.scot, and for .gov.wales and .llyw.cymru — both of which live inside .wales and .cymru, the two generic top-level domains Nominet holds in its own right. Nominet runs those TLDs. Somebody else runs the governments\u0026rsquo; corner of them.\nSo even inside one country\u0026rsquo;s namespace the authority is split, and it is split by history and convention rather than by anybody\u0026rsquo;s design.\nSo one organisation sat in four different relationships to the name system at once. Outside ICANN\u0026rsquo;s reach entirely for .uk. Inside it, under contract and continuously measured, for the gTLDs. The instrument ICANN reached for when it needed to take a top-level domain off somebody else. And the resolver deciding what a government department was allowed to look up.\nNone of that is a contradiction. It is what the structure actually looks like once you stop reading org charts and start reading contracts. Authority here attaches to individual delegations, not to companies.\nTwo Regimes, One Platform Now the bit you only see from inside, and it is the reason the estate matters rather than being trivia.\nFor the gTLDs, Nominet signs the ICANN registry agreement like anybody else. Specification 6 applies — no wildcards, no synthesised answers. So does Specification 10, and ICANN measures compliance with it continuously from outside: DNS resolution, the shared registration system, the Whois service, escrow deposits and correctly signed zones are all monitored against thresholds, and falling through one is a compliance event rather than merely an outage.\nCount the zones. On one side, .uk, with no contract. On the other, several dozen gTLDs, every one under a registry agreement and a monitoring probe. The un-contracted zone was outnumbered on its own platform by something like thirty to one.\nTwo contractual regimes on one platform, and the strict one wins One platform, two contractual regimes Nominet — one registry platform, one set of runbooks .uk 1 zone no registry agreement no Specification 6 nobody monitoring generic top-level domains ≈30 zones .cymru\u0026#160;· .wales\u0026#160;· MMX\u0026#160;×28\u0026#160;· .bbc\u0026#160;· .bentley ICANN registry agreement Spec 6, Spec 10, monitored from outside the strict regime becomes the house standard Nobody maintains two operating standards to preserve an exemption for one zone. .uk gets ICANN's rules anyway\u0026#160;— no contract, no consultation, nobody in the UK asked. One platform carrying both regimes. The un-contracted zone is outnumbered about thirty to one, so the rules written for the gTLDs become the way everything is run — including the zone nobody has any authority over. Running two operational regimes on one platform is painful, so you do not do it. Two escrow processes, two monitoring regimes, two sets of runbooks, two on-call procedures, two answers to the same question depending on which zone the ticket happens to be about — that is how mistakes get made at three in the morning. When ICANN mandates something for the gTLDs, you do not build it twice. You build it once and run everything on it.\nWhich means ICANN\u0026rsquo;s requirements landed on .uk as well. Not because ICANN had any authority over .uk — it had none — but because the cheapest safe way to satisfy a rule binding most of your estate is to apply it to all of it, and because deliberately maintaining a split so the ccTLD could do things the gTLDs may not buys you nothing except a second way for the platform to break.\nNo contract was signed for that. No consultation happened. Nobody in the UK was asked. A requirement written in Los Angeles for generic top-level domains shaped how the country\u0026rsquo;s own registry ran, by way of a build decision.\nFor the record, that also made ICANN a real presence on the on-call rota rather than a line in a policy paper. For part of the estate it was a counterparty with a contract, a probe pointed at our infrastructure, and an escalation path.\nSelling .uk to Pay for the Rest The commercial logic of all this was the thing members eventually objected to.\n.uk is a monopoly. There is exactly one place to buy a .uk domain, the demand is close to inelastic, and the margin funds whatever the board decides to fund. From 2016 The Register was making the argument out loud that .uk registrants were being overcharged to subsidise the rest of the operation.\nUnder CEO Russell Haworth, appointed in 2015, Nominet pushed hard into cyber security — the NCSC contract among other things — on the reasoning that a registry sitting on national DNS infrastructure was well placed to sell security services.\nThat was the direction of travel for the whole of my time there. The revolt did not come out of nowhere in 2021. The conditions for it were laid down years earlier, in plain view of anyone working in the building.\nThe Charity Went First In January 2018, while I was there, Nominet withdrew from its own charitable foundation.\nThe Nominet Trust had been funded by the registry since 2008 — £44 million over that period, £4 million in 2016, £5.4 million in 2017. It gave money to technology-for-good projects. It was, in a fairly direct sense, the public benefit in a public benefit company. It went independent and in May 2018 became Social Tech Trust.\nHaworth\u0026rsquo;s stated reason was that \u0026ldquo;the grant-giving, single funder model we set up in 2008 was not the most effective route to greatest impact.\u0026rdquo;\nWhat was announced alongside it was a Cyber Advisory Panel — chaired by Haworth — targeting government and enterprise business, backed by a marketing programme and, in Nominet\u0026rsquo;s own words, potentially an acquisition.\nPut those two next to each other, because they were announced together. The money leaving the building for an arm\u0026rsquo;s-length charity stopped. A commercial venture chaired by the chief executive started. The members were consulted about neither.\nOne of them, Andrew Bennett, asked the obvious question at the time: where are all future operating profits going to be spent?\nThree years later the membership answered it.\nThis is also why the paragraphs below are not just my impression. The single largest movement of public benefit money in Nominet\u0026rsquo;s history went from an independent trust with its own governance to a panel the chief executive chaired, in one announcement, without asking the owners. Whatever you make of the intent, the direction of the money is on the record.\n\u0026ldquo;Profit With a Purpose\u0026rdquo; That was the phrase. It was the line for the whole strategy and we heard it a great deal.\nThe trouble with it was not that it was ambitious. It was that the two halves had come apart. The profit was real, growing, and came from a captive market that had nowhere else to buy a .uk. The purpose was a slide in a deck.\nProfit up, purpose withdrawn — the two halves of the slogan moving apart \"Profit with a purpose\", in the accounts Purpose — giving to the Nominet Trust £4.0m 2016 £5.4m 2017 withdrawn, January 2018 Trust cut loose to find its own funders 2018 onwards £44m given in total between 2008 and 2018 — then nothing. Over the same years the price of a .uk domain went up by more than 50%. The profit half of the slogan was working exactly as intended. The two halves of the slogan, moving in opposite directions. Giving figures from the Nominet Trust\u0026rsquo;s funding history; the price rise is the members\u0026rsquo; own complaint in 2021. You can hold a company to a slogan like that, and eventually the members did. Inside, it mostly produced the particular tiredness that comes of being told at every all-hands that the commercial push is the public benefit, while the actual public benefit line goes down.\nWhere the Money Went Three things ran while I was there, and none of them is DNS for the United Kingdom.\nA radio spectrum registry. Nominet built a TV White Space database — a register of which radio frequencies are free to use in a given place at a given time — and got itself approved by the FCC as a database administrator in the United States. The argument was that a registry is a registry, and a company good at one kind of lookup could sell another.\nA DNS security product. NTX, threat detection built on inspecting DNS traffic, sold to governments and enterprises. In other words a competitor to OpenDNS, which is to say a competitor to Cisco, entered by a British domain registry.\nDriverless cars. Along with drones and the internet of things, held up as the next great wave of things that would need naming and registering.\nWhat happened to them is the answer to whether they were a strategy. The spectrum business was sold to RED Technologies when Nominet refocused. The cyber work produced the technology behind the NCSC contract, which Nominet then lost in 2024. The driverless cars never arrived and neither did the domains they were going to need.\nBidding for Australia There was a fourth, and it says most about the ambition. In 2018 Nominet bid to run Australia\u0026rsquo;s registry.\nauDA, which administers .au, had put the registry operations out to tender. Nine bids came in from around the world, three were shortlisted, and Afilias took over on 1 July 2018, ending sixteen years of AusRegistry running it. auDA never published who else bid, so you will not find this in the record — but Nominet was in it, and I was there while we went for it.\nHold that against the same year\u0026rsquo;s other news. In January 2018 Nominet withdrew from the charitable foundation it had given £44 million, on the grounds that single-funder grant-giving was not the most effective route to impact. In the same twelve months it was bidding to run the domain registry of a country on the other side of the world.\nIt is also worth noticing what the .au tender is, because it is exactly what most people assume .uk must be. auDA can put its registry out to competitive tender and hand it to somebody else, and in 2018 it did. There is no equivalent for .uk. Nominet\u0026rsquo;s position is not a contract that comes up for renewal — which is a stronger position than Afilias won in Australia, and as such it is worth remembering when weighing how much pressure the members\u0026rsquo; vote actually represented. It was the only lever there was.\nAnd here is what was happening to auDA while Nominet was bidding to work for it.\nAlongside the tender, the Australian government was running a review of auDA itself. In April 2018 it reported that auDA\u0026rsquo;s management and governance framework was \u0026ldquo;no longer fit-for-purpose\u0026rdquo;, issued 29 required reforms as new terms of endorsement, and put a senior officer from the Department of Communications onto auDA\u0026rsquo;s board to watch the work. It also said, plainly, that it would transition the delegation for .au to another provider if auDA could not deliver.\nThat is a national government stating in writing that it will move its country\u0026rsquo;s domain if the body holding it does not sort itself out. auDA\u0026rsquo;s own members were mutinying at the same time — a petition for a special general meeting to remove four of its leadership, three years before Nominet\u0026rsquo;s members did the same thing to five of theirs.\nTwo national registries, two not-for-profits holding a country\u0026rsquo;s namespace, both accused of governance failure inside four years of each other. It is not a Nominet quirk. It is what this structure does when nobody is watching it closely enough.\nWhat the Members Saw As to why those and not others — I will be careful, because I can tell you what it looked like and not what was in anybody\u0026rsquo;s head. The view widely held among staff at the time was that funding tracked the enthusiasms of the people approving it. I cannot show you a ledger and I am not going to pretend I can.\nWhat I can point at is that three years later the membership\u0026rsquo;s formal complaint was a more evidenced version of the same suspicion — that a body with no shareholders and a public benefit purpose was spending the surplus from a national monopoly on things that suited the people running it, and that the giving which was supposed to justify the whole arrangement had been cut while it happened.\nWhere the .uk surplus went, and how each one ended Where the surplus went .uk one buyer, no rival prices up 50%+ Nominet Trust £44m of giving, 2008 to 2018 stopped, Jan 2018 TV White Space spectrum database FCC-approved in the United States sold off NTX cyber security £12.6m revenue in its last full year £2.4m loss Bid to run Australia's registry nine bidders, three shortlisted lost to Afilias Driverless cars, drones, the internet of things the next great wave of things needing names never arrived The one line that was doing what the company existed to do is the one that got cut. Everything below it was paid for by the people buying .uk domains. Five destinations for the money from a monopoly. Four of them ended in a sale, a loss, a lost bid or nothing at all — and the fifth, the one the constitution existed to fund, was the one that got stopped. The members\u0026rsquo; complaint was not that diversification is wrong. It was arithmetic. .uk prices rose more than 50 per cent. Charitable and public benefit giving fell, while the organisation was making monopoly margins on a national asset. Operating performance was going backwards despite cost-cutting. Executive pay and bonuses went up through all of it. And member feedback, offered through the channels the constitution provides, went nowhere for years.\nA company with no shareholders had started behaving like one with impatient shareholders, and the surplus from the national namespace was paying for it.\nThe Members Revolt The campaign was PublicBenefit.uk, organised by Simon Blackler of the hosting company Krystal. Its resolution was blunt: remove named directors.\nNominet\u0026rsquo;s response is the part worth recording, because it tells you what the organisation had become.\nIt campaigned against its own members using the members\u0026rsquo; own money — emails, phone calls and mailings urging rejection. It refused to engage with the campaign\u0026rsquo;s substance. When the campaign exercised its right to members\u0026rsquo; contact details in order to make its case, Nominet did not send a spreadsheet. It sent a physical package with the details printed across more than 500 sheets of paper, with the email addresses left out.\nIt also blocked a second resolution that would have installed two qualified caretaker directors, on the argument that it was not legal — and then criticised the campaign for having no succession plan.\nOn 22 March 2021 it went to a vote. Turnout was 53%. The resolution carried with 52.7%, and five of the eleven board members went:\nMark Wood Chairman Russell Haworth Chief Executive Eleanor Bradley Managing Director, Registry Ben Hill Chief Financial Officer Jane Tozer Non-executive Director Haworth resigned hours before the vote rather than lose it. Rob Binns became acting chair.\nHere is what I think that pair amounted to, and I am flagging it as opinion because that is what it is.\nWood and Haworth ran Nominet the way a venture capital firm runs a portfolio company. Not as a registry that happens to produce a surplus, but as a balance sheet with an underworked asset bolted to it — a captive monopoly throwing off cash that could be put to better use somewhere with more upside.\nEvery documented thing above is consistent with that. You take the reliable income and raise its price, because the customers have nowhere else to go. You stop the outgoing that produces no return, which is the charity. You put the difference into a portfolio — spectrum, cyber, a foreign registry, driverless cars — on the theory that one of them comes good. You pay the people directing it at the rate that sort of work commands. And you keep going until somebody with the standing to stop you does.\nThere is nothing unusual about running a company that way. It is no way to run a public benefit body holding a national asset in trust, because that surplus was never capital in search of a return. It was the thing the whole arrangement existed to produce, and the constitution said so.\nThe members eventually said the same, in the only language that was available to them.\nIt is worth being honest about the margin. 52.7% on a 53% turnout is not a landslide, it is a narrow win on a split membership, and the organisation fought it with every resource it had. It still lost.\nWhat Happened Next Nominet pulled back from the commercial cyber direction and turned towards the registry and public benefit work. Then the numbers arrived.\nThe core business peaked the year I left. .uk domains under management topped out at 13,348,378 in 2019. By January 2023 that was 11,045,559. By January 2024, 10,688,932. A fifth of the base, gone, and still falling.\nThe diversification lost money. In its last full year before the reckoning the cyber business unit turned over £12.6 million and posted a £2.4 million loss. That is the answer to whether the projects the surplus paid for were an investment or an indulgence, and it is Nominet\u0026rsquo;s own figure.\nThen the contract went. In 2024 the NCSC retendered Protective DNS — a service handling around half a trillion queries a year — and Nominet lost it. The work went to Cloudflare with Accenture from September 2024, on a deal reported at about £30 million. Chief executive Paul Fletcher said the government had picked a cheaper competitor.\nAnd the backend book churns. In April 2021, weeks after the EGM, MMX sold its portfolio to GoDaddy Registry for $120 million — and the 28 strings that had put Nominet in the top tier of registry operators went with the buyer. .blog had already left for CentralNic in 2019.\nIt is only fair to say this runs both ways. Amazon moved the bulk of its 54 gTLDs onto Nominet\u0026rsquo;s platform in 2019, and Microsoft later shifted .skype and .office across from GoDaddy. Nominet wins portfolios as well as losing them.\nBut that is the point about the business, not a defence of it. Backend registry services is revenue you do not control. It arrives and departs on somebody else\u0026rsquo;s corporate transaction — MMX did not leave because Nominet ran the platform badly, it left because MMX was sold. Building the finances of a public benefit body on a book of business that can walk out on a Tuesday because its owner took $120 million is a strategic choice, and it was made.\nAnd in the same year, it won .gov.uk.\n.gov.uk had been run by Jisc for years, pro bono, under a legacy memorandum of understanding — an arrangement the government eventually concluded was not meeting internationally recognised standards. The Central Digital and Data Office ran a procurement through Crown Commercial Service, Nominet won it in November 2023, and the transition completed on 26 June 2024, timed a week before the general election to keep voter registration out of harm\u0026rsquo;s way. Jisc\u0026rsquo;s own registry page now records the date flatly — as of 26 June 2024 it no longer manages the .gov.uk namespace, though it stayed on as a registrar for some customers. Twenty-eight years running a national namespace, ending in a line on a web page.\nThe stated requirements were resilience, compliance with ICANN\u0026rsquo;s DNS standards, and meeting the NCSC\u0026rsquo;s Cyber Assessment Framework.\nRead that against the argument earlier in this post. The reason Nominet was a credible bidder for the government\u0026rsquo;s own namespace is that it already ran to ICANN\u0026rsquo;s standards — standards it had adopted because most of its estate was contractually bound by them, and which reached .uk because nobody maintains two regimes on one platform. The discipline that arrived sideways, through the gTLD contracts, is what qualified it to run .gov.uk.\nSo 2024 was not simply a bad year. It lost the biggest contract it had and won the one with the government\u0026rsquo;s name on it.\nTwo things about that win are worth noticing, because they are the difference between operating a namespace and holding it.\nGovernment kept the authority. Nominet runs the registry. The Central Digital and Data Office still manages and approves the applications — who is allowed a .gov.uk remains a government decision, not the registry\u0026rsquo;s. That is not how .co.uk works, where accredited registrars sell to whoever turns up. The technical operation was outsourced. The say over the namespace was not.\nAnd it is a contract. .gov.uk was procured through Crown Commercial Service, which means it has a term and an end date, and the government has already demonstrated exactly once what it does when it decides the arrangement is not good enough: it moved the whole thing off Jisc after twenty-odd years.\nSo Nominet holds .gov.uk on terms it does not hold .uk on. One can be taken back at the end of a contract by an official deciding not to renew. The other has no contract, no term and no renewal — and the only people who have ever managed to discipline it had to organise a members\u0026rsquo; vote to do it.\nIn March 2024, with PDNS gone and .uk shrinking, Fletcher announced a restructure with up to 70 roles at risk.\nAnd then the line that closes the circle. Fletcher noted that domain pricing \u0026ldquo;cannot be held at the level set in January 2020 indefinitely.\u0026rdquo;\n.uk price rises were among the things the members revolted over. Five years on, with the diversification wound down at a loss, the flagship contract lost to a cheaper bidder and the registrations falling, the answer being prepared is to put the price of .uk up again.\nAnd the View From Inside Has Not Recovered The members got their board change. Whether they got the organisation back is a separate question.\nAs of 25 August 2026 Nominet\u0026rsquo;s Glassdoor rating stands at 2.9 out of 5 across 97 reviews, with 31% saying they would recommend it, 22% positive about the business outlook and 29% approving of the chief executive. The lowest-scoring category is senior management, at 2.3. Reviewers rate their colleagues and the work highly. What they do not rate is the layer above them.\nThat is a mediocre score rather than a damning one, and it should be read as what it is: a self-selecting sample on a public site. But the shape of the complaint is recognisably the one the members made in 2021 — strategy nobody can explain, reward flowing upward, and the people running the national registry not being asked.\nDifferent chief executive. Same complaint.\nWhat It Says About Who Governs a ccTLD Come back to the question at the top.\nICANN could not have done any of this. It had no contract over .uk, no standing, and no mechanism beyond writing a letter. Everything in the ICANN post about Specification 6, compliance and monitoring applied to Nominet\u0026rsquo;s gTLDs and not to the zone that actually matters to the country.\nThe UK government did not do it either, and it is worth being clear about why, because the common assumption is wrong. .uk is not awarded on a government contract. There is no tender, no renewal date and no competitor waiting to bid for it. Nominet holds the delegation from IANA, the same way every other country-code registry holds theirs, and nobody hands it out every few years.\nWhat the government does have is a reserve power, and almost nobody knows it is there.\nSections 19 to 21 of the Digital Economy Act 2010 let the Secretary of State act where there is a \u0026ldquo;serious relevant failure\u0026rdquo; at a qualifying internet domain registry — a failure adversely affecting the availability or reputation of UK communications, or the interests of consumers or the public. Nominet is that registry. After notice and a chance to make representations, the Secretary of State may appoint a manager over the registry, or apply to court to alter its constitution.\nThat is a good deal more than ICANN has ever had over any ccTLD. And here is the part worth sitting with: those two sections sat dormant for fourteen years, and were commenced on 6 April 2024 — the same year Nominet lost the NCSC contract and announced 70 redundancies.\nI am not going to claim those facts are connected, because I do not know that they are. What is on the record is that the power to put a manager into the .uk registry became live in 2024, and had not been for the fourteen years before it.\nAnd that is only the domestic route. There is a second one, and it is the reason no ccTLD operator anywhere holds its delegation by right.\nA country-code delegation can be moved. IANA redelegates ccTLDs, under RFC 1591, ICP-1 and the GAC Principles, and it has done so repeatedly — .kz, .iq, .za, .gd, .gw and others. The GAC Principles hold that each government carries the ultimate responsibility within its own territory for national public policy, and in practice IANA treats the view of the recognised government as a major consideration in any transfer of its country\u0026rsquo;s domain.\nAustralia, as above, put that in writing in 2018 — reform or the delegation moves.\nSo a UK government that decided .uk needed to be somewhere else would not be blocked. It would not be instant, it is not a power exercised by announcement, and the local internet community would be consulted. But the machinery exists, the precedents exist, and the government\u0026rsquo;s opinion is the heaviest single input to it.\nWhich puts .uk in a position worth stating plainly. Communications is one of the UK\u0026rsquo;s critical national infrastructure sectors, and the state treats DNS accordingly — the NCSC has been buying protective DNS for the public sector for years. A registry running critical national infrastructure, whose government can appoint a manager over it domestically and whose word would carry the day in a redelegation internationally, does not hold its delegation as property. It holds it on sufferance, and the sufferance is conditional on not embarrassing anybody.\nWhat disciplined it was a membership vote, narrowly, six years into the problem, after a campaign run by one hosting company that had to be handed its own members\u0026rsquo; details on 500 sheets of paper to make its case.\nThat is a better accountability mechanism than anything ICANN has. It removed the chief executive and the chairman of a national registry, which no ICANN process has ever done to anybody. It is also slow, adversarial, dependent on someone deciding to spend a year of their life on it, and it came within a couple of percentage points of failing.\nWhich is roughly where DNS governance is everywhere. The mechanisms that work are the ones somebody with resources chooses to operate. .cm wildcarded a whole top-level domain for years because nobody with standing objected. Freenom stopped only when Meta sued. .uk changed course because a hosting company organised a vote.\nNone of that is governance in the sense the word implies. It is who happened to turn up.\nA Free Thought to Finish Everything above is either sourced or marked as my own experience. This last part is neither. It is what I think, and you are free to disagree with it.\nI started at Nominet in 2017, which is nine years ago. Russell Haworth took over in 2015, which is eleven. That is long enough for an organisation to have learned something.\nHas it?\nOn paper, yes. The board that was voted out is gone. The commercial cyber ambitions are wound down. The charitable arm went independent and is still going under its own name. The members used the one mechanism they had and it worked.\nThen look at what is actually in front of you. .uk is a fifth smaller than it was in 2019 and still shrinking. The contract that the diversification eventually produced has gone to a cheaper bidder. Seventy roles went with it. The chief executive is briefing that .uk pricing cannot be held at 2020 levels indefinitely — which is the same lever that helped start the trouble in the first place. And the staff rate senior management 2.3 out of 5, with 29% approving of the chief executive. A different chief executive, and recognisably the same complaint.\nSo my honest answer is that the people were changed and I am not convinced the organisation was.\nIt is worth noting where Haworth went afterwards, because it is not a criticism and that is exactly why it matters. Before Nominet he had spent fourteen years at Thomson Reuters, in financial data. Afterwards he went to NBS, then Byggfakta, then Acclaro — subscription businesses with professional customers, recurring revenue and pricing power. NBS turned over £45m at £23.6m EBITDA in 2023. Those are good numbers, and getting them is what a commercial chief executive is for.\nWhich is rather the point, and you do not have to take my word for the characterisation. His own professional profile describes him as a chief executive for private equity backed B2B businesses. That is the job he does, he says so himself, and he is clearly good at it.\nHe is a commercial operator and he did commercial things. The error was never his temperament — it was putting that temperament in charge of an organisation whose surplus was not capital and was never supposed to be, and then leaving nobody in a position to check it for six years.\nAlthough NBS is worth a second look, because the shape of it is familiar.\nNBS sells Chorus, the specification tool most UK architectural practices work in. It is wired into Revit and ISO 19650 workflows, and there is no serious alternative — architects describe it as a vital tool they simply have to keep paying for. One practitioner\u0026rsquo;s licence went from £1,385 in 2015 to £7,350, which the Architects\u0026rsquo; Journal reported as a 400% rise over a decade. There is no small practice rate, so a sole practitioner pays the same per seat as a 200-person firm. The words in the trade press are \u0026ldquo;rip-off increases\u0026rdquo; and \u0026ldquo;well above inflation\u0026rdquo;, alongside the observation that there is little alternative to paying.\nBe careful with the attribution, because the dates do not line up the way you might want them to. Chorus launched in 2018. RIBA sold NBS between 2018 and 2020 for around £172 million, and a former RIBA president has said it should never have ceded control — a sentence with a certain echo to it. Haworth ran the business from October 2021 to October 2024. The sharpest increases reported, and the loudest complaints, are from 2024 and 2025, after he had left, and the chief executive publicly defending them now is somebody else.\nSo this is not evidence about a man. It is evidence about a shape.\nA professional body sells the tool its members cannot work without. The buyer discovers the customers cannot leave. The price rises well beyond inflation, year after year, and the complaints go nowhere because there is nowhere for them to go. That is .uk, and it is what .org nearly became when ICANN lifted the price caps four months before a private equity vehicle bid $1.135 billion for it.\nWhich is the actual lesson, and it is not a comfortable one for anybody hoping this was about one executive. Captive professional customers, an essential tool, no alternative supplier: that arrangement produces the same outcome whoever is running it, unless summat is standing in the way. At Nominet the members eventually were. At NBS there is nobody — the professional body that might have played that part sold its stake and walked away with £172 million.\nI do not think that is really about any individual, either. Put anybody in charge of a member-owned monopoly running a national asset with no regulator watching, and give them a surplus with no obvious owner, and the pull is always going to be towards spending it on something more interesting than the thing that earns it. Running .uk well is not a story you can tell at a conference. Bidding for Australia is.\nThe correction, when it comes, has to come from members — which means somebody has to give up a year of their life to organise a vote against an incumbent spending those same members\u0026rsquo; money to resist it. That happened once, and it carried by 2.7 percentage points on a 53% turnout. Nobody would design an accountability mechanism that way.\nThe two real backstops sat unused throughout. The Digital Economy Act powers only came into force in 2024. A redelegation has never been seriously suggested by anyone.\nNine years on, the thing I would like to know, and cannot tell from outside, is whether somebody at Nominet today could say out loud in a meeting — we are a non-profit, should we really be spending money on this? — and still be working there a year later.\nIf the answer is yes, it has learned something. If it is no, then all that changed in 2021 was the names.\n","permalink":"https://blogs.damiendye.uk/en/dns/what-happened-at-nominet/","summary":"The .uk registry is owned by its members, and in March 2021 they voted half the board out. What the estate actually looked like from inside, why running .uk alongside dozens of gTLDs shaped how it was run, and how a registry with no regulator ended up disciplined by the only people who could.","title":"What Happened at Nominet"},{"content":"Part 1 of 8. This series is written plain, with short lines and everyday words.\nWhat this post covers What changed in how you buy technology. Why it changed. Why leaving a supplier got harder. Why open source did not put a stop to it. Where the whole thing has ended up. The change Twenty years back, you bought software.\nYou paid once. That copy was yours. You could keep using it as long as it worked.\nNow most software is rented. You pay every year just to carry on using it.\nStop paying, and it stops working. You own nowt.\nThat is what a subscription is. A subscription is a payment you make over and over to keep hold of a service.\nThe same thing happened to computing itself. You used to buy servers. Now a lot of outfits rent their computing from a cloud provider.\nA cloud provider is a company that runs the computers for you, in its own buildings.\nWhy it changed Let us be fair about this, because it matters. It changed because renting was often the better deal.\nBuying servers costs a lot up front. Renting costs very little up front.\nCloud services were reliable and quick to set up. Small teams got tools that only big companies could afford before.\nSubscriptions brought regular updates too. No more planning a great upheaval every few years.\nSo most outfits moved, and they had good reason. Nobody was forced.\nThe question this series asks comes later. It is about what happens once going elsewhere has become difficult.\nWhy leaving got harder Here is the part that matters.\nEvery step in this change shifted where it was hard to leave.\nAt first it sat in the file format. A file format is the way a program saves your work. If only one program could open your files, you kept buying that program.\nOpen file formats sorted that. Most documents today open in more than one program.\nSo the hard part moved. It went to the interface.\nAn interface is the way one piece of software talks to another. Build all your systems round one supplier\u0026rsquo;s interface, and changing supplier means rewriting the lot.\nThen it moved again, to your data.\nShifting a small amount of data is easy enough. Shifting years of it is slow, and some providers charge you for taking it out.\nEach move made the software itself matter less. As such, what counts now is how much work it would take you to leave.\nWhere the difficulty of leaving has sat, over time Then Later Now Your data File format Interface The program Your data File format Interface The program Your data File format Interface The program The blue box is the part that is hard to change. Each time one layer was opened up, the difficulty moved to the next. The same layers, 3 times. The hard-to-change part has climbed: file format first, then the interface, then the data. Why open source did not put a stop to it Open source software is software whose code anyone can read, use and change.\nPlenty of people expected it to solve this. It solved a real part of it, but not the part you would think.\nOpen source won. Most of the internet runs on it. The code for most of the important pieces is there to read.\nThe hard part of leaving did not go away, though, because it had already moved on.\nYou can read every line of a database. You still cannot easily shift 5 years of data out of the managed service running it.\nA managed service is where a provider runs the software for you, so you do not have to.\nThe code is open. The service is not.\nThere is a second point here, and it is a fair one.\nA lot of open source is kept going by volunteers who are not paid. Big companies build products on that work.\nIn 2024 a widely used compression tool was found to have harmful code hidden in it. That tool was maintained by 1 unpaid person. Bruce Schneier and others wrote it up in detail.\nThe lesson is not that open source is risky. It is that shared work everyone leans on needs paying for, and often is not.\nWhere it has ended up Put those steps together and you get where we are now.\nA small number of companies provide the services most outfits depend on.\nMost of those companies sit in 1 country, the United States.\nResearch gathered by the EuroStack project puts non-European providers at around 85% of the European cloud market.\nNo single decision did this. It is the result of a great many sensible choices, made one at a time, over 20 years.\nThat is what makes it a trend and not an event.\nWhat it means for you None of this says American software is poor. Much of it is very good, which is exactly how it got everywhere.\nIt does mean 1 question matters more than it used to.\nHow long would it take you to move to a different supplier?\nThat number decides a fair bit, and part 2 explains why.\nFirst published: 2026-08-25. Last updated: 2026-08-25.\n","permalink":"https://blogs.damiendye.uk/en/random/technology-you-rent/","summary":"Twenty years back you bought software and the copy was yours. Now you rent it, and the supplier sets the terms. Part 1 of 8 walks through how that happened, and why every step made it harder to go elsewhere.","title":"How Technology Became Something You Rent"},{"content":"Part 2 of 8. Part 1 covered how renting replaced buying.\nWhat this post covers How the price gets set once leaving is hard. Where the profit is taxed. Whose law applies to your data. How trade decisions reach your kit. What happens if a service simply stops. 1. The price follows your exit cost Start with the one most organisations have already had.\nIn 2023 Broadcom bought VMware. VMware makes the software used to run many virtual servers on 1 physical server.\nBroadcom then stopped selling permanent licences. Customers had to shift to subscriptions.\nPrices went up sharply for a lot of them. AT\u0026amp;T said in court papers its costs were set to rise by about 1,050%.\nCISPE is a trade body for European cloud providers. It has complained to the European Commission about the changes, using figures of a similar size.\nIn March 2026 CISPE complained again, after the European partner programme was shut. The Register reported what providers made of it.\nNothing unlawful has been proven here. The complaints are still being looked at.\nBut the pattern is plain enough to learn from.\nA renewal is a negotiation. Your position in it rests on 1 thing: how easily you could walk.\nIf walking would take you 3 years, you have very little room to say no.\nThis is not just about American suppliers, mind. Any supplier in that position has the same advantage.\n2. Where the profit is taxed The second cost is harder to see.\nPlenty of large technology companies sell to customers in 1 country and book the profit in another.\nThe local company is often treated as a service provider to its parent. Most of the profit goes elsewhere.\nSo you get big sales in a country and a small tax bill in that country.\nTaxWatch is a research group in the United Kingdom. It reckoned that 7 large technology groups made close to £15 billion of profit from UK customers in 2021. It estimated their arrangements knocked around £2 billion off the UK corporation tax due.\nEurope has taken these arrangements on. The results are mixed, and it is only fair to show both sides.\nApple lost. In September 2024 the Court of Justice of the European Union confirmed that Irish tax arrangements were unlawful state aid. Ireland got back around €14 billion. Amazon won. In December 2023 the same court threw out the European Commission\u0026rsquo;s appeal in a similar case about Luxembourg. There is a global agreement as well, called Pillar Two. It sets a minimum tax rate of 15%.\nIn January 2026 the Organisation for Economic Co-operation and Development published a side-by-side arrangement. Under it, companies headquartered in the United States follow United States rules instead of most of the global ones.\nSome countries tried their own digital services taxes. Canada passed one, then dropped it in June 2025 after the United States called off trade talks over it.\nSo the money is earned in one place. Where it gets taxed is settled somewhere else.\n3. Whose law applies This one gets misunderstood more than any other, so it is worth being exact.\nPlenty of people believe that keeping data in Europe keeps it under European law and nothing else.\nThat is not quite right. What counts is who controls the company holding it.\nThe CLOUD Act is a United States law from 2018. It requires a United States provider to hand over data it controls, wherever that data sits. The Department of Justice sets out the detail.\nSo a data centre in Frankfurt, run by a company owned in the United States, can still be reached by a United States legal order.\n\u0026ldquo;Our data stays in Europe\u0026rdquo; and \u0026ldquo;our data is outside United States law\u0026rdquo; are 2 different statements. Careful suppliers only make the first.\nA legal order follows who controls the company, not where the building is Data centre in Europe Your data is stored here The site follows EU law Parent company Based in another country Controls the data operates the site A legal order arrives here The order does not need to go to the building. It goes to whoever controls the data. The building is in Europe. The company running it is owned elsewhere. A legal order goes to the owner, not the building. The rules here are shifting too, in both directions.\nSection 702 is a United States surveillance law. It ran out on 2026-06-12 when Congress did not renew it.\nThat does not mean collection stopped. Approvals already granted carry on until they expire, expected to be around March 2027. The Brennan Center keeps track.\nThere is also the EU–US Data Privacy Framework. It lets personal data move from Europe to the United States.\nA legal challenge to it was thrown out in September 2025. That decision is now under appeal.\nThe framework is good law today. It is also the third of its kind, because the 2 before it were struck down.\nIf a transfer arrangement fails, the cost lands on the European organisation using it. In 2023 the Irish Data Protection Commission fined Meta €1.2 billion over transfers to the United States.\nThere is a calm, practical answer to all this, and it is worth knowing.\nIf your provider holds the keys to your data, your provider can answer a legal order.\nHold the keys yourself and the order has to come to you instead. As such, the decision sits under your own country\u0026rsquo;s law.\n4. Trade decisions reach your kit Technology is part of trade policy now.\nIn January 2026 the United States put a 25% tariff on a narrow group of advanced computer chips and the products containing them. A second stage has been signalled.\nWhether your own kit is caught changes over time. Your supplier is the right one to ask.\nTrade relationships can turn quickly, too. Talks between the United States and Canada broke down on 2026-08-22, with new tariffs of 50%. Canada announced measures in reply.\nThe Canadian Prime Minister set out the reasons publicly.\nThe point for you is a narrow one, and it does not rest on anybody\u0026rsquo;s politics.\nThe cost and availability of kit can change because of a decision taken in another country, at short notice.\nWorth allowing for in a 3-year plan.\n5. A service can stop for reasons that are not yours This last one is unlikely for most organisations. It is here because it works differently from the rest.\nIn February 2025 the United States put sanctions on the Prosecutor of the International Criminal Court.\nThe Prosecutor then lost the use of his Microsoft email account, and moved to a Swiss provider.\nMicrosoft\u0026rsquo;s president said publicly that the company did not close the account. Exactly what happened is still disputed, and it is only fair to say so.\nWhat is not disputed is the effect. The German technology publication heise reported it as a turning point for digital sovereignty in Europe. It was raised in the European Parliament.\nBy late 2025 the Court had moved to openDesk, an open source alternative.\nIt is the mechanism that matters here.\nNo unpaid bill. No rule broken. Nowt wrong with the service.\nA government took a decision, and a supplier had to work out what it could lawfully carry on providing.\nYour agreement is with your supplier. Your supplier\u0026rsquo;s legal duties are to its own government.\nYou cannot monitor for this. There is no warning light on a dashboard.\nFor most readers the risk is small, and accepting it is reasonable. It is only reasonable if you have thought about it.\nThe pattern across all 5 These 5 costs are very different from one another.\nYou will most likely meet the first. You may never meet the last.\nBut they share 1 thing.\nEvery one of them gets worse the harder it is for you to leave.\nA price rise is a nuisance if you have somewhere to go. It is serious if you have not.\nThat is the thread running through the lot, and it points somewhere useful.\nThe question worth working on is not which country your supplier sits in. It is how quickly you could change your mind.\nPart 3 looks at who else pays for the software you use.\nFirst published: 2026-08-25. Last updated: 2026-08-25.\n","permalink":"https://blogs.damiendye.uk/en/random/what-renting-costs/","summary":"Once moving supplier is hard, the costs change shape. The price follows your exit cost. Profit is declared in one country and earned in another. The law follows who owns the company, not where the building stands. Part 2 of 8, in plain language.","title":"What Renting Your Technology Costs"},{"content":"Part 3 of 8. Part 2 set out what dependency costs you.\nWhat this post covers The years when open source was fought. Why the fight stopped. Who keeps the software going today. What happened when the makers tried to charge. What happens to companies after they are bought. What it means for you. The years when open source was fought Open source is software whose code anyone can read, use and change.\nIt is ordinary now. For about 15 years it was treated as a threat.\nIn October 1998 an internal Microsoft memo leaked. Eric Raymond published it with notes, and it became known as the Halloween documents.\nThe memo was honest about the quality of open source. It called the way thousands of people work together on it remarkable.\nThen it set out what to do about it. One line explains a great deal:\nBy extending these protocols and developing new protocols, we can deny OSS projects entry into the market.\nA protocol is an agreed way for 2 systems to talk to each other.\nSo the plan was not to build a better product. The plan was to change the joins between products, so a competitor could not be dropped in.\nSeveral other steps followed over the next 10 years.\nIn 2003 a company called SCO claimed Linux contained code it owned. The case dragged on for years and SCO did not win. While it ran, plenty of organisations were unsure whether Linux was safe to take on.\nMicrosoft also signed patent agreements with mobile phone makers. For several years it took a payment on every handset sold running Android, an operating system it had not written.\nIn 2008 Microsoft\u0026rsquo;s document formats were approved as an international standard. Several national standards bodies objected to how that was handled.\nEurope pushed back on some of it. The European Commission fined Microsoft €497 million in 2004 over information competitors needed to work with its products.\nIt fined the company €561 million again in 2013, because an agreed remedy had not been put in place.\nThat second fine is worth a moment. The problem was not a new one. The agreed fix simply had not been done.\nWhy the fight stopped It stopped around 2014, because it had not worked.\nLinux had become the software most servers run on. It also runs most cloud services and most mobile phones.\nYou cannot take something out through the courts once everything is built on top of it.\nSo the approach turned right round.\nMicrosoft joined the Linux Foundation. It released some of its own tools as open source. In 2018 it bought GitHub, the site where much of the world\u0026rsquo;s open source gets written.\nAmazon and Google built very large businesses on open source, and all 3 companies now put a great deal of work back into it.\nThat work is real, and the software is better for it. It would be daft to pretend otherwise.\nBut 1 thing did not change.\nThe companies that once tried to slow open source down are now among its biggest funders. They are also still the biggest companies in the market.\nOpen source won the technical argument. It did not change who holds the strongest hand.\nWho keeps the software going today Here is the bit that surprises folk outside software.\nA great deal of widely used open source is kept going by very small teams. Some of it by 1 person, in their own time, for nowt.\nThose same pieces then sit inside products sold by very large companies.\nIn December 2021 a fault turned up in Log4j, a small tool Java programs use to record what they are doing.\nThe fault let attackers run their own code on affected systems. It hit an enormous number of organisations at once. The UK National Cyber Security Centre put out guidance on it.\nThe tool was maintained by a small group of volunteers.\nThe lesson is not that open source is risky. Most of it is very good, and open to inspection in a way closed software never is.\nThe lesson is about paying for things. Shared work that many businesses lean on needs funding, and often does not get it.\nThis one is fixable, and some organisations do fix it. Paying a maintainer, or funding a foundation, is usually a very small cost next to what the software saves you.\nWhat happened when the makers tried to charge Some companies build open source and sell a paid service round it. That worked well enough for a long while.\nIt got harder when cloud providers started offering the same software as a service of their own.\nThe provider took the revenue. The company that wrote the software did not.\nSeveral of them responded by changing their licence, so others could not offer their software as a competing service.\nMongoDB changed its licence in 2018. Elastic followed in 2021. HashiCorp changed in 2023. Redis changed in 2024. The Open Source Initiative sets the accepted definition of open source. It ruled that 1 of these new licences did not meet it.\nPlenty of Linux distributions then dropped the affected software.\nThe wider community took copies of the last open versions and carried them on separately. A copy like that is called a fork.\nOpenSearch carries on the earlier Elasticsearch code. OpenTofu carries on the earlier Terraform code. Valkey carries on the earlier Redis code. Elastic and Redis have both since gone back to open licences.\nIt is worth looking at how that ended, because nobody involved got what they were after.\nA company built useful software. A bigger company made more from it than the maker did. The maker restricted the licence to survive. The community objected, rightly, that this was no longer open source. A fork appeared, often backed by the bigger companies.\nAt the end of it the software is still free to use. The forks are mostly steered by the biggest players. The business that paid for the original work is weaker than when it set off.\nWhat happens to companies after they are bought The second way value leaves a business has nowt to do with software licences.\nIt is about how a company gets bought.\nHere is the pattern, plainly.\nA buyer borrows most of the purchase price. The loan is then secured against the company being bought. So the company ends up carrying the debt used to buy it.\nThe buyer may then sell the company\u0026rsquo;s buildings and rent them back. The money raised can be paid out to the new owners. The company now pays rent on premises it used to own.\nCosts that do not show up this year get cut. That usually means maintenance, staff numbers, research and product development.\nThen the company is sold on.\nThe Bank of England has looked at the financial stability side of this. Its review of private equity notes that heavy borrowing on buyouts makes those companies more likely to default, and leaves their lenders exposed to losses. It has kept watching the sector since.\nNot every buyout works this way, and plenty of owners invest rather than strip. This is a pattern to recognise, not a description of all buyers.\nBut where it does happen, the outcome is the same every time.\nThe business was usually working fine. What did not work was the business plus the debt taken on to buy it.\nThe same business before and after a buyout funded by borrowing Before After The work Same staff, same customers Owns its buildings No debt from being bought Funds maintenance Funds research Rent paid: none The work Same staff, same customers Rents the same buildings Carries the loan used to buy it Maintenance reduced Research reduced Rent paid: every month bought with borrowed money The business still does the same work for the same customers. What changed is what it owns, and what it now has to pay out. The same business, before and after. The work and the customers stay put. The buildings and the borrowing change hands, and the running costs go up. The same pattern in software This reached software some years back.\nThe asset being worked is not a building. It is the customers who cannot easily leave.\nA mature product with long-standing customers can be repriced. Those customers cannot move quickly, and as such most of them pay.\nAt the same time, spending on engineering can be cut. That takes years to show, and by then the sale has gone through.\nPart 2 covered what this looked like for VMware customers.\nThere is a version funded by investors rather than debt, too. It has been described often enough to earn a name: enshittification, a word used by the Canadian writer Cory Doctorow from 2022.\nIt runs in 3 stages.\nThe service is sold below what it costs to run, paid for by investors. It is cheap and good, so people move to it. Once people cannot easily leave, it is changed to suit the paying business customers instead. Once both sides are committed, terms change again to raise profit. Stage 1 is the one that matters for this series.\nA company selling below cost for years is not winning because it is better. It is being funded to take the market.\nSmaller businesses that have to cover their own costs cannot match that price. Plenty shut.\nWhen prices later climb to something sustainable, the choice that used to exist has often gone.\nWhat it means for you Two practical things come out of this, and both are easy enough to do.\nCheck how many people keep your dependencies going.\nLook at the tools your systems could not run without. Find out how many people actively maintain each one.\nIf the answer is 1 or 2, that is worth knowing. You might want to fund that work, keep your own copy of the source, or plan what you would do if it stopped.\nWatch who owns your suppliers.\nYour terms follow the owner, not the product. A change of ownership can shift your renewal more than any change in the software.\nWhen you sign, check what you keep if you stop paying. Check whether you can still run what you have already installed, and whether you still get security updates.\nNeither takes long. Both are far easier before a renewal than during one.\nPart 4 looks at what courts and regulators have already decided.\nFirst published: 2026-08-25. Last updated: 2026-08-25.\n","permalink":"https://blogs.damiendye.uk/en/random/who-pays-for-the-software/","summary":"The first 2 posts looked at what dependency costs you. This one looks at who else foots the bill: the volunteers keeping software going for very large companies, and the businesses bought, borrowed against and stripped back. Part 3 of 8, in plain language.","title":"Who Pays for the Software You Use"},{"content":"Part 4 of 8. Part 3 looked at who pays for the software you use.\nWhat this post covers Why the record is worth reading. What has been decided about Google. What has been decided about Apple. What has been decided about Amazon. What has been decided about Meta. The pattern across the lot. Why the record is worth reading Picking a supplier is a prediction. You are deciding how they will behave in 3 or 5 years.\nCompany values statements are not much use for that. Everyone has one, and they all say much the same.\nFindings of fact are better. These are conclusions a court or a regulator reached after hearing evidence.\nThat is what this post uses. Where a European decision exists, it comes first.\nTwo things before the list, because they keep it honest.\nNot every case went against the companies. Some went against the regulators, and those are in here too.\nAnd this is not a claim that European companies behave better. Volkswagen and Wirecard settle that well enough. The pattern follows from size and market position, not from where a company comes from.\nGoogle The European Commission got there first, and the cases took a long time to finish.\nIn 2017 it fined Google €2.42 billion. It found Google had pushed its own shopping comparison service up the search results and rivals down.\nGoogle appealed, and the Court of Justice of the European Union upheld the decision in September 2024.\nThat is 7 years from the fine, and longer still from the behaviour.\nIn 2018 the Commission fined Google €4.34 billion over the conditions placed on phone makers who wanted the Play Store. The fine was later cut to about €4.1 billion, and the final appeal was dismissed in 2026.\nIn 2019 the Commission fined Google €1.49 billion over advertising contracts. The General Court annulled that one in 2024, finding the case had not been made. The regulator lost it.\nIn the United States, a court found in 2024 that Google had unlawfully held on to a monopoly in search. In September 2025 the same court set out the remedies. It required changes to default agreements and some data sharing with competitors. It declined to make Google sell Chrome or Android.\nA separate United States ruling in April 2025 found Google had unlawfully tied 2 of its advertising products together. That remedy is still being decided.\nApple In 2021 a United States court ordered Apple to let app developers tell customers about other ways to pay.\nIn April 2025 the same court found Apple had not complied.\nIt also found that a senior Apple executive had given evidence that was untrue, and that the company\u0026rsquo;s own internal documents said otherwise. CNBC reported the ruling at the time.\nThe court referred the matter to prosecutors to consider criminal contempt proceedings. Apple said it would appeal.\nThis one matters for a different reason from the rest.\nA company can take a firm view of its legal position and still deal straight with a court. A finding that evidence was untrue is another thing altogether.\nYour relationship with a supplier is a set of promises. This is direct evidence of how promises get treated once keeping them turns expensive.\nAmazon In September 2025 Amazon agreed to pay $2.5 billion to settle a case brought by the United States Federal Trade Commission. The Commission published the detail.\nIt was $1 billion in penalties and $1.5 billion back to customers.\nThe case was about the design of Prime sign-up and cancellation. The allegation was that the screens were built to enrol people who had not chosen to enrol, and to make cancelling hard work. The Commission said about 35 million people were affected.\nAmazon agreed to change the sign-up and cancellation process as part of the settlement.\nInterface designs at that scale get tested and measured before release — that is normal, sensible practice. As such, the effect of a design is usually known before it ships.\nMeta In 2019 the United States Federal Trade Commission imposed a $5 billion penalty on Facebook over privacy practices, after the Cambridge Analytica case.\nThe bigger item is not a fine.\nIn 2021 a former employee handed internal company research to journalists, then gave evidence to a United States Senate committee. The BBC reported her central claim, which was that the company had put profit ahead of safety, again and again.\nSome of that research looked at the effect of Instagram on teenage users. Parts of it came back worrying.\nMeta disagreed with how the research was described. It said the findings were more mixed than reported, and that the former employee had not worked on those teams. The wider science on this is not settled, and it is only fair to say so.\nOne point is not in dispute, and it is the one to keep.\nThe company studied whether its product was harming a group of users. Some of that work came back worrying. The results reached the public because somebody carried them out of the building.\nResearch that only gets out that way is nowt like independent oversight.\nThe pattern across the lot Put the cases side by side and 4 things stand out. These matter more for planning than any single case.\n1. The penalties are small next to the gain.\nA few billion is a small figure against revenues in the hundreds of billions. It also lands years after the money was made.\nWhen a penalty comes to less than the benefit, it works like a cost of doing business rather than a deterrent.\n2. The remedies come late.\nThe gap between the behaviour and an enforceable remedy runs from about 7 to 12 years in these cases.\nCompetitors rarely last that long. By the time a remedy bites, the market it was meant to protect has usually already changed shape.\n3. Individuals are rarely touched.\nFines are paid by the company — so by its shareholders — while the people who approved the decisions usually keep their pay and their jobs.\nThat matters because it shapes the incentive. Where the cost falls on the company and the reward falls on the individual, the decision looks worth making to the person making it.\n4. Problems are known inside before they are known outside.\nIn several of these cases the outfit had the information first. It reached the public through a leak, a lawsuit or a regulator instead.\nBe careful what you take from that, mind.\nThese are not usually cases of one person setting out to do harm. A team improved a sign-up screen against a target. A ranking system was tuned to increase use. A research finding went unpublished.\nThe decisions are spread across a lot of people, and each step looks reasonable on its own. That is what makes the pattern so steady, and why asking companies to try harder is unlikely to shift it.\nWhat to do with this Use it for the question it actually answers.\nWhere a supplier has a commercial interest and you have no real alternative, the record says the outcome tends to go the supplier\u0026rsquo;s way.\nAny correction usually arrives years later, and usually costs the supplier less than the behaviour earned.\nThat is not a reason to avoid a supplier over the country it sits in.\nIt is a good reason not to depend on any supplier you could not leave.\nPart 5 looks at how one country\u0026rsquo;s law can reach companies in another.\nFirst published: 2026-08-25. Last updated: 2026-08-25.\n","permalink":"https://blogs.damiendye.uk/en/random/what-the-record-shows/","summary":"Picking a supplier means guessing how they will behave later. The fairest basis for that guess is what courts and regulators have already decided. Part 4 of 8 sets out the findings, includes the cases regulators lost, and draws the pattern.","title":"What the Record Shows"},{"content":"Part 5 of 8. Part 4 looked at what courts and regulators have decided.\nWhat this post covers Law that reaches companies in other countries. What France made of it. Older worries about commercial information. Pressure applied through trade. Assumptions that travel with a product. Law that reaches companies in other countries Part 2 explained how one country\u0026rsquo;s law can reach data held by a company it controls.\nThe same principle reaches companies themselves, and the effects can be a good deal bigger.\nThe best documented example in Europe is Alstom, a French engineering company.\nIn 2014 Alstom pleaded guilty in the United States and agreed to pay $772 million. The case came under the Foreign Corrupt Practices Act, an American anti-bribery law. The United States Department of Justice published the detail.\nThe bribery was real. This is not a tale about an innocent company.\nAn Alstom executive was also arrested while travelling through the United States, and spent time in prison.\nAround the same period Alstom sold its power business to General Electric, an American company.\nThe executive has since argued that the unresolved case worked as pressure during that sale. United States prosecutors have denied acting to help an American buyer.\nThat disagreement has never been settled, and as such this post cannot settle it either.\nWhat France made of it What can be shown is what the French state concluded afterwards.\nIn June 2019 a report went to the French Prime Minister. It is known as the Gauvain report, and its title is about restoring French and European sovereignty and protecting companies from laws with extraterritorial reach.\nExtraterritorial means a law that applies beyond the borders of the country that made it.\nThe report found 3 things worth repeating here.\nVery large penalties, running to tens of billions of dollars, had been imposed on French, European and other non-American companies. Much of the conduct had little direct connection to United States territory. French companies had no effective legal tools to defend themselves. It also noted that American companies were rarely the target.\nFrance then updated its blocking statute, a law limiting what information French companies may hand to foreign authorities. An earlier parliamentary inquiry had looked at the same subject in 2016.\nYou do not have to swallow every conclusion in that report. It is enough that a national parliament looked at the question and decided it was a matter of sovereignty.\nFor your own planning the point is short.\nLegal exposure to another country is a commercial exposure. Not just a compliance one.\nOlder worries about commercial information This worry is not new, and it is worth knowing how far back it runs.\nIn 2001 the European Parliament published a report on a global communications interception system, then known as ECHELON. The Parliament has since published a study revisiting that work.\nThe report concluded such a system existed. It also examined claims that intercepted information had been used for commercial advantage, including European companies.\nThose claims were contested at the time and remain hard to prove.\nThe reason to mention it is not to settle them. It is that the European Parliament took the question seriously enough to run a formal inquiry 25 years back, and it has not gone away since.\nPressure applied through trade Part 2 covered tariffs on kit, and Canada dropping its digital services tax after trade talks were called off.\nThere is one more step worth recording on its own, because it applies to people rather than goods.\nIn January 2026 the United States State Department imposed visa restrictions on 5 European officials. They had worked on the Digital Markets Act and the Digital Services Act, the European laws governing large online platforms.\nIt came alongside tariff threats connected to European enforcement action.\nThe Centre for Strategic and International Studies has described this as using trade measures to discourage digital regulation.\nThe Istituto Affari Internazionali, an Italian institute, has looked at whether enforcement of the Digital Markets Act is becoming negotiable as a result.\nThat is the open question. If European technology rules can be shifted by trade pressure, the protection they give you is less certain than the text suggests.\nIt is only fair to add that trade pressure is a normal tool of statecraft, used by plenty of countries including European ones. The point here is what it is being used on.\nAssumptions that travel with a product The last item is quieter than the rest — no court case, no fine — and it affects you every day.\nSoftware gets built for the market its makers know best. That market\u0026rsquo;s assumptions travel with it.\nSome you will recognise straight away.\nPrivacy settings often default to sharing, because that is the common approach in the United States. European law starts from asking permission. Staff management tools often assume employment can end at short notice, and that there is no works council to consult. Standard contract terms often want disputes heard in a United States court, under United States law, even for a European customer. Content rules follow 1 country\u0026rsquo;s approach to free expression, then get applied worldwide. Support for smaller languages turns up last, and sometimes not at all. None of this is done to cause bother — it is the ordinary result of building for a home market and selling the result everywhere else.\nIt does explain something about the wider argument, though.\nThe General Data Protection Regulation, the Digital Markets Act and the Digital Services Act are mostly Europe stating that it is a separate legal area with its own settled rules.\nThe answer to that has included the trade measures above.\nWhat it means for you Three practical points fall out of this.\nRead the governing law clause. Check whose courts would hear a dispute. On a large contract that is worth negotiating.\nAsk where your supplier\u0026rsquo;s parent company is registered. That answer decides more than the address of the data centre.\nCheck a product fits your legal setting before you buy. Consent, employment rules and record keeping are the usual places an imported default does not match local law.\nNowt there needs a view on any government. It is ordinary supplier diligence, applied to a question most procurement checklists still miss out.\nPart 6 looks at security, and at the standard applied when a supplier is shut out on national security grounds.\nFirst published: 2026-08-25. Last updated: 2026-08-25.\n","permalink":"https://blogs.damiendye.uk/en/random/how-far-the-law-reaches/","summary":"Part 2 covered the law that reaches your data. Part 5 of 8 covers the law that reaches your company: large penalties under American law, a French parliamentary report on whether that works as a commercial weapon, trade pressure applied to tax law and to regulators, and products carrying one market\u0026rsquo;s assumptions everywhere.","title":"How Far One Country's Law Reaches"},{"content":"Part 6 of 8. Part 5 looked at how far one country\u0026rsquo;s law reaches.\nWhat this post covers The reason given for shutting some suppliers out. What has been established about weakened products. What happened in 2024. Why this lands on Europe too. What you can do about it. The reason given for shutting some suppliers out Several governments have shut Chinese suppliers out of telephone and internet networks, and restricted some Chinese apps on official devices.\nThe reason given is usually the same: a company can be made to help its own government. So the kit carries a risk, whatever the code does.\nThat reasoning is sound. It is also the same reasoning used in part 2 about whose law applies to your data.\nLet us be straight about one thing before going on.\nCyber operations by the Chinese state are real. European national security agencies document them as well as American ones, and they include large-scale theft of commercial information. Nowt in this post is a defence of that.\nBut if the principle is that a supplier can be made to help its own government, then it applies to every supplier that has a government.\nApply it evenly and the record on the European and American side is not empty. In places it is better established than the accusations, because parliaments have investigated it.\nWhat has been established about weakened products Three items, each with a source, in order of how firmly they stand up.\nAn intelligence service owned an encryption company.\nCrypto AG was a Swiss company that sold cipher machines to more than 120 governments from the 1950s onwards.\nIt was owned by the American CIA and the German BND. The machines were altered so those services could read the messages of the governments buying them.\nIt rests on more than journalism. After the story broke, the Swiss Parliament\u0026rsquo;s own intelligence oversight body investigated and published a report in November 2020.\nSwiss public broadcaster SWI has covered the findings, including that Swiss intelligence had known since 1993.\nCustomers included European governments. It ran for roughly 50 years.\nA cryptographic standard was withdrawn.\nA random number generator called Dual_EC_DRBG was published as a United States standard — random numbers are what encryption uses to make keys unpredictable.\nResearchers showed its design let whoever picked certain values inside it predict its output.\nThe standards body advised against using it in 2013 and removed it in 2014. The Register reported on the related commercial arrangements at the time.\nKit has been intercepted in transit.\nIn December 2013 the German magazine Der Spiegel published a catalogue of interception tools covering routers, firewalls, servers and storage firmware from well-known manufacturers.\nThe reporting also described kit being diverted in transit, altered, repackaged and sent on to the customer.\nThat last one is worth a moment from anybody who buys hardware.\nIt means the trust boundary is not just your supplier — it is your supplier and everything that handles the delivery.\nAn honest supplier cannot give you an assurance about the second part.\nWhat happened in 2024 This is the event that answers the engineering question, and it is the most useful thing in this post.\nIn 1994 the United States passed a law requiring telephone companies to build their networks so communications could be intercepted on a lawful request.\nSo the way in was permanent, built in, and governed by legal process.\nIn 2024 a group linked to the Chinese state, publicly named Salt Typhoon, was found to have got into at least 9 major American telephone companies.\nAmong the systems reached were the interception systems themselves. The Register reported on the response from lawmakers.\nThe way in that one government required to be built became the way another government got in.\nThat is not bad luck. It follows from how such a thing works.\nA built-in way in is a capability, not a rule. The law governing it only binds people who accept that law. Somebody who has broken in does not.\nEvery argument for built-in lawful access assumes the access can be kept to the intended user. This is the clearest evidence going that the assumption does not hold.\nWhy this lands on Europe too If that reasoning is right, it is right everywhere, and as such it lands on European proposals too.\nProposals to scan messages on the device before they are encrypted create the same sort of built-in way in.\nThe United Kingdom has a power to require companies to provide technical capabilities. It has been used to require changes to an encryption feature, and the supplier pulled that feature in the United Kingdom rather than change it.\nBoth are pursued by democracies, with oversight, for serious reasons such as protecting children.\nBoth also build exactly the kind of capability that 2024 showed cannot reliably be kept to its intended user.\nA European way in is no safer than any other, because an intruder does not care how accountable the institution is. What they care about is whether the thing exists.\nWhich is why the useful version of this argument is a technical one, not a political one.\n\u0026ldquo;Do not trust suppliers from country X\u0026rdquo; has to be reopened every time the politics shift.\n\u0026ldquo;Assume any built-in way in will eventually be used by somebody it was not built for\u0026rdquo; holds regardless.\nWhat you can do about it The practical answer is not mainly about who you buy from. It is about designing so the question matters less.\nTreat the network as untrusted. Encrypt traffic end to end, and do not let one system trust another just because of where it sits on the network.\nHold your own encryption keys where you can. Part 2 explained why that changes who can be asked for your data. It also limits what an intruder finds worth having.\nSend less. Data you never collected cannot be intercepted, requested or lost. Least fashionable control on the list, and often the most effective.\nCheck what your kit runs before you trust it. Verified boot and firmware checking are worth turning on where your hardware supports them.\nOn genuinely sensitive systems, check deliveries. Tamper-evident packaging and recorded serial numbers are simple enough.\nNone of these need you to decide which government to worry about most.\nThat is exactly why they are the better answer. A control that only works if your guess about politics is right is not really a control.\nPart 7 steps outside technology, and looks at how other industries handled the same problem.\nFirst published: 2026-08-25. Last updated: 2026-08-25.\n","permalink":"https://blogs.damiendye.uk/en/random/backdoors-and-who-is-accused/","summary":"Shutting a supplier out on security grounds needs a standard, applied evenly. Part 6 of 8 sets out what has been established about weakened products and interception, including a Swiss parliamentary inquiry, and shows why a built-in way in belongs to whoever reaches it.","title":"Backdoors, and Who Gets Accused of Them"},{"content":"Part 7 of 8. Part 6 looked at security and built-in ways in.\nWhat this post covers Why the comparison is useful. Seeds, and licence terms on living things. Supermarkets, and the rules built for them. Why technology has nowt like it. What it tells you. Why the comparison is useful There is a common view that technology is a new kind of power and needs a new kind of thinking.\nIt is a comforting view, and it is mostly wrong. Believing it is one reason the answer has been so slow coming.\nThe problem in this series is an old one. A supplier becomes essential. The customer\u0026rsquo;s alternatives disappear. Terms stop being agreed and start being announced.\nThat has happened in industries with no software in them at all.\nThe comparison helps 2 ways.\nIt shows what tends to come next, because those industries are further down the road.\nAnd it shows what a serious answer looks like, because some of those answers already exist and work.\nSeeds, and licence terms on living things A farmer buying seed is in much the same spot as an organisation buying core software.\nFour companies, Bayer, Corteva, Syngenta and BASF, hold most of the global seed market and a similar share of pesticides. The Swiss organisation Public Eye publishes analysis of it.\nFor some individual crops it is tighter still, and a handful of firms hold most of the relevant patents.\nThe licensing will look familiar if you have ever read a software agreement.\nA patented seed is not simply sold. It is licensed, and the terms can stop the oldest practice in farming: keeping part of this year\u0026rsquo;s harvest to plant next year.\nSo the farmer buys the use of the seed for 1 season. The right to reproduce it, which is the whole point of a seed, stops with the supplier.\nThat is a permanent purchase turned into a recurring one. It happened to farming before it happened to software.\nThere is a second familiar detail.\nThe seed is often built to work with a particular pesticide, and the same company sells both. Buying one makes buying the other easier than switching.\nAgricultural economists have tracked the results: fewer varieties about, higher input costs, and less ability to change supplier.\nSupermarkets, and the rules built for them The other half of food shows the same problem from the buying side. This is the example worth copying.\nA small number of retailers handle most grocery sales in the United Kingdom.\nSo a supplier faces very few possible customers. Losing one can finish the business. For the retailer, replacing a supplier is a morning\u0026rsquo;s work.\nThat imbalance got documented well enough that Parliament acted. The Groceries Code Adjudicator was set up in 2013. It oversees whether large retailers follow the Groceries Supply Code of Practice.\nLook at what that code restricts, then think about your last software renewal.\nChanging agreed terms after the fact, without agreement. Charging a supplier for the right to carry on doing business. Moving costs and risks onto the supplier without compensation. Using removal from sale as leverage in a dispute. The European Union built its own version. Directive 2019/633 on unfair trading practices covers the agricultural and food supply chain.\nIt bans late payment, cancelling orders at short notice, changing terms one-sidedly, and threatening commercial retaliation. Every member state has an authority to enforce it.\nNeither regime is perfect. The Adjudicator\u0026rsquo;s powers are limited, it does not cover pricing itself, and it has used its strongest powers sparingly.\nBut the principle inside them is the one missing from technology.\nWhere bargaining power is very unequal, freedom of contract does not really exist. So certain practices get banned outright, rather than left to negotiation.\nNobody has suggested a software supplier should be stopped from changing agreed terms after the fact.\nNobody treats a charge for taking your own data out as shifting a cost onto the weaker party.\nIn food, both would be spotted straight away. In software we call it a licensing model.\nWhy technology has nowt like it Four reasons, and they are worth understanding because they point at what would have to change.\nSpeed. Grocery concentration took about 100 years. Cloud concentration took about 15. By the time you could see it clearly, it had happened.\nFree services. Competition law in most countries grew up round prices paid by consumers. A service priced at zero looks harmless under that test, even when the customer is not the user.\nThe claim of novelty. The industry argued, successfully, that its economics were different, and that being hard to leave was just a property of complex systems rather than a design choice.\nOrganisation. Farmers have unions, cooperatives and a long habit of bargaining together. They have used it to win specific legal protections.\nTechnology buyers have user groups and a conference. When VMware customers were hit with large increases, the real pushback came from a trade body of cloud providers, and it went through competition law. Part 2 covered how long that takes.\nWhat it tells you Two conclusions, and both are worth having.\nThe problem is structural, not national.\nBayer is German. The biggest supermarkets are British, French and German. The buyout model in part 3 is used keenly across Europe.\nReplace every American supplier in this series tomorrow with 4 European ones and the behaviour comes back, because it follows from the structure.\nAs such, the advice here is about being able to change supplier, rather than about picking a country.\nA working answer already exists.\nWe do not need to invent a new theory of platform regulation.\nThere is a model where certain practices are banned because bargaining power is unequal, an adjudicator hears complaints, and the weaker party does not carry the whole burden of proof.\nIt was built for cabbages. It would work on cloud contracts.\nPart 8 looks at what is being built in Europe now, and how to check where you stand.\nFirst published: 2026-08-25. Last updated: 2026-08-25.\n","permalink":"https://blogs.damiendye.uk/en/random/the-same-pattern-elsewhere/","summary":"None of this is new, and it is not really about technology. Four companies control most of the world\u0026rsquo;s seed supply. Ten retailers move most of Britain\u0026rsquo;s food. Both produced a legal answer software has nowt like. Part 7 of 8 compares them, and asks why technology got away with it.","title":"The Same Pattern in Other Industries"},{"content":"Part 8 of 8. Part 7 compared the pattern to other industries.\nWhat this post covers What Europe has actually built. What is realistic today, and what is not. 3 questions to work out where you stand. An honest word about this website. The trend has turned For a lot of years, European digital independence was mostly a conversation.\nSince 2025 that has changed. There are working products now, and outfits using them every day.\nThis is the most useful trend in the whole series, because it is the one you can act on.\nOffice and collaboration software openDesk is a set of open source tools for everyday office work. Documents, email, calendars, files, chat and video meetings.\nIt is looked after by ZenDiS, the German centre for digital sovereignty.\nIt is built from tools that already existed: Nextcloud for files, Collabora for documents, Open-Xchange for email and calendars, Element for chat, Jitsi for video.\nIt is in real use. The German armed forces signed a multi-year agreement. The Robert Koch Institute, a German public health body, runs it for several thousand people. The International Criminal Court picked it in 2025.\nThere are commercial options now too.\noffice.eu launched in The Hague in March 2026. It is a European-owned hosted alternative to the big American office suites.\nEuro-Office followed in June 2026. It is a shared document editor built jointly by several European companies, including IONOS, Nextcloud, XWiki, OpenProject, Open-Xchange and office.eu.\nWorking together like that matters. It means each company is not building the same editor over again on its own.\nGovernments actually moving Some governments have already moved, which gives the rest of us real experience to learn from.\nThe Danish Ministry of Digitalisation moved from Microsoft 365 to LibreOffice from July 2025. LibreOffice is a free office suite.\nThe German state of Schleswig-Holstein is moving tens of thousands of staff computers to open source.\nFrance and Germany are building shared tools together as well, rather than each paying for its own.\nIt is only fair to say this is not always plain sailing.\nMunich moved its city council to Linux over several years, then moved back to Windows in 2017. The Register reported on the cost at the time.\nThe main lesson from Munich was not a technical one. The software mostly worked. What did not last was the political backing, across changes of administration, on a project running more than a decade.\nSo these things need steady support over many years. That is a real requirement, and worth planning for honestly.\nPayments Payments follow the same trend, and plenty of people do not realise how concentrated they are.\nAnalysis of European Central Bank figures suggests Visa and Mastercard handled around 47% of card payment value in the euro area in 2025. In several countries the share is a good deal higher.\nSeveral European countries do have their own card systems. Trouble is they do not work across borders, so the one that works everywhere is the American one.\nWero is the European answer. It is run by the European Payments Initiative.\nIt started with payments between people in Belgium, France and Germany. It now has tens of millions of registered users. Luxembourg joined in 2026, and the Netherlands is moving its existing iDEAL system onto it. The European Payments Council publishes progress.\nPaying in shops and online is the next stage. That is the stage that decides whether Wero becomes a real alternative.\nThe digital euro sits behind it, on a longer timescale of around 2029.\nNew rules The European Commission published its Cloud and AI Development Act proposal on 2026-06-03.\nIt is the first serious go at turning cloud sovereignty from a voluntary label into a legal requirement.\nIt grades cloud providers in tiers, on where the infrastructure sits, how independently it can be run, and who owns it. Public sector buyers would need providers to meet at least the lowest tier.\nIt is a proposal. It is not law yet, and it may change on the way through.\nWhat is realistic, and what is not A hopeful list can easily go wrong here, so let us be plain.\nRealistic today. Office documents, email, calendars, files, chat and video. Working European and open source options exist, and outfits are running them in production right now.\nGetting there. Payments, and some cloud services. The parts exist. The coverage is still filling in.\nNot yet. Large-scale cloud infrastructure. European capacity is a long way short of demand.\nWhat can be replaced with a European or open source option today How replaceable each layer is today Office work Documents, email, calendars, files, chat, video Replaceable now Payments Wero is growing; shops and online are the next stage Partly ready Cloud infrastructure European capacity is much smaller than demand Not yet Computer chips Real European strengths, but no full substitute today Not this decade The further up the stack you look, the more choice you have today. Not this decade. The chips themselves. Europe has real strengths, including ASML in the Netherlands and Infineon and NXP in chip design. None of that replaces a data centre graphics processor today.\nThe Bertelsmann Stiftung, a German foundation, has reckoned that building a full European stack would take around 10 years and about €300 billion.\nSo full independence is not on the table, and anybody offering it is overselling what they have.\nAs such, full independence was never the useful goal.\n3 questions to work out where you stand Here is a simple way to look at your own systems. It works for any supplier, in any country.\nAsk these 3 questions about every service you depend on.\n1. What stops working if this service stops?\nBe specific. Name the systems and the people affected. \u0026ldquo;It is important\u0026rdquo; is not something you can plan with.\n2. How many weeks of work would it take to move?\nThat number sets your position at every renewal for as long as you keep the supplier.\nIf you cannot estimate it, that is already worth knowing.\n3. Could your organisation actually run the alternative?\nNot whether an alternative exists. Whether your team, at the size it really is, could run it.\nLibreOffice is a genuine answer for a ministry with a training budget. It is a harder answer for a small company with no IT staff.\nRank your services on those 3 answers. The list is usually short, and often not the one people expect.\nFor plenty of outfits, email and staff accounts come out above cloud computing. Moving them is rarely tested, and would take a long while.\nThen ask the same 3 questions about any European alternative you are weighing up.\nA single European supplier you cannot leave is not much better than a single American one you cannot leave.\nThe real advantage of open formats and open source is not the price. It is that they make question 2 answerable, and question 3 possible.\nAn honest word about this website It would not be fair to write all this without turning it on myself.\nThis site is built with Hugo and served by Cloudflare, an American company. Every point in this series applies to it.\nI picked Cloudflare because it works well and costs me nowt. If the account stopped tomorrow I would change some DNS settings and publish somewhere else.\nSo the exposure is real and the consequence is small. That is a fair trade, and I made it on purpose.\nThe rest of my own kit is mixed as well. Proxmox, which I write about often, is Austrian. It runs on processors designed in the United States, on firmware I cannot inspect.\nThere is no setup that gets to zero. That was never the goal.\nThe short version Technology went from something you bought to something you rent.\nRenting is often better, which is exactly why it happened.\nEvery cost of renting grows the harder it is for you to leave.\nSince 2025 real alternatives have landed for everyday office work, and are landing for payments.\nSo the practical goal is not dodging any particular country. It is cutting down the number of services you could not leave inside 3 months.\nStart with the ones you have never tested. That is usually where the surprise is waiting.\nFirst published: 2026-08-25. Last updated: 2026-08-25.\n","permalink":"https://blogs.damiendye.uk/en/random/choice-coming-back/","summary":"Since 2025 the European answer has turned from plans into things you can install. Part 8 of 8 covers what has landed, what is honestly still years off, and gives 3 questions to work out where you stand.","title":"Where Your Choices Are Coming Back"},{"content":"Every machine in an Active Directory domain finds its domain controller by asking DNS. Not from a configured list. It asks for an SRV record, and it goes wherever the answer sends it. Which makes a handful of records under _msdcs the most security-critical thing in the zone, and raises two questions worth answering properly: how do you sign them, and who is allowed to write them. Neither gets asked often, and both get answered badly by default.\nThis post answers both for a Samba4 domain controller. BIND with dlz_bind9 serving the directory, inline signing, a hidden primary that no client ever reaches, a delegation inside a public zone so the trust chain runs to the root, and client dynamic updates kept well away from the locators.\nBut signing is only worth anything if you know what it protects, and that means starting one level down — at how a name gets resolved at all. So the first third of this is the walk down from the root. If you already have that cold, skip to Samba\u0026rsquo;s two DNS backends.\nA Resolver Starts Knowing Almost Nothing It is worth being clear about how little a resolver is born with, because everything else in this post follows from it.\nA freshly installed recursive resolver knows two things. It knows the addresses of the root servers — a hints file, thirteen names from a.root-servers.net to m.root-servers.net, served by anycast from a great many more instances than thirteen. And, if it validates, it knows one public key: the root\u0026rsquo;s.\nThat is the whole of the built-in configuration. Every other fact it will ever serve — which servers are authoritative for uk, where your domain lives, what address your mail server has — is learned at runtime by asking, and then cached until its TTL runs out.\nThis is a good design. Nobody has to ship a list of the internet\u0026rsquo;s name servers, and no central party has to approve a change to your zone. But it has a consequence that the rest of this post is about: a resolver believes what it is told, by whichever server the previous answer pointed it at. Take away signatures and the whole structure is a chain of assertions, each one authenticated only by the fact that it arrived from the address the last answer named. That is a lot of weight to hang on a return address.\nThe Walk Down the Tree A lookup is not one question. It is a series of referrals down the name tree, and each step is a separate query to a different server.\nSay a client wants dc01.ad.example.co.uk. A resolver with a cold cache does this:\nAsk a root server. The root does not know the answer and does not pretend to. It returns a referral: an empty answer section, and in the authority section the NS records for uk, with the addresses of those name servers as glue in the additional section. Ask a uk server. Another referral, this time to the name servers for example.co.uk. Ask an example.co.uk server. If ad.example.co.uk is delegated — and in the design later in this post it is — one more referral. Ask an ad.example.co.uk server. This one is authoritative for the name, so it answers with the A record and sets the AA (authoritative answer) bit. A single lookup walked down the delegation tree, one referral at a time Recursive resolver Born knowing only: \u0026#183; the root hints \u0026#183; the root\u0026#8217;s public key 1 \u0026#160; Root server a\u0026#8211;m.root-servers.net Referral the NS records for uk. \u0026#8212; plus their addresses as glue 2 \u0026#160; uk name server authoritative for uk. Referral the NS records for example.co.uk. 3 \u0026#160; example.co.uk holds the delegation Referral ad.example.co.uk. is delegated \u0026#8212; ask its servers 4 \u0026#160; ad.example.co.uk authoritative for the name Answer the A record for dc01, with the AA bit set Every step is cached until its TTL expires, so a warm resolver starts at step 3 or 4. Which is exactly what makes one poisoned cache entry worth so much. One lookup, four servers. Each step returns not the answer but the name of somewhere better to ask, and the resolver caches every step on the way down. Three things about that walk matter later.\nThe referral is the delegation. A zone\u0026rsquo;s parent does not hold its contents. It holds NS records saying \u0026ldquo;ask over there\u0026rdquo;, and — once signing enters the picture — a DS record saying \u0026ldquo;and here is the fingerprint of the key you should expect when you do\u0026rdquo;. Delegation is the only structural mechanism DNS has, and it is the thing you will use to keep an internal zone internal.\nNearly all of this is cached. The root and TLD referrals have long TTLs, so a warm resolver skips straight to step three or four. This is why a resolver\u0026rsquo;s cache is worth attacking: poison one entry and you have redirected everything under it for as long as the TTL you chose.\nThe full name is not sent to every server. Under QNAME minimisation, a resolver asks the root only about uk, not about dc01.ad.example.co.uk. Worth knowing if you ever go looking for your internal names in a query log upstream. They should not be there.\nStub, Recursive, Authoritative Three words that get used interchangeably and mean quite different jobs. The distinction is load-bearing for the Active Directory half of this post, so it is worth nailing down.\nWhat it does Walks the tree? Holds zone data? Stub resolver The library or local service your applications call. Asks one configured server and takes the answer. No No Forwarder Passes queries to another resolver, caches the replies. No No Recursive resolver Does the walk above, caches every step, optionally validates signatures. Yes No Authoritative server Answers for zones it has been given, and only those. Says \u0026ldquo;I don\u0026rsquo;t know\u0026rdquo; about everything else. No Yes The important line is the last one. An authoritative server has no business doing recursion, and a recursive resolver has no business being authoritative. Combining them is how you get a machine that will both accept arbitrary questions from clients and hold data those clients depend on — which is exactly the machine you do not want your domain controller to be.\nOn a modern Linux desktop there is another layer worth being aware of. systemd-resolved runs a stub listener on the loopback address and writes a resolv.conf pointing at itself, with options edns0 trust-ad. That trust-ad is the one to notice: it tells the stub to believe the AD (authenticated data) bit in replies from that server. The AD bit is not proof of anything on its own — it is the upstream resolver\u0026rsquo;s claim that it validated. Trusting it is reasonable when the upstream is yours and the hop to it is trustworthy, and meaningless otherwise.\nHow Services Are Actually Found An A record answers one question: what address does this name have. It says nothing about which port the service is on, which of several servers to prefer, or what to do when the first one is down.\nSRV records answer all three. From RFC 2782, the shape is:\n_service._proto.name. TTL IN SRV priority weight port target The parts of an SRV record and what each one decides the transport the domain the service which kind of server _ldap. _tcp. dc._msdcs. ad.example.com. 600 IN SRV 0 100 389 dc01.ad.example.com. one owner name \u0026#8212; underscores keep it out of host space TTL record type priority weight port target Lower priority is tried first. Weight shares load between the servers that sit at the same priority. The target must be a host name with A or AAAA records \u0026#8212; never a CNAME. A single target of \".\" means the service is deliberately not offered here. An SRV record dissected. The owner name encodes the service and the transport; the record body carries the failover order, the load share within an order, the port, and a target that must be a real host name. Each field earns its place:\nThe underscore prefixes keep the service labels in a namespace of their own. _tcp can never collide with a host actually called tcp, because a host name may not begin with an underscore. Priority is the failover order — lower is tried first. Servers at a higher priority number are only used when everything below is unreachable. Weight shares load within one priority. A pair at weight 100 and weight 300 gets roughly a quarter and three quarters of the clients. It is a proportional draw, not round-robin. Port frees the service from a well-known number, which is how a client finds an LDAP server on 3268 without anyone hard-coding it. Target must be a name with A or AAAA records. RFC 2782 is explicit that it must not be a CNAME, and that a single target of . means \u0026ldquo;this service is deliberately not offered here\u0026rdquo; — a useful thing to publish on purpose. This is what makes a domain locatable rather than configured. A domain-joined machine does not have a list of your domain controllers. It has a domain name, and it asks.\nThe Names Active Directory Publishes An AD domain is, from a client\u0026rsquo;s point of view, mostly a set of SRV records. The tree under _msdcs is the interesting part, because it is how clients distinguish \u0026ldquo;an LDAP server\u0026rdquo; from \u0026ldquo;a domain controller for this domain\u0026rdquo; from \u0026ldquo;a global catalogue for this forest\u0026rdquo;:\nName What asks for it _ldap._tcp.\u0026lt;domain\u0026gt; Anything wanting LDAP in the domain _ldap._tcp.dc._msdcs.\u0026lt;domain\u0026gt; Domain controller location — the big one _ldap._tcp.pdc._msdcs.\u0026lt;domain\u0026gt; The PDC emulator specifically _ldap._tcp.gc._msdcs.\u0026lt;forest\u0026gt; Global catalogue _kerberos._tcp.\u0026lt;domain\u0026gt;, _kerberos._udp.\u0026lt;domain\u0026gt; KDC location, before any ticket exists _kpasswd._tcp.\u0026lt;domain\u0026gt;, _kpasswd._udp.\u0026lt;domain\u0026gt; Password changes _ldap._tcp.\u0026lt;site\u0026gt;._sites.dc._msdcs.\u0026lt;domain\u0026gt; Site-aware location — find a DC that is near Look at what that list is for. Kerberos authentication cannot begin until the client has found a KDC, and it finds the KDC by DNS. Site-aware location means the answer also decides which datacentre a client authenticates against.\nThe Same Idea, Modernised SRV has a successor worth knowing about. SVCB and HTTPS records (RFC 9460) generalise the pattern: service parameters in DNS, including ALPN, port, and address hints, in one record. That is how a browser learns to go straight to HTTP/3 without a redirect. The HTTPS record is already widely deployed. The mechanism is the same in spirit as SRV: the client is told, by DNS, how to reach the service. Which means the argument in the next section applies to it just as much.\nEverything After the Lookup Trusts the Lookup Here is the pivot, and it is the reason a DNS post has a security section rather than the other way round.\nA domain-joined machine boots and asks where a domain controller is. It gets a name and a port. It connects there, and then it starts doing all the things we normally think of as the security layer: Kerberos, LDAP signing, channel binding, certificate validation.\nSo what happens if the answer was a lie?\nBeing fair about this matters, because the answer is not \u0026ldquo;instant catastrophe\u0026rdquo;. Kerberos was designed with mutual authentication exactly so that a client is not at the mercy of its name lookup: a host that cannot produce a service ticket for the name the client asked for cannot complete the exchange. Get an SRV answer pointing at a machine with no key in the domain and, for a client configured strictly, it fails.\nThe realistic damage is subtler, and it is all in the gaps around that:\nDowngrade. A client that falls back to NTLM when Kerberos does not work has just been handed to whoever answered. Relay and coercion. The attacker does not need to be the DC. Being the name a client connects to is enough to start relaying that authentication somewhere it is useful. Denial of service that looks like a fault. Point _ldap._tcp.dc._msdcs at something that does not answer and the domain is intermittently broken in a way nobody diagnoses as DNS for a good while. Everything with no mutual authentication at all. Time sync, syslog, monitoring, backup agents, that internal HTTP thing with verify=False. Plenty of estates have more of these than they would like to admit. The general principle is the one worth taking away: DNS is a discovery mechanism, not an authorisation mechanism. It is entirely reasonable to find a service by name. It is not reasonable to grant anything on the strength of a name, and the number of systems that quietly do — PTR-based allowlists, host-name ACLs, \u0026ldquo;it\u0026rsquo;s on the internal network so it must be ours\u0026rdquo; — is the actual attack surface.\nWhat DNSSEC Fixes, and What It Does Not DNSSEC exists to answer one question: did this answer really come from the zone that owns the name, unmodified?\nIt works (RFC 4033 and the two that follow it) by signing, and by chaining the signatures to something you already trust:\nEvery RRset in a signed zone has an RRSIG — a signature over that set. The zone\u0026rsquo;s public keys are published as DNSKEY. The parent zone publishes a DS record: a hash of the child\u0026rsquo;s key. That DS is itself signed by the parent, whose key is hashed in its parent\u0026rsquo;s DS, all the way up to the root — and the root\u0026rsquo;s key is the one thing your resolver was born knowing. A chain of trust from the root down to an internal zone . \u0026#8212; the root the one key your resolver already has Where every chain starts. Shipped with the resolver, rolled rarely, trusted implicitly. signed DS for com. com. or uk., or whatever you sit under A parent vouches for its child's key. The DS is a hash of the child's key, signed by the parent. signed DS for example.com. example.com. public, signed, DS lodged with the parent The zone you already own and already sign. It holds the delegation for the internal zone, and the DS that authenticates it. That is all it holds of the inside. signed DS for ad.example.com. above the line: published to the internet below it: served only inside the estate ad.example.com. the AD-derived view \u0026#8212; internal servers only Signed inside, trusted from outside. Your resolvers validate it root \u0026#8594; com \u0026#8594; example.com \u0026#8594; here, with no local trust anchor and nothing to distribute. The data never leaves. Only the trust comes in, down the ordinary public chain. The chain a validating resolver builds. Each parent vouches for its child\u0026rsquo;s key with a signed DS record, so one built-in trust anchor at the root authenticates every zone below it — including a delegated internal zone whose contents never leave the building. Denial is signed too, which is easy to overlook and matters more than it sounds. Without it, \u0026ldquo;no such name\u0026rdquo; is an unauthenticated answer an attacker can forge to make something disappear. NSEC and NSEC3 give an authenticated \u0026ldquo;nothing exists between these two names\u0026rdquo;.\nThat is a real fix for a real problem. Cache poisoning is not theoretical: the Kaminsky attack made off-path spoofing cheap enough to force emergency patching across the entire internet, and the mitigations that followed — source port randomisation, 0x20 case mixing — are all attempts to make guessing harder, not to make answers verifiable. Side-channel work since has repeatedly chipped away at them. Signatures are the only thing in that list that changes the game rather than raising the price.\nNow the honest limits, because DNSSEC is oversold in both directions.\nIt is not confidentiality. Signing is public-key authentication; every query and answer is still in clear text on the wire. Privacy on the client hop is DoT, DoH or DoQ, and those are a different mechanism solving a different problem. Encrypting the hop to a resolver that does not validate buys you a private conversation with something that can still be lied to.\nIt does not validate meaning, only origin. DNSSEC proves the zone\u0026rsquo;s owner published this. If the record is wrong, or malicious, or was written by someone the zone unwisely allowed to write it, the signature is applied to that just as faithfully. Signing a zone that untrusted machines can write does not make the contents trustworthy — it authenticates the lie. Hold onto that sentence; the last third of this post is essentially its consequence.\nIt only helps if somebody validates. If your resolver does not validate, signatures on the zones you query are decoration. And if validation happens at a resolver across the network from the client, the client is trusting the AD bit and the path — see trust-ad above.\nAnd it needs operating. An expired signature is not a degraded answer, it is SERVFAIL — the name goes dark. A DS in the parent that no longer matches the child\u0026rsquo;s key does the same. This is the honest cost, and as such the automation in the next sections matters more than the initial signing.\nWhy a Made-Up Internal TLD Cannot Be Signed This is where naming decisions made years ago come back with a bill, and it is worth spelling out because it is the reason this post\u0026rsquo;s design uses a real, owned domain for internal names.\nLook again at how the chain is built: a zone\u0026rsquo;s key is vouched for by a DS record in its parent. So a validating resolver can only authenticate an internal zone if that zone has a parent willing and able to publish a DS for it.\nad.corp.local has no such parent. Neither does .lan, .home, or anything else invented on the day the domain was provisioned:\n.local is reserved for mDNS. Using it for unicast DNS is not merely unsigned, it is a documented collision with how every operating system in the building is entitled to behave. It gets its own section below, because it is in a different league to the other three. .lan, .corp, .home are unregistered strings. There is no parent to hold a DS, and the applied-for versions were withheld from delegation because of the name-collision mess in private estates. home.arpa is properly reserved for home networks, which sounds ideal — but its delegation is deliberately insecure. There is no signed path down to it, by design. .internal has been set aside by ICANN for exactly this use, and it solves the collision problem. It does not solve this one: no delegation means no DS, so no chain. Your options with an unchained internal zone are to leave it unsigned, or to configure a local trust anchor on every validating resolver in the estate and become the root of your own private island — a key you now have to distribute, monitor and roll by hand, on every resolver, forever, with SERVFAIL across the entire domain as the failure mode.\nThere is a much easier answer, and it is free: make the internal zone a delegation inside a public zone you already own and already sign. ad.example.com, delegated from example.com. The DS goes in the public parent. Every validating resolver in the estate then authenticates internal names with the trust anchor it already has, and you distribute nothing.\nThe contents stay inside. Only the trust comes from outside. That distinction is the design in the rest of this post.\n.local Is Not a Style Choice One of those four deserves its own section, because it is the one people defend, and because the defence is always the same: it works, we have used it for years, what is the problem.\nThe problem is that .local is not an unregistered string somebody might one day sell. It is a namespace that already has an owner and a defined behaviour, and the behaviour is not \u0026ldquo;ask the DNS server\u0026rdquo;. RFC 6762 §3 is not gentle about it:\nAny DNS query for a name ending with \u0026ldquo;.local.\u0026rdquo; MUST be sent to the mDNS IPv4 link-local multicast address 224.0.0.251 (or its IPv6 equivalent FF02::FB).\nRead what that MUST actually governs. It is not a claim about who owns the string. It is an instruction about where the query goes — and the answer is a multicast group on the local link, not your domain controller. The same section says names under .local are \u0026ldquo;meaningful only on the link where they originate\u0026rdquo;, the DNS equivalent of a 169.254 address.\nSo a client that follows the specification correctly will never ask your DNS server for a .local name. It shouts on the wire and takes whatever answers.\nThat produces a set of failures with a very particular flavour:\nResolution depends on the client, not on your DNS. macOS resolves .local through Bonjour and always has; systemd-resolved routes the local domain to mDNS by default; Windows has done mDNS since Windows 10. Three stacks, three sets of rules, and your zone file is not consulted by any of them. It stops at the first router. mDNS is link-local by design. A name that resolves at a desk on the same VLAN does not resolve from another floor, another site, or the VPN — which is the \u0026ldquo;works in the office, breaks at home\u0026rdquo; ticket that gets closed as \u0026ldquo;network issue\u0026rdquo; four times before somebody reads the spec. The behaviour changes underneath you. Whether a given machine tries multicast first, unicast first, or both in parallel depends on the resolver stack and its version. Estates that ran on .local \u0026ldquo;fine for years\u0026rdquo; are usually estates where a distro update has not yet changed the order. The evidence is missing from where you look for it. The DC\u0026rsquo;s query log shows nothing, because nothing arrived. People spend days on the server, and the query never left the client. And it cannot be signed, which is the point of this section. No parent, no DS, no chain, no way to publish one. Now put that under an Active Directory domain. The realm is derived from the domain name. The service principals are derived from the realm. The _msdcs locators — the records this entire post is about — sit beneath a suffix that a conforming client is required to resolve by shouting at the local segment. You are building Kerberos on top of a namespace that half your estate resolves with a protocol designed for finding printers.\nAnd the exit is expensive, which is what makes the original decision worth being unsentimental about. There is no in-place rename on Samba. The documented path is samba-tool domain backup rename followed by a restore: you take a renamed copy of the database, re-seed a new DC from it, and re-add every other DC from scratch, with no overlap permitted between old-name and new-name DCs. It is not Windows\u0026rsquo; rendom, which at least walks the DCs one at a time. It is a rebuild with a nicer name.\nSo the honest summary. .local for a unicast domain is not a preference, a convention, or a harmless bit of legacy. It is a documented conflict with a protocol that ships enabled on every operating system in the building, and the bill arrives years later as intermittent, unreproducible resolution failures — and, when you finally want to sign the zone, as a domain rebuild.\nThe Same Mistake, Wearing a Different Hat While we are here: port 5353.\nmDNS listens on UDP 5353, and it is not a spare port that happens to be free. Standing a normal unicast DNS service on it, or pointing resolvers at :5353 because 53 was busy or needed root, does the same damage from the other direction — every mDNS-aware machine on that segment is now talking to your name service, and your name service is now fielding multicast service discovery it was never meant to answer. The name and the port are two halves of one namespace, and both of them already belong to somebody else.\n.local on 53, or unicast DNS on 5353. Same misunderstanding, same class of intermittent failure, same weeks of somebody else\u0026rsquo;s time.\nAnd Be Honest About What That Tells You Microsoft recommended .local in the Small Business Server era, which is why so many estates still carry it, and stopped recommending it a long time ago. Inheriting a .local domain is bad luck. Most people reading this who have one did not choose it, and the section above is a migration plan rather than an accusation.\nDeploying one is a different matter entirely, and I am going to be blunt about it, because softening this has not helped anybody.\nAnybody who deploys .local for Active Directory or for standard unicast DNS, or who stands a DNS service on port 5353, is not an IT professional. They may hold the job title. They are not doing the job. RFC 6762 has been published since 2013, it is readable in twenty minutes, and the sentence that settles the entire question is in section 3. Building a company\u0026rsquo;s name service on a namespace the specification reserves for link-local multicast is not a defensible engineering decision. It is somebody guessing in the one part of the stack where guessing produces failures nobody can reproduce and everybody blames on the network.\nThey should be removed from your IT department. Not moved sideways, not given DNS to look after with supervision. Removed from the function. The role exists to know which packet goes where and on whose authority. Somebody who has not read the document governing the namespace they chose for every machine in the building is not performing that role, and keeping them in it means the next decision of this size gets made the same way.\nAnd every system they touched should be audited. This is the part people skip, and it is the part that matters most. A decision like this is never isolated. Somebody who did not check RFC 6762 before naming the domain did not check anything else either — so go and look at what else they built. Expect to find:\nDNS servers open to the world, forwarders answering anybody who asks, and no ACLs anywhere. Dynamic update left wide open, and zone containers with permissions nobody has reviewed since the domain was created. Certificates and Kerberos built on top of names that were never going to resolve consistently, with the failures papered over by hosts files. Hosts files. Everywhere. Whole estates have been held together by them exactly because .local never worked properly and somebody found a workaround instead of a cause. Firewall rules and service accounts created to make the symptoms go away, still in place, still granting more than anybody now remembers. That is not vindictiveness. It is what a competence signal is for. When you find one decision that was made without reading the spec, the correct response is to assume the rest were made the same way and go and check — because the same person configured your authentication, your certificates and your access control, and you now have direct evidence of how they approach a problem they do not fully understand.\nInheriting this mess costs you a migration. Employing the person who is still creating it costs you considerably more.\nSamba\u0026rsquo;s Two DNS Backends A Samba AD domain controller stores its DNS data in the directory itself, and there are two ways to serve it.\nSAMBA_INTERNAL is Samba\u0026rsquo;s own DNS server, built into the AD DC. It handles the AD zones and Kerberos-authenticated dynamic updates, and it hands anything else to a forwarder. Samba describe it as supporting \u0026ldquo;the basic feature required in an AD\u0026rdquo; and recommend it \u0026ldquo;for simple DNS setups\u0026rdquo;, which is fair and worth taking literally. It gets its own section below, because what it does not do is longer and more interesting than what it does.\nBIND9_DLZ runs BIND as the DNS server, with Samba\u0026rsquo;s dlz_bind9 module loaded into it. DLZ — Dynamically Loadable Zones — is a BIND interface for backing a zone with something other than a zone file. The module answers BIND\u0026rsquo;s queries out of sam.ldb directly, so there is no copy, no export step, and no synchronisation to go wrong: BIND is reading the directory as it serves.\nThe Internal DNS Server Does Not Do Recursion — It Cannot Start with the thing that is almost always described wrongly, including by people who run it.\n\u0026ldquo;The DC is our DNS server, it does recursion for the clients\u0026rdquo; is not what is happening, because the internal DNS server cannot do recursion at all. Samba\u0026rsquo;s own feature list says it plainly. The internal DNS does not support:\nacting as a caching resolver recursive queries (but it can forward to another recursive DNS nameserver) shared-key transaction signature (TSIG) stub zones zone transfers Round Robin load balancing among DCs with scavenging and conditional forwarders listed as not implemented either.\nRead the first two together, because that is the whole story. It cannot resolve, and it cannot cache. What dns forwarder actually buys you is a relay: a client asks the DC for windowsupdate.com, the DC asks a real resolver, the answer comes back through the DC, and then the DC forgets it entirely. The next client asks the same question and the whole thing happens again.\nLook back at the table earlier in this post and note what that is: a forwarder without the forwarder\u0026rsquo;s cache. It has the role\u0026rsquo;s costs — an extra hop, a dependency, a thing to be down — and none of its benefit.\nSo a DC with SAMBA_INTERNAL serving the estate is not a DNS server in the sense people mean. It is an uncached proxy for a real resolver, and you have placed it on the machine that holds your directory.\nWhy That Is a Risk and Not Just an Inefficiency The inefficiency is easy to see: every external lookup in the building becomes a round trip through the AD DC, permanently, with no cache to take the edge off. On a quiet domain nobody notices. That is exactly why it survives.\nThe security argument is the one worth making, and it has three parts.\nThe listener is in the wrong process. The internal DNS is a service task of the AD DC itself, running with the directory\u0026rsquo;s privileges — not a separate daemon under its own account the way named is. So the thing parsing unauthenticated UDP from anything that can reach port 53 is running inside the process that serves LDAP and Kerberos and owns sam.ldb. BIND has thirty years of hostile attention, a dedicated user, and a habit of being run in a jail, exactly because a DNS listener is a rough neighbourhood. Samba\u0026rsquo;s internal server is a convenience feature that happens to live in the crown jewels.\nServing clients means being reachable by clients. To be the estate\u0026rsquo;s DNS server, it must accept queries from every workstation, every printer, every contractor\u0026rsquo;s laptop on the guest VLAN somebody bridged by accident. That is a large, permanently open, unauthenticated attack surface on the most valuable host you own, and you are running it to save the cost of a resolver that a Raspberry Pi could host.\nAnd a forwarder open to the world is somebody else\u0026rsquo;s weapon. A DC forwarding for anything that asks is an open forwarder. Once it is reachable off your network it becomes a reflection and amplification participant, which means the traffic and the abuse reports both arrive at your domain controller. There is no rate limiting to reach for, because the internal server has none.\nNone of that requires a Samba vulnerability to be a bad idea. It is a bad idea on shape alone: it puts an unauthenticated, internet-facing-by-accident, parser-heavy service in the same process as your directory, to do a job it is documented as not being able to do properly.\nIt Will Fall Over Sooner, and Take More With It Worth being clear about what this argument is not. It is not a claim that Samba\u0026rsquo;s DNS code has more bugs than BIND\u0026rsquo;s. BIND has a long CVE list, mostly because it is the most examined DNS implementation in existence, and counting advisories would be a poor way to choose between them.\nThe comparison that matters is structural, and it comes down to three questions with three uncomfortable answers.\nWho can send it a malformed packet? With SAMBA_INTERNAL serving the estate: every workstation, every phone on the wireless, everything that can route to port 53 on that box. With the design in this post: the signer. One host, one TSIG key, allow-query limited to it. That is not a small difference in degree. It is the difference between an exposed service and one that is effectively unreachable, and it dwarfs any difference in code quality between the two implementations.\nWhat falls over when it does? This is the one that decides the severity. named is a separate daemon under its own account; when it dies, DNS stops and the domain controller carries on authenticating. Samba\u0026rsquo;s internal DNS is a service task inside the AD DC, so anything that wedges, exhausts or crashes it is happening inside the process that serves LDAP and Kerberos. A DNS problem becomes a directory outage. And systemctl restart named costs a second, whereas restarting a DC is a different sort of morning.\nWhat can you do about it while it is happening? BIND has response rate limiting, allow-query, allow-recursion, blackhole, per-view policy, and the option of not answering clients at all. The internal server has dns forwarder and a log file. When something starts hammering it, there is no dial to turn.\nThen add the missing cache. Every client query is a fresh outbound round trip, so a query flood costs the DC an upstream lookup per packet rather than a cache hit — and it costs it in the same process that is trying to issue Kerberos tickets. You do not need an exploit for that; you need a busy morning, a misbehaving application, or somebody pointing a scanner at the wrong VLAN. Sustained, it presents as authentication being slow and nobody thinking to look at DNS.\nSo yes — even with DLZ in the picture, BIND is the safer place for this. The DLZ module does give named access to Samba\u0026rsquo;s data, and that is a real consideration this post has already made a point of. But a named crash is a DNS outage rather than a directory outage, and in this design that named is not taking queries from the estate in the first place. Bugs are a fact of every codebase. Blast radius and reachability are things you choose.\nAnd It Cannot Build the Design in This Post There is a simpler, more final reason it is not the backend here.\nNo zone transfers. The pipeline in the next section — hidden primary, signer, standalone authoritative servers — starts with an AXFR out of the DC. SAMBA_INTERNAL has nothing to transfer with. No TSIG, either, so even the authentication that transfer would need is absent. And no signing, and no validation of anything it forwards.\nSo the honest summary of SAMBA_INTERNAL: it is the backend that lets you stand up an AD domain in a lab on a Sunday afternoon without configuring BIND, and it is very good at that. Samba say \u0026ldquo;simple DNS setups\u0026rdquo; and they mean it. It is not a resolver, it was never built to be the DNS service for an estate, and the moment you want signing, transfers, views, ACLs, caching or rate limiting, the answer is not to tune it. It does not have those dials. The answer is BIND.\nIf you are running it today with every client pointed at the DC, the fix is not urgent but it is not optional either: give the clients a real validating resolver, and set dns forwarder on the DC to point at that, so the DC is answering for its own zones and nothing else.\nAnd this is the part I want to be clear about, because \u0026ldquo;don\u0026rsquo;t use DLZ\u0026rdquo; gets repeated as though it were a hardening rule: DLZ is not the exposure. It is the extraction mechanism. What matters is not which module is loaded into BIND — it is who is allowed to talk to that BIND, and what happens to the zone next. A DLZ back-ended BIND that answers only a transfer request from your signer is not an attack surface in any interesting sense. A SAMBA_INTERNAL DC fielding every name lookup from four hundred laptops very much is. Only one of those two turns up on hardening checklists, and it is not the one that matters.\nTwo things about DLZ are true and worth planning around rather than fearing:\nThe module is version-coupled to BIND. Samba ships a separate .so per BIND version, and named.conf names one specifically. A BIND major upgrade means the matching module must be in place, or named does not start. It is a packaging dependency to test in advance, not a security property. named needs access to Samba\u0026rsquo;s data, which is why Samba keeps a dedicated directory for the bits BIND needs rather than granting it the whole private directory. The grant is meant to be narrow — worth verifying it still is on your DCs, since it is the one place where DLZ does widen what a compromise of named would reach. The reason DLZ is the right choice here is what it enables: BIND\u0026rsquo;s DLZ interface supports enumerating an entire zone, which is what makes a zone transfer out of a DLZ-backed zone possible at all. That transfer is the first hop of the pipeline, and it is BIND doing it, which means the rest of the pipeline is ordinary BIND configuration rather than anything exotic.\nThere is one constraint that shapes everything downstream, and it is worth stating plainly because it is easy to assume otherwise. A DLZ zone cannot itself be signed. ISC are explicit about it in the BIND ARM: DLZ \u0026ldquo;is unable to handle DNSSEC-signed data due to its limited API\u0026rdquo;. You cannot hang a dnssec-policy off the dlz statement and be done.\nWhat you can do — and what ISC suggests for DLZ in the same breath — is run it as a hidden primary, with the signing done by a normal BIND zone that transfers the data in. That is the next section, and the constraint is the reason it has the shape it does.\nIt is worth knowing that this is a Samba constraint rather than a law of Active Directory. Microsoft\u0026rsquo;s DNS server has done online signing of dynamic, AD-integrated zones since Windows Server 2012. The zone is signed in place, the private keys replicate to the Key Masters through AD replication itself, and dynamic updates keep working. On Windows, \u0026ldquo;sign the AD partitions\u0026rdquo; is a real option and the answer to the churn objection is built in. (On Server 2008 R2 it was not: you could sign an AD-integrated zone, but not one accepting dynamic updates, and every change meant re-signing by hand — which is where the folklore about AD zones being unsignable comes from.)\nSamba has no equivalent. Neither backend signs: the internal server has no DNSSEC at all, and DLZ cannot carry signed data. So on Samba the transfer-and-sign pipeline is not one design among several. It is the way.\nThe Publishing Pipeline Now the shape of it. Four roles, and the DC is at the back where nothing can reach it.\nPublishing an Active Directory zone from a hidden-primary domain controller Domain controller \u0026#8212; hidden primary BIND with dlz_bind9, authoritative from the directory recursion off \u0026#183; transfers to the signer only, with TSIG The directory host takes no questions. The highest-value machine in the estate is not reachable by the machines that depend on it. AXFR on the SOA refresh a poll \u0026#8212; DLZ cannot notify Signing instance \u0026#8212; a normal secondary transfers the zone in and serves it signed a second named on the DC, or its own host The DLZ zone cannot be signed itself. So signing is a normal secondary holding the copy \u0026#8212; and with clients out of the zone, there is no churn. transfer out the signed zone Authoritative servers plain secondaries of the signed zone no keys, no route to the directory A compromise gets a copy, not a forgery. With no key on the servers that face the estate, no new record can be made that will validate. queries answered with signatures Validating resolvers what client resolv.conf points at internal zones forwarded, public tree walked Validation happens next to the client. The chain runs to the root trust anchor, because the internal zone is a delegation in a public one. no client-facing DNS on the DC Publication runs one way, from the directory outwards. Nothing a client sends ever arrives at the top of it. The DC is a hidden primary: it serves the AD zone by transfer and answers nothing else. The signing instance is an ordinary secondary zone holding the transferred copy — the DLZ zone itself cannot be signed, so signing always happens one hop along, whether that hop is a second named on the DC or a separate host. The standalone authoritative servers are the only thing clients ever see, and they are secondaries of the signed zone. 1. The domain controller — hidden primary. BIND with dlz_bind9, authoritative for the AD zones out of the directory. Recursion off. No client-facing service. allow-transfer restricted to the signer alone, with TSIG. From the network\u0026rsquo;s point of view the DC does not serve DNS at all, and the only thing that ever queries it is the next box along.\n2. The signing instance — a normal BIND zone that transfers the data in. This is where inline-signing and the dnssec-policy live, on a zone of the same name that holds the transferred copy. BIND keeps the unsigned copy it received and the signed copy it publishes as separate things, and re-signs as new transfers arrive. Because it is a transfer rather than a shared file, this instance can sit on the DC itself — a second named on its own address — or on a separate host, and the configuration is nearly identical either way.\nThe trade is the one it looks like: on the DC is one fewer machine to run, while a separate host keeps private keys off the box that holds the directory. Both are defensible, and the choice does not change anything else in the pipeline. What matters is that signing happens once, at a defined point, under a key policy — the difference between DNSSEC you operate and DNSSEC that expires at three in the morning.\n3. The authoritative servers — the only thing clients see. Plain secondaries of the signed zone. They hold no keys, do no signing, and have no path to the directory. If one is compromised, the attacker has a copy of a zone and no ability to forge a new record that will validate.\n4. The resolvers. Validating recursive resolvers for the estate, which conditionally forward the internal zones to those authoritative servers and walk the public tree for everything else. These are what client resolv.conf points at.\nTwo properties fall out of this arrangement that are worth stating on their own, because they are the whole point:\nThe machine holding the directory is not reachable by the machines using it. A domain controller is the highest-value host in the estate. Giving it a client-facing network service — one that answers unauthenticated UDP from any workstation — is a poor trade for a service that other machines can perform. Signing happens once, at a defined point. The usual objection to signing an AD zone is churn — that the content changes too often for signatures to keep up. That objection is really an objection to the default layout, not to signing. With client registration moved out, as the next section argues it must be, the AD zone changes when a domain controller is promoted or demoted and roughly never otherwise. A zone written only by DCs is a stable zone, and a stable zone is an unremarkable thing to sign. Keeping clients out is not only a security control; it is what makes the zone quiet enough to sign cleanly. How the Signing Is Actually Wired Up The shape above is the important part, but \u0026ldquo;a normal BIND zone that transfers the data in\u0026rdquo; deserves showing rather than describing, because the first attempt at this usually founders on trying to sign the DLZ directly.\nOn the DC, the DLZ side stays deliberately dull. It serves the directory, it hands the zone to exactly one peer, and it does nothing else:\nkey \u0026#34;transfer-to-signer\u0026#34; { algorithm hmac-sha256; secret \u0026#34;...\u0026#34;; }; options { recursion no; allow-query { key transfer-to-signer; localhost; }; allow-transfer { key transfer-to-signer; }; notify no; }; dlz \u0026#34;AD DNS Zone\u0026#34; { database \u0026#34;dlopen /usr/lib64/samba/bind9/dlz_bind9_18.so\u0026#34;; }; Note the .so carries the BIND major version in its name. That is the version coupling from the previous section made concrete — a BIND upgrade needs the matching Samba module in place before named will start.\nOn the signing side, the zone is an ordinary secondary with the signing options attached. This is the part that cannot live on the DLZ:\ndnssec-policy \u0026#34;ad-internal\u0026#34; { keys { ksk lifetime P365D algorithm ecdsa256; zsk lifetime P90D algorithm ecdsa256; }; }; zone \u0026#34;ad.example.com\u0026#34; { type secondary; primaries { 192.0.2.10 key transfer-to-signer; }; file \u0026#34;ad.example.com.axfr\u0026#34;; inline-signing yes; dnssec-policy \u0026#34;ad-internal\u0026#34;; allow-transfer { key transfer-to-public; }; also-notify { 192.0.2.20; 192.0.2.21; }; }; That this works on a secondary is the load-bearing detail, and the ARM states it directly:\nIf yes, BIND 9 maintains a separate signed version of the zone. An unsigned zone is transferred in or loaded from disk and the signed version of the zone is served with, possibly, a different serial number.\nSo BIND keeps two copies — the unsigned one it received, written to file, and the signed one it serves, written alongside it with a .signed extension. A transfer arrives, the signed version is regenerated, and the downstream servers get a notify, because at this point it is an entirely ordinary zone. inline-signing yes is in fact the default once a dnssec-policy is attached; it is written out above because a config that says what it does is worth the line.\nKey rollover comes with the policy rather than with a cron job. ZSK rollovers need no input at all; KSK rollovers need the new DS getting to the parent, which is the CDS/CDNSKEY automation from later in this post. rndc dnssec -status ad.example.com tells you where each key is in its lifetime.\nWhether that block runs on the DC or on its own host is a matter of which address primaries points at. On the DC it is a second named instance on a second address, transferring from the first over the loopback or a management interface. That is what \u0026ldquo;BIND with DLZ can be the signer\u0026rdquo; means in practice: same software, same box if you like, but the signing happens on the transferred copy rather than on the DLZ zone.\nThe one operational catch is notify — and it is on the DLZ side. ISC\u0026rsquo;s manual is blunt about it: DLZ \u0026ldquo;has no built-in support for DNS notify\u0026rdquo;, so secondary servers are not automatically informed of changes to the zones in the database. Samba can change a record in the directory and the DLZ instance has no idea it should tell anybody.\nSo the first hop is a poll, not a push. The signer refreshes on the SOA timer of the zone it is transferring, which means:\nPropagation from a DC change to a signed, published record is bounded by that refresh interval, not by seconds. Promote a DC and its new _msdcs records appear downstream up to one refresh later. That interval is the dial to turn if the delay matters. It is a trade against how often you want the DLZ queried, which brings up the other thing ISC says about DLZ: it does real-time database lookups with no caching and is \u0026ldquo;not recommended for use on high-volume servers\u0026rdquo;. Both of those are arguments for this topology rather than against it. The only client the DLZ instance ever has is the signer, asking once per refresh. The estate\u0026rsquo;s actual query load lands on the standalone authoritative servers, which are serving a plain signed zone file at full speed. Everything downstream of the signer is conventional: the authoritative servers are secondaries of the signed zone, they get a proper notify, and they never touch the directory or a key.\nClients Must Not Write the Zone That Holds the Locators This is the important one, and it is a design rule rather than a setting.\nActive Directory registers records dynamically. A machine joins, and it registers itself; a DC starts, and it registers the SRV records that advertise its services. \u0026ldquo;Secure\u0026rdquo; dynamic update means those updates are authenticated — the machine proves it is who it says it is with its own credentials, and there is a per-record ACL so a machine can generally only modify a record it created.\nRead that carefully, because the guarantee is narrower than it first appears. Secure dynamic update authenticates who is writing. It does not evaluate what the record means. And the set of accounts entitled to write is far wider than people assume: in a default AD-integrated zone the Authenticated Users group holds Create All Child Objects on the zone container in the directory, because ADIDNS stores every record as an AD object under CN=MicrosoftDNS,DC=DomainDnsZones. Not just machine accounts — any authenticated account in the domain, including the one belonging to whoever opened the invoice attachment this morning.\nWhich produces two failure modes in a zone that holds both client records and service locators:\nNames that do not exist yet belong to nobody. Per-record ACLs protect a record that already has an owner. A name that has never been registered has no object to enforce an ACL on, so the first account to create it gets it. That is the mechanism behind the whole family of ADIDNS attacks, the sharpest being a wildcard: create * and every name in the zone that nobody has explicitly claimed — typos, decommissioned hosts, wpad — resolves to the attacker. Existing records are untouched, which is exactly why it goes unnoticed.\nThe blast radius includes the locators. The records under _msdcs are how every client in the domain finds a domain controller and a KDC. They are the most security-critical records you own, and in a default deployment they sit in the same zone that four hundred laptops write to every time they get a DHCP lease.\nWho may write which zone: one combined zone against a split by writer Default \u0026#8212; one zone ad.example.com locators and client records together _ldap._tcp.dc._msdcs. SRV _kerberos._udp SRV dc01 A laptop-042 A * A \u0026#8592; unclaimed a name nobody has registered has no ACL to enforce every joined machine may write all of it One compromised machine account can rewrite the records that locate a domain controller. Split by who writes ad.example.com \u0026#8212; DCs only no client has an update path into this zone _ldap._tcp.dc._msdcs. SRV _kerberos._udp SRV dc01 A NS delegation dyn.ad.example.com or a separate Samba zone with its own ACL laptop-042 A every joined machine may write no path to locators Who may write what. On the left, one zone, and every joined machine is an authorised writer in the zone that holds the DC locators. On the right, the AD zone is written only by DCs, and client registrations land in a separate sub-zone where the worst a compromised machine can do is lie about itself. So the rule: the zone that holds the locators is written by domain controllers, and by nothing else. Client dynamic registration goes somewhere else.\nSomewhere else can be either of two things, and both are fine:\nA delegated sub-zone served elsewhere. The AD zone holds an NS delegation for, say, dyn.ad.example.com, and the records land on a separate server that accepts the updates. Samba\u0026rsquo;s partitions never take a client write at all. A separate zone in Samba with its own update ACL. Still in the directory, but its own zone, so a client write has no path to _msdcs or to a DC\u0026rsquo;s own records. The first gives the harder separation; the second is less to run. What fails the rule is neither of those. It is the default, where the two live together. Which is how most domains are still running, because nobody chose it and nobody has been back to look.\nAnd there is a second dividend, the one that ties this back to the pipeline. A zone that only domain controllers write is a zone that hardly ever changes: a DC promotion, a DC demotion, and otherwise silence. All the churn in a default AD zone is client registration. Take that out and the objection to signing the AD partitions goes with it — there is no stream of updates for the signatures to chase, so signing it is routine rather than a fight. The write discipline and the signing are the same decision seen twice.\nAlongside either of those, there is a permission to go and look at — and on a Microsoft DC it is the direct fix. Because ADIDNS records are directory objects, the entitlement comes from an AD ACL rather than from anything in the DNS protocol: Authenticated Users holding Create All Child Objects on the zone container. Tightening that is the cleanest mitigation for the unclaimed-name and wildcard problem, and in many estates the permission can be removed outright once you know what actually needs to self-register. Samba\u0026rsquo;s AD DNS stores its records in the directory the same way, so the same question applies — go and look at what your zone container actually grants.\nBut Hang On — Should Clients Be Registering At All in 2026? Everything above assumes client dynamic update is a thing you need and are trying to make safe. Before accepting that, it is worth asking the question nobody asks, because the answer has changed since this behaviour was designed.\nDynamic DNS registration was built for a desktop. A box under a desk, one network cable, one address it kept for years. In that world a machine registering its own name was tidy and basically true.\nNow look at what a client is in 2026. It wakes up on home wifi. It comes into the office and joins the corporate wireless. It goes into a docking station and picks up a wired address as well. Somebody starts the VPN and a tunnel adapter appears with a third address. They go to a coffee shop, tether to a phone, and the VPN comes back on a fourth. That is one machine, one name, and half a dozen addresses in a working day — and by default it will try to register a good number of them.\nSo the zone fills up with claims that were true once.\nWindows registers every adapter it has, unless somebody has been round and unticked Register this connection\u0026rsquo;s addresses in DNS per interface. A docked laptop on the VPN is a machine with three live adapters and an opinion about all of them. Multiple A records for one name is not an error state, it is the normal result. A lookup returns them all, clients try them in whatever order they fancy, and connections to that name fail in proportion to how many of the addresses are dead. This is the mechanism behind \u0026ldquo;remote support can see the machine one minute and not the next\u0026rdquo;. VPN addresses are the worst of them, because a tunnel address is valid for an hour and the record outlives it. The tunnel drops, the pool address goes to somebody else, and the name now points at a colleague. Docking stations muddy the identity itself. Unless MAC address pass-through is configured, the lease belongs to the dock rather than the laptop, so a hot-desk estate has names, leases and machines drifting apart from each other daily. And ownership makes it permanent. A record can only be updated by the account that created it. When a record was made by DHCP under one credential and the machine later tries to update it under its own, the update fails, quietly, and the stale address stays exactly where it is. The clean-up story is not the rescue it sounds like either. Samba has had scavenging since 4.9, but it is off by default (dns zone scavenging = yes, with samba-tool dns zoneoptions --aging=1), and Samba themselves say it \u0026ldquo;should only be enabled on new zones or new installations\u0026rdquo;, because older versions marked dynamic records as static and static ones as dynamic. On the estates most likely to be full of rubbish — the ones that have been running for years — the tool for clearing it up is the one you are advised not to switch on. It has also had a CVE of its own.\nSo put the question directly: what actually consumes a laptop\u0026rsquo;s A record?\nIn most outfits, very little. Users connect to servers; servers do not connect to laptops. The genuine consumers are remote support tools, RDP to a named workstation, and inventory or monitoring — and nearly all of that tooling maintains its own inventory and works from an agent checking in, because it could never rely on DNS for mobile clients in the first place.\nWhich suggests inverting the default:\nServers and infrastructure get records from provisioning. With NetBox and Ansible already in the picture, the record is created by the same thing that created the machine, it is correct by construction, and it is removed when the machine is. The stable wired estate can take records from DHCP if something actually needs them, with one credential owning them so updates do not fail. Mobile clients register nothing at all. They are consumers of DNS, not publishers of it. If something needs to reach a laptop, it needs an agent, not an A record. You end up in the same place the security argument put you, from a completely different direction. Fewer writers means a smaller ADIDNS surface, a zone that is not full of expired claims, and — back to the pipeline — a zone quiet enough to sign without thinking about it.\nThe security case says clients must not write the zone that holds the locators. The operational case asks why they are writing DNS at all. In 2026, for a fleet that changes address five times a day, \u0026ldquo;they are not\u0026rdquo; is a perfectly good answer, and considerably less work than making their mess safe.\nSplit Views, and Where the Trust Comes From The last piece ties the two halves of the post together.\nThere are two views of the namespace. A public zone, published to the internet, holding the handful of names the world needs. And an internal view — the AD-derived content, every joined host, every service locator, the site topology — which the world has no business seeing. That content is a map of the estate, and it should be unreachable and untransferable from outside.\nBut the trust for the internal view comes from the public side, and that is what makes this design better than the usual internal-DNS island:\nexample.com is public and signed, with its DS in the parent and a chain to the root. ad.example.com is delegated from it. The public parent publishes the delegation and a DS for the internal zone\u0026rsquo;s key. The internal authoritative servers serve the signed ad.example.com. The internal resolvers validate it — root → com → example.com → ad.example.com — using nothing but the root trust anchor they already had. No local trust anchor. No island. No hand-distributed key. The data never leaves the building, and the validation path is the ordinary public one. When you add a resolver, it validates internal names correctly with no DNSSEC configuration at all.\nBeing straight about the trade, because there is one. Publishing a delegation and a DS in the public zone means the existence of ad.example.com, and the names of its name servers, are public. The contents are not, and never are — but you have told the world that the zone exists. In exchange, every resolver you own validates internal names against the real root. That is a good trade for most estates, and it should be a deliberate one rather than a surprise.\nTwo things to get right alongside it:\nKeep the internal view unenumerable and untransferable. allow-transfer on the internal authoritative servers is for the signer and your own secondaries, nothing else. And bear in mind that NSEC authenticated denial lets anyone who can query the zone walk it end to end; NSEC3 raises that cost, but the real control is that outsiders cannot reach the servers at all. Automate the DS. A DS in the parent that stops matching the child\u0026rsquo;s key takes the whole internal domain to SERVFAIL. CDS/CDNSKEY exist so the child can signal a key change and the parent can pick it up without a human editing a record during a rollover. If the parent zone is at a registrar or provider that supports it, use it; if not, the rollover procedure needs writing down before the first roll, not during it. Making Windows and Linux Actually Validate Everything up to here has been about publishing a zone that can be verified. None of it does anything until something on the client side insists on verifying it. A perfectly signed zone and a client that never checks a signature produce exactly the same experience as an unsigned zone, right up until the day they don\u0026rsquo;t.\nThere are only two places validation can happen, and the difference between them is the difference between a security control and a polite suggestion.\nValidate at the resolver, and trust the AD bit. The client asks a resolver, the resolver does the cryptography, and it reports the result by setting one bit — AD, authenticated data — in the reply. The client believes the bit. This is the model Windows uses, and it is only as strong as the path between the client and the resolver, because anything that can answer as the resolver can set that bit.\nValidate on the client itself. The machine runs its own validating resolver, so the \u0026ldquo;path to the resolver\u0026rdquo; is a loopback socket inside the machine and there is nothing left to spoof. This is stronger, and on Linux it is entirely achievable.\nThe Default Is Nothing Before configuring anything, it is worth seeing what a mainstream Linux workstation does out of the box. This is a Fedora 44 machine, systemd 259, untouched:\n$ resolvectl status | head -3 Global Protocols: LLMNR=resolve -mDNS -DNSOverTLS DNSSEC=no/unsupported resolv.conf mode: stub $ grep options /etc/resolv.conf options edns0 trust-ad $ dig +dnssec cloudflare.com A | grep flags ;; flags: qr rd ra; QUERY: 1, ANSWER: 3, AUTHORITY: 0, ADDITIONAL: 1 Read those three together, because they tell a small story.\nThe stub is configured with trust-ad — it has been told to believe the AD bit. systemd-resolved reports DNSSEC=no/unsupported, so it is not validating anything itself. And the reply for a signed zone comes back with flags qr rd ra and no ad — nothing anywhere in that path claimed to have validated.\nThat is not a misconfiguration. That is the default. A client can be told to trust an assertion that nothing in the chain is making. Nowt is checking it. Worth sitting with for a minute before configuring anything else.\nLinux Three options, in increasing order of how little you have to trust the network.\n1. systemd-resolved, validating locally. A drop-in rather than editing the shipped file:\n# /etc/systemd/resolved.conf.d/dnssec.conf [Resolve] DNSSEC=yes DNSOverTLS=opportunistic Then systemctl restart systemd-resolved and confirm with resolvectl status that the line now reads DNSSEC=yes.\nThe setting to be careful about is the middle one. DNSSEC=allow-downgrade looks like a sensible compromise and is not a security control — resolved.conf(5) says so itself:\nNote that this mode makes DNSSEC validation vulnerable to \u0026ldquo;downgrade\u0026rdquo; attacks, where an attacker might be able to trigger a downgrade to non-DNSSEC mode by synthesizing a DNS response that suggests DNSSEC was not supported.\nAn attacker who can forge answers is exactly the attacker DNSSEC exists to stop, so a mode they can switch off by forging an answer buys you nothing against them. It is yes or it is decoration.\n2. A real validating resolver on the host. systemd-resolved\u0026rsquo;s validator is convenient rather than thorough. Where it matters, run Unbound or BIND on the loopback and point the stub at it:\n# unbound: validate against the root anchor, refuse to be stripped server: module-config: \u0026#34;validator iterator\u0026#34; auto-trust-anchor-file: \u0026#34;/var/lib/unbound/root.key\u0026#34; harden-dnssec-stripped: yes val-permissive-mode: no # the internal zone is reached like any other name — no local anchor needed forward-zone: name: \u0026#34;ad.example.com.\u0026#34; forward-addr: 192.0.2.53 BIND\u0026rsquo;s equivalent is one line — dnssec-validation auto; — which uses its built-in copy of the root anchor and manages the rollover for you.\n3. Estate-wide, at the resolvers you already run. These are stage four of the pipeline earlier in this post. Validation happens there, clients trust the AD bit, and the hop between them is the thing you have to protect — with DoT, or with a network you are willing to make that assumption about.\nAnd note what is not in any of those configurations: a trust anchor for the internal zone. Because ad.example.com is a delegation inside a publicly signed zone, every one of these validates internal names through the ordinary chain from the root. That is the design from the previous section paying for itself. The alternative is pushing a local anchor to every client and resolver in the estate, and re-pushing it at every rollover.\nWindows The important thing first, because it is routinely misunderstood: the Windows DNS Client does not validate DNSSEC. It performs no cryptography, checks no signature, and holds no trust anchor. It is a stub resolver, and it always has been.\nWhat you can do is force it to refuse answers that were not validated by the server on its behalf. That is the Name Resolution Policy Table, and it is per-namespace rather than global:\n# Require validated answers for the internal zone Add-DnsClientNrptRule -Namespace \u0026#34;.ad.example.com\u0026#34; ` -DnsSecEnable -DnsSecValidationRequired # What is actually in force on this machine, including from Group Policy Get-DnsClientNrptPolicy -Effective Get-DnsClientNrptRule For the estate, the same thing lives in Group Policy under Computer Configuration → Policies → Windows Settings → Name Resolution Policy: create a rule for the namespace, tick the DNSSEC option, and tick the requirement that the client check the data was validated by the DNS server.\nTwo things follow from that, and both matter.\nSomething upstream still has to do the validating. The NRPT rule makes the client demand the AD bit; it does not create one. The resolver those clients point at must be a validating resolver, or every name in that namespace fails.\nAnd this is why the NRPT rule has IPsec options next to it. Microsoft put them there for the reason set out at the top of this section: requiring a bit that any on-path attacker can set is not much of a requirement. If you are relying on the resolver-validates model on an untrusted network, the last hop needs protecting — IPsec between client and resolver, or DoT where the resolver supports it.\nIt Fails Closed, So Roll It Out in That Order Forcing validation converts a class of silent compromise into a class of loud outage. That is the correct trade, and it is still an outage: an expired RRSIG, a DS in the parent that no longer matches after a rollover, or a resolver that cannot reach the parent zone all produce SERVFAIL, and SERVFAIL for _ldap._tcp.dc._msdcs means the domain is down rather than degraded.\nSo do it in this order:\nTurn validation on at the resolvers first, and leave the clients alone. Watch for SERVFAIL in the resolver logs for a couple of weeks — this is where you find the zone that has been quietly broken for a year. Automate the DS before forcing anything, per the previous section. Most self-inflicted DNSSEC outages are a rollover where the parent was never updated. Then force the clients, one namespace at a time, starting with your own workstation and a test OU rather than the whole estate. The failure you are engineering against is a client being handed a forged domain controller. The failure you are risking is a client being handed nothing at all. The second one is recoverable and obvious; the first is neither. You will hear about the outage inside a minute. You would never have heard about the other one at all.\nEncrypting the Last Hop: DoT and DoH on the Internal Resolvers The validation section left one thing hanging. In the resolver-validates model — the one Windows gives you — the client is trusting a single bit set by the resolver, and that bit is only worth the path it travelled over. Something has to protect that path.\nBIND does support both encrypted transports natively, so this is a configuration job rather than a procurement one:\nDNS over TLS — a tls block referenced from listen-on, conventionally on port 853. DNS over HTTPS — the same tls block plus an http block, on 443. Outgoing DoT, because forwarders takes a TLS transport per address or for the whole list. Zone transfers over TLS, because a type secondary zone\u0026rsquo;s primaries statement takes one too — which is directly useful to the pipeline earlier in this post. First, be clear about what this buys, because DoT and DNSSEC get conflated constantly and they are not alternatives. DNSSEC authenticates the data, all the way back to the zone that published it. DoT protects the conversation with the resolver. One survives a hostile resolver and a hostile network between resolvers; the other stops the machine on your wifi reading and rewriting what your laptop asked. You want both, and neither substitutes for the other. Encrypting the hop to a resolver that does not validate is a private conversation with something that can still be lied to.\nServing It tls internal-resolver { key-file \u0026#34;/etc/pki/dns/resolver.key\u0026#34;; cert-file \u0026#34;/etc/pki/dns/resolver.pem\u0026#34;; protocols { TLSv1.3; }; }; http internal-doh { endpoints { \u0026#34;/dns-query\u0026#34;; }; }; options { dnssec-validation auto; listen-on port 53 { 192.0.2.53; }; listen-on port 853 tls internal-resolver { 192.0.2.53; }; listen-on port 443 tls internal-resolver http internal-doh { 192.0.2.53; }; listen-on-v6 port 853 tls internal-resolver { 2001:db8::53; }; }; The certificate is the actual work, and it is the part that gets skipped. A client that verifies — which is the entire point — needs a certificate valid for the name it was configured with, issued by something it already trusts. That means your internal CA and your existing certificate automation, not the ephemeral keyword. ephemeral generates a throwaway self-signed certificate; it is there so you can prove the listener works, and it is worthless to any client actually checking.\nForwarding Over It If those resolvers forward anywhere rather than walking the tree themselves, the upstream hop can be encrypted too — and there is a distinction here worth getting right:\ntls upstream { ca-file \u0026#34;/etc/pki/tls/certs/ca-bundle.crt\u0026#34;; remote-hostname \u0026#34;dns.example.net\u0026#34;; }; options { forwarders port 853 tls upstream { 192.0.2.1; }; }; Without remote-hostname, you get encryption without authentication: the traffic is unreadable to a passive observer, and an active attacker who can intercept the connection simply presents their own certificate. With remote-hostname and ca-file, BIND verifies who it is talking to. The first is worth something. Only the second is worth calling a control.\nThe Client Half Is Not Symmetrical This is where a mixed estate gets awkward, and it is the reason to configure both transports rather than picking one.\nLinux does DoT properly. systemd-resolved takes DNSOverTLS=yes for strict mode, and the server can be given the name to verify against:\n[Resolve] DNS=192.0.2.53#resolver.ad.example.com DNSOverTLS=yes DNSSEC=yes As with DNSSEC=, the middle setting is the trap: DNSOverTLS=opportunistic falls back to clear text when TLS is unavailable, which an attacker who can interfere with the connection can arrange.\nWindows does DoH, and not DoT. Client DoH support shipped in Windows 11 and Server 2022, configured per server with a template:\n$doh = \u0026#34;https://resolver.ad.example.com/dns-query\u0026#34; netsh dnsclient add encryption server=192.0.2.53 dohtemplate=$doh DoT, at the time of writing, has only appeared in Insider builds. So on released Windows the encrypted option is DoH or nothing, which is exactly why the NRPT section earlier reached for IPsec instead.\nHence serving both from the same BIND instance. DoT for the Linux fleet and anything else that speaks it, DoH for Windows, one resolver, one certificate.\nWhich One, Where For an internal resolver, DoT is the better transport and DoH is the compatibility answer.\nDoT sits on its own port. You can see it, allow it, deny it, and alert on anything doing DNS that is not doing it. DoH\u0026rsquo;s advantage — indistinguishable from ordinary web traffic on 443 — is a genuine benefit on a hostile network and a nuisance on your own, where being able to tell what is DNS is a feature you paid for. On the estate you control, prefer the transport you can observe, and run DoH because Windows leaves you no choice rather than because it is better.\nWhile You Are There: Encrypt the Transfers The pipeline earlier in this post moves the AD zone by AXFR, and the split-view section made the point that its contents are a map of the estate. TSIG authenticates those transfers; it does not conceal them. Since a secondary\u0026rsquo;s primaries statement accepts a TLS configuration, the transfer can run over TLS as well — RFC 9103 if you want the standard:\nzone \u0026#34;ad.example.com\u0026#34; { type secondary; primaries { 192.0.2.10 port 853 tls xfr-tls key transfer-to-signer; }; ... }; Authenticated by the key, encrypted by the transport. If any hop in that pipeline crosses a site link, a hypervisor you share, or anything you would not happily put a hub on, it is worth the twenty minutes.\nWhat It Does Not Fix It is not validation. Covered above, and worth repeating because vendors sell \u0026ldquo;secure DNS\u0026rdquo; meaning encryption alone. It does not hide anything from the resolver. The resolver sees every query in full. Encryption protects the path, not the privacy of the lookup from the operator — which is fine when the operator is you. It does nothing for a client that does not verify the certificate, and opportunistic modes are downgradeable by exactly the attacker you are worried about. And it is not a reason to put a listener on the domain controller. Encrypted DNS on the DC would be solving the wrong problem beautifully. The DC still answers nobody. How to Check What You Have Commands to run against your own estate. The outputs are the interesting part, and a couple of them tend to be uncomfortable reading the first time.\n# Walk the tree yourself, one delegation at a time dig +trace dc01.ad.example.com # Validate, and show the chain being built delv +rtrace +vtrace ad.example.com SOA # Is the resolver you are pointed at actually validating? # A deliberately broken test name must come back SERVFAIL, not an address dig @\u0026lt;resolver\u0026gt; dnssec-failed.org A # What does the estate advertise as a domain controller? dig SRV _ldap._tcp.dc._msdcs.\u0026lt;domain\u0026gt; dig SRV _kerberos._udp.\u0026lt;domain\u0026gt; # Is a DC answering for names it has no business answering? # Ask it for something it is not authoritative for. \u0026#34;recursion requested # but not available\u0026#34; is the answer you want. An actual address means it is # serving the estate — recursing if it is BIND, relaying to the forwarder # if it is SAMBA_INTERNAL. Either way it should not be doing that. dig @\u0026lt;dc\u0026gt; www.example.org A # Is the DC configured as the estate\u0026#39;s DNS relay? grep -E \u0026#39;dns forwarder|server services\u0026#39; /etc/samba/smb.conf # Will a DC hand its zone to anybody who asks? dig @\u0026lt;dc\u0026gt; AXFR ad.example.com # Which backend is this DC running, and does named have the module? grep -r dlz /etc/named.conf /var/lib/samba/bind-dns/ 2\u0026gt;/dev/null samba-tool dns query \u0026lt;dc\u0026gt; \u0026lt;domain\u0026gt; @ ALL # Is the internal zone chained to the public parent? dig DS ad.example.com @\u0026lt;public-authoritative-for-example.com\u0026gt; # What does the local stub actually do with the AD bit, and is the # hop to the resolver encrypted? resolvectl status # DNSSEC= and DNSOverTLS= per link grep options /etc/resolv.conf # trust-ad, trusting whom exactly? # Did anything in the path claim to have validated? Look for \u0026#34;ad\u0026#34; in the flags dig +dnssec ad.example.com SOA | grep flags # Validate independently of whatever the local resolver believes delv ad.example.com SOA # \u0026#34;fully validated\u0026#34; is the line you want # Is the resolver actually listening for DoT, and does its certificate # match the name clients are configured with? kdig +tls @192.0.2.53 ad.example.com SOA openssl s_client -connect 192.0.2.53:853 \\ -servername resolver.ad.example.com \u0026lt;/dev/null 2\u0026gt;/dev/null \\ | openssl x509 -noout -subject -dates And on a Windows client, to see whether it is demanding anything at all:\nGet-DnsClientNrptPolicy -Effective # the rules actually in force Resolve-DnsName ad.example.com -DnssecOk Get-DnsClientDohServerAddress # is the hop to the resolver encrypted? The two that most often produce a surprise are the recursion check and the AXFR attempt. If a DC answers either of them for an arbitrary client, the pipeline in this post is not in place, whatever the diagram on the wiki says.\nThe third is resolvectl status on a machine nobody has touched. DNSSEC=no/unsupported alongside trust-ad in resolv.conf is the normal state of a Linux desktop, and it means the signing work described above is currently being checked by nobody.\nThe Short Version A resolver is born knowing the root servers and one key, and learns everything else by being told. Names are found by walking down delegations, and services are found by asking for an SRV record — so by the time a client starts doing Kerberos with a domain controller, the identity of that domain controller came from a DNS answer. Discovery by DNS is correct and fine. Authorisation by DNS is not, and a surprising amount of infrastructure quietly does it anyway.\nDNSSEC is what makes those answers verifiable: signatures on every set, a DS in each parent, a chain to a single trust anchor at the root, and authenticated denial so a name cannot be made to disappear. It buys origin authentication and integrity — not privacy, and not correctness. It signs whatever the zone says, which is why it cannot rescue a zone that untrusted machines are allowed to write. And it needs a parent: an invented internal TLD has nowhere to put a DS, so .local, .lan, .internal and home.arpa all leave you either unsigned or running a private island of hand-distributed keys.\nFor an Active Directory domain, that produces a design rather than a list of settings. Use a delegation inside a public zone you own, so the trust comes down the ordinary chain from the root while the data never leaves. Serve the AD partitions with BIND and dlz_bind9 — DLZ is not the risk, it is how you get the zone out of the directory — and let the DC be a hidden primary that transfers out and answers nobody else. A DLZ zone cannot itself be signed, so the signing is a normal secondary zone holding the transferred copy, with inline-signing and a dnssec-policy on it: a second named on the DC if you want fewer machines, a separate host if you want private keys off the directory. Either way it is signed once, at a defined point, under a key policy — and the first hop is a poll rather than a push, because DLZ cannot send a notify. Publish from standalone authoritative servers that hold no keys and have no route to the directory.\nAnd keep clients out of the zone that matters. Secure dynamic update authenticates the writer, not the meaning, and in a default AD-integrated zone the writers are Authenticated Users — every account, not just every machine — so a zone holding both laptop records and _msdcs locators is one phished user away from a client being told, with a perfectly valid signature, that the domain controller is somewhere else. Client registrations belong in a sub-zone, delegated out or separate in Samba, where the worst a compromised account can do is lie about itself.\nThough the better question is whether clients should be registering at all. A 2026 laptop has an address on home wifi, another on the office wireless, another through the dock and another on the VPN, and it will cheerfully publish most of them. What consumes a laptop\u0026rsquo;s A record is almost nothing — the tools that need to reach a workstation keep their own inventory, because DNS was never reliable for mobile clients anyway. Records for servers should come from provisioning, and the fleet should register nowt.\nThen make something check the signatures, because none of the above is worth anything until a client refuses an answer. On Linux that means DNSSEC=yes in systemd-resolved, or a real validating resolver on the loopback — never allow-downgrade, which an attacker capable of forging answers can simply switch off by forging an answer. On Windows it means accepting that the DNS client never validates anything itself, and using an NRPT rule to make it require an answer the resolver validated, with the last hop protected because that requirement is a single bit. Turn it on at the resolvers first and watch for SERVFAIL, automate the DS, then force the clients. It fails closed, which is the right way round and still an outage.\nNone of this is exotic. It is delegation, transfer and signing — the three things DNS has always done — arranged so that the machine holding your directory is not the machine taking questions from the car park, and so that when something does lie to a client, the client notices.\n","permalink":"https://blogs.damiendye.uk/en/dns/samba4-securing-ad-records-with-dnssec/","summary":"Every domain-joined machine finds its domain controller by asking DNS for an SRV record, so the _msdcs locators are the most security-critical records you own. This is how to publish and sign them properly from a Samba4 DC: BIND with dlz_bind9 reading the directory, inline signing, a hidden primary that clients never reach, and client dynamic updates kept out of the zone that holds the locators. Then how to force Windows and Linux clients to actually check the signatures, because a signed zone nobody validates behaves exactly like an unsigned one.","title":"Samba4 and Securing AD Records Using DNSSEC"},{"content":"The Bay That Takes Anything The pitch for a tri-mode adapter is genuinely good, and it is worth stating properly before pulling it apart.\nBuy a chassis with a U.3 backplane and a tri-mode controller, and every drive bay becomes universal. Slot 0 can hold a 24G SAS drive, slot 1 a cheap SATA boot device, slot 2 a Gen4 NVMe SSD, and the adapter negotiates whatever turns up. Broadcom calls the silicon Tri-Mode SerDes; the bay standard is SFF-TA-1001, known as U.3, which defines a common connector for SAS x1/x2, SATA, and NVMe at x1, x2 or x4. The management side is SFF-TA-1005, Universal Backplane Management, which is how the enclosure works out what it is actually talking to and drives the right activity LEDs.\nFor anyone specifying servers, that solves a real and annoying problem. You no longer have to decide the storage protocol at purchase order time, or keep two chassis SKUs, or discover that the NVMe-capable bays are the four on the left and your drives went into the other twenty. One part number covers the fleet, and a SAS estate can move to NVMe a drive at a time instead of a chassis at a time.\nNone of that is marketing. It is the reason these adapters sell, and it would be daft to pretend otherwise.\nBut the flexibility is not free, and the bill is not paid in pounds. It is paid in queues.\nWhat Actually Happens to the Drive An NVMe SSD is a PCIe endpoint. In a direct-attached server, its four lanes run to the CPU\u0026rsquo;s root complex — through a retimer or a PCIe switch, but electrically and logically it is a device on the PCIe bus. The kernel enumerates it, binds the nvme driver, and from that point the driver talks to the drive\u0026rsquo;s registers directly.\nPut the same drive behind a tri-mode adapter and that stops being true.\nThe drive\u0026rsquo;s lanes now terminate at the controller. Broadcom\u0026rsquo;s documentation calls the relevant block the PCIe device bridge, and the word bridge is doing a lot of work: this is not a transparent switch that forwards your CPU\u0026rsquo;s transactions to a drive it can still see. The adapter is the PCIe endpoint your host enumerates. The drive is a target hanging off the far side of it, and the controller\u0026rsquo;s firmware re-originates every I/O.\nDirect-attached NVMe against the same drives behind a tri-mode adapter Direct-attached Behind a tri-mode adapter CPU root complex CPU root complex x4 x4 x4 x4 four independent links about 7 GB/s each, in parallel x8 Gen4 everything below shares this Tri-mode controller the only PCIe endpoint the host enumerates NVMe NVMe NVMe NVMe nvme0n1 nvme1n1 nvme2n1 nvme3n1 NVMe NVMe NVMe NVMe sda sdb sdc sdd nvme driver one queue pair per CPU core, per drive bandwidth grows as you add drives mpt3sas driver \u0026#8212; the drives are SCSI targets queue depth 128 each, one tag pool between them bandwidth stops at the adapter The same four drives, wired two ways. On the left each drive owns four lanes to the root complex. On the right the lanes stop at the adapter, and everything downstream shares one x8 uplink and one controller. So the adapter is not passing your NVMe commands through. It is terminating them, and speaking to the drive on your behalf.\nWhich raises the question of what protocol it speaks to you.\nThe OS Never Sees an NVMe Drive It speaks SCSI.\nPlug an NVMe SSD into a Broadcom tri-mode HBA and it does not appear as /dev/nvme0n1. It appears as /dev/sdb, bound to mpt3sas — the same driver that has been running LSI SAS controllers for over a decade. nvme list returns nothing. lsblk -o NAME,TRAN reports the transport as sas. As far as every layer of the storage stack above the driver is concerned, you bought a SAS disk.\nThis is not a bug or a firmware limitation waiting to be fixed. It is the design. Presenting everything as a SCSI target is exactly how one adapter serves three protocols: the controller normalises SAS, SATA and NVMe into a single device model, and the host gets one driver, one enumeration path, one set of tooling. The flexibility in the marketing and the SCSI presentation in dmesg are the same architectural decision looked at from either end.\nThe 9500 and 9600 generations do add a passthrough mechanism so vendor tooling can reach a drive\u0026rsquo;s NVMe admin commands, and Broadcom\u0026rsquo;s newer parts are much better at surfacing drive health than the 9400 was. But that is a management side channel. The data path — every read and write your workload issues — still runs down the SCSI stack.\nAnd the SCSI stack has a queue model that predates flash by twenty years.\nThe Queue Model You Just Gave Up This is the part that actually costs you performance, and it is worth being precise about, because \u0026ldquo;NVMe is faster than SAS\u0026rdquo; is not the reason.\nNVMe\u0026rsquo;s central design decision was not a faster wire. It was to stop pretending that a storage device is a single serialised thing.\nThe specification allows up to 65,535 I/O queue pairs, and that figure gets quoted in every NVMe explainer going. It is the wrong number to reach for. No drive implements anything close to it, so anyone who has actually looked at a running system can wave the comparison away — and they would be right to. The real number is smaller, unglamorous, and makes the point better.\nSo here is a real drive. Not an enterprise part: a 256 GB SK Hynix OEM SSD, the kind soldered into a mid-range laptop, in a 16-core machine.\n$ nproc 16 $ cat /sys/class/nvme/nvme0/queue_count 17 $ ls /sys/block/nvme0n1/mq | wc -l 16 $ cat /sys/block/nvme0n1/queue/nr_requests 1023 Seventeen queues: one admin queue, and sixteen I/O queues for sixteen cores. Each 1023 commands deep. The mapping is one-to-one — every hardware queue is bound to exactly one CPU:\n$ cd /sys/block/nvme0n1/mq \u0026amp;\u0026amp; grep -H . */cpu_list 0/cpu_list:1 1/cpu_list:9 2/cpu_list:3 3/cpu_list:11 ... One CPU per queue, all the way down — hardware queue 0 serves core 1 and nothing else.\nThat is what \u0026ldquo;NVMe has lots of queues\u0026rdquo; actually means in practice. Not 65,535 — one per core, however many cores you have. Linux creates a queue pair per CPU up to whatever the controller will grant, and controllers grant far more than a typical server has cores, so in practice the core count is the number. Put this drive in a 64-core box and you get 64.\nThat per-core split is where the performance comes from:\nA core submits into its own queue. No lock, because no other core touches it. Each queue gets its own MSI-X vector, affinitised to that core. The completion interrupt lands back on the core that issued the I/O, where the relevant cache lines already are. All sixteen cores can be in flight at once without ever contending on a shared structure. Parallelism scales with your core count, and no core\u0026rsquo;s work is ever queued behind another\u0026rsquo;s. A cheap consumer drive does this. It is table stakes.\nNow look at what the drive gets behind the adapter. The numbers below are not estimates — they are constants in the mainline mpt3sas driver.\nThe per-device queue depth is set from ioc-\u0026gt;max_nvme_qd, which the driver takes from what the controller firmware reports and otherwise falls back to a compile-time default in drivers/scsi/mpt3sas/mpt3sas_base.h:\n#define MPT3SAS_SATA_QUEUE_DEPTH\t32 #define MPT3SAS_SAS_QUEUE_DEPTH\t254 #define MPT3SAS_RAID_QUEUE_DEPTH\t128 #define MPT3SAS_NVME_QUEUE_DEPTH\t128 128. A device capable of tens of thousands of outstanding commands is given a queue depth of 128 — and notice it is shallower than the SAS default of 254 sitting two lines above it. The drive\u0026rsquo;s own capabilities never enter into the decision. The number comes from the controller.\nThe hardware queue count is worse, and the driver is candid about it. From mpt3sas_scsih.c:\nshost-\u0026gt;nr_hw_queues = 1; if (shost-\u0026gt;host_tagset) { shost-\u0026gt;nr_hw_queues = ioc-\u0026gt;reply_queue_count - ioc-\u0026gt;high_iops_queues; ... dev_info(\u0026amp;ioc-\u0026gt;pdev-\u0026gt;dev, \u0026#34;Max SCSIIO MPT commands: %d shared with nr_hw_queues = %d\\n\u0026#34;, shost-\u0026gt;can_queue, shost-\u0026gt;nr_hw_queues); } Read that carefully, because three separate things are going on.\nThe default is one hardware queue. nr_hw_queues = 1. Multi-queue only happens on gen35 controllers with the host_tagset feature enabled, and even then it is what the driver\u0026rsquo;s own commit history describes as simulated multiple hardware queues — the I/O controller hardware is a single submission queue with multiple reply queues, and blk-mq is being fitted over the top of that.\nThe queue count comes from the controller, not the core count. It is reply_queue_count minus the high-IOPS queues — the adapter\u0026rsquo;s MSI-X vector allocation. It has nothing to do with how many CPUs you have, and it does not grow when you add drives. This is the exact inversion of the drive above, where the number of queues was the core count.\nAnd the tag pool is shared. host_tagset means exactly what it says: one tag pool for the whole host adapter, and that log line says \u0026ldquo;shared\u0026rdquo; out loud. Every drive on the card draws from the same set of command slots. A twenty-four-bay chassis has twenty-four drives competing for one controller\u0026rsquo;s tags.\nPut the two side by side. That laptop SSD had sixteen private queues of 1023, one per core, answering only to itself. The same drive behind the adapter gets a share of the card\u0026rsquo;s reply queues, 128 outstanding commands, and twenty-three neighbours drawing on the same pool.\nPer-core NVMe queue pairs against one shared adapter tag pool Queues multiply with cores Drives divide one pool core 0 core 1 core 2 core 3 SQ + CQ SQ + CQ SQ + CQ SQ + CQ 1023 deep 1023 deep 1023 deep 1023 deep NVMe SSD /dev/nvme0n1 one submission and completion pair per core own MSI-X vector, completions land on that core no lock, no cross-core contention core 0 core 1 core 2 core 3 Tri-mode controller reply queues come from the card's MSI-X vectors, not from your core count one shared tag pool for every drive sda sdb sdc sdd qd 128 qd 128 qd 128 qd 128 and 20 more bays drawing on the same pool. 128 outstanding per device, whatever the drive can do adding drives divides a fixed resource Left: one queue pair per core, private and 1023 deep, with the completion interrupt landing back on the submitting core — measured on the drive above. Right: every core funnelled into the controller\u0026rsquo;s reply queues, drawing on one shared tag pool, with each drive capped at 128. So the loss is not that SCSI is slow. Modern SCSI on blk-mq is fine. The loss is structural:\nQueues belong to the adapter, not the drive. Adding drives divides a fixed resource rather than adding to it. The tag pool is shared host-wide. One drive under heavy load can starve the others in a way that simply cannot happen when each drive has its own queues. Per-device depth is capped at 128, regardless of what the drive can sustain. Interrupt locality is weakened. Completions arrive on whichever reply queue the controller used, not necessarily the core that submitted. For a queue depth of 1 or 2 — a single-threaded process doing occasional reads — none of this registers. You will measure the same latency either way, within noise. The penalty appears exactly where you bought NVMe to help: many cores issuing many concurrent I/Os. The deeper the workload, the more of the drive you have paid for and cannot reach.\nThe Uplink Is One x8 Slot The queue model is the subtle problem. The bandwidth ceiling is the obvious one, and you can read it off Broadcom\u0026rsquo;s own product briefs without needing a benchmark.\nThe 9500 series HBA is an x8 PCIe Gen 4.0 card. Broadcom\u0026rsquo;s published figures for it are 13,700 MB/s at 256K sequential read and 3M IOPS at 4K random read. The same brief says it supports up to 32 NVMe devices.\nPut those two numbers next to each other and the question answers itself: how many drives does it take to run out of adapter?\nNot many, and fewer every year. A Gen4 x4 SSD does roughly 7 GB/s. A Gen5 x4 SSD does roughly 14. Both are ordinary parts in 2026 — Gen4 is what the used U.2 market is full of, and Gen5 is what you get buying new.\nCeiling Gen4 drives to reach it Gen5 drives to reach it HBA 9500 — 13,700 MB/s sequential 2 1 HBA 9500 — 3M IOPS (4K RR) 3 1–2 eHBA 9600 — 6.4M IOPS (4K RR) ~6 ~3 MegaRAID 9600 — 1.1M RAID 5 IOPS (4K RW) ~1 ~1 Read the top row again. A single Gen5 SSD meets the entire sequential ceiling of an HBA 9500. One drive, in a card rated for thirty-two. Everything after that is capacity. Not performance.\nAnd the bottom row is the one that should stop a purchase order: on the current-generation MegaRAID, a full shelf of NVMe in RAID 5 delivers roughly what one mainstream drive does on its own.\nThe Drives Do Not Even Link at x4 There is a second throttle underneath the shared uplink, easy to miss because it sits in a specifications table rather than a headline.\nDell\u0026rsquo;s PERC 12 User\u0026rsquo;s Guide, covering the H965i tri-mode controllers, says:\nSupports drive speeds for NVMe drives are 8 GT/s (Gen 3) and 16 GT/s (Gen 4) at maximum x2 lane width.\nEach NVMe drive gets two lanes, not four. So before any contention for the uplink, before the tag pool, before the SCSI translation, a Gen4 drive is already down to about 3.5 GB/s — half of what it can do. Put a Gen5 drive in that bay and it negotiates down to Gen4 x2 and delivers roughly a quarter of its rated bandwidth.\nIt is worth being precise about what this does and does not change. It does not mean the adapter goes further. It takes about four x2-limited drives to fill the 9500\u0026rsquo;s uplink instead of two, but only because each drive is contributing half as much. The bottleneck has moved from the uplink down to the drive link. The total you can extract has not improved.\nDirect-attached, those same thirty-two drives would each have their own x4 path to the root complex, at whatever generation the drive and the CPU can negotiate.\nThe RAID 5 row in that table deserves its own look, because it is Broadcom\u0026rsquo;s own number and it is published without spin. From the 9600 series brief:\n900K to 1.1M RAID 5 IOPS (4K RW)\nParity RAID in controller firmware is the most expensive thing you can ask a tri-mode card to do, and this is the current generation doing it. Worth reading before someone specifies RAID 5 across twenty-four NVMe drives and expects twenty-four drives\u0026rsquo; worth of performance.\nDrive count against two tri-mode ceilings, an x8 Gen4 card and an x16 Gen5 card 0 15 30 45 60 75 90 aggregate GB/s 1 2 3 4 5 6 NVMe drives Gen5 direct \u0026#8212; about 14 GB/s each Gen4 direct \u0026#8212; about 7 GB/s each out of reach of either card PERC13, Gen5 x16 \u0026#8212; 52.5 GB/s measured HBA 9500, Gen4 x8 \u0026#8212; 13.7 GB/s 2 Gen4 drives reach the 9500 \u0026#8212; 4 Gen5 drives reach even a PERC13 Vendor and review figures, not measured here. The 9500 is rated for 32 NVMe devices, the PERC13 for 16. Both cards additionally link each drive at x2, which these direct-attach lines do not. How many drives it takes to run out of adapter, against two ceilings. Two Gen4 drives reach the HBA 9500; four Gen5 drives reach even a PERC13. The cards are rated for thirty-two and sixteen devices respectively. What About an x16 Card? The obvious objection to everything above is that the 9500 is an x8 Gen4 card, and the ceiling is an artefact of a narrow host link. Give the adapter sixteen lanes of Gen5 and the problem goes away.\nIt is a fair objection, and it deserves the strongest example rather than a straw man. So take Dell\u0026rsquo;s PERC13 H975i — the current generation, and about as good as tri-mode gets. Its user\u0026rsquo;s guide specifies \u0026ldquo;Gen 4 and Gen 5 PCIe x16 host interfaces\u0026rdquo;, and StorageReview measured 52.5 GB/s and 12.5M IOPS per controller, against up to sixteen NVMe drives.\nThose are serious numbers, and they change the picture substantially. Against the 9500\u0026rsquo;s 13,700 MB/s and 3M IOPS, that is roughly four times the bandwidth and four times the IOPS, spread over half as many drives. Dell did not just widen the pipe — they also halved the fan-out, and the oversubscription ratio improved as a result. On IOPS in particular, 12.5M across sixteen drives is about 780K per drive, which is close to what a mainstream drive delivers on its own. At that point the controller is genuinely not the thing holding you back.\nCredit where it is due, then: a modern x16 Gen5 tri-mode card is a much better piece of engineering than an x8 Gen4 one, and if bandwidth was your only objection, x16 largely answers it.\nThree things it does not fix.\nThe drives still link at x2. This is the one that surprised me. The PERC13 guide, describing a Gen5 x16 controller, still says:\nSupports drive speeds for NVMe drives are 8 GT/s (Gen 3), 16 GT/s (Gen 4), and 32 GT/s (Gen 5) at maximum x2 lane width.\nA wider host link does not widen the downstream drive links. Every NVMe drive on the newest, fastest tri-mode RAID controller Dell sells is still connected by two lanes instead of four, and still gives up half its bandwidth before anything else happens.\nThe queue model is completely untouched. Nothing in this post\u0026rsquo;s queue section is a function of host link width. nr_hw_queues comes from the controller\u0026rsquo;s MSI-X reply queue allocation; the per-device depth of 128 is a driver and firmware constant; the tag pool is shared host-wide because host_tagset says so. Widen the host link to x16, x32, whatever you like — the drives are still SCSI targets sharing the card\u0026rsquo;s queues, there is still no /dev/nvme0n1, and you still cannot pass a drive to a VM.\nAnd x16 does not manufacture bandwidth — it fans out lanes you already had. This is the argument that actually settles it. Sixteen Gen5 lanes into a PERC13 buys you 52.5 GB/s shared across sixteen bays. Those same sixteen lanes wired directly to four Gen5 drives at x4 buy you roughly 56 GB/s across four bays — the same bandwidth from the same lanes, except each drive gets its full x4, its own queue pair per core, and a real nvme device node.\nSo the honest way to describe an x16 tri-mode card is not \u0026ldquo;a faster adapter\u0026rdquo;. It is a lane multiplexer: it converts a fixed lane budget into more drive bays, and charges you the queue model for the conversion. Whether that deal is any good depends on one thing. Bays or parallelism.\nWhat Happens When You Fill All the Bays Which brings us to the case that actually matters, because nobody buys a 24-bay chassis to put four drives in it.\nPast the saturation point, the aggregate line is flat. Adding drives adds capacity, and nothing else — so per-drive performance falls as 1/N. That arithmetic is unforgiving at realistic populations:\nDrives on an HBA 9500 Aggregate Per drive Fraction of a Gen4 drive 2 13.7 GB/s 6.9 GB/s 98% 12 13.7 GB/s 1.14 GB/s 16% 24 13.7 GB/s 0.57 GB/s 8% Look at the bottom row. Twenty-four NVMe drives behind an HBA 9500 deliver about 570 MB/s each. A SATA SSD does around 550. You have bought twenty-four NVMe drives, paid for a tri-mode controller to attach them, and arrived at SATA-class per-drive bandwidth.\nThe IOPS arithmetic is the same shape: 3M spread over twenty-four drives is 125K each, against the 1M a mainstream Gen4 drive manages alone — about an eighth of what you own.\nThe x16 card improves this considerably but does not escape it. A PERC13 at its full sixteen drives is 52.5 GB/s ÷ 16 = 3.3 GB/s per drive, or roughly 23% of a Gen5 drive — and that is before the x2 link halves it again.\nTwo effects at high drive counts are worse than the division suggests:\nTag starvation is cross-device. The shared host tag pool means a single drive under heavy load can consume slots that other drives need. Twenty-four devices each nominally allowed 128 outstanding commands want 3,072 between them, drawn from one controller\u0026rsquo;s can_queue. Head-of-line blocking between separate drives is a failure mode that simply does not exist when each drive owns its queues. Rebuilds hit everything. A parity rebuild across a populated shelf saturates the one shared uplink, so foreground I/O to every other drive on the card degrades at the same time. With drives on independent lanes and software RAID, the rebuild competes for CPU, not for a single pipe. When None of This Matters There is an important counterweight, and it is the reason plenty of 24-bay tri-mode servers run perfectly happily.\nThe adapter ceiling only bites if something downstream can consume more than it delivers. A server with 2 × 25GbE has 6.2 GB/s of network — it cannot fill even an HBA 9500. If those twenty-four drives are a capacity tier serving files over that link, the adapter is nowhere near the bottleneck and the per-drive arithmetic above is irrelevant.\nThe moment it starts to matter is when the consumer gets faster than the card: 100GbE (12.5 GB/s) puts you level with a 9500\u0026rsquo;s entire sequential ceiling on its own, and local workloads — databases, compilation, analytics, virtualisation hosts with busy guests — have no network in the path at all.\nSo the question to ask about a populated shelf is not \u0026ldquo;is the adapter slow\u0026rdquo; but \u0026ldquo;what is going to consume this, and can it consume more than the card can deliver?\u0026rdquo; If the answer is a 25GbE link, stop worrying. If the answer is 100GbE, NVMe-oF, or a local database, the card is your bottleneck and the drive count is making it worse.\nWhat Else Goes Missing Beyond throughput, presenting an NVMe drive as a SCSI disk means the NVMe-specific parts of your toolkit stop working:\nDirect-attached Behind a tri-mode adapter Device node /dev/nvme0n1 /dev/sdb Driver nvme mpt3sas / mpi3mr nvme-cli Works Nothing to talk to Health data NVMe SMART log pages Translated SCSI log pages Namespace management Yes No Firmware updates nvme fw-download Vendor tool via the controller Format / sanitize NVMe Format NVM SCSI equivalents, if implemented Hardware queues One pair per core (16 on the machine above) The card\u0026rsquo;s reply queues, shared by every drive Queue depth 1023 per queue 128 per device One consequence catches people out often enough to call out separately: you cannot pass an individual drive through to a virtual machine. PCIe passthrough needs the drive to be a PCIe endpoint with its own IOMMU group, and behind a tri-mode adapter it is not one — the only PCIe device present is the controller. You can pass the entire adapter through, with every drive attached to it, or nothing. If your plan involved handing specific NVMe drives to specific guests, the backplane decision has already made that call for you.\nSo Who Actually Wants This in 2026? Here is where the pitch at the top of this post has to face a harder question, because the world it was designed for has mostly gone.\nTri-mode was conceived when NVMe was the expensive tier you added to a SAS estate. In 2026 that is backwards: NVMe is the default, U.2 enterprise drives are abundant and cheap on the used market, and \u0026ldquo;mixed SAS, SATA and NVMe in one chassis\u0026rdquo; describes fewer and fewer real deployments. So who is actually buying it?\nMostly nobody — deliberately. The honest answer is that most tri-mode controllers were not chosen. They arrived, because the server vendor ships one, and the vendor ships one because a single U.3 backplane SKU lets them sell SAS, SATA and NVMe configurations out of the same chassis. That is a supply-chain win for the OEM. It does nowt for your performance, and as such it was never sold as doing so.\nThree of the classic justifications no longer hold up well:\n\u0026ldquo;I need mixed drive types.\u0026rdquo; Rarely in the same chassis, and even when you do, tri-mode is not the only way. A plain SAS HBA for the spinning disks plus NVMe wired to the root complex gets you both, with neither one penalised. Mixed estate does not imply mixed controller.\n\u0026ldquo;I do not have enough PCIe lanes.\u0026rdquo; This was the real argument in 2019, on 40-lane platforms with 24 bays. A single-socket Genoa or Turin Epyc has 128 lanes. Twenty-four drives at x4 is 96. The scarcity that justified aggregating drives behind one controller has all but gone, and where it has not, a PCIe switch does the job without terminating the protocol.\n\u0026ldquo;The bays need to be universal.\u0026rdquo; This one is worth separating carefully, because it is the argument most often used to justify the wrong component. U.3 is a backplane standard, not a controller requirement. A U.3 backplane can be cabled straight to the CPU\u0026rsquo;s PCIe lanes instead of through a tri-mode controller, and vendors document both topologies. You can keep the universal bays and delete the tax. If you have inherited a tri-mode server, the single most valuable thing you can check is whether the backplane can be re-cabled direct.\nWhat is actually left:\nHardware RAID at density, where policy or platform requires it — an audit requirement, a support matrix, a Windows or ESXi deployment with no software layer to do the job. This is the real remaining market, it is the one case where you are buying the RAID engine rather than the connectivity, and on current silicon it is a capable product: sixteen NVMe drives in hardware RAID 5 with a supercap-protected cache, out of sixteen lanes, is something direct attachment cannot offer at all. Bulk SAS HDD capacity, where £/TB still belongs decisively to spinning disks. But that is a plain SAS HBA\u0026rsquo;s job, and a cheaper one. Very large bay counts and external enclosures, where SAS expanders reach further and wider than PCIe will. Workloads that never go deep. If your queue depths sit in single digits, none of this registers. Plenty of real systems live here quite happily. There is also a plain operational argument — one enclosure type, one driver, one spare on the shelf — and for a general-purpose virtualisation host that is worth something real. Just price it honestly against the fact that, per the table above, one Gen5 drive can meet the whole card\u0026rsquo;s sequential ceiling.\nWhen It Is the Wrong Tool The trade turns bad in proportion to how much concurrency your workload has.\nCeph is the clearest case. A storage node runs one OSD per drive, each with its own thread pools, all issuing I/O at once — and then a whole cluster\u0026rsquo;s worth of clients drives them concurrently. That is the shared-tag-pool worst case: twenty-four daemons contending for one controller\u0026rsquo;s command slots, each drive capped at 128 outstanding, everything funnelled through one x8 uplink. Put those drives straight on the root complex and each OSD gets its own queues, its own tags and its own lanes. An all-NVMe Ceph node should not have a tri-mode adapter in the data path.\nThe same logic applies to NVMe-oF targets, where you are re-exporting drives and every layer of serialisation compounds; to databases with deep asynchronous I/O; and to anything built on io_uring or SPDK, which exist specifically to exploit per-core queues that the adapter has just taken away.\nThe general rule: the more parallelism your software was written to exploit, the more a tri-mode adapter charges you for it.\nHow to Tell What You Have If you have inherited a server and want to know which side of this you are on:\n# What is the transport? \u0026#34;nvme\u0026#34; is direct, \u0026#34;sas\u0026#34; means it went through a controller lsblk -o NAME,TRAN,MODEL,SIZE # Is there a tri-mode controller in the machine at all? lspci -nn | grep -Ei \u0026#39;sas|megaraid|serial attached\u0026#39; # Which driver claimed the disk? ls -l /sys/block/sdb/device/driver # Per-device queue depth — 128 is the mpt3sas NVMe default cat /sys/block/sdb/device/queue_depth # How many hardware queues does this device actually get? ls /sys/block/sdb/mq/ | wc -l ls /sys/block/nvme0n1/mq/ | wc -l # compare against a direct-attached drive # On a direct-attached drive, what did the controller actually grant? # One admin queue plus one I/O queue per core, so expect nproc + 1 cat /sys/class/nvme/nvme0/queue_count nproc # The driver says it out loud at load time dmesg | grep -i \u0026#39;nr_hw_queues\u0026#39; That last one prints the Max SCSIIO MPT commands: N shared with nr_hw_queues = M line quoted earlier. If nvme list is empty on a machine you were told is all-flash NVMe, the adapter is why.\nThe Short Version A tri-mode adapter converts your NVMe drives into SCSI disks. That conversion is not a side effect. It is how one card serves three protocols, and it is what you are buying.\nWhat you give up is specific and measurable: per-core queue pairs replaced by a controller\u0026rsquo;s shared reply queues, a per-device depth of 128, a tag pool divided among every drive on the card, an x2 link where the drive wanted x4, and one shared uplink where each drive previously had its own path to the root complex. The adapter stops being a connection and becomes the bottleneck, and on current hardware it becomes one quickly: two Gen4 drives reach an HBA 9500, four Gen5 drives reach even a PERC13.\nA wider host link does help — an x16 Gen5 card has roughly four times the bandwidth and IOPS of an x8 Gen4 one — but it does not change the shape. It buys bays, not parallelism: the same sixteen lanes wired straight to four drives deliver the same bandwidth with none of the queue tax. And it does not rescue a full shelf. Twenty-four drives behind a 9500 get about 570 MB/s each, which is what a SATA SSD does.\nIn 2019, when NVMe was the tier you added to a SAS estate and platforms were short of lanes, that was a reasonable trade. In 2026 it usually is not. NVMe is the default, used U.2 drives are cheap, a single-socket Epyc has lanes to spare, and the one benefit that still stands — universal drive bays — belongs to the U.3 backplane, not to the controller. You can very often keep the bays and delete the tax by cabling the backplane straight to the CPU.\nSo buy a tri-mode adapter if you are buying its RAID engine and you need one. Do not buy it for the flexibility, and if you have inherited one in a chassis full of NVMe, go and find out how that backplane is cabled.\nThere is no sense paying twice for drives you then cannot use properly.\n","permalink":"https://blogs.damiendye.uk/en/hardware/tri-mode-adapters-nvme-as-sas/","summary":"A tri-mode adapter lets any bay take SAS, SATA or NVMe — real flexibility, and the reason U.3 backplanes exist. What it does not advertise is that your NVMe drives stop being NVMe drives: they arrive in Linux as SCSI disks on mpt3sas, queue depth 128, sharing one tag pool and one x8 uplink. Two Gen4 drives saturate the card. A single Gen5 drive is already past it. Whether that trade still makes sense in 2026, and why the universal bays belong to the backplane rather than the controller.","title":"Tri-Mode Adapters Buy Flexibility With Your NVMe Queues"},{"content":"The Assumption Ansible Usually Gets to Make Almost every Ansible module you have used works like this: Ansible connects to the host named in the inventory, copies a small Python program to it, runs it, and reads back the result. The host is the thing being changed and the thing doing the work.\nCreating a virtual machine breaks that in the most basic way possible. The host you are building does not exist. It has no IP, no SSH daemon, no Python, and no operating system. There is nowt to connect to.\nSo community.proxmox is not a configuration agent. It is an API client that happens to be shipped as an Ansible collection. As such, everything in this post follows from that one fact — where the tasks run, how you get credentials in, why re-running does not do what you expect, and why --check is not telling you the truth.\nThe examples here are cut down from a playbook that builds Windows and Linux VMs from NetBox records: proxmox-create-vms.yml. I have genericised the node, storage and bridge names for readability — the real thing is in that repository.\nEverything I claim about module behaviour below was checked against community.proxmox 1.6.0, which is the version I have installed:\n$ ansible-galaxy collection list community.proxmox # /home/damien/.ansible/collections/ansible_collections Collection Version ----------------- ------- community.proxmox 1.6.0 First: The Collection Moved If you are reading an older playbook or an older answer, the modules were called community.general.proxmox_kvm. They now live in a dedicated collection, and that is where the development is happening. Version 1.6.0 ships 47 modules, covering Ceph, SDN, firewall, HA rules and cluster join, none of which existed in the community.general era.\nansible-galaxy collection install community.proxmox pip install \u0026#39;proxmoxer\u0026gt;=2.0\u0026#39; requests The Python dependency is not optional and not bundled: the collection declares requirements: [\u0026quot;proxmoxer \u0026gt;= 2.0\u0026quot;, \u0026quot;requests\u0026quot;], and those need to be installed wherever the module actually executes — which, as the next section explains, is not the Proxmox node.\nrequirements.yml, if you would rather pin it:\n--- collections: - name: community.proxmox version: \u0026#34;\u0026gt;=1.6.0\u0026#34; The collection is tested against ansible-core 2.17 through 2.20. Renaming your community.general.proxmox_* tasks to community.proxmox.proxmox_* is most of the migration.\nEvery Proxmox Task Runs on localhost Two lines at the top of the play do the heavy lifting, and both look like they are disabling something useful:\n- name: Create Proxmox virtual machines hosts: \u0026#34;{{ target_hosts | default(\u0026#39;cluster_pve:\u0026amp;status_planned\u0026#39;) }}\u0026#34; gather_facts: false serial: 1 gather_facts: false is not an optimisation. Fact gathering connects to the inventory host, and the inventory host is a VM that has not been built yet. Leave it on and the play fails before the first task.\nThen every Proxmox task carries delegate_to: localhost:\n- name: Create Proxmox VM delegate_to: localhost register: created_vm community.proxmox.proxmox_kvm: api_user: \u0026#34;{{ proxmox_user }}\u0026#34; api_password: \u0026#34;{{ proxmox_password }}\u0026#34; api_host: \u0026#34;{{ proxmox_api_ip }}\u0026#34; name: \u0026#34;{{ inventory_hostname }}\u0026#34; node: \u0026#34;{{ proxmox_api_host }}\u0026#34; ... The inventory host is now just a name and a bag of variables. inventory_hostname becomes the VM\u0026rsquo;s name; its variables describe the machine you want. Nothing connects to it. The task runs on the control node, which opens an HTTPS session to api_host and posts a VM definition.\nYou will also see this written as local_action:, which is the older syntax for the same thing. The teardown playbook in that repository uses it throughout. They are equivalent. delegate_to is the current spelling.\nThe Escape Hatch Points at the Node Some things genuinely have to happen on a Proxmox host, and those tasks delegate somewhere else entirely:\n- name: Fail if no ISO file exists for the OS delegate_to: \u0026#34;{{ proxmox_api_host }}\u0026#34; ansible.builtin.stat: path: \u0026#34;{{ iso | replace(\u0026#39;isos:\u0026#39;, \u0026#39;/mnt/pve/isos/template/\u0026#39;) }}\u0026#34; register: iso_file failed_when: not iso_file.stat.exists That is a real SSH connection to a real node, checking a real path on shared storage, because the API will happily accept an ISO reference that does not resolve to a file and you would rather find out now than at boot. Note the string surgery translating a PVE storage reference (isos:iso/debian.iso) into a filesystem path. The storage abstraction is not available to stat.\nSo a single play has tasks executing in three different places, and getting them mixed up is the most common way these playbooks fail:\nWhere each task in a Proxmox build playbook actually executes Ansible control node delegate_to: localhost proxmox_kvm proxmox_vm_info proxmox_disk proxmox_access_acl every one of them is an HTTPS client, not an agent proxmoxer \u0026#8805; 2.0 + requests installed here, not on the node gather_facts: false Proxmox node \u0026#8212; pve1 pvedaemon, REST API on port 8006 qm, /etc/pve, storage the VM definition lands here the VM you are creating no IP \u0026#183; no SSH \u0026#183; no Python no operating system in the inventory it is only a name and a bag of vars API SSH qm set stat nothing to connect to not until a later play Three execution contexts in one play. The Proxmox modules never touch the node or the guest — they are HTTPS clients running beside the playbook. The qm escape hatch is the only part that needs SSH to a hypervisor. Credentials, and a Default That Is About to Change The auth options are shared by every module in the collection through a documentation fragment, so they are the same everywhere: api_host, api_user, and then either api_password or the pair api_token_id / api_token_secret. All of them fall back to environment variables — PROXMOX_HOST, PROXMOX_USER, PROXMOX_PASSWORD, PROXMOX_TOKEN_ID, PROXMOX_TOKEN_SECRET, PROXMOX_VALIDATE_CERTS — which is the cleanest way to keep secrets out of the play entirely.\nAn API token is the better default. It is scoped, it is revocable without changing a human\u0026rsquo;s password, and it can be given exactly the privileges the playbook needs rather than the ones a person happens to have:\n- name: Create Proxmox VM delegate_to: localhost community.proxmox.proxmox_kvm: api_host: \u0026#34;{{ proxmox_api_ip }}\u0026#34; api_user: ansible@pve api_token_id: automation api_token_secret: \u0026#34;{{ proxmox_token_secret }}\u0026#34; validate_certs: true ca_path: /etc/ssl/certs/pve-cluster-ca.pem Set validate_certs explicitly, today. The collection\u0026rsquo;s own documentation says it plainly:\nCurrently defaults to false and changes default to true with community.proxmox 2.0.0.\nWhich means a playbook that never mentions it is not validating TLS right now, and will start validating — and therefore start failing against the self-signed certificate every fresh Proxmox install ships with — the moment someone runs --upgrade. Better to make that decision on purpose than to have it land in the middle of a build. If you are keeping the self-signed certificate, say validate_certs: false and take the finding; if you have a proper chain, point ca_path at it. Either way it is written down.\nThe Minimum Viable Create Strip the production task down to what actually defines a machine and it is readable:\n- name: Create Proxmox VM delegate_to: localhost register: created_vm community.proxmox.proxmox_kvm: api_host: \u0026#34;{{ proxmox_api_ip }}\u0026#34; api_user: ansible@pve api_token_id: automation api_token_secret: \u0026#34;{{ proxmox_token_secret }}\u0026#34; validate_certs: true node: pve1 name: \u0026#34;{{ inventory_hostname }}\u0026#34; cores: \u0026#34;{{ vcpus | int }}\u0026#34; memory: \u0026#34;{{ memory }}\u0026#34; machine: q35 bios: ovmf ostype: l26 scsihw: virtio-scsi-single scsi: scsi0: \u0026#34;vmdata:32,format=qcow2,discard=on,ssd=1\u0026#34; sata: sata0: \u0026#34;isos:iso/debian-13-netinst.iso,media=cdrom\u0026#34; net: net0: \u0026#34;virtio,bridge=vmbr0\u0026#34; efidisk0: storage: vmdata format: raw efitype: 4m pre_enrolled_keys: true boot: \u0026#34;order=scsi0;sata0\u0026#34; agent: \u0026#34;enabled=1,fstrim_cloned_disks=1\u0026#34; onboot: true tags: - production A few things about that shape are worth knowing before you write your own.\nThe device options are PVE syntax inside YAML. scsi, sata, net, virtio, ide are all typed dict, keyed scsi0, net0 and so on, and the values are the comma-separated option strings straight out of man qm — \u0026lt;storage\u0026gt;:\u0026lt;size\u0026gt;,option=value for a disk, [model=]\u0026lt;enum\u0026gt;,option=value for a NIC. The module does not model them; it forwards them. When something is rejected, the answer is in the PVE options reference, not in the Ansible docs.\nYou will also see these written as a JSON string rather than a YAML mapping:\nnet: \u0026#39;{\u0026#34;net0\u0026#34;:\u0026#34;virtio,bridge={{ vlan_bridge }}\u0026#34;}\u0026#39; Both work. Ansible coerces the string for a dict-typed parameter. The JSON form exists because it is easier to template a whole structure in one Jinja expression. The mapping form is easier to read six months later.\nboot has two generations of syntax. The module accepts the legacy letters, where boot: \u0026quot;cdn\u0026quot; means \u0026ldquo;try disk, then CD-ROM, then network\u0026rdquo;. Current PVE wants an explicit ordered list — boot: \u0026quot;order=scsi0;sata0;net0\u0026quot; — which is unambiguous about which disk. The legacy form still works. The explicit form is what you want in new work. One genuine gotcha buried in the module docs: network boot requires setting rng0 since PVE 8.3.5.\nnuma and numa_enabled are different parameters. numa_enabled is the boolean that turns NUMA on. numa is a dict describing a topology (cpus, hostnodes, memory, policy). Setting numa: true is a type error, and it is an easy hour to lose. If you care why any of this matters, NUMA alignment on Proxmox covers the underlying problem.\nmachine: q35 and bios: ovmf are the right defaults, not decoration — that argument in full.\nOmitting vmid means the module asks the API for the next free ID. Convenient, and the direct cause of the next thing.\nWhy serial: 1 \u0026ldquo;Fetch the next available ID, then create a VM with it\u0026rdquo; is two API calls with a gap in the middle. Two workers doing that concurrently can read the same free ID, and the loser gets an error or, worse, a surprise.\nserial: 1 makes the create phase one host at a time. It is not fast and it does not need to be. The expensive part of building a VM happens after this playbook hands off. The counterpart play that starts the finished VMs uses serial: 5, because starting has no shared counter to race on.\nIf you would rather have the parallelism, allocate the VMID yourself from your source of truth and pass it explicitly. Then there is no read-modify-write and no race.\nIdempotency Is Not What You Expect This is the section to read twice, because proxmox_kvm does not behave like ansible.builtin.package.\nname is not an identity. VM names are not unique across a Proxmox cluster, and the module says so. With state: present and no vmid, if a VM with that name already exists, the module exits changed=false with msg: \u0026quot;VM with name \u0026lt;x\u0026gt; already exists\u0026quot; and does nowt. It does not compare your parameters to reality. It does not converge. It declines.\nupdate defaults to false. So editing memory: in your playbook and re-running is a no-op. The VM keeps the memory it was built with, the task reports success, and nothing anywhere tells you the two have diverged.\nupdate: true still refuses the interesting parameters. From the module documentation:\nBecause of the operations of the API and security reasons, I have disabled the update of the following parameters net, virtio, ide, sata, scsi. Per example updating net update the MAC address and virtio create always new disk…\nupdate_unsafe: true lifts that restriction, and the warning is not for show:\nUse this option with caution because an improper configuration might result in a permanent loss of data (for example disk recreated).\nSo the disk you thought you were resizing can be replaced with a new empty one. Do not reach for this — the refused parameters have their own modules, and that is the next section.\nAnd --check does not cover any of this. The collection declares check-mode support per module, and it is inconsistent in exactly the wrong direction:\nModule check_mode diff_mode proxmox_kvm none none proxmox_disk none none proxmox_template none none proxmox_snap full none proxmox_nic full none proxmox_pool full none The split is not \u0026ldquo;read-only modules can, write modules cannot\u0026rdquo; — proxmox_nic creates and deletes interfaces and honours check mode perfectly well. It is that the three modules dealing in storage and VM lifecycle do not. A --check run of a build playbook skips the VM creation silently and then reports on a world where the VM was never made, so every task after it is reasoning about the wrong state. On a build playbook, --check is not a safety net, and treating it as one is worse than not running it.\nWhat a second run of proxmox_kvm actually does second run \u0026#8212; state: present, and a VM of that name exists the module never compares your parameters to the running VM update: false the default changed = false \u0026#8220;VM with name \u0026lt;x\u0026gt; already exists\u0026#8221; edit memory in the play, re-run, and nothing anywhere tells you update: true converges most of it cores, memory, tags, agent, onboot applied net, virtio, ide, sata, scsi, efidisk0, tpmstate0 refused by design use proxmox_disk and proxmox_nic for those update_unsafe: true converges all of it the refused parameters are applied too a disk parameter can recreate the disk permanent data loss is the documented risk --check tells you nothing check_mode: none the task is skipped every later task then reasons about a world where the VM was never created Two of the four converge anything, and the one that covers disks is the one that can destroy them. So gate on existence yourself, and treat creation as a one-time event. What a second run actually does. Two of the four paths converge anything at all, and the only one that covers disks is the one that can destroy them. So Make Existence the Gate Given all that, the workable pattern is to stop asking the module to be idempotent and decide for yourself whether to build. The production playbook does it like this:\n- name: Check if VM is present or manually built delegate_to: localhost community.proxmox.proxmox_vm_info: api_host: \u0026#34;{{ proxmox_api_ip }}\u0026#34; api_user: \u0026#34;{{ proxmox_user }}\u0026#34; api_password: \u0026#34;{{ proxmox_password }}\u0026#34; name: \u0026#34;{{ inventory_hostname }}\u0026#34; config: current register: existing_vm ignore_errors: true failed_when: (existing_vm.proxmox_vms | length) == 0 - name: Configure Proxmox VM when: existing_vm is failed block: - name: Create Proxmox VM ... proxmox_vm_info with config: current returns the VM and its live configuration, or an empty list. failed_when turns \u0026ldquo;empty list\u0026rdquo; into a failure, ignore_errors: true stops that failure ending the play, and when: existing_vm is failed becomes \u0026ldquo;the VM is not there, build it\u0026rdquo;.\nUsing a deliberately failed task as a boolean reads badly, and I am not going to pretend otherwise. The alternative is when: (existing_vm.proxmox_vms | default([]) | length) == 0, which is honest about being a length check and does not need ignore_errors. Both work. The version above is what is in production, and its one real advantage is that the registered result carries the existing configuration for later tasks to read.\nThe important part is the shape, not the spelling: check, then branch, and treat creation as a one-time event. A VM\u0026rsquo;s ongoing configuration is a different problem from a VM\u0026rsquo;s existence, and this module is only good at the second one.\nDisks and NICs Have Their Own Modules Here is the thing I glossed over above, and it changes the whole picture: the parameters proxmox_kvm refuses to update are not a gap in the collection. They are delegated. community.proxmox.proxmox_disk and community.proxmox.proxmox_nic add, change and remove exactly the things the create module will not touch — keyed on the same scsi0 and net0 names you used when you built the VM.\nBoth are better-behaved than proxmox_kvm, and one of them is the only module in this workflow that can be dry-run.\nproxmox_nic — Add, Retag or Remove an Interface - name: Move the VM\u0026#39;s primary NIC to a new bridge and VLAN delegate_to: localhost community.proxmox.proxmox_nic: api_host: \u0026#34;{{ proxmox_api_ip }}\u0026#34; api_user: ansible@pve api_token_id: automation api_token_secret: \u0026#34;{{ proxmox_token_secret }}\u0026#34; vmid: \u0026#34;{{ created_vm.vmid }}\u0026#34; interface: net0 bridge: vmbr1 tag: 120 model: virtio mtu: 1 queues: 4 firewall: true state: present interface is the only required option beyond auth — net[n] where n is 0 to 31 — and state: present or absent gives you add and remove. model defaults to virtio, which is the right answer unless a guest cannot cope.\nThe reason this module exists is the MAC address. Recall why proxmox_kvm refuses to update net: \u0026ldquo;updating net update the MAC address\u0026rdquo;. proxmox_nic fixes that explicitly:\nWhen not specified this module will keep the MAC address the same when changing an existing interface.\nSo you can retag a VLAN, move a bridge, change the MTU or turn the firewall on without the guest\u0026rsquo;s NIC identity changing underneath it. That matters more than it sounds: a new MAC invalidates DHCP reservations, breaks anything licensed to a NIC, and desynchronises the NetBox interface record the create playbook so carefully wrote. This is the module that lets a day-2 network change be boring.\nA few options worth knowing before you need them:\nrate is in MBps — MegaBytes per second, not bits. The documentation is explicit and the factor-of-eight error is very easy to make. link_down: true disconnects the interface, described in the docs as \u0026ldquo;like pulling the plug\u0026rdquo;. A clean way to isolate a suspect VM without stopping it or touching the guest. trunks takes a list of VLAN IDs to pass through, for a guest that does its own tagging. mtu: 1 is not a typo and not a 1-byte MTU — it means \u0026ldquo;inherit the bridge MTU\u0026rdquo;, and it only applies to virtio. queues sets multiqueue, 0 to 16. Worth matching to vCPU count on anything pushing real traffic. And it supports check mode fully. --check on a proxmox_nic task tells you the truth, which makes it the one part of this workflow you can safely rehearse. Its messages are properly idempotent too. An unchanged interface reports Nic net0 unchanged on VM with vmid 103 rather than claiming a change.\nproxmox_disk — The Whole Disk Lifecycle proxmox_disk is the largest module of the three, and its state is doing five different jobs:\nstate What happens Reversible? present create the disk, or update options on an existing one n/a resized grow it — PVE cannot shrink, and the docs say do that manually no detached becomes unused[n]; the volume and its data stay yes moved change backing storage, or hand the disk to another VM original kept unless delete_moved absent removed from backing storage no The gap between detached and absent is the safety net proxmox_kvm never gives you. Detaching is a config change; deleting destroys data. Two different words, two different consequences.\nAdding a second disk to a VM that already exists:\n- name: Add a data disk delegate_to: localhost community.proxmox.proxmox_disk: api_host: \u0026#34;{{ proxmox_api_ip }}\u0026#34; api_user: ansible@pve api_token_id: automation api_token_secret: \u0026#34;{{ proxmox_token_secret }}\u0026#34; vmid: \u0026#34;{{ created_vm.vmid }}\u0026#34; disk: scsi1 storage: vmdata size: 200 format: qcow2 iothread: true aio: io_uring discard: \u0026#34;on\u0026#34; ssd: true backup: true state: present create is the knob proxmox_kvm should have had. It controls what state: present is allowed to do:\nregular (the default) — create the disk if missing, otherwise update its options. disabled — update options only, and never create. This is the one to reach for when you are changing cache or iothread on a disk that must already exist. It cannot surprise you by conjuring a new volume because a key was misspelled. forced — always create. An existing disk is detached and left unused, not deleted. That last behaviour is the important detail. create: forced is the destructive-looking option, and it still does not destroy anything: the old volume survives as unusedN and you can re-attach it. Compare that with proxmox_kvm plus update_unsafe, whose documented failure mode is a disk recreated. Same rough operation, much better blast radius.\nGrowing a disk:\n- name: Grow the data disk by 100 GiB delegate_to: localhost community.proxmox.proxmox_disk: api_host: \u0026#34;{{ proxmox_api_ip }}\u0026#34; api_user: ansible@pve api_token_id: automation api_token_secret: \u0026#34;{{ proxmox_token_secret }}\u0026#34; vmid: \u0026#34;{{ created_vm.vmid }}\u0026#34; disk: scsi1 size: \u0026#34;+100G\u0026#34; state: resized Watch the units, because size changes meaning with state. With state: present it is GiB as a bare number (size: 200). With state: resized it takes a suffix — +100G to add to the current size, or 500G as an absolute target. One parameter, two conventions, and the failure is silent if you guess wrong.\nMoving a disk to different storage, which is the live-migration-of-one-volume case:\n- name: Move the disk to NVMe storage delegate_to: localhost community.proxmox.proxmox_disk: api_host: \u0026#34;{{ proxmox_api_ip }}\u0026#34; api_user: ansible@pve api_token_id: automation api_token_secret: \u0026#34;{{ proxmox_token_secret }}\u0026#34; vmid: \u0026#34;{{ created_vm.vmid }}\u0026#34; disk: scsi1 target_storage: nvme-pool bwlimit: 200000 delete_moved: true timeout: 3600 state: moved target_storage moves within one VM; target_vmid hands the disk to a different VM and requires the same storage on both. They are mutually exclusive. delete_moved defaults to false, so by default you finish with two copies and the original sitting there as unused — safe, and a good way to fill a storage pool if you never come back to it.\ntimeout defaults to 600 here, against 30 in proxmox_kvm. Same-looking parameter, twentyfold difference, because these operations copy data. Raise it for large images or slow storage — the docs say so for both moved and import_from.\nWhich brings up the option that makes this module the V2V and cloud-image path:\nimport_from: \u0026#34;vmdata:9000/base-debian13.qcow2\u0026#34; import_from builds the disk from an existing volume rather than allocating an empty one — \u0026lt;STORAGE\u0026gt;:\u0026lt;VMID\u0026gt;/\u0026lt;NAME\u0026gt;, or \u0026lt;STORAGE\u0026gt;:import/\u0026lt;NAME\u0026gt; using the storage import directory on PVE 9.x and later. It is mutually exclusive with size, and only root can use absolute filesystem paths.\nThe rest of the parameter list is the reason to attach disks with this module rather than inline in the create call: cache, aio, iothread, discard, ssd, backup, detect_zeroes, and the full throttling family — iops, iops_rd, iops_wr, their _max and _max_length variants, and the bps_*_max_length burst controls. None of that is reachable through proxmox_kvm after creation.\nTwo caveats, both from the module\u0026rsquo;s own documentation:\nSome option changes need a reboot. \u0026ldquo;Some updates on options (like cache) are not being applied instantly and require VM restart.\u0026rdquo; A green task means the config was written, not that the running VM is behaving differently. It does not support check mode. check_mode: none, same as proxmox_kvm. So the collection splits down the middle: NIC changes can be rehearsed with --check, disk changes cannot. The Division of Labour To do this Use Create the VM proxmox_kvm, once, gated on existence Change cores, memory, tags, agent, onboot proxmox_kvm with update: true Add, retag, disconnect or remove a NIC proxmox_nic Add, grow, move, detach or remove a disk proxmox_disk Snapshot proxmox_snap (also full check mode) Anything none of them expose qm set over SSH Change a disk or NIC through proxmox_kvm nothing — this is what update_unsafe is for, and it is why you should not use it Build the VM with a minimal proxmox_kvm call, then attach the disks and interfaces with their own modules. It is more tasks, and it is the version where day-2 changes have a route that does not involve an option whose documented risk is losing a disk.\nWhere the Module Stops proxmox_kvm has a huge parameter list and still does not cover everything qm can do. Rather than wait, the production playbook drops to the CLI on the node:\n- name: Set RNG source and better SPICE quality delegate_to: \u0026#34;{{ proxmox_api_host }}\u0026#34; become: true ansible.builtin.command: cmd: \u0026gt;- /usr/sbin/qm set {{ created_vm.vmid }} --rng0 source=/dev/urandom --spice_enhancements videostreaming=all There is nothing wrong with this. It is not idempotent in any meaningful sense — qm set is a write, and it will report changed every run — but it is explicit, it is readable, and it does not pretend. If a module gains the parameter later, you delete the task.\nOther things need a second pass through the module with update: true, because they cannot be set in the same call that creates the VM:\n- name: Add SPICE-compatible USB device delegate_to: localhost community.proxmox.proxmox_kvm: api_host: \u0026#34;{{ proxmox_api_ip }}\u0026#34; api_user: \u0026#34;{{ proxmox_user }}\u0026#34; api_password: \u0026#34;{{ proxmox_password }}\u0026#34; node: \u0026#34;{{ proxmox_api_host }}\u0026#34; vmid: \u0026#34;{{ created_vm.vmid }}\u0026#34; usb: usb0: \u0026#34;spice,usb3=1\u0026#34; update: true when: spice_usb | default(false) Note it passes vmid, not name. Once you have the ID, use it. It is the only identifier the API treats as unique.\nAnd for the things QEMU can do that PVE has no option for, there is args, which is passed to the QEMU command line verbatim:\nargs: \u0026gt;- -global scsi-hd.physical_block_size=4k -global scsi-hd.logical_block_size=4096 That one presents the virtual disk as 4Kn rather than 512e, which matters more than it sounds like it does — block sizes, 4Kn and 512e. The module labels args \u0026ldquo;for experts only\u0026rdquo;, and the reason is that PVE does not validate it and a bad flag stops the VM booting with an error that comes from QEMU rather than from Proxmox.\nThe Cluster Is Not Instantly Consistent - name: Let registration complete on cluster ansible.builtin.pause: seconds: 5 when: created_vm.changed A pause in a playbook is usually a smell, and this one is load-bearing. The create call returns when the API has accepted the definition, which is not the same as every node agreeing the VM exists — and the very next task wants to set an ACL on /vms/\u0026lt;vmid\u0026gt;. Five seconds of patience is cheaper than a retry loop around an error that only shows up under load.\nThe teardown playbook has the same shape for the same reason: stop, wait, then delete.\nReading Back What You Built proxmox_kvm documents three return values: vmid, status and msg. In practice you will want a fourth, and it is not in the documentation.\n- name: Get MAC address of VM ansible.builtin.set_fact: primary_mac_addr: \u0026#34;{{ created_vm.mac.net0 }}\u0026#34; created_vm.mac is real — the module builds it in get_vminfo() and splats it into the result — but it is absent from the documented RETURN block, which means nothing promises it will keep working. Worth knowing exactly how it behaves, because there are two traps in it:\nIt only appears when the module actually created the VM. mac is only assembled on the create-and-deploy path. Take the \u0026ldquo;already exists\u0026rdquo; branch and the result has vmid and msg and nothing else. It only contains the interfaces you passed in. The code walks the parameters you supplied and picks out the ones matching net[0-9], then reads each one\u0026rsquo;s stored config back from the API. No net parameter, no mac key. Which is why the production playbook needs both halves, and the second one is ugly:\n- name: Get MAC address of VM ansible.builtin.set_fact: primary_mac_addr: \u0026gt;- {{ created_vm.mac.net0 if created_vm is defined and created_vm.changed else ((existing_vm.proxmox_vms[0].config.net0 | split(\u0026#39;,\u0026#39;))[0] | split(\u0026#39;=\u0026#39;))[1] }} When the VM already existed, there is no mac, so the MAC has to be dug out of the raw config string. net0 comes back from the API as virtio=AE:AE:5C:A8:89:85,bridge=vmbr0, so: split on commas, take the first field, split on =, take the second half. It is string surgery on an API response, and it is the honest cost of a module whose return shape depends on which branch it took.\nIf you need the MAC reliably in both cases, get it from proxmox_vm_info unconditionally and parse one shape rather than two.\nDriving It From a Source of Truth Look again at the line that opens the play:\nhosts: \u0026#34;{{ target_hosts | default(\u0026#39;cluster_pve:\u0026amp;status_planned\u0026#39;) }}\u0026#34; That is the actual architecture, and it is worth stating plainly: the VMs to build are not a list in a vars file. They are the hosts in your inventory whose recorded status says they should exist and do not yet.\nThe inventory here is NetBox. A VM is requested by creating a NetBox record with status planned, carrying its CPU, memory, disk, VLAN, owner and platform. The playbook selects planned machines, builds them, allocates an IP, writes DNS, and then sets the record to staged — at which point that host no longer matches the play\u0026rsquo;s host pattern, and a handler refreshes the inventory so the next play sees the new state:\nhandlers: - name: Refresh inventory ansible.builtin.meta: refresh_inventory The status field is a state machine, the playbook is one transition in it, and the whole thing is re-runnable because a host that has already moved on is no longer selected. That is a much better property than any amount of module-level idempotency, and it is the reason the create task can get away with being a one-shot.\ncommunity.proxmox also ships its own inventory plugin, which builds an inventory from the cluster — the right choice when Proxmox is the source of truth. Here it is the other way round: NetBox is authoritative and Proxmox is where its intent gets realised. That is a whole post of its own and I will write it separately.\nTaking It Away Again Creation without teardown is half a lifecycle, and the removal path has its own trap — you cannot delete a running VM:\n- name: Force stop the VM if it is running delegate_to: localhost community.proxmox.proxmox_kvm: api_host: \u0026#34;{{ proxmox_api_ip }}\u0026#34; api_user: \u0026#34;{{ proxmox_user }}\u0026#34; api_password: \u0026#34;{{ proxmox_password }}\u0026#34; name: \u0026#34;{{ inventory_hostname }}\u0026#34; state: stopped force: true timeout: 10 - name: Allow the cluster to stop the VM before removing it ansible.builtin.pause: seconds: 10 - name: Remove the VM from the cluster delegate_to: localhost community.proxmox.proxmox_kvm: api_host: \u0026#34;{{ proxmox_api_ip }}\u0026#34; api_user: \u0026#34;{{ proxmox_user }}\u0026#34; api_password: \u0026#34;{{ proxmox_password }}\u0026#34; name: \u0026#34;{{ inventory_hostname }}\u0026#34; state: absent force: true timeout: 10 state: stopped is a graceful shutdown, and the interaction with timeout is documented and worth memorising: if the timeout is reached with force: true the VM is powered off hard; with force: false the task fails instead. A ten-second graceful window followed by a pull of the plug is a reasonable policy for a machine that is being torn down, and a terrible one for anything else.\nThe full teardown (proxmox-remove-vms.yml) then unwinds the rest of the record: DNS, the allocated IP, the NetBox interfaces, the NetBox VM, and the stale entries in known_hosts. It wraps the block in ignore_errors: true, which is defensible in a teardown. You are removing things that may already be gone, and a half-deleted machine is worse than a noisy log.\nWhat I Would Change in a Fresh Build Having read the module source rather than just its documentation, four things:\nUse an API token, not api_user plus api_password. Scoped, revocable, and it never belongs to a person. Set validate_certs explicitly, before 2.0.0 changes it under you. Allocate the VMID yourself from the source of truth. It removes the read-modify-write race, lets you drop serial: 1, and gives every later task a stable identifier instead of a name that is not unique. Create the VM bare, then attach its disks and NICs with proxmox_disk and proxmox_nic. More tasks, but every disk and interface then has a module that can change it later — including create: disabled for option-only edits and state: detached instead of deletion — rather than a config that can only be changed through update_unsafe. Get the MAC from proxmox_vm_info in one place, so there is one shape to parse instead of a conditional across a documented and an undocumented return value. And do not reach for --check on a build playbook. The module that matters cannot honour it.\nA dry run that always says yes is worse than no dry run at all, because you will believe it.\nReferences community.proxmox collection docs — the full 47-module surface community.proxmox on GitHub — where the source above lives; plugins/modules/proxmox_kvm.py is the file to read when the docs are ambiguous proxmox_kvm module documentation — the parameter list, and the update / update_unsafe warnings quoted above proxmox_vm_info module documentation — config: current and config: pending PVE qm options reference — the real specification for every scsi[n], net[n] and boot string you pass through proxmoxer — the Python client the collection is built on damo2929/ansible-example — the playbooks these excerpts come from, including the NetBox-driven inventory, DHCP generation and hypervisor build ","permalink":"https://blogs.damiendye.uk/en/ansible/proxmox-create-vms-community-proxmox/","summary":"The community.proxmox collection is an API client, not a configuration agent, and that changes the shape of every playbook that uses it. Where the tasks actually run, why proxmox_kvm declines to converge rather than updating, why proxmox_disk and proxmox_nic are where disk and NIC changes belong, and the undocumented return value you will end up needing.","title":"Creating Proxmox VMs with Ansible — The Host You Are Building Does Not Exist Yet"},{"content":"The Question Every Proxmox Evaluation Starts With \u0026ldquo;Is Proxmox actually enterprise grade?\u0026rdquo;\nIt comes up in nearly every migration conversation, and the worry underneath it is almost never the web interface. Nobody seriously fears that a browser dashboard will corrupt their data. What people are asking is whether the thing standing between a virtual machine and the hardware — the component that has to keep one tenant out of another tenant\u0026rsquo;s memory, forever, without a single mistake — is a serious piece of engineering or a community project that got popular.\nThat is exactly the right thing to be nervous about. It is just aimed at the wrong layer, because Proxmox VE does not contain a hypervisor.\nThe hypervisor is KVM. It is part of Linux, it has been there since 2007, and if your organisation uses EC2, Google Cloud, Oracle Cloud, Alibaba Cloud, DigitalOcean or Nutanix, you are already running it in production today — you have just never had to think about it, because somebody else owned the layers above.\nThis post is about what that shared foundation actually means. Both halves of it: the part of the argument that genuinely holds, and the part that gets overclaimed in vendor slides.\nProxmox VE Is a Management Layer Virtualisation on Linux is four distinct layers, built and maintained by four different sets of people.\n1. Hardware virtualisation extensions. Intel VT-x with EPT, or AMD-V with NPT. Silicon. This is what makes a guest able to run its own kernel at native speed with its own page tables, without anything emulating instructions.\n2. KVM — the hypervisor. Kernel modules: kvm.ko for the architecture-independent core, plus kvm-intel.ko or kvm-amd.ko for the vendor extensions. This is the component that owns the isolation boundary. It sets up the guest\u0026rsquo;s virtual machine control structures, handles VM exits, manages the second-level page tables, and delivers interrupts.\n3. The VMM — the virtual machine monitor, in user space. On Proxmox VE this is QEMU. It builds the virtual motherboard: chipset, PCIe topology, disks, NICs, serial ports, firmware. KVM runs the CPU; QEMU decides what hardware the guest thinks it has.\n4. The management layer. This is Proxmox VE: pve-manager and pveproxy for the API and interface, qemu-server to turn a VM config file into a QEMU command line, pve-container for LXC, pmxcfs on top of Corosync for the replicated cluster configuration, pve-ha-manager for fencing and restart, plus the firewall and SDN stack.\nThose four layers exist on the platform you are migrating from too, doing the same four jobs — vCenter is layer 4, VMkernel is layer 2 — and I will come back to that comparison once the pieces are on the table.\nThe same four layers on Proxmox VE and on vSphere Proxmox VE VMware vSphere 4 \u0026#183; management Proxmox VE pve-manager \u0026#183; pveproxy \u0026#183; qemu-server pmxcfs on Corosync \u0026#183; pve-ha-manager \u0026#183; SDN vCenter Server a separate appliance to size, licence, patch and back up 3 \u0026#183; device model pve-qemu, in user space the virtual motherboard: chipset, PCIe, disks, NICs upstream QEMU + 78 Proxmox patches 2 \u0026#183; hypervisor \u0026#183; the isolation boundary KVM \u0026#8212; kvm.ko + kvm-intel.ko / kvm-amd.ko vCPU entry and exit \u0026#183; EPT/NPT \u0026#183; interrupts upstream Linux, in a Proxmox kernel build the VMX userworld, one per VM device I/O, snapshots, remote console user space \u0026#8212; not inside the kernel VMkernel, plus a VMM per vCPU guest instructions and memory VMware's own words for VMkernel: \u0026#8220;a POSIX-like operating system\u0026#8221; 1 \u0026#183; silicon Intel VT-x + EPT \u0026#183; AMD-V + NPT the same silicon The same four layers on both sides. Proxmox writes layer 4, patches layer 3 heavily, builds layer 2\u0026rsquo;s kernel and takes KVM itself from upstream. VMware writes all four and lets you read none of them. What Proxmox Actually Maintains Proxmox VE owns layer 4 outright. It would be wrong to say it just packages layers 2 and 3, though, and that is the most common misreading of what the company does.\npve-qemu carries 78 patches against upstream QEMU in its series file at the time of writing:\nWhat Proxmox adds to QEMU: 78 patches, to scale extra/ bitmap-mirror/ pve/ 78 patches, to scale 26 6 46 Backported upstream fixes. A striking number are security fixes in the device model: qxl and virtio-gpu stride validation, a guest-triggerable intel_iommu abort, re-entrant DMA in lsi53c895a and virtio-net. Dirty-bitmap sync modes for drive-mirror. Proxmox's own work: savevm-async, the VMA format, the PBS block driver, pbs-restore, alloc-track, backup fleecing. Drawn to scale from debian/patches/series. Most of the queue is Proxmox\u0026rsquo;s own engineering, and the whole of it sits in layer 3 — none of these patches touch kvm.ko. They build their own kernel as well, and maintain packaging or patches for most of the surrounding stack — pve-edk2-firmware for OVMF, plus lxc, zfsonlinux, openvswitch, libiscsi, corosync-pve, lvm and ceph.\nSo the real boundary is: all of layer 4, substantial engineering inside layer 3, and a kernel they compile. What they have not done is write a hypervisor. Layer 2\u0026rsquo;s KVM code is upstream Linux.\nAnd none of this is a criticism of Proxmox. As such, it is the reason a company of Proxmox\u0026rsquo;s size can be trusted with the job at all. A small company in Vienna did not sit down and write a hypervisor from scratch; they built on one that Intel, AMD, Red Hat, Google, Amazon and IBM were already paying engineers to maintain, and spent their own effort on the layer above it and the integration work that layer needs. That is where a team that size can make a difference, and the patch queue shows them making it.\nWhat KVM Actually Is KVM stands for Kernel-based Virtual Machine, and the name is exact: it is a kernel feature, not a program.\nLoad the modules and Linux gains a character device, /dev/kvm, plus a small set of ioctl calls on it — KVM_CREATE_VM, KVM_CREATE_VCPU, KVM_SET_USER_MEMORY_REGION, KVM_RUN. That interface is the entire hypervisor API. Anything that can open a file descriptor and call ioctl can create virtual machines.\nIt was written by Avi Kivity at Qumranet and merged into Linux 2.6.20, released in February 2007 — nineteen years ago. Red Hat acquired Qumranet in 2008, and KVM has shipped in every kernel release since, on the kernel\u0026rsquo;s usual nine-to-ten-week cadence. It has been ported well beyond x86: arm64, POWER, s390 on IBM Z, and RISC-V.\nThe important architectural point is why it is small.\nKVM does not have a scheduler, because Linux has one. A vCPU is an ordinary host thread, and the completely fair scheduler puts it on a core like any other thread. It does not have a memory manager, because Linux has one. Guest RAM is a normal user-space mapping, so it can be paged, backed by hugepages, or placed on a NUMA node using the same machinery as any other process. It does not have a driver stack, a block layer, a network stack, or a filesystem, because Linux already had all of them and they are the ones your hardware vendor is testing against.\nThat is the actual maturity argument, and it is much stronger than a version number. Every NUMA balancing improvement, every io_uring change, every network driver, every new CPU errata workaround that lands in Linux lands underneath your virtual machines, because there is no separate hypervisor kernel for anyone to port it to.\nThe Type 1 Argument Draws the Line in the Wrong Place The objection that follows this, reliably, is that KVM is \u0026ldquo;only a type 2 hypervisor\u0026rdquo;. It runs on a host OS, unlike ESXi, which runs on bare metal.\nThat taxonomy is older than hardware virtualisation by three decades, and the thing it was drawing a line around is no longer where anyone thinks it is.\nWhat Actually Happens When a vCPU Runs QEMU calls ioctl(vcpu_fd, KVM_RUN). Control passes into kvm.ko, which loads the guest CPU state and executes VMLAUNCH. From that instruction until the next VM exit, the guest is executing directly on the physical core, in guest mode, with its own page tables active through EPT, at full hardware speed. There is nowt beneath it interpreting anything. The \u0026ldquo;host OS\u0026rdquo; is not in the path. It is not even running on that core.\nWhen the guest does something that needs handling, the CPU exits to the host — and lands in kvm.ko, in the kernel, at exactly the privilege level ESXi\u0026rsquo;s VMkernel occupies. Most exits are resolved right there and re-entered without user space ever being involved.\nThe path a vCPU takes from QEMU into guest mode, and where its exits stop user space kernel \u0026#8212; ring 0 host mode guest mode \u0026#8212; VMX non-root QEMU \u0026#8212; the vCPU thread one ordinary host thread per vCPU kvm.ko VM entry \u0026#183; exit handling \u0026#183; EPT \u0026#183; interrupts the guest, on the physical core its own kernel, its own page tables through EPT running at full hardware speed ioctl(KVM_RUN) VMLAUNCH VM exit resolved here exit to user space Nothing is interpreting the guest, and the host is not running on that core. Where an exit stops Resolved in the kernel \u0026#183; EPT/NPT violation \u0026#183; local APIC write \u0026#8212; with APICv, often no exit \u0026#183; inter-processor interrupt, posted in hardware \u0026#183; virtio-net doorbell, taken by vhost-net Reaches QEMU \u0026#183; register access on an emulated e1000 or IDE \u0026#183; PCI configuration space write \u0026#183; anything QEMU alone knows how to answer Only the second list pays a user-space round trip, and it is the list you shrink by using virtio devices and a machine type that carries no legacy hardware. The path a vCPU takes. Everything above the dashed line is a host thread; everything below it is the guest on bare silicon. The only exits that reach QEMU are the ones QEMU has to answer. ESXi Has the Same Split Now look at the platform that supposedly proves the distinction.\nVMware\u0026rsquo;s own architecture documentation describes VMkernel as \u0026ldquo;a POSIX-like operating system\u0026rdquo; which provides \u0026ldquo;process creation and control, signals, file system, and process threads\u0026rdquo;. That is an operating system, by its author\u0026rsquo;s description. And a running VM on ESXi is not a thing inside the kernel — it is a group of userworld processes: a VMM per virtual CPU, which virtualises the guest\u0026rsquo;s instructions and manages its memory, and a VMX process per VM, which handles I/O to the devices that are not performance-critical and talks to the snapshot manager and the remote console.\nRead that with QEMU in mind. Per-vCPU execution context in the kernel, per-VM user-space process doing device emulation and management. VMware split it for the same reason everyone else did.\nHyper-V is no different. Microsoft\u0026rsquo;s documentation is explicit that the Virtual Machine Worker Process, vmwp.exe, is \u0026ldquo;a user mode component of the virtualization stack\u0026rdquo;, spawned per VM, and that all emulated devices are implemented in it — running in the root partition, which is Windows. Even Xen, the architecture the taxonomy fits best, needs a general-purpose Linux in dom0 to function, and gets its device model for fully virtualised guests from QEMU.\nSo Where Is the Line? The criterion was never \u0026ldquo;does it use user space\u0026rdquo;. Goldberg\u0026rsquo;s taxonomy, from the early 1970s, asks whether the hypervisor is an application running on a pre-existing operating system that already owns the hardware and does the scheduling. That is a type 2, hosted hypervisor: VMware Workstation, VirtualBox, Parallels Desktop, plain QEMU with no acceleration. You install a general-purpose OS, then you install a program onto it, and that program asks the OS for memory and CPU time like any other program.\nThat is not what KVM is. kvm.ko is not a program on top of Linux. It is part of Linux, executing at the same privilege level as the code that owns the hardware, and when a guest exits it lands there directly. There is no host OS underneath the hypervisor. The kernel is the hypervisor. Proxmox VE ships that kernel as the system, exactly the way ESXi ships VMkernel as the system.\nAnd the \u0026ldquo;it needs user space, so it is type 2\u0026rdquo; version cannot be rescued, because applied consistently it catches everything. No VM runs on ESXi without its VMX process, none on Hyper-V without vmwp.exe, none on Xen as an HVM guest without QEMU. A test that puts every shipping hypervisor in one bucket is not distinguishing anything.\nKernel that owns CPU and memory virtualisation Per-VM user-space device model Needs a pre-existing host OS? VMware ESXi VMkernel VMX per VM No Microsoft Hyper-V hypervisor plus the Windows root partition vmwp.exe No Xen Xen hypervisor plus dom0 Linux QEMU, in dom0 or a stub domain No KVM kvm.ko QEMU No VirtualBox, VMware Workstation the host\u0026rsquo;s kernel, via an installed driver the application itself Yes Four products, one shape — and then a fifth row that is genuinely different. That last row is what \u0026ldquo;type 2\u0026rdquo; was coined to describe, and it is the only one where something else was already in charge of the hardware.\nSo the taxonomy does still draw a line. It just does not draw it anywhere near where the argument assumes: KVM and ESXi are on the same side of it. Calling KVM type 2 borrows a word from the VirtualBox category and applies it to something architecturally in the ESXi category.\nWhich leaves the label doing no useful work in an evaluation, because both of the products you are choosing between sit in the same box. What differs is not the type number. It is that one vendor also wrote the kernel and will not let you read it — a licensing and transparency distinction wearing an architecture diagram\u0026rsquo;s clothes. Nutanix shows the point commercially: it ships the same KVM code Proxmox does and describes AHV as a bare-metal type 1 hypervisor. Same code, opposite label, different marketing department.\nThere is a real concern hiding inside the accusation, and it deserves a better name: a general-purpose kernel is doing a thousand jobs a purpose-built one is not, which is more code and more attack surface beside the isolation boundary. That is legitimate and measurable, and I come back to it near the end.\nWhere the Work Actually Happens The one place the \u0026ldquo;it\u0026rsquo;s on a host OS\u0026rdquo; instinct has a real point is exit handling, so it is worth being clear about which exits go where.\nThe guest does this Handled by Cost Touches a page not yet mapped in EPT/NPT kvm.ko, in kernel One exit, microseconds Writes to its local APIC The CPU itself, via APICv/AVIC Often no exit at all Sends an inter-processor interrupt Posted interrupts in hardware Often no exit Transmits on a virtio-net queue with vhost-net Kernel thread, no user-space hop One doorbell Reads a register on an emulated e1000 or IDE controller All the way out to QEMU Exit plus a user-space round trip Only the last row looks anything like the type 2 caricature — and it is also the row you engineer away, by using virtio devices and not presenting emulated legacy hardware you do not need. That is the same reasoning behind choosing Q35 over i440fx: fewer trapped register accesses, fewer legacy devices to walk.\nThe Host OS Is a Feature, Not Baggage The other half of that ledger never makes it into the argument, so here it is: a Proxmox node is a machine you can actually work on. It is Debian, so the whole Debian archive is one apt install away.\nMonitoring you already run — a Prometheus node exporter, smartmontools, your existing agent — rather than whatever the appliance chooses to expose. Diagnostics when something is slow: fio, iperf3, nvme-cli, perf, bpftrace. Backup agents from any vendor shipping a Linux binary. Configuration management, so the hypervisor sits in the same Ansible inventory as everything else instead of being a special case. fwupd for firmware, on hardware whose vendor supports LVFS. None of that needs a plugin format, a signed bundle, or vendor blessing. Compare it with the ESXi model, where the shell is deliberately restricted, third-party code arrives as a VIB, and there is no package manager to reach for at all.\nIt Reaches Into the Storage Stack Monitoring agents are the boring version of this. Proxmox\u0026rsquo;s own storage documentation lists the native plugins — dir, NFS, CIFS, CephFS, ZFS, BTRFS, LVM, LVM-thin, iSCSI, FC/SAS, RBD, ZFS-over-iSCSI, PBS — and then adds a sentence worth taking literally:\nyou may use all storage technologies available for Debian Linux\nThat is a statement about where the boundary is, and the mechanism behind it is generic: get a block device onto every node, put LVM on it, add it as an LVM storage with shared enabled. That is exactly how the supported Fibre Channel and iSCSI paths work, so anything that can produce a shared block device can use the same route.\nFour transports, one route to shared Proxmox storage the transport what Linux gives you what Proxmox sees Fibre Channel / SAS native plugin iSCSI native plugin NVMe/TCP or RDMA no plugin \u0026#8212; nvme-cli ATA over Ethernet no plugin \u0026#8212; aoetools a block device on every node /dev/sdX \u0026#183; /dev/nvme0n1 LVM volume group shared: 1 storage type: lvm live migration across the cluster Proxmox VE only has to recognise the two boxes on the right. Everything to their left is the Linux block layer, and it does not care how the device arrived. Two of these transports have a Proxmox plugin and two have nothing at all, and it makes no difference past the second box. The LVM layer is where a shared LUN becomes cluster storage, whatever delivered it. NVMe over TCP or RDMA is the case worth knowing about, because it is fast, current, and absent from that plugin list. nvme-tcp and nvme-rdma are in-tree Linux host drivers — NVMe/TCP has been in mainline since 5.0 — so there is nothing to compile. nvme-cli is a Debian package (nvme discover, then nvme connect), and nvmetcli configures the in-kernel nvmet target at the other end. A connected namespace appears as /dev/nvmeXnY, and from there it is an ordinary block device.\nATA over Ethernet makes the same point from the opposite end of the spectrum — ancient, obscure, equally unsupported as a plugin. The aoe driver is in mainline; aoetools gives you aoe-discover and aoe-stat; vblade turns any file or block device on another machine into a target. Same route, same outcome.\nNeither is a plugin API, an SDK, or a certification programme. It is what happens when the hypervisor\u0026rsquo;s storage layer is the Linux block layer.\nThe honest caveats, because this reads as a party trick until it is 3 a.m. Neither transport is a tested Proxmox storage type, so the integration and its failure modes — reconnect behaviour, multipath, timeouts under load — are yours to own and yours to test before anything important lives on it. And AoE is a bare Layer 2 protocol, non-routable and with no authentication, so it belongs on an isolated storage VLAN and nowhere else. NVMe/TCP at least has a discovery model and can be routed, which is much of why it is the one to reach for now.\nWho Else Runs KVM Here is where the \u0026ldquo;already running it\u0026rdquo; claim comes from. Every platform below runs the same kernel module.\nPlatform Where you meet it The KVM part The user-space VMM Amazon EC2 (Nitro) Public cloud KVM core module Custom — QEMU removed, device model in the Nitro cards AWS Lambda, Fargate Serverless /dev/kvm Firecracker — a minimal microVM monitor in Rust Google Compute Engine Public cloud KVM since launch Google\u0026rsquo;s own VMM, deliberately not QEMU Alibaba Cloud ECS Public cloud Simplified KVM (X-Dragon) Custom, with net and storage offloaded to a MoC card Oracle Cloud (OCI) Public cloud Oracle Linux KVM — the same stack Oracle ships on-premises QEMU lineage DigitalOcean, Linode/Akamai, Vultr, Hetzner, OVHcloud, Scaleway, UpCloud Public cloud Stock KVM QEMU Nutanix AHV On-premises HCI Stock KVM QEMU with libvirt and Open vSwitch OpenStack (Nova) Private cloud Stock KVM QEMU via libvirt — the default and best-tested driver Apache CloudStack, OpenNebula, oVirt On-premises Stock KVM QEMU via libvirt OpenShift Virtualisation, SUSE Harvester Kubernetes Stock KVM QEMU inside a pod, via KubeVirt Proxmox VE On-premises Stock KVM QEMU, with LXC alongside for containers Everyone Keeps the Kernel Half and Rewrites the User-Space Half The hyperscalers did not fork KVM. They forked QEMU\u0026rsquo;s job.\nAWS moved EC2 off Xen onto the Nitro hypervisor, which is built on the KVM core kernel module with QEMU thrown out entirely. The device model lives in dedicated Nitro cards instead, which is how they get performance indistinguishable from bare metal. Every current-generation EC2 instance type runs this. Separately, Lambda and Fargate run Firecracker, a purpose-built VMM in Rust that talks to the same /dev/kvm.\nGoogle has run every Compute Engine VM on KVM since Compute Engine launched, and wrote its own user-space VMM rather than using QEMU, explicitly to avoid QEMU\u0026rsquo;s huge matrix of guests, devices and modes. They also went the other way and hardened the kernel module upstream, removing emulated devices nobody needed and narrowing the set of emulated instructions. That work is in the KVM you are running.\nAlibaba did the same shape of thing with X-Dragon: a stripped-down KVM hypervisor with the network and storage virtualisation offloaded onto an FPGA-based MoC card.\nNutanix AHV is KVM plus libvirt plus QEMU plus Open vSwitch plus Nutanix\u0026rsquo;s orchestration — of every commercial product on that list, the closest relative Proxmox VE has. A different layer 4, and a very different invoice.\nFive platforms, five device models, one KVM Amazon EC2Nitro Google CloudCompute Engine Alibaba CloudECS, X-Dragon NutanixAHV Proxmox VEon Debian management device model hypervisor hardware EC2 control plane console and API per instance-hour GCE control plane console and API per instance-hour ECS console and API per instance-hour Prism and AOS on every node per node, per year pve-manager pveproxy, HA, SDN on every node optional subscription custom VMM, no QEMU at all devices live on the Nitro cards Google's own user-space VMM, deliberately not QEMU simplified VMM, network and storage offloaded to a MoC card QEMU, libvirt and Open vSwitch QEMU, with LXC alongside it KVM \u0026#8212; kvm.ko, kvm-intel.ko / kvm-amd.ko the isolation boundary, and it is the same code in every column Intel VT-x + EPT \u0026#183; AMD-V + NPT The same kernel module in every column. What changes going up the stack is the device model, then the management layer, then the licensing — which is also the order in which these platforms actually differ. Proxmox VE Keeps QEMU On Purpose It is tempting to read the table as a ranking, with AWS and Google at the top for having replaced QEMU. That is the wrong reading, because their constraint is not yours.\nAWS and Google run one hardware profile, at a scale where a single device-emulation bug is a fleet-wide event, and they control every guest image boundary they care about. In that world QEMU\u0026rsquo;s breadth is nearly all liability, so deleting it is obviously correct.\nYou are not in that world. You have an appliance image from 2013 that wants an e1000. You have a Windows VM whose machine type must stay pinned for the rest of its life. You have a GPU to pass through, an emulated SAS controller to satisfy an installer, a UEFI variable store to preserve. QEMU is exactly what lets a general-purpose platform say yes to all of that.\nWhat AWS calls attack surface is what you call a compatibility matrix. Both descriptions are accurate; the difference is whether you get to choose your workloads.\nIt also explains the shape of that patch queue. AWS and Google solved backup and snapshots outside the VMM, in their own storage services. Proxmox had no storage service to solve it in, so they put it in QEMU — which is why savevm-async and the PBS block driver exist as patches rather than as products.\nAnd that is worth knowing for a practical reason: the parts of Proxmox VE you would miss most are the parts that are not upstream. Your VM configs are plain text and your disk images are standard formats, so a machine will move. But a snapshot including RAM state, and a Proxmox Backup Server incremental chain, depend on Proxmox\u0026rsquo;s QEMU fork. That is a much lighter dependency than a proprietary hypervisor — the fork is public, AGPL, and you can read every patch in it — but it is not zero, and \u0026ldquo;no lock-in at the software level\u0026rdquo; should carry that footnote.\nSo — Is It Enterprise Grade? The lazy form of this argument does not work, and it is worth saying so plainly. \u0026ldquo;AWS uses KVM, therefore Proxmox VE is enterprise grade\u0026rdquo; is a non-sequitur: it takes a claim about one layer and quietly applies it to an entire product.\nHere is the version that does hold.\nWhat is shared is the layer that is hardest to get right and most dangerous to get wrong. CPU and memory virtualisation, and the isolation boundary between tenants, is the part where a bug is a breach rather than an outage. That code is reviewed by engineers paid by Amazon, Google, Red Hat, Intel, AMD, IBM and Alibaba, and its bugs are found by the outfits running the largest fleets in existence — usually before the kernel reaches you. When a guest-escape vulnerability does land, the fix arrives through the normal kernel update you were going to apply anyway. You are not waiting on one vendor\u0026rsquo;s release cycle for a hypervisor only that vendor can see.\nWhat is not shared is everything above the boundary. Which means the question collapses into two much more answerable ones:\nIs the management layer good enough for the way you operate? Can you buy support for it with terms you can live with? Both can be tested in a proof of concept and written into a contract. Neither requires faith in a hypervisor.\nThat is a much better position than the one the original question assumes — that you are being asked to trust a novel hypervisor from a small vendor. You are not. The Proxmox-specific code is a management layer mostly in Perl and increasingly in Rust, plus that patch queue against QEMU, and note where the patches land: the device model and the backup path, not the isolation boundary.\nIt is also worth knowing what the Proxmox-specific layer failing costs you. The QEMU processes are ordinary independent processes on the host, so pveproxy falling over does not stop a single virtual machine. That is a very different blast radius from losing the component that owns the isolation boundary.\nWhat \u0026ldquo;The Same Hypervisor\u0026rdquo; Does Not Buy You This is where vendor confidence documents tend to stop, which tells you something about who they are written for. It is the more useful half.\nIt does not buy you AWS\u0026rsquo;s reliability. Nitro\u0026rsquo;s availability has very little to do with KVM. It comes from the control plane, the network fabric, the storage service, the capacity management and the operational practice around it. Your cluster\u0026rsquo;s reliability will come from your Corosync quorum design, your fencing configuration, your storage choice and your network redundancy. KVM has no opinion on any of those. Sharing a hypervisor with a hyperscaler does not inherit their operations.\nIt does not buy you Nitro\u0026rsquo;s attack surface, and this is where the legitimate half of the type 2 accusation lands. You are running QEMU on a general-purpose kernel, and historically QEMU\u0026rsquo;s device model is where the memorable VM escapes lived — VENOM, in an emulated floppy controller nobody was using, being the canonical example. Google could note at the time that Compute Engine was unaffected exactly because it does not run QEMU. You do run it, so the compensating controls are yours: prefer virtio over emulated hardware, do not present devices you do not need, leave the AppArmor profiles alone, and patch QEMU on the same discipline as the kernel. Proxmox\u0026rsquo;s extra/ patches are them doing exactly that on your behalf, which is a reasonable thing to check they are still doing.\nIt does not make VM behaviour portable between platforms. CPU model selection, machine type versioning, live migration compatibility and clock behaviour are all decided in layers 3 and 4, and they differ everywhere. A cluster of mixed CPU generations will still punish you for setting the CPU type to host, and guest clocks still drift whatever the logo on the platform is. Same hypervisor is not same behaviour.\nIt does not answer the support question — which is the one procurement actually cares about, and rightly so. Proxmox VE is AGPLv3 and free to run in production. The subscription buys the enterprise-tested repository and vendor support, starting at €120 per socket per year. Proxmox\u0026rsquo;s own support is delivered during Austrian business hours, so round-the-clock coverage comes from partners rather than from Vienna. croit, where I work, is one of those partners, and covers 24/7, 365 days a year. That is a commercial negotiation, not a technical risk — and having it be a commercial negotiation is the point of everything above.\nWhat To Actually Evaluate Instead If the hypervisor is settled, a proof of concept should spend its time on the layer that is genuinely specific to Proxmox VE:\nFencing and HA. Pull the power on a node with running HA guests and time the restart. Then do it to two nodes and confirm the remaining ones behave the way you expect when quorum is lost. Live migration across CPU generations. With the CPU model you actually intend to standardise on, not host. Backup and, more importantly, restore. Restore times under load, not backup times. Nobody has ever been thanked for a fast backup. The permissions model. Whether you can hand an application team console and power control over their own VMs and nothing else, granularly enough to satisfy an auditor. The API. Everything the web interface does is an API call; if your automation cannot drive it, the platform will not fit how you work. The upgrade path. A major-version upgrade on the PoC cluster, before you have 400 VMs on it. None of those are questions about KVM. That is rather the point.\nThe hypervisor is the settled part. Spend the proof of concept on the parts that are not.\nReferences KVM itself\nKVM API documentation — the /dev/kvm ioctl interface, which is the whole hypervisor contract Linux 2.6.20 release notes — the release KVM was merged into, February 2007 Some KVM developments — LWN, January 2007, on KVM in the days after it landed in mainline How the other hypervisors are built\nInterpreting virtual machine monitor and executable failures — Broadcom\u0026rsquo;s own description of a running ESXi VM as \u0026ldquo;several processes or userworlds\u0026rdquo;, with one VMM per vCPU and one VMX per VM The Architecture of VMware ESXi (PDF) — the whitepaper calling VMkernel \u0026ldquo;a POSIX-like operating system\u0026rdquo; with processes, signals, a file system and threads. Linked via a mirror because the original VMware URL did not survive the Broadcom reorganisation, which is its own small comment on vendor continuity Hyper-V architecture — Microsoft on the Virtual Machine Worker Process as a user-mode component, one per VM, where all emulated devices live The Nutanix Bible — AHV architecture — AHV described as KVM with libvirt, QEMU and Open vSwitch Who runs KVM, and how they changed it\n7 ways we harden our KVM hypervisor at Google Cloud — Google on running KVM without QEMU, and the upstream hardening they did The AWS Nitro System — which EC2 instance types run the Nitro hypervisor AWS EC2 Virtualization 2017: Introducing Nitro — Brendan Gregg\u0026rsquo;s write-up of the Xen-to-KVM transition, still the clearest account of it Firecracker — AWS\u0026rsquo;s minimal KVM-based VMM behind Lambda and Fargate Alibaba Cloud\u0026rsquo;s sixth-generation ECS instances — the X-Dragon hypervisor and MoC offload KubeVirt — the QEMU/KVM-in-a-pod model behind OpenShift Virtualisation and Harvester OpenStack Nova hypervisor support matrix — KVM as the reference driver What Proxmox maintains\nThe Proxmox GitHub organisation — 89 repositories, including pve-qemu, pve-kernel, pve-edk2-firmware and packaging for lxc, zfsonlinux, openvswitch, libiscsi, corosync-pve and ceph pve-qemu patch series — the 78 patches counted above, and the fastest way to see exactly what Proxmox adds to QEMU Proxmox VE storage documentation — the native plugin list, which storage types are shared, and the line about using all storage technologies available for Debian Linux nvme-cli and nvmetcli — NVMe-oF host and target tooling; aoetools and vblade for the ATA over Ethernet equivalent. All in Debian trixie, which is what Proxmox VE 9 is built on Proxmox VE pricing and subscription tiers — what the subscription covers Disclosure: I work for croit, a Proxmox Gold Partner. The technical claims above are sourced and checkable; the commercial paragraph is the part where I have an interest, so treat it accordingly.\n","permalink":"https://blogs.damiendye.uk/en/proxmox/kvm-the-hypervisor-inside-proxmox-ve/","summary":"\u0026ldquo;Is Proxmox enterprise grade?\u0026rdquo; is a question about the hypervisor, and Proxmox VE does not contain one. The hypervisor is KVM — the same kernel module underneath EC2, Google Cloud, Alibaba Cloud, Nutanix AHV and OpenShift Virtualisation. What Proxmox really maintains, why the type 1 argument is aimed at the wrong line, and what the shared foundation does not buy you.","title":"Proxmox VE Is Not a Hypervisor — KVM Is, and You Are Already Running It"},{"content":"Why HDDs Are Back on the Menu For years the advice was simple enough to fit on a sticky note: put it all on SSDs.\nThat advice was right, and it was also cheap. Neither is quite true now. NAND supply has tightened, AI demand has pushed flash pricing up, and high-capacity NVMe has become hard to justify for a cost-conscious Proxmox deployment — the kind that needs to be operational, reliable, and as inexpensive as honesty allows.\nEnterprise HDDs are interesting again for the same reason they always were: capacity per pound, and nothing else comes close.\nThe catch is that everybody who has run VMs on spinning disks knows how badly it can go. That experience is real, and it is usually the architecture\u0026rsquo;s fault rather than the disks'.\nBlaming the disks is easier, mind. It is also wrong, and it is expensive, because it talks people into buying flash they did not need.\nEverything below assumes Proxmox with hyper-converged Ceph OSDs, because that is where the layout decisions actually bite.\nThe Problem Is Latency, Not Throughput A modern enterprise HDD moves large sequential data perfectly well. That was never the bottleneck.\nA 7.2k RPM enterprise drive delivers somewhere around 80 to 100 IOPS of random work, because every random operation waits on a head seek and a platter rotation. That number has barely moved in twenty years. It is mechanics, not electronics. An entry-level enterprise SSD is three or four orders of magnitude away from it.\nGive the same drive sequential work and it is a useful device. The operation rate roughly doubles, to about 200 IOPS, but each of those operations now carries a large block because the head is already in the right place and can keep streaming. So you get real throughput, 150 to 250 MB/s of it, at a price per terabyte nothing else touches.\nThat is the important asymmetry, and it is not \u0026ldquo;HDDs are slow\u0026rdquo;. A spinning disk is good at sequential work and hopeless at random work, and the two numbers are close together only because the bounded thing is the operation count, not the bytes.\nWhich sets the whole design brief. You are not trying to make the disk faster. You cannot. You are trying to arrange for it to spend its time in the mode where it is already competent, and to keep the small random traffic away from it. As such, everything that follows is that one idea applied twice.\nThe problem is that virtual machines do not generate the workload HDDs are good at. They generate metadata updates, journal writes, small synchronous flushes, logging, filesystem housekeeping, and bursts of random activity from a dozen guests at once that arrive interleaved at the disk.\nHopeless at random work, genuinely good at sequential — which mode the disk runs in is the whole design Operations per second, one device — bars not to scale, because three orders of magnitude will not fit 7.2k HDD — random 80–100 every operation waits on a seek and a rotation — hopeless 7.2k HDD — sequential ~200 and each one carries a big block: 150–250 MB/s of real throughput Enterprise NVMe 100k+ no moving parts, so no seek to pay for What a hyper-converged cluster actually sends the disk guest metadata updates journal writes and fsync logging and housekeeping RocksDB metadata random bursts, many guests large sequential transfers Five of those six are small and random. One of them is what the spindle is good at. This is not \"HDDs are slow\". A spinning disk is genuinely good at sequential work and hopeless at random work, and the two figures sit close together only because the bounded quantity is the operation count, not the bytes each one carries. So the job is not to make the disk faster. It is to keep it in the mode it is already competent in. Random and sequential operation rates sit surprisingly close together, because what is bounded is the number of operations rather than the bytes each one carries. Sequential is where the disk earns its money; a hyper-converged cluster natively sends it the other kind of work. Ceph adds its own layer to this. Every write carries RocksDB metadata updates alongside the data. If all of that lands on raw spindles, the disks spend their time seeking instead of serving, and you get the symptom every administrator recognises: high I/O wait, sluggish guests, inconsistent responsiveness, and a cluster that feels far slower than its capacity sheet suggests.\nCeph\u0026rsquo;s own hardware guidance is blunt about where this leads. It warns to \u0026ldquo;consider carefully the ostensible cost-per-gigabyte advantage of larger HDDs, and the concomitant limitations of IOPS per TB\u0026rdquo;, and says drives above 8 TB \u0026ldquo;may be best suited for storage of large files / objects that are not at all performance-sensitive.\u0026rdquo;\nWorth doing that arithmetic, because the drives that make this economic in the first place are large ones. 20 TB and up. A 20 TB drive still does its 80 to 100 random IOPS, because capacity buys you no operations whatsoever. So it offers about 4 to 5 random IOPS per terabyte, where an 8 TB drive manages roughly eleven, and a 2 TB drive forty.\nBy Ceph\u0026rsquo;s own measure, then, a 20 TB spindle is well inside the territory it tells you not to put performance-sensitive workloads on.\nThat is not an argument against buying them. It is the reason the rest of this post exists. The design below is exactly what makes a 20 TB spindle viable for VM storage, by arranging for the small random work never to reach it.\nMove the Metadata Off the Spindle The highest-value change to an HDD-backed cluster is to stop making the spindle store Ceph\u0026rsquo;s bookkeeping.\nBlueStore keeps three things per OSD: the data itself on block, the RocksDB metadata database on block.db, and the write-ahead log on block.wal. By default all three live on the same device. Put block.db on NVMe instead and a large amount of small random I/O leaves the disk entirely.\nCeph\u0026rsquo;s hardware documentation puts it plainly: \u0026ldquo;HDD OSDs may see a significant write latency improvement by offloading WAL+DB onto an SSD.\u0026rdquo;\nBlueStore OSD layout: everything on the spindle, versus metadata on NVMe Default: one device holds all three object data writes RocksDB metadata One 7.2k HDD block block.db .wal Both streams share one 80-IOPS budget. The disk seeks between data and metadata for every write. block.db on NVMe: the metadata leaves object data writes RocksDB metadata 7.2k HDD block — data only Enterprise NVMe block.db .wal The spindle now does one job. Name a DB device and the WAL goes with it. Size\u0026#160;block.db\u0026#160;at 1–2% of\u0026#160;block\u0026#160;for RBD workloads, or 2.5% to be comfortable. Undersize it and RocksDB does not fail — it spills back onto the spindle, which is the one outcome that wastes the flash you just bought. Useful sizes step at roughly 3 GB, 30 GB and 300 GB, because a DB partition can only fully use sizes matching the sums of RocksDB's levels. Rounding up to the next step is usually cheap; rounding down buys nothing. Left, everything on one spindle: data writes and RocksDB updates compete for the same 80-odd IOPS. Right, block.db on NVMe: the metadata traffic leaves the disk, and the spindle is left doing the one thing it is good at. Use proper enterprise NVMe here, not consumer flash. The reason is not headline benchmark numbers. It is stable latency under sustained load, and power-loss protection. Ceph\u0026rsquo;s documentation is direct: \u0026ldquo;Enterprise-class SSDs are best for Ceph: they feature power loss protection (PLP)\u0026rdquo;, and \u0026ldquo;bargain client-class or off-brand SSDs are a false economy.\u0026rdquo;\nSamsung PM9A3 and Micron 7450 Pro class drives are the sort of thing that belongs here. Under write pressure Ceph cares about consistency far more than peak numbers, and that is exactly what separates an enterprise drive from a fast consumer one.\nSizing block.db, and What Spillover Costs Get the size wrong and the benefit quietly evaporates, because when block.db fills, RocksDB does not fail — it spills back onto the slow device, exactly where it would have been without the NVMe.\nThat is the worst of both worlds: you have bought the flash and you are still seeking on the spindle.\nThe current Ceph guidance:\nWorkload block.db as a fraction of block RBD / block — VM disks 1% to 2% RGW / object at least 4% General recommendation when offloading WAL+DB at least 2.5% Proxmox VM storage is the RBD row, so 1–2% is the honest planning figure, and 2.5% is a comfortable place to sit.\nRun that against a 20 TB drive and the numbers get your attention:\nblock.db for one 20 TB OSD At 1% 200 GB At 2% 400 GB At 2.5% 500 GB Take that with the RocksDB level steps below and 300 GB per OSD is the sensible landing spot. The useful step nearest the middle of the range.\nWhich changes what the shared metadata device actually is. Eight 20 TB OSDs want 2.4 TB of NVMe between them; fifteen of them, Ceph\u0026rsquo;s stated maximum behind one NVMe, want 4.5 TB. On drives this large it is DB capacity that decides how many OSDs sit behind one device, not the ratio. You will run out of gigabytes long before you run out of Ceph\u0026rsquo;s blessing.\nThere is one more wrinkle worth knowing, because it makes intuitive sizing wrong. RocksDB\u0026rsquo;s level structure means a DB partition can only fully use sizes that correspond to the sums of its levels — historically producing useful steps around 3 GB, 30 GB and 300 GB, with sizes in between offering nothing over the step below. Rounding 18 GB up to 30 GB is often free in practice; rounding it down to 20 GB buys you nothing over 3 GB in the worst case.\nCheck for spillover after the cluster has been running a while, not on day one. It is a slow-onset problem.\nYou Do Not Need a Separate WAL Device This one saves a partition and a lot of fiddling.\nIf you specify a DB device and no explicit WAL device, Ceph puts the WAL on the fast device with the DB automatically. The documentation states that \u0026ldquo;whenever a DB device is specified but an explicit WAL device is not, the WAL will be implicitly colocated with the DB on the faster device.\u0026rdquo;\nSo \u0026ldquo;put the DB and WAL on NVMe\u0026rdquo; is the right goal, but it is one argument, not two. A separate block.wal only makes sense if you have a third, faster tier to put it on.\nTurn Off the HDD Write Cache Small, unglamorous, and easy to miss.\nCeph\u0026rsquo;s documentation notes that OSD performance \u0026ldquo;may be dramatically increased \u0026hellip; by disabling this write cache\u0026rdquo; on HDDs. The volatile cache in the drive reorders and delays writes in ways that fight Ceph\u0026rsquo;s flush semantics, and turning it off usually makes things faster rather than slower.\nIt becomes mandatory rather than advisable once bcache is involved, which is the next section.\nWhile you are in there: you do not need a RAID-capable HBA. Ceph says so directly: \u0026ldquo;You do not need an RoC (RAID-capable) HBA.\u0026rdquo; Give the OSDs the disks.\nOne Optane Per Spindle Moving the metadata off the disk fixes the steady state. It does not fix bursts, because a burst of small synchronous writes still has to be acknowledged, and the spindle still acknowledges at spindle speed.\nThat is what bcache is for, and what makes it work here is Optane.\nWhat matters is not capacity. It is behaviour under write load. Against NAND, Optane offers very low latency, very high endurance, and none of the garbage-collection cliffs that make an SSD\u0026rsquo;s write-cache performance unpredictable once it has been in service a while. Absorbing small, bursty, write-heavy traffic is the thing it is best in the world at.\nThe layout that matters: one Optane per HDD, each pair its own bcache device. Not one Optane shared across a shelf of disks.\nThree tiers, two sharing models: paired write cache, shared metadata device One Proxmox node, three OSDs Guest VM writes — small, synchronous, bursty Ceph RocksDB metadata bcache0 Optane write cache HDD osd.0 block bcache1 Optane write cache HDD osd.1 block bcache2 Optane write cache HDD osd.2 block One Optane per spindle — paired, never shared A cache failure costs exactly one OSD, which Ceph rebuilds routinely. Endurance needed here: 30–100 DWPD. Optane only. Shared enterprise NVMe block.db + implicit WAL, all three OSDs power-loss protected Shared, because metadata is small Ceph says 4–5 HDDs per SATA SSD, no more than 15 per NVMe. Lose this one and every OSD behind it goes with it. Enterprise NAND is fine — but a fast NVMe, not a cheap one. Both tiers are required, and they are not the same purchase. The Optane shortens write acknowledgement; it does nothing for the metadata reads Ceph issues constantly, and only defers the metadata writes. Skip the NVMe and RocksDB is back on the platter. Every write to an OSD crosses its cache device, which is why that one must be Optane. The metadata device need not be — but it must still be a fast NVMe: small write volume, and yet it carries the metadata latency of every OSD behind it. Three tiers, two different sharing models. The DB/WAL device is shared across several OSDs because RocksDB write volume is small, though it still has to be a fast NVMe because it carries the metadata latency of every one of them. The Optane write cache is paired one-to-one with its spindle, so a cache failure can only ever cost one OSD. The pairing is a deliberate choice, and it buys two things.\nNo contention. Each spindle gets a whole Optane\u0026rsquo;s latency and queue depth to itself, rather than queueing behind seven other OSDs\u0026rsquo; bursts.\nA failure domain Ceph already knows how to survive. A shared write cache holds dirty data for every OSD behind it, so losing it loses all of them at once; paired one-to-one, losing an Optane costs exactly one OSD. That distinction is the single most important consequence of this design, and it gets its own treatment in what you are trading away.\nThis Has To Be Optane. Not NAND. The word \u0026ldquo;Optane\u0026rdquo; is not a brand preference or a nice-to-have here. Substituting a NAND NVMe — however expensive — builds something that wears out, and the reason is endurance.\nLook at where the cache device sits. Because this design turns bcache\u0026rsquo;s sequential bypass off — the sequential_cutoff=0 in the rule below — every single write to that OSD passes through it: not just the bursts, all of it. That makes the cache the highest-duty-cycle device in the node, asked to absorb the write volume of an entire spinning disk for the life of the machine.\nEndurance is quoted in drive writes per day, and the gap is not incremental:\nDevice Rated endurance Optane P5800X 100 DWPD Optane P4800X 30 DWPD Write-intensive enterprise NAND, top of the range around 10 DWPD Mixed-use enterprise NAND 3 DWPD or less Mainstream enterprise NAND — PM9A3, 7450 Pro class around 1 DWPD Optane is between ten and a hundred times the endurance of the NAND you would otherwise put there. Put a 1 DWPD drive in the one position that receives every write in the cluster and you have not built a cache, you have built a consumable.\nIt is worse than the table suggests, because NAND suffers write amplification internally and Optane does not. A NAND drive\u0026rsquo;s rated DWPD is what it takes at the interface; what its cells actually absorb is larger. Optane\u0026rsquo;s number needs no such asterisk.\nRebuilds are where that gets tested. Backfill pushes terabytes of writes through the surviving nodes in a compressed window, every byte of it across their cache devices — an endurance event, not merely a latency one, arriving exactly when you least want a second device to fail.\nSo the two flash tiers want different drives, for different reasons:\nTier Write volume What the device needs DB/WAL small — RocksDB only A fast enterprise NVMe. Low latency, high random IOPS, PLP. Enterprise NAND is the right technology; a PM9A3 or 7450 Pro is the right buy Write cache all of it PLP and endurance in the tens of DWPD. NAND is the wrong technology at any price Buying one kind of drive for both jobs is the mistake this section exists to prevent — but do not read the first row as permission to economise, because low write volume is not low demand.\nThe DB/WAL device still has to be genuinely fast NVMe, for three reasons that have nothing to do with how many bytes cross it.\nIt serves every OSD behind it at once, so the queue it sees is four or fifteen sets of metadata traffic, not one disk\u0026rsquo;s worth.\nIts latency lands directly on your clients. RocksDB lookups sit on the critical path for finding objects, for peering and for scrub, and they are small random reads. The workload where the gap between a fast NVMe and a mediocre one is widest. Every microsecond is multiplied by every OSD depending on it.\nAnd compaction is bursty. RocksDB periodically rewrites its levels, turning steady churn into a concentrated spike of reads and writes. A device that copes with the average and stalls on the spike stalls every OSD behind it at the same moment.\nCeph\u0026rsquo;s own ratios are the tell. It allows three times as many OSDs behind an NVMe as behind a SATA SSD, and that is the interface and the device class talking, not the capacity. Use NVMe.\nSo: the cache device must be Optane, the DB/WAL device may be NAND, and neither position is where you save money.\nWhich Optane: A 32 GB M10 Would Often Do None of that means the biggest Optane is the right one. The device you need is decided by your write workload — though as it turns out, the second-hand market may decide it for you regardless. The arithmetic still matters, because it tells you what you are over- or under-buying.\nStart from what the cache is for. It holds a burst, not a disk. At writeback_percent=40 a 32 GB device gives roughly 13 GB of dirty buffer, which is a huge number of small synchronous writes.\nSo do not be put off by the ratio. A 32 GB module in front of a 20 TB disk is 0.16% of it, which sounds absurd until you remember it is a burst absorber rather than a working-set cache trying to hold hot data. A 375 GB device in front of the same disk is 1.9%, which is simply more headroom than you will use.\nWhich puts the cheap end of the range in play. Here is the little consumer module against the datacentre part, because the gap is not where people expect:\nOptane Memory M10 32 GB Optane DC P4800X 375 GB Random read, 4K 240,000 IOPS 550,000 IOPS Random write, 4K 65,000 IOPS comparable to read Sequential read 1,200 MB/s 2,400 MB/s Sequential write 290 MB/s 2,000 MB/s Typical latency — under 10 µs Endurance 365 TBW 12.3 PBW — 30 DWPD Interface PCIe 3.0 x2, M.2 2280 PCIe 3.0 x4, U.2 or AIC Look at the write IOPS row first. The little M10 does 65,000 random writes a second; the spindle behind it does 80. Even the cheapest Optane is roughly 650 times the disk it is caching, so on the performance axis the argument is over before it starts. The DC part\u0026rsquo;s extra IOPS have nowhere to go when there is one 7.2k disk downstream.\nThe row that decides the purchase is endurance, where the gap is 34×. Turned into a daily budget over a five-year life:\nDevice Sustainable writes per day, over five years M10 32 GB — 365 TBW about 200 GB/day P4800X 375 GB — 12.3 PBW about 6.7 TB/day Under 200 GB of writes a day and a 32 GB M10 sees out the build. At twice that you get two and a half years. Properly write-heavy and the DC part earns its price — not for the IOPS, which you cannot use, but for the thirty-odd times more writing it tolerates.\nSo measure rather than guess. nvme smart-log on an existing cache device gives you data_units_written; sample it a week apart, divide, and compare against the rated TBW of whatever you are considering. bcache keeps its own totals under /sys/block/bcacheN/bcache/stats_total/ if you would rather read it there.\nTwo caveats on the M10. Its random figures are quoted over an 8 GB span rather than the full device, so on a cache you intend to run substantially full, treat them as the optimistic end. And it is a consumer part carrying no enterprise PLP claim — Optane\u0026rsquo;s media is written in place with no DRAM buffer in the write path, which is why these modules behave far better on sudden power loss than consumer NAND would, but \u0026ldquo;behaves well in practice\u0026rdquo; is not a specification. If you want the guarantee in writing, buy a DC-series part.\nThe Floor Is Bandwidth, Not Capacity — Skip the 16 GB Modules There is a second constraint, independent of everything above, and it disqualifies the cheapest part in the range.\nThe cache has to be faster than the disk sequentially, or it is a throttle. The 16 GB M10 is not, and it is the module you will be most tempted by because it is nearly free:\nSequential write Optane Memory M10 16 GB 150 MB/s Optane Memory M10 32 GB 290 MB/s Optane DC P4800X 375 GB 2,000 MB/s A 7.2k enterprise HDD 150–250 MB/s Read the first and last rows together. A 16 GB module writes sequentially at about the same speed as the spinning disk it is supposed to be accelerating, and slower than a good one. It remains tremendously faster for random work — 35,000 random write IOPS against the disk\u0026rsquo;s 80 — but on a sequential stream it is at best a wash and at worst a ceiling below what the bare disk managed unaided.\nAnd this design guarantees you meet that ceiling, because turning off the sequential bypass sends everything through the cache, making the cache\u0026rsquo;s own sequential write bandwidth the hard ceiling for the whole OSD. Recovery is where it bites hardest: backfilling a 20 TB OSD is about as sequential as this workload gets, and capping it at 150 MB/s makes an already slow rebuild slower.\nSo the M10 line has a floor at 32 GB, and it is a bandwidth floor rather than a capacity one. The capacity arithmetic said 32 GB was ample; the bandwidth arithmetic rules out 16 GB regardless of how little you needed to store. Both tests have to pass, and the cheap part fails the one people do not check.\nWhatever you are considering, put its sequential write figure next to 250 MB/s before you buy it.\nBuying a Product Nobody Makes Any More Optane is discontinued. Intel wound the business down in 2022, writing off $559 million of inventory and ending development. There is no new production and no restock; the total supply only goes down from here.\nWhich raises an obvious question, because there is a surprising amount of it for sale. Understanding where it comes from tells you what you are buying.\nIt is OEM service-spares inventory being liquidated. The big three server makers stocked Optane as spare parts to support the platforms they sold it in. Those platforms have gone end-of-life and dropped off support, so the spares backing them became dead stock overnight — warehouses of parts for machines nobody is contractually obliged to fix any more. That inventory is what is flowing onto AliExpress and the refurbishers.\nTwo consequences, and both are good news.\nA lot of it is unused rather than pulled. Service spares sat on a shelf waiting for a failure that never came, so the wear figure on arrival is often effectively zero — not \u0026ldquo;a few years of light duty\u0026rdquo; but never written. Verify rather than trust: nvme smart-log gives you percentage_used and data_units_written, and on genuine spares stock those should be startlingly low. Anything showing real wear is a pull being sold as something else, though even then the endurance headroom means a used Optane can have more life left than a brand-new NAND drive of the same capacity.\nIt explains which capacities you will find. OEM spares were stocked for servers, which means datacentre parts. Here is the full range, and note where the flagship starts:\nPart Capacities Form Optane Memory M10 16, 32, 64 GB M.2 2280, consumer Optane SSD 800P 58, 118 GB M.2 2280, consumer Optane SSD P1600X 58, 118 GB M.2 2280, datacentre boot and cache part Optane SSD DC P4801X 100, 200, 375 GB M.2 110 mm or U.2 Optane SSD DC P4800X 375 GB at its smallest, 750 GB, 1.5 TB U.2 or add-in card Optane SSD P5800X 400 GB, 800 GB, 1.6 TB U.2 or add-in card In theory the small M.2 parts are the elegant answer — the M10, the 800P, the P1600X and the 100 GB P4801X all do the job, and cheaply. In practice the market does not have them, because nobody warehoused desktop accelerator modules as server spares. What is listed is 375 GB and upwards, P4800X-class hardware.\nSo plan on buying more capacity than the role needs, because that is what is for sale. It is not a bad outcome. Overbuying a cache device lands you on 30 DWPD and sub-10 µs latency when your workload demanded a fraction of either, and for a 20 TB spindle you intend to keep for years, that is the right way to err. It does mean the entry price is higher than the arithmetic suggests, and that the sizing above becomes a check against under-buying rather than a shopping list.\nThree things to plan for, since this is a liquidation rather than a supply chain:\nBuy your spares with the build. A finite pool is being cleared. When a device fails in three years you will not be ordering a replacement, you will be hunting one — so cost the spares in now, while the stock is there.\nExpect OEM firmware and check the namespace. Parts from a vendor\u0026rsquo;s spares programme often carry that vendor\u0026rsquo;s firmware and may arrive formatted to an LBA size or with metadata settings suiting whatever they were stocked for. Confirm with nvme id-ns before you build on it, and be ready to nvme format to a plain 4096-byte format.\nThere is no warranty, no support, and no more firmware updates. Intel\u0026rsquo;s own support notice covers what the wind-down means for devices already in service, which is the position anything you buy is already in. Verifying that the part which arrived is the part in the listing is on you.\nbcache Does Not Replace the DB/WAL Offload Worth stating explicitly, because it is the obvious place to try to save money and it does not work: you still need block.db on NVMe. Both layers, not one or the other.\nThe Optane is a write cache. That is all it is. It shortens the acknowledgement path for writes heading towards the disk, and it does nothing else.\nRocksDB does not only write. Ceph reads its metadata constantly — to find objects, to service peering, to answer scrubs — and a write cache offers a read exactly nothing once the data has been flushed out of it. Nor does the cache remove the metadata traffic; it defers and batches it, so every RocksDB update still arrives at the spindle eventually, competing for the same seeks. And compaction turns modest churn into far more device traffic than the writes that caused it.\nSo the two changes fix different problems and neither substitutes for the other:\nWhat it fixes What it does not block.db on NVMe metadata lives on flash — reads and writes both, permanently off the spindle nothing for a burst of guest writes Optane via bcache burst acknowledgement latency for data writes nothing for metadata reads; only defers metadata writes Skip the DB offload and keep the Optane, and RocksDB is back on the platter with its reads served at 80 IOPS. Skip the Optane and keep the DB offload, and the steady state is decent but bursts still stall at spindle speed.\nThe udev Rule, Line by Line bcache\u0026rsquo;s tunables live in sysfs, and sysfs resets them every time the device is registered — which is every boot. So they belong in a udev rule rather than a script somebody has to remember to run:\n# /etc/udev/rules.d/99-bcache.rules ACTION==\u0026#34;add|change\u0026#34;, SUBSYSTEM==\u0026#34;block\u0026#34;, KERNEL==\u0026#34;bcache*\u0026#34;, \\ ATTR{bcache/cache_mode}=\u0026#34;writeback\u0026#34;, \\ ATTR{bcache/sequential_cutoff}=\u0026#34;0\u0026#34;, \\ ATTR{bcache/congested_read_threshold_us}=\u0026#34;0\u0026#34;, \\ ATTR{bcache/writeback_rate}=\u0026#34;81920\u0026#34;, \\ ATTR{bcache/writeback_rate_minimum}=\u0026#34;20480\u0026#34;, \\ ATTR{bcache/writeback_percent}=\u0026#34;40\u0026#34; KERNEL==\u0026quot;bcache*\u0026quot; matches every bcache device on the node, so one rule covers all the pairs. ACTION==\u0026quot;add|change\u0026quot; means it reapplies whenever a device appears or is re-attached, not just at boot.\nWhat each setting is doing, and why:\ncache_mode=writeback — the entire point. In the default writethrough, a write is not acknowledged until it reaches the HDD, so the cache does nothing for write latency. In writeback, the Optane acknowledges and the spindle catches up later.\nsequential_cutoff=0 — by default bcache detects sequential I/O and routes it straight past the cache once it passes 4 MB, on the theory that the backing disk handles sequential fine. Zero disables that bypass so everything is cached. On a hyper-converged node that is the right call: what looks sequential to one bcache device stops being sequential at the platter once several OSDs interleave, and you want every write acknowledged at Optane speed regardless. It does carry one obligation, though — with no bypass, the cache device\u0026rsquo;s own sequential write bandwidth becomes the ceiling for the whole OSD, which is why the 16 GB modules are disqualified.\ncongested_read_threshold_us=0 — bcache watches its own cache-device latency and starts bypassing the cache when it judges it congested, defaulting to 2000 µs for reads. Optane does not get congested the way NAND does, so this is bcache second-guessing a device it has mismeasured. Zero switches the tracking off.\nwriteback_rate=81920 and writeback_rate_minimum=20480 — the background flush rate, in sectors per second, so roughly 40 MB/s target with a 10 MB/s floor. A PD controller moves the actual rate between them. The target is set near what a spindle can absorb sequentially, and the floor stops the controller throttling flush towards zero and letting dirty data pile up indefinitely. Both are starting points rather than constants — how to change them is below.\nwriteback_percent=40 — how much of the cache bcache will let sit dirty before pushing back hard, against a default of 10. Forty gives you a much deeper burst buffer. It also means up to 40% of that Optane holds the only copy of data in the system, which is the trade-off, and it is the reason the one-per-spindle layout matters. This is the value most worth revisiting once you have watched a real workload.\nChanging the Fill and Flush Values Later The 40% and the two rates above are the values running here, not universal constants. They are the first thing you will want to move once you have watched a real workload, so it is worth knowing that there are two places to change them, doing two different things.\nsysfs changes it now. The udev rule changes it next boot. You want both, and in that order.\nChange it live, on one device:\necho 30 \u0026gt; /sys/block/bcache0/bcache/writeback_percent Or across every pair on the node:\nfor d in /sys/block/bcache*/bcache; do echo 30 \u0026gt; \u0026#34;$d/writeback_percent\u0026#34; done The flush rate works the same way. Both values are in sectors per second, so these halve the target and the floor:\nfor d in /sys/block/bcache*/bcache; do echo 40960 \u0026gt; \u0026#34;$d/writeback_rate\u0026#34; echo 10240 \u0026gt; \u0026#34;$d/writeback_rate_minimum\u0026#34; done That takes effect immediately and survives exactly until the device is re-registered. The kernel documentation is explicit that these settings \u0026ldquo;do not persist across reboot\u0026rdquo; — which is the entire reason the udev rule exists.\nOnce you are happy with a value, edit the rule and reload it without rebooting:\nudevadm control --reload udevadm trigger --subsystem-match=block --action=change This is where ACTION==\u0026quot;add|change\u0026quot; earns its keep. The trigger fires a change event at devices that are already present, so the rule reapplies to a running node instead of waiting for the next boot. Had the rule matched add alone, that command would do nothing.\nThen read the values back, because a typo in a udev rule fails silently:\ngrep . /sys/block/bcache*/bcache/writeback_percent Knowing Which Way To Move Them Do not tune this from first principles — bcache exposes what you need under the same sysfs directory.\ndirty_data is the one to watch: how much data is currently sitting in the cache and nowhere else. The docs describe it as \u0026ldquo;continuously updated unlike the cache set\u0026rsquo;s version, but may be slightly off\u0026rdquo;, which is fine for this purpose. Sample it through a normal working day and a backup window.\ncache_hits, cache_misses and cache_hit_ratio tell you whether the cache is being used, with the caveat that \u0026ldquo;a partial hit is counted as a miss\u0026rdquo;. bypassed counts IO that went past the cache entirely — with sequential_cutoff=0 that should be close to flat, so a growing number means something is still routing around it.\nAll of those come as running totals plus versions that decay over the past day, hour and five minutes, which makes the short-window ones far more useful for spotting a problem than the lifetime figure.\nFrom there:\nSymptom Knob Direction Bursts stall — writes hit spindle latency mid-burst writeback_percent up, for a deeper buffer More data exposed on the cache than you are comfortable with writeback_percent down dirty_data pinned at the ceiling during ordinary work writeback_rate up — the buffer is not the problem, drainage is Background flush competing with guest reads on the spindle writeback_rate down dirty_data creeping up over days rather than hours writeback_rate_minimum up, so the controller cannot throttle to a crawl If dirty_data sits pinned at the ceiling no matter what you do, neither knob is the answer — the cluster is writing faster than the spindles can absorb, and the honest fixes are more spindles or fewer writes.\nOne file to leave alone: writeback_running. Setting it off stops writeback entirely, and the documentation says it is \u0026ldquo;only meant for benchmarking\u0026rdquo;. On a production OSD it means dirty data accumulates until the cache is full and never drains.\nWhat You Are Trading Away Most write-ups of this design stop at the good news. These are the parts that will actually hurt you, and every one of them is worth knowing before you build rather than after.\nThe two added devices fail very differently The shared DB/WAL NVMe dies block.db for osd.0–2 gone osd.0 lost osd.1 lost osd.2 lost Three OSDs, one event. A correlated failure across the node, which is exactly what replication does not protect you from. Choose the ratio for the rebuild you are willing to sit through, not the price per OSD. One paired Optane dies Optane for osd.0 gone Optane for osd.1 fine Optane for osd.2 fine osd.0 lost osd.1 serving osd.2 serving One OSD, and Ceph does this every day. The blast radius is a single disk's worth of data, which is the failure a replicated pool exists to absorb. This is the entire argument for buying several small devices instead of one large one. Same class of loss on both sides — each device holds state that exists nowhere else, so the OSD is destroyed rather than stopped and has to be re-created and backfilled. Only the count differs, and that is what the one-per-spindle pairing is buying. Both added devices hold state that exists nowhere else, so losing either destroys the OSDs that depend on it. The difference is only how many that is: the shared DB/WAL NVMe takes down every OSD behind it, while an Optane paired one-to-one takes down exactly one. That is what the extra devices buy. Losing either flash device destroys the OSD. Not stops it — destroys it. Both devices hold state that exists nowhere else: block.db holds the RocksDB that makes sense of block, and a writeback cache holds every write not yet flushed. The kernel documentation does not hedge on the second one: \u0026ldquo;In writeback mode you\u0026rsquo;ll lose data if something happens to your SSD.\u0026rdquo; Either way the OSD does not come back, it gets re-created and backfilled.\nSo these are the same class of risk, and it is worth being clear about that rather than treating the cache as the scary one. The Optane is not a more dangerous device than the NVMe. Two things separate them, and neither is severity per OSD.\nThe first is blast radius, and it is the entire justification for the pairing. One Optane per spindle means a cache failure costs one OSD — a single-disk loss, which is exactly the event a replicated pool exists to absorb, and which Ceph handles without anyone being paged. The DB/WAL device is shared, so losing it costs every OSD behind it at once, which is a correlated failure replication does not protect you from. Same failure, one device versus five.\nThe second is how the failure presents, and this one is an operational trap. When a cache dies in writeback the backing device stops and returns I/O errors, which is fine — Ceph marks the OSD down and gets on with it. The bad case is a reboot where the backing device comes up without its cache attached. It then looks like a mountable filesystem that is simply missing every dirty write, which is corrupt rather than merely stale. A missing block.db fails loudly and the OSD refuses to start; a missing cache can fail quietly and let you mount the wreckage. Treat a bcache backing device as unmountable without its cache, and never let anything helpfully mount it for you.\nbcache does not guarantee power-safe writeback by itself. This is where the HDD write cache stops being an optimisation and becomes a requirement — turn it off, and run a kernel recent enough to have the FUA handling, so synchronous writes are honoured through the stack rather than acknowledged early somewhere in the middle. A cache device without power-loss protection compounds the problem, because it can lose data it has already reported safe.\nWhich makes the DB/WAL ratio the number that matters. Ceph allows 4–5 HDD OSDs per SATA SSD and no more than 15 per NVMe, and warns about \u0026ldquo;balancing the risk of reducing costs by placing too many responsibilities into too few failure domains.\u0026rdquo; Unlike the cache tier there is no pairing available to contain this one — sharing is the point of the device.\nOn 20 TB drives that is the whole conversation, because the rebuild is huge. One failed OSD is 20 TB to backfill; at the 150–250 MB/s a spindle sustains, throttled so recovery does not starve the guests, you are looking at well over a day of degraded operation for a single disk. Lose a shared DB device carrying five of them and there is 100 TB to move.\nSo the ratio is not really a cost decision. It is a decision about how long you are prepared to run degraded, and how much rebuild traffic the cluster can carry while still serving VMs.\nThe design depends on a device nobody makes any more. This is the one with no technical answer. Optane is the right part for the cache position and nothing current replaces it — NAND cannot take the write volume, and the CXL-based memory Intel pivoted towards is not a drop-in for a block cache. So the mitigation is commercial rather than clever, it is covered above, and the honest summary is that this architecture has an end date somewhere out in the future.\nNone of this makes an HDD cluster an all-flash cluster. Sustained random write traffic that outruns the flush rate will fill the cache, and once it is full you are writing at spindle speed with extra layers in the path. This design absorbs bursts and removes metadata overhead. It does not manufacture IOPS.\nWhat It Adds Up To Three tiers, each doing the one thing it is best at — and all three are required:\nOptane absorbs random write bursts, one device per spindle. It has to be Optane, because every write crosses it and NAND endurance is wrong for that position A fast enterprise NVMe holds RocksDB and the WAL, shared across a handful of OSDs at a ratio you chose deliberately rather than accepted HDDs provide bulk capacity, freed of both metadata and bursts The point is not to pretend spinning disks are flash. It is to stop sending them the work they are worst at, so the capacity you actually paid for is usable.\nThat is the whole trick, and there is nowt clever about it. Put each job on the device that is good at it, and stop paying twice for the ones that are not.\nFor a hyper-converged Proxmox cluster that needs multi-terabyte capacity without all-flash pricing, that is a defensible design. Built properly it feels far faster than raw HDDs, it fails in ways Ceph is built to handle, and the money goes where it changes the outcome.\nTwo related things worth reading alongside this: the sector-size work in 4Kn, 512e and 512n matters a great deal for what the spindles do with the writes that reach them, and if you are building this on direct-attached nodes, the switchless mesh covers the network side of a small Ceph cluster.\nReferences Ceph — Hardware Recommendations — the IOPS-per-TB warning on large HDDs, \u0026ldquo;HDD OSDs may see a significant write latency improvement by offloading WAL+DB onto an SSD\u0026rdquo;, the 4–5 HDD per SATA SSD and ≤15 per NVMe ratios, power-loss protection on enterprise SSDs, disabling the HDD write cache, and not needing an RoC HBA Ceph — BlueStore Configuration Reference — block.db at 1–2% of block for RBD and at least 4% for RGW, the 2.5% general recommendation, spillover back onto the primary device, the RocksDB level sizes behind the 3/30/300 GB steps, and the WAL being implicitly colocated with the DB Linux kernel — bcache admin guide — the cache modes, sequential_cutoff and its 4 MB default, the 2000 µs read congestion default, writeback_rate in sectors per second, the writeback_percent PD controller, the dirty_data / cache_hit_ratio / bypassed counters and their decaying day, hour and five-minute versions, the warning that writeback_running is \u0026ldquo;only meant for benchmarking\u0026rdquo;, the statement that these settings \u0026ldquo;do not persist across reboot\u0026rdquo;, and \u0026ldquo;in writeback mode you\u0026rsquo;ll lose data if something happens to your SSD\u0026rdquo; Intel — Optane Memory M10 32 GB specifications — 365 TBW, 240,000 random read and 65,000 random write IOPS at 4K over an 8 GB span, 1200/290 MB/s sequential, PCIe 3.0 x2 Intel — Optane Memory M10 16 GB specifications — the 150 MB/s sequential write and 35,000 random write IOPS that put this part below a spinning disk on sequential throughput PC Perspective — Optane SSD DC P4800X performance — the 375 GB part\u0026rsquo;s 550K random 4K read IOPS, 2400/2000 MB/s sequential, sub-10 µs typical latency, and 12.3 PBW at 30 DWPD ServeTheHome — Optane DC P4801X 100GB M.2 review — the small-capacity datacentre M.2 parts, for the capacity ladder StorageReview — Intel Optane SSD P5800X — the 100 DWPD rating, the P4800X\u0026rsquo;s 30 DWPD, and the comparison against NAND enterprise drives topping out around 10 DWPD for write-intensive parts and 3 or less for mixed-use Bcache — ArchWiki — the practical failure modes, including a backing device coming up without its cache after a reboot, and the power-safety requirements around the HDD write cache and FUA Intel — Optane business update — the wind-down, and what it means for warranty and support on devices already in service, which is the position anything you buy on the recycler market is already in ","permalink":"https://blogs.damiendye.uk/en/proxmox/hdd-backed-ceph-bcache-optane/","summary":"Flash pricing has made all-NVMe hard to justify, and enterprise HDDs are worth another look. Getting acceptable latency out of them on hyper-converged Proxmox Ceph means moving RocksDB off the spindle, pairing each disk with its own Optane through bcache — Optane specifically, because NAND endurance is wrong for that job — and being honest about the failure modes it buys you.","title":"Making HDD-Backed Proxmox Ceph Clusters Fast — NVMe Metadata, One Optane Per Spindle, and What It Costs You"},{"content":"Why This Has Been Expensive Until Now VDI — Virtual Desktop Infrastructure — delivers a full desktop from a data centre or cloud instead of from the device in front of you. A user signs in from a laptop, a thin client or their own machine, and gets a familiar Windows or Linux desktop whose apps, files and processing all happen centrally.\nThe appeal for an IT team is that patching, access control and data protection all happen in one place, and people can reach the same desktop from anywhere.\nThe obstacle has never been the hypervisor. It has been the GPU.\nSharing one physical GPU between several desktops has been NVIDIA\u0026rsquo;s territory, and NVIDIA charges for the vGPU software that does the slicing. That licence is why VDI with hardware-accelerated graphics has mostly been the preserve of outfits with corporate budgets. Paying a yearly fee to switch on something the silicon can already do is a hard thing to be keen about.\nIntel changed the arithmetic with the Arc Pro line. These cards support hardware splitting through standard SR-IOV, which is a PCIe feature rather than a product. As such, there is no licence server, no subscription, and no separate vGPU software to buy.\nSo I got hold of an Intel Arc Pro B50 to see how cleanly it goes together with Proxmox. The short answer: the mechanism works exactly as advertised, and then firmware got in the way. Both halves are below.\nHow one card becomes several: physical function on the host, virtual functions to guests Intel Arc Pro Bxx 03:00.0 — physical function host driver: xe 03:00.1 vfio-pci 03:00.2 vfio-pci 03:00.3 created where supported 03:00.4 created where supported Windows desktop 1 hardware accelerated Windows desktop 2 hardware accelerated Windows desktop 3 hardware accelerated Windows desktop 4 hardware accelerated No vGPU licence is involved at any point. SR-IOV is a PCIe capability the card either advertises or does not — this one advertises it at\u0026#160;[320]. How many functions it will create is set in firmware, because each one is given a fixed slice of the card's memory: a bigger slice means fewer functions. A 16 GB B50 gives two at 8 GB each; the 24 and 32 GB cards report seven. The physical function stays with the host\u0026rsquo;s xe driver. Each virtual function is a PCIe device in its own right, bound to vfio-pci and handed to a guest — the same passthrough machinery as a whole card, just several times over. The dashed pair depends on the card: how many functions each one will create is further down. The Elephant: You Need Windows First Before any of this works, the card wants its firmware updating. Intel ships that update inside the Windows driver installer.\nWhich is awkward, because the reason you bought the card is to run it under Proxmox.\nThere is a way round it that needs nothing but Proxmox: build a Windows VM, pass the whole card through to it, let Windows update the firmware, then hand the card back to the host. It is a bootstrap loop, and it is worth knowing about before you plan the build rather than after.\nBreaking the Windows-first bootstrap loop with a temporary VM The catch: SR-IOV needs current firmware, and the firmware ships inside a Windows driver installer. 1 Card in the host lspci → 03:00.0 2 Whole card → Windows VM hostpci0, Q35 + OVMF 3 Install Intel driver firmware updates with it 4 Reboot host card reinitialises card handed back to Proxmox 5 SR-IOV present lspci -v → [320] 6 tmpfiles.d at boot numvfs, unbind, bind 7 VFs to guests as many as firmware allows Steps 2 and 3 exist only to update firmware. The Windows VM is temporary — once the card is back on the host it plays no further part, and nothing about the finished setup depends on Windows running on the hypervisor. The card cannot do SR-IOV until its firmware is current, and the firmware update arrives as a Windows driver. One temporary Windows VM with the whole card attached breaks the loop. Finding the Card On the Proxmox host, lspci to locate it:\nlspci Here it is at 03:00.0, reported as Battlemage G21, the silicon behind the Arc Pro B50. Note the address down. You will need it several times, and there is a separate audio function at 04:00.0 that comes with it.\nPassing the Whole Card to a Windows VM Add the card to a Windows VM as a raw PCI device. In the VM\u0026rsquo;s hardware, that is a PCI Device entry: 0000:03:00 with pcie=1:\nWorth noticing what else that VM is, because none of it is accidental: Q35 machine type, OVMF firmware, VirtIO SCSI single, and a TPM for Windows 11. Q35 in particular is not optional for this. A passed-through GPU on i440fx appears as a legacy PCI device, which is the wrong shape for a modern graphics driver. That is covered in Always Use Q35, Not i440fx.\nBoot the VM and check Windows sees the card:\nIt appears as a Microsoft Basic Display Adapter because no driver is installed yet. That is the expected state, and it is enough. Windows has found the hardware.\nUpdating the Firmware Get the current driver from Intel\u0026rsquo;s Arc Pro B50 download page. At the time of writing that was version 32.0.101.8306 (Q4.25), for Windows 11 and Windows 10 22H2:\nInstall it, and let the firmware update run as part of the process rather than cancelling out early. That firmware step is the entire reason for this detour.\nWhen it finishes, give the card back to the host and reboot, so it reinitialises fully under Proxmox.\nConfirming SR-IOV Is There Now ask the card what it can do:\nlspci -v The line that matters:\nCapabilities: [320] Single Root I/O Virtualization (SR-IOV) That is the whole proposition in one line of lspci output, with no licence attached to it.\nThree other things in that output are worth reading while you are there:\nKernel driver in use: xe — the card is on Intel\u0026rsquo;s newer xe driver rather than i915, which is what will be creating the virtual functions. [420] Physical Resizable BAR and [220] Virtual Resizable BAR — the card supports resizable BARs, and so do its virtual functions. Worth knowing what that costs you in address space if you are passing several of them through: see PCIe Resizable BAR and Modern GPUs. IOMMU group 13 — the card is in a group of its own, which is what you want for clean passthrough. Why that matters is in the IOMMU tax article. Creating the Virtual Functions at Boot Virtual functions are not persistent. Asking for them is a write to sysfs, so it has to happen on every boot.\ntmpfiles.d is a tidy way to do that declaratively, rather than bolting a script onto a unit file:\nThe file does three jobs in order.\nCreate the virtual functions, by writing the count to the physical function\u0026rsquo;s sriov_numvfs:\nw /sys/devices/pci0000:00/0000:00:01.1/0000:01:00.0/0000:02:01.0/0000:03:00.0/sriov_numvfs - - - - 4 Unbind each new function from xe, because the host driver claims them as they appear and a guest cannot have a device the host is holding:\nw /sys/bus/pci/drivers/xe/unbind - - - - 0000:03:00.1 w /sys/bus/pci/drivers/xe/unbind - - - - 0000:03:00.2 w /sys/bus/pci/drivers/xe/unbind - - - - 0000:03:00.3 w /sys/bus/pci/drivers/xe/unbind - - - - 0000:03:00.4 Bind them to vfio-pci, which is what makes them available to pass through:\nw /sys/bus/pci/drivers/vfio-pci/bind - - - - 0000:03:00.1 w /sys/bus/pci/drivers/vfio-pci/bind - - - - 0000:03:00.2 w /sys/bus/pci/drivers/vfio-pci/bind - - - - 0000:03:00.3 w /sys/bus/pci/drivers/vfio-pci/bind - - - - 0000:03:00.4 Note the addresses: the physical function is 03:00.0 and the virtual functions come up as .1 through .4.\nOne detail about w that explains the ordering, and that will bite you if you get it wrong. systemd documents it as: \u0026ldquo;Write the argument parameter to a file, if the file exists.\u0026rdquo; sriov_numvfs only exists once a driver has bound to the physical function, and the virtual function paths only exist once that write has happened. So the sequence in the file is not stylistic. Each line depends on the one before it having taken effect.\nWhy Two, and Not Four Driver 32.0.101.8306 — the one installed above — carries graphics firmware BMG__21,1162, and that is the release where Intel first officially enabled SR-IOV on Arc Pro. Intel\u0026rsquo;s stated default for the B50 in that release is two virtual functions, each with an 8 GB VF Local Memory BAR.\nWhich makes the number arithmetic rather than policy. The B50 has 16 GB. At 8 GB per virtual function, two is all that fits.\nThere is a wrinkle worth knowing if you go looking for a workaround. Before official support existed, some people ran older firmware that exposed 12 virtual functions on a B50, and going back to driver 32.0.101.6979 restores that count. Those 12 shared the same 16 GB, so each got a fraction of the memory mine get. Intel\u0026rsquo;s position is that two was chosen deliberately to give each function enough compute, capacity and bandwidth to behave predictably.\nSo the cap can be moved, but not by you. The maximum VF count and the VF Local Memory BAR size live in the IFWI, there is no public tool to change either, and the supported answer on a current stack is two.\nHow Many Desktops Each Card Gives You The B50 is the small card in the family, and its two functions are the family\u0026rsquo;s low water mark. If the seat count is what you care about, buy further up the range.\nThe whole Battlemage Arc Pro line does SR-IOV. What differs is how many functions the firmware will carve out, and that follows the memory:\nCard Memory VFs on the current supported stack Seen elsewhere Arc Pro B50 16 GB 2, at 8 GB VF BAR each — Intel\u0026rsquo;s documented default 12 on pre-official firmware, via driver 32.0.101.6979 Arc Pro B60 24 GB 7 reported 24 on an early ASRock firmware, cut to 7 by a later one Arc Pro B60 Dual 2 × 24 GB 7 per GPU — two GPUs, so 14 from one slot as above; the two halves are independent Arc Pro B65 32 GB no published count found — Arc Pro B70 32 GB 7 reported, on firmware 8517 4 on earlier firmware Only the B50 row is Intel-documented. The B60 and B70 numbers are what people report from lspci, and they have moved more than once. The B60 in particular went from 24 down to 7 in a firmware update, which is the same kind of narrowing the B50 saw. Nobody appears to have published a VF count for the B65 at all, so treat that row as unknown rather than as zero.\nTwo things follow from that table, and both matter more than any single number in it.\nThe VF count is a memory division, not a die feature. Intel\u0026rsquo;s rule is that a larger VF Local Memory BAR means fewer functions. That is why the 16 GB card gives two and the 32 GB cards give seven: nothing about the GPU\u0026rsquo;s shaders decides it.\nCheck the card you are about to buy, not the family. SR-IOV presence has varied between board vendors on the same chip — Sparkle\u0026rsquo;s B60 Blower initially shipped without the capability visible at all and only gained it after an igsc firmware update. Ask for lspci -v output from the exact model, or budget for a firmware update before you count on any of this.\nThe Dual B60 Is Two Cards Wearing One Bracket Maxsun\u0026rsquo;s Arc Pro B60 Dual 48G Turbo is the interesting one for seat count, and the thing to understand is that the 48 GB is not a pool.\nIt is two B60 GPUs — two BMG-G21 dies — on one board, with 24 GB of GDDR6 wired to each, and no PCIe bridge chip between them. Both dies hang straight off the x16 gold fingers at PCIe 5.0 x8 each.\nWhich means the host has to split the slot for you. The card needs the primary x16 slot bifurcated to x8/x8, and most consumer boards do not enable that by default. It is a firmware setting you go looking for, in the same category as the IOMMU and ACS settings any of this needs.\nGet that right and the operating system sees two separate GPUs, each with its own physical function and its own SR-IOV capability. So you get two lots of virtual functions from one slot — 14 seats if each die behaves like a single B60 — and the tmpfiles.d file above doubles up, one sriov_numvfs write per die.\nGet it wrong and you see one GPU and half the card is invisible.\nWorth being plain about what the 48 GB is not: a guest attached to a virtual function on the first die cannot reach the second die\u0026rsquo;s memory. This is two 24 GB cards in one physical space, which is exactly what you want for VDI seats and exactly what you do not want for one large model.\nSo: if two seats are enough, the B50 is a 70 W card that will do it. If you want seven, plan around a B70. If you want fourteen and have a board that will bifurcate, the dual B60 gets you there in one slot.\nCan You Run AI On a Virtual Function? Short answer: treat it as unsupported. Longer answer, because the reason matters and it is not the one you would guess.\nThese are marketed as AI cards and they are not pretending. The B50\u0026rsquo;s 128 XMX engines are rated at 170 peak TOPS, the B70 at 367, and Intel\u0026rsquo;s software story is real — vLLM serves models from 8B up to 120B on Arc Pro B-series, and IPEX-LLM and llama.cpp\u0026rsquo;s SYCL backend both run on them.\nBut look at how every one of those results is produced. vLLM\u0026rsquo;s own Arc Pro numbers come from a Docker container on bare metal, on systems with four and eight whole B60 cards doing tensor parallelism. Intel\u0026rsquo;s post does not mention SR-IOV or virtual functions once.\nThat pattern holds everywhere I looked. Intel scopes the SR-IOV use cases to virtualised remote desktop, guest OS graphics acceleration, and media encode and decode. Compute is not on that list, and I could not find a single published case of anyone running LLM inference inside a VM attached to a virtual function.\nWhat people actually do is telling: they run the model in Docker on the host, and hand virtual functions to VMs for desktops. One person doing both at once reports simply that \u0026ldquo;VRAM gets pretty tight.\u0026rdquo;\nWhich is the real problem, and it is arithmetic rather than driver support.\nA virtual function gets a fixed slice of local memory — 8 GB on the B50, set in firmware. That slice is the hard ceiling for weights plus KV cache in that guest, and it does not grow because the card has more. An 8B model at FP16 is around 16 GB of weights before you add any context at all, so it does not fit in a B50 virtual function on any driver. Quantise to Q4 and an 8B fits in about 4 GB, leaving a few GB for context — which works, but is a long way from what the card can do undivided.\nSo the two workloads compete for the same memory, and the split is decided in firmware before either of them starts.\nIf AI is the job, do not divide the card. Pass the whole thing through to one VM — the same hostpci0 passthrough used for the firmware update earlier in this post — or run the container on the host and skip virtualisation for that workload. Both give the model all 16 GB and the full XMX array.\nIf VDI is the job, virtual functions are right, and expect desktop graphics rather than an inference server behind each one. Hardware-accelerated desktops, video playback and encode work. That is what the mechanism is documented for.\nWorth saying plainly: absence of published evidence is not proof it fails. The xe driver exposes compute through Level Zero and OpenCL, and it is entirely possible a VF-backed guest brings those up fine. But nothing from Intel says it is validated, no one appears to have shown it working, and the memory ceiling limits the payoff even if it does. That is not something to build a plan on.\nWhat It Is Still Worth Two virtual functions is two hardware-accelerated Windows desktops from one card, with no vGPU licence, no subscription and no licence server. On a hypervisor that costs nothing to run. That is enough to prove the approach works, which is the honest job for a B50. It is the bottom of the range.\nFor an actual VDI deployment I would be specifying the dual B60.\nFourteen functions from one slot — if each die behaves like a single B60 — puts it in the same seat-count territory as the NVIDIA cards sold for this workload, at a far lower price, and with nothing to license per user. That last part is the one that compounds. NVIDIA\u0026rsquo;s vApps, vPC and RTX vWS are all licensed per concurrent user, either as an annual subscription or as a perpetual licence that has to be bought alongside a five-year support and maintenance subscription. Every seat is a line item, and it comes round again. On the Intel side there is no equivalent line. You buy the card.\nAnd the mechanism is the part that matters long term. SR-IOV on the GPU is a PCIe capability, not a product tier, so the tmpfiles.d file just grows to match whatever the card allows.\nThe shape is one sriov_numvfs write, then an unbind and a bind for each function — so two functions is five lines, and the nine above are four functions asked for on a card that delivers two. Seven functions is fifteen lines. A dual B60 is thirty, because each die is its own physical function and gets its own sriov_numvfs write.\nNowt else changes as you scale it. No licence server appears at any point in that file.\nReferences Intel support — why the latest Arc Pro B50 firmware shows 2 SR-IOV VFs — the authoritative statement: SR-IOV officially enabled from graphics firmware BMG__21,1162 in driver 32.0.101.8306, two VFs at an 8 GB VF Local Memory BAR each on the B50, and the maximum VF count and BAR size set at IFWI level with no public tool to change them Intel Community — \u0026ldquo;Why did the latest Intel Arc Pro B50 firmware nerf SR-IOV VFs from 12 to 2?\u0026rdquo; — the 12-VF pre-official firmware, the 32.0.101.6979 rollback that restores it, and Intel\u0026rsquo;s reasoning for the lower default Level1Techs — B60 SR-IOV support in the Arc Pro drivers — field lspci reports for the B60, the igsc firmware update that exposed the capability, and where the 24-then-7 figures come from Level1Techs — B50, B60 or B70 for SR-IOV — the reported VF counts per card and per firmware, source for the B70 rows ASRock — Intel Arc Pro B65 Creator 32GB — the B65\u0026rsquo;s specifications: 32 GB GDDR6, 20 compute units, 160 XMX engines, 256-bit, PCIe 5.0 MAXSUN — Arc Pro B60 Dual 48G Turbo — the vendor\u0026rsquo;s own statement that the card \u0026ldquo;uses PCIe 5.0 x8 + x8 interfaces and runs efficiently on consumer platforms that support PCIe x16 lane bifurcation\u0026rdquo; vLLM — Fast and affordable LLM serving on Intel Arc Pro B-Series — the AI story on these cards, and the fact that it is a Docker-on-bare-metal story across four and eight whole B60s, with no mention of SR-IOV or virtual functions Linux kernel — Intel Xe driver — the driver in use on the card, per lspci -v tmpfiles.d(5) — the w line type, and its \u0026ldquo;if the file exists\u0026rdquo; condition that dictates the ordering above NVIDIA Virtual GPU Software Packaging, Pricing and Licensing Guide — the licensed alternative this design avoids: vApps, vPC and RTX vWS all sold per concurrent user, as an annual subscription or as a perpetual licence bundled with five years of support and maintenance Proxmox VE — PCI(e) Passthrough — host-side passthrough requirements ","permalink":"https://blogs.damiendye.uk/en/proxmox/licence-free-vdi-intel-arc-pro-sriov/","summary":"Intel\u0026rsquo;s Arc Pro cards do SR-IOV natively, with no vGPU licence to buy — which makes a licence-free Windows VDI on Proxmox genuinely possible. The whole build on a B50, the Windows-first bootstrap problem, why the firmware caps it at two virtual functions, and how many each card in the B-series gives you.","title":"Licence-Free Windows VDI on Proxmox with an Intel Arc Pro B50 — and the Firmware Limit That Stopped It"},{"content":"The Switch You Do Not Buy A three-node Proxmox cluster with Ceph wants a fast network between the nodes. The usual answer is a 100 Gbit switch, and the usual objection is what one costs.\nThere is another answer for three nodes: cable them straight to each other in a triangle and route across it. No switch in the storage path at all.\nThe cheapest component in any design is the one you do not buy, and it is the only one that never fails.\nThat buys four things.\nThe switch cost disappears. You need network cards and three cables, not a 100 or 200 Gbit switch with the port count to match.\nThe data path has no single point of failure. A switch that fails or reboots takes the whole cluster\u0026rsquo;s east-west traffic with it. Nodes wired directly to each other do not care what happens to a switch elsewhere in the building.\nThe heavy traffic is on its own wires. Live migration and Ceph replication stay on the mesh rather than competing with office and management traffic.\nGrowing it is a wiring job, not a port-budget job. There is no central box acting as a bandwidth or port-count ceiling, so the fabric grows as long as each server has a free PCIe slot. OpenFabric routes across new links on its own.\nAnd one honest limit, because it matters more than the four benefits. A full mesh needs a cable between every pair of nodes. Three nodes is three cables. Four is six. Five is ten. Wiring grows faster than the node count, and this design does not survive past a small cluster without moving to something switched, like leaf-spine.\nThree nodes is exactly where a full mesh makes sense. Beyond that, with two mesh ports per node, what you can still build is a ring — and that needs one thing this build otherwise avoids, covered below.\nWhat Gets Built Three Proxmox VE hosts, each with two dedicated 100 Gbit interfaces, cabled into a triangle so every node has two direct neighbours.\nProxmox SDN runs the routing. OpenFabric is the protocol: it works out the best path across the mesh, and when a cable is pulled or a link drops it reroutes through the remaining path. Nobody logs in.\nThe fabric lives in 10.10.10.0/24, reserved for the mesh and nothing else — not management, not guests, not storage addressing, not anything external.\nNode Mesh address Mesh interfaces mesh1 10.10.10.1/32 nic1, nic2 mesh2 10.10.10.2/32 nic1, nic2 mesh3 10.10.10.3/32 nic1, nic2 Each node gets a /32, not a slice of the subnet. That is the point of a routed mesh rather than a bridged one. The address identifies the node, OpenFabric advertises it, and the two physical links are just paths to reach it. Proxmox creates a dummy loopback interface to hold it.\nThe three-node triangle, and which NIC each cable lands on 2.5 Gbit switch management + client mesh1 10.10.10.1/32 mesh2 10.10.10.2/32 mesh3 10.10.10.3/32 DAC-01 nic1 ↔ nic1 DAC-03 nic2 ↔ nic2 DAC-02 mesh2 nic2 ↔ mesh3 nic1 nic0 nic0 nic0 Each node reaches the other two directly, so every mesh hop is one hop. The dashed nic0 links are the separate 2.5 Gbit management path — never part of the fabric, and the reason you can still log in if the mesh is broken. Three cables, six ports, and two paths to every node. Pull any single cable and every node is still reachable — the remaining two links form a chain that OpenFabric will route around. Wiring It Active fibre DAC cables, in a triangle, with jumbo frames enabled on both mesh interfaces on every node.\nCable From To DAC-01 mesh1 nic1 mesh2 nic1 DAC-02 mesh2 nic2 mesh3 nic1 DAC-03 mesh3 nic2 mesh1 nic2 Active fibre rather than copper DAC, for two reasons that are about the rack rather than the network:\nCable management. Active fibre DACs are thinner and far more flexible than copper, so they route cleanly and do not pile up behind the servers. Airflow. Less cable bulk behind the chassis means less disruption to front-to-back airflow, which matters when several high-speed links land in the same few rack units. Two Networks, Not One The mesh is not the only network, and it must not become one.\nEach node has nic0 on a 2.5 Gbit switch, presented to Proxmox as the vmbr0 bridge. That carries the web interface, admin access and client traffic. It is also the path you use while building the fabric, which is why it has to be independent of it.\nnic0 is deliberately left out of the fabric. Only nic1 and nic2 are selected when the fabric nodes are created.\nThe mesh carries three things:\nCeph. Node-to-node replication, recovery and backfill for the hyper-converged storage. Client virtual networking. VXLAN-backed VNets stretch guest networks across all three nodes, using the routed mesh as the underlay. Corosync, as a second path. Cluster membership traffic runs over the management network and the mesh, so quorum does not depend on either one surviving alone. What rides the management network, what rides the mesh, and the one thing on both 2.5 Gbit management network nic0 → vmbr0 → switch Proxmox web interface Administrative access Client-facing traffic 100 Gbit routed mesh nic1 + nic2 → direct DAC, no switch Ceph replication, recovery, backfill VXLAN client virtual networks Live migration Corosync — both paths Cluster membership does not depend on either network surviving alone. The separation is the design. Management stays reachable whatever state the 100 Gbit interfaces are in, which is what makes it safe to build — and to roll back — the fabric from the web interface. Corosync is the only thing on both. Everything else has exactly one home — and management access is the one that has to keep working while you are changing the other. The rule worth writing on the change ticket: management and client access stay available through the 2.5 Gbit switch at all times, whatever state the 100 Gbit interfaces are in.\nEverything Through the Web Interface This build is done entirely in the Proxmox web interface. That is a deliberate choice, not a limit of the tooling.\nNot used, on purpose:\nEditing /etc/network/interfaces by hand. Editing FRR configuration files by hand. vtysh as a build method. Proxmox generates the underlying network and routing configuration from the SDN objects you define. If something genuinely cannot be set in the interface, that is worth calling out as a prerequisite rather than quietly fixing on the command line. The next person to open the web interface will not know you did.\nCommand-line output appears below only as evidence, never as a build step.\nThe Build 1. Open the Proxmox web interface and go to Datacentre → SDN → Fabrics.\n2. Add the fabric. Name it, give it the mesh prefix, and set the timers.\nHello and CSNP intervals of 1 update state as fast as possible when something changes, which is what you want on a fabric this small. The cost is more control-plane chatter. Irrelevant on three nodes with two links each, worth a second thought if the fabric ever grows.\n3. Add each node with Add node. Give it an address from the mesh range and tick the interfaces that take part.\nNote what is not ticked: nic0 stays out, and vmbr0 keeps the management address. Use Create another for the first two nodes and Create on the last.\n4. Check the result before applying it.\nThree nodes, three addresses, nic1, nic2 on each, all marked new — nothing has been written yet.\n5. Apply the SDN configuration.\nThere is a Dry-Run next to Apply if you would rather see what it intends to do first.\nKnowing It Worked The status view should show both zone and fabric entries ok on all three nodes, and no pending changes left to apply.\nThen check the fabric from a node\u0026rsquo;s own point of view. Routes first — each node should have a /32 to each of the others, and the Via column tells you which neighbour it is using.\nNeighbours next. Two, both Up, on a three-node triangle.\nThen the interfaces, which is where the shape of the thing shows up: dummy_Mesh as the loopback holding the router address, and nic1 and nic2 as Point-To-Point rather than broadcast segments.\nFinally, prove it end to end.\nNo loss, and averages of 0.134 ms and 0.141 ms. Both neighbours are one direct hop away, which is what a triangle gives you.\nThe check worth doing that no screenshot can show: pull one cable and confirm everything is still reachable. That is the whole reason for choosing a routed mesh over a pair of point-to-point links. As such, it is the only test that matters.\nBeyond Three Nodes: The Ring A triangle is a full mesh. Every node has a direct cable to every other node, every hop is one hop, and no node ever carries traffic that is not its own.\nThat property is what two mesh ports per node buys you at three nodes, and it is exactly what you lose at four. Not some of it. All of it. A full mesh of four nodes needs three ports each. With two, the most you can wire is a ring.\nA ring changes the traffic model. Adjacent nodes still have a direct cable, but nodes on opposite sides of the ring do not — their traffic has to cross an intermediate node. And that node has to be willing to forward packets between its two mesh interfaces, which Linux will not do by default. The kernel documents ip_forward as \u0026ldquo;Forward Packets between interfaces\u0026rdquo; with a \u0026ldquo;Default: 0 (disabled)\u0026rdquo;.\nA four-node ring: opposite nodes have no cable, so one node forwards for them mesh1 10.10.10.1 mesh2 forwards mesh3 10.10.10.3 mesh4 10.10.10.4 mesh1 → mesh3 no cable between them so it crosses mesh2 transit node forwards between nic1 and nic2 Full mesh at 4 nodes 6 cables, 3 ports per node no transit, no forwarding Ring: 4 cables, 2 ports — which is why you are here Two of the six node pairs have no direct cable. Their traffic is carried by a neighbour, which means that neighbour's links carry other nodes' Ceph replication as well as their own — the cost the triangle does not have. mesh1 to mesh3 has no cable. Its traffic crosses mesh2 or mesh4, and that node only forwards it because forwarding is enabled on the two interfaces it arrives on. Turn it on for the mesh interfaces, and only those:\n# /etc/sysctl.d/99-mesh-forwarding.conf net.ipv4.conf.nic1.forwarding = 1 net.ipv4.conf.nic2.forwarding = 1 sysctl --system Read them back rather than assuming:\nsysctl net.ipv4.conf.nic1.forwarding net.ipv4.conf.nic2.forwarding The per-interface setting is the right scope here, and it works on its own. The global net.ipv4.ip_forward is not a prerequisite. The kernel\u0026rsquo;s forwarding decision reads the receiving interface\u0026rsquo;s own value:\n#define IN_DEV_FORWARD(in_dev) IN_DEV_CONF_GET((in_dev), FORWARDING) IN_DEV_CONF_GET returns that device\u0026rsquo;s setting, not an AND with the global one. So nic1 and nic2 forward transit traffic for the fabric while nic0 and vmbr0 stay exactly what they should be: host interfaces that do not route. Turning on the global switch would make every interface on the box a router, including the one facing your office network. That is nowt this design needs.\nOne thing to know about the global IPv4 switch even though you are not setting it: net.ipv4.ip_forward is a bulk setter, which is why the kernel documentation warns that changing it \u0026ldquo;resets all configuration parameters to their default state\u0026rdquo;. If anything else on the host ever writes it, it overwrites these per-interface values. Worth knowing before you spend an afternoon on why transit stopped working.\nIPv6 is the exception, and it is the only place the global switch belongs. The kernel documentation says so directly under conf/all/forwarding:\nEnable global IPv6 forwarding between all interfaces. IPv4 and IPv6 work differently here; the force_forwarding flag must be used to control which interfaces may forward packets.\nSo there is no IPv6 equivalent of the tidy per-interface approach above. If the fabric carries IPv6 you enable forwarding globally and then scope it with force_forwarding, documented as \u0026ldquo;Enable forwarding on this interface only — regardless of the setting on conf/all/forwarding\u0026rdquo;. Note the same clobbering behaviour applies in reverse: setting conf.all.forwarding to 0 resets force_forwarding on every interface.\nThe build above leaves the fabric\u0026rsquo;s IPv6 prefix empty, so none of that applies here — it only matters if you add one.\nNote what forwarding does not change: OpenFabric was already advertising each node\u0026rsquo;s /32 and already working out the path across the ring. Forwarding is the missing permission, not the missing brains. The routing table was right all along. The kernel was simply declining to act as a router.\nWhat the ring costs, compared with the triangle:\nTransit traffic. On a four-node ring, the two diagonal pairs cross an intermediate node, so their traffic consumes that node\u0026rsquo;s link bandwidth as well as its own. Ceph notices this first, because replication is all-to-all rather than neighbour-to-neighbour. An extra hop of latency on those paths, on top of the switchless design\u0026rsquo;s usual sub-millisecond figures. Less headroom for failure. One broken link turns a ring into a chain: still fully connected, but with longer paths and more transit. A second break partitions the cluster. A triangle tolerates one break with no transit at all. And it gets worse in a way that is easier to draw than to describe. Add a fifth node and half of every node pair in the cluster depends on somebody else forwarding:\nA five-node ring: half the node pairs now depend on somebody else forwarding mesh1 mesh2 mesh3 mesh4 mesh5 cable — direct, one hop no cable — needs a neighbour to forward Five nodes, two ports each 10 node pairs in total 5 have a direct cable 5 do not, and transit a neighbour Every node now forwards traffic that is not its own. A full mesh instead? 10 cables, and 4 ports per node — two more NICs in every server, which is where the design stops. At three nodes nothing transits. At four, two pairs do. At five, half of them do — and a single break lengthens every path. Ten node pairs, five cables. The dashed lines are pairs with no cable between them — every one of those is Ceph traffic riding through a third node\u0026rsquo;s links. The full mesh that would avoid it wants four ports per server. That is the real ceiling on this design, and it is not the routing protocol. OpenFabric copes fine. It is that ports per node is fixed, so past three nodes every new box converts more of your traffic into somebody else\u0026rsquo;s transit.\nAnd the honest note, because it matters for a build that has been web-interface-only up to here: a sysctl is not a web interface action. By the rule set out earlier, that makes it a prerequisite to call out rather than something to fix quietly on the command line — write it into the runbook, because the next person to open the SDN panel will see a healthy fabric and no hint that a ring depends on a file in /etc/sysctl.d.\nBacking It Out Removal is the build in reverse, in the same interface: delete the SDN objects created for the mesh, apply the configuration, and confirm management access is untouched.\nThat last step is why nic0 and the 2.5 Gbit switch exist. If rolling back the mesh could cost you the web interface, the design was wrong before you started.\nBefore You Start Management access is genuinely independent of the mesh. Check it, do not assume it. All three nodes are healthy before any SDN change. Node names, interface names and cabling are written down, because nic1 on one host being nic2 on another is a bad afternoon. Jumbo frames are set on both mesh interfaces, and the MTU accounts for VXLAN\u0026rsquo;s overhead on the underlay. The mesh range is reserved and not in use anywhere else. References Proxmox VE — Software-Defined Network — the SDN Fabrics documentation. Fabrics \u0026ldquo;provide automated routing between nodes in a cluster\u0026rdquo;, OpenFabric is \u0026ldquo;based on IS-IS and optimized for the spine-leaf topology common in data centers\u0026rdquo;, each node needs a unique Router-ID, and \u0026ldquo;a dummy \u0026rsquo;loopback\u0026rsquo; interface with the router-id is automatically created\u0026rdquo; Proxmox VE — Cluster Manager — Corosync networking and redundant links, behind the two-path membership design Proxmox VE — Deploy Hyper-Converged Ceph Cluster — the network expectations for a hyper-converged cluster Linux kernel — IP sysctl documentation — ip_forward and its default of 0, the per-interface forwarding control, and the warning that changing the global switch resets per-interface configuration ","permalink":"https://blogs.damiendye.uk/en/proxmox/proxmox-routed-mesh-sdn-openfabric/","summary":"Three Proxmox nodes wired directly to each other in a triangle, with OpenFabric routing over the mesh and Ceph plus VXLAN client networks riding on top. Built entirely in the web interface, and honest about where the design stops scaling.","title":"A Switchless Proxmox Mesh with SDN OpenFabric — Ceph and Client Networks Without a 100G Switch"},{"content":"The Problem With Copied Boot Lines Search for Proxmox tuning and you will find a single long GRUB_CMDLINE_LINUX line, presented as a unit, with no indication of which flags apply where.\nThat matters more than it sounds. The hypervisor and the guest are solving opposite problems.\nThe host wants deterministic access to real hardware: IOMMU behaviour, PCIe link states, physical idle states. The guest wants to stop pretending it has hardware at all — its timers are approximations, its idle states are fiction, and its stalls are usually somebody else\u0026rsquo;s scheduler. As such, the same flag can be correct on one side, pointless on the other, and occasionally harmful.\nBelow is where each one actually belongs.\nWhere each kernel boot flag belongs Host only real hardware lives here iommu=pt amd_iommu=pgtbl_v2 pcie_acs_override=… pcie_aspm=off pci=pcie_bus_perf → pcie_bus_safe if USB4 processor.max_cstate=1 + intel_idle.max_cstate on Intel amd_pstate=disable set a governor too The guest has no PCIe links, no C-states and no cpufreq. Guest only stop pretending it is hardware cpuidle.off=1 guest idle states are fiction nmi_watchdog=0 a descheduled vCPU trips it softlockup_panic=0 the stall was the host's fault Leave kvm-clock alone. Do not force tsc or hpet — the counters are not yours. On the host these three all mean something different, and two of them cost you something real. Both — different reasons same flag, separate decision mitigations=off host: guest-to-host leaks guest: process isolation default_hugepagesz + hugepages host: backs guest RAM guest: backs an application pick one layer, not both watchdog_thresh / nowatchdog host: noise you accepted guest: never trustworthy \"Both\" does not mean you should set it in both places. It means the decision has to be made twice. Neither — these do nothing amd_iommu=on not a valid option; the kernel logs \"Unknown option - 'on'\" and carries on consoleblank=0 already the kernel default If a copied boot line contains either of these, it was never verified against /proc/cmdline — which is the cheapest check available, and the one that catches the wrong-bootloader mistake as well. The flags in circulation, sorted. Two of the popular ones do nothing on either side, and the watchdog row is a genuine judgement call rather than a rule. First: Are You Even Editing The Right File? A Proxmox install on ZFS root boots with systemd-boot, where /etc/default/grub is read by nobody. Editing it and rebooting produces no change and no error. That is a frustrating hour.\nproxmox-boot-tool status # tells you which bootloader is in use # systemd-boot: edit /etc/kernel/cmdline, then proxmox-boot-tool refresh # GRUB: edit /etc/default/grub, then update-grub Either way, verify rather than assume:\ncat /proc/cmdline And on the GRUB side, use GRUB_CMDLINE_LINUX_DEFAULT, not GRUB_CMDLINE_LINUX. The latter applies to every boot entry including recovery — and recovery is exactly when you want stock behaviour, not mitigations disabled and C-states pinned.\nThe Host Line IOMMU and passthrough iommu=pt amd_iommu=pgtbl_v2 pcie_acs_override=downstream,multifunction iommu=pt puts the IOMMU in passthrough mode: devices assigned to VMs get translated, host-native devices bypass translation. It is real and it is handled in arch/x86/kernel/pci-dma.c, which calls iommu_set_default_passthrough(true). The kernel documents it as equivalent to iommu.passthrough=1.\namd_iommu=on is not a thing. This is the most widely copied non-existent parameter in Proxmox guides. The kernel\u0026rsquo;s parse_amd_iommu_options() accepts fullflush, force_enable, off, force_isolation, pgtbl_v1, pgtbl_v2, irtcachedis, nohugepages and v2_pgsizes_only. Anything else lands here:\npr_notice(\u0026#34;Unknown option - \u0026#39;%s\u0026#39;\\n\u0026#34;, str); AMD-Vi is enabled by default when the firmware advertises it. Check your own log and you will find the parameter was never doing the job it was credited with:\ndmesg | grep -i \u0026#34;AMD-Vi\\|Unknown option\u0026#34; amd_iommu=pgtbl_v2 is valid — it selects the v2 DMA page-table format, which shares the CPU page-table structure rather than using AMD\u0026rsquo;s own. Two things to know: the documentation scopes it to the DMA-API, meaning the host\u0026rsquo;s own device domains rather than the VFIO domains used for passthrough; and it fails safe with a log line you should check for:\nif (amd_iommu_pgtable == PD_MODE_V2) { if (!amd_iommu_v2_pgtbl_supported()) { pr_warn(\u0026#34;Cannot enable v2 page table for DMA-API. Fallback to v1.\\n\u0026#34;); amd_iommu_pgtable = PD_MODE_V1; } } So it is worth measuring on a node with heavy host-side IO, and worth verifying you actually got it.\npcie_acs_override=downstream,multifunction is Proxmox\u0026rsquo;s out-of-tree patch. It splits IOMMU groups by asserting isolation the hardware does not advertise, which is what makes passthrough possible on consumer boards. It is also, exactly, telling the kernel something untrue about the topology. Fine on a box whose guests you trust as much as the host. Not fine otherwise. There is more on why in the IOMMU tax article.\nLatency and jitter pcie_aspm=off processor.max_cstate=1 amd_pstate=disable pcie_aspm=off keeps PCIe links out of low-power states so an arriving IO never waits for one to wake. It costs a few watts per link and removes a latency tail that is hard to diagnose. See PCIe ASPM and passthrough.\nprocessor.max_cstate=1 caps ACPI idle at C1. Note the driver: this is the processor/acpi_idle knob, so on Intel you need intel_idle.max_cstate=1 as well, because intel_idle takes precedence. On AMD this is the right one.\nThere is a real counter-argument. Deep sleep on idle cores is what gives the package thermal and power headroom to boost the busy ones, so pinning everything at C1 can lower your peak single-thread frequency while raising idle draw. On a latency-sensitive host that trade is usually worth it. On a host chasing throughput it might not be. Measure it rather than inheriting it.\namd_pstate=disable falls back to acpi-cpufreq. Worth knowing the documented alternatives before reaching for it: passive (driver requests a performance level), active (the EPP driver, biasing towards performance or efficiency), and guided. active with a performance bias, or passive plus the performance governor, often gets the same latency while keeping CPPC\u0026rsquo;s finer control. And if you do disable it, set a governor deliberately — landing on acpi-cpufreq with schedutil may be a step back.\nMemory default_hugepagesz=1G hugepages=64 default_hugepagesz=1G on its own reserves nothing. The kernel documents it as setting \u0026ldquo;the size of the default HugeTLB page… the default hugetlb size used for shmget(), mmap() and mounting hugetlbfs\u0026rdquo; — a unit, not an allocation. The allocation comes from hugepages=, documented as \u0026ldquo;Number of HugeTLB pages to allocate at boot\u0026rdquo;.\nThat matters far more for 1 GiB than 2 MiB pages, because contiguous 1 GiB regions are effectively unobtainable once the host has been up and fragmented memory. Boot is your only reliable chance.\nThen the guest has to opt in (hugepages: 1024 in the VM config). Reserved pages nothing uses are just memory you cannot have back, and you lose ballooning and KSM on the VMs that use them.\nThe security trade mitigations=off This is not one switch. The kernel expands it to a list, and on a hypervisor these are the entries that matter:\nl1tf=off mds=off mmio_stale_data=off kvm.nx_huge_pages=off gather_data_sampling=off retbleed=off spec_rstack_overflow=off nospectre_v2 nopti indirect_target_selection=off The kernel\u0026rsquo;s own summary is \u0026ldquo;improves system performance, but it may also expose users to several CPU vulnerabilities\u0026rdquo;. L1TF, MDS and MMIO stale data are specifically guest-to-host and guest-to-guest leak paths, and kvm.nx_huge_pages is the iTLB-multihit mitigation inside KVM itself.\nDefensible on a single-tenant box where every guest is as trusted as the host. Not defensible where guests are untrusted or belong to different tenants. And note it stacks with pcie_acs_override: two independent isolation guarantees removed on the same line. Worth doing on purpose rather than by inheritance.\nPCIe Over USB4 Changes Two Of These If your PCIe devices arrive over USB4 or Thunderbolt — an external GPU or NVMe enclosure — two of the answers above change.\nMPS tuning stops being free On fixed slots, pci=pcie_bus_perf is a small free win. The kernel describes it as:\nSet device MPS to the largest allowable MPS based on its parent bus. Also set MRRS (Max Read Request Size) to the largest supported value… for best performance.\nThe catch is that it configures bridges at boot, from the topology present at boot. Over USB4 hot-plug is the normal case, and a device added later may support a smaller MPS than the bridge was already set to.\nThe kernel says the quiet part out loud while advertising a different policy:\npcie_bus_peer2peer — Set every device\u0026rsquo;s MPS to 128B, which every device is guaranteed to support… This also guarantees that hot-added devices will work.\nOnly one policy carries that guarantee, and it is the one that pins everything to 128 bytes — exactly what MaxPayloadSize tuning sets out to escape. For a hot-plug topology, pcie_bus_safe (the largest value supported by all devices below the root complex) or simply leaving tune_off is the safer starting point. The tunnel\u0026rsquo;s own switch caps the achievable MPS anyway, so the ceiling was never yours to raise.\nWhy hot-plug changes the MPS policy answer At boot,\u0026#160;pcie_bus_perf\u0026#160;configures the bridge from what it can see bridge — MPS 512 device A · 512 device B · 512 present at boot hot-added · 256 only mismatch nothing renegotiated In a chassis that never happens — the topology at boot is the topology forever. On a USB4 port it is the normal case. The four policies, and which one the kernel says is hot-plug safe pcie_bus_tune_off leave the BIOS values alone pcie_bus_safe largest value all devices below the root complex support pcie_bus_perf largest the parent bus allows, per device — plus MRRS pcie_bus_peer2peer 128 B everywhere — \"guarantees that hot-added devices will work\" Only one policy carries that guarantee, and it is the one that throws away the payload size you were tuning for. The bridge is configured once, at boot, from the devices present then. Everything after that has to live with the decision — which is fine in a chassis and not fine on a port. The security stack-up gets serious External PCIe means someone can plug a DMA-capable device into your hypervisor. The kernel\u0026rsquo;s Thunderbolt documentation is direct about it:\n…the connected devices can be DMA masters and thus read contents of the host memory without CPU and OS knowing about it. There are ways to prevent this by setting up an IOMMU but it is not always available for various reasons.\nThe IOMMU is the defence. Now count what the host line does to it. iommu=pt gives host-owned devices untranslated identity domains, pcie_acs_override asserts isolation that is not there, and mitigations=off disables the guest-isolation mitigations. Each is defensible alone. Together, on a machine with a physically reachable USB4 port, they stack up.\nCheck where you stand:\ncat /sys/bus/thunderbolt/devices/domain*/security # none | user | secure | dponly | usbonly none alongside that boot line is an open door. If the ports are reachable by people you would not give root to, iommu=pt is the first thing I would reconsider.\nThe Guest Line These are the ones that belong inside the VM, and three of them mean something different here than they would on the host.\nnmi_watchdog=0 softlockup_panic=0 cpuidle.off=1 cpuidle.off=1 disables the cpuidle subsystem. In a guest that is close to free: the guest\u0026rsquo;s idle states are emulation, and there is no physical core to put to sleep, so all the framework buys you is wake-up latency. On the host the same flag is a real power and boost-headroom trade, and it overlaps processor.max_cstate=1. Guest side only.\nsoftlockup_panic=0 stops a soft lockup panicking the guest. This is genuinely protective in a VM, because a soft lockup there is frequently not the guest\u0026rsquo;s fault. A descheduled vCPU looks exactly like a task that refused to yield. That is the same mechanism behind why a VM\u0026rsquo;s clock cannot be trusted. Check whether you need it, though. It is 0 by default on most builds.\nsysctl kernel.softlockup_panic nmi_watchdog=0 is the interesting one, and it deserves more than a rule.\nThe Watchdog Question There are two detectors sharing one threshold:\nwatchdog_thresh= — Set the hard lockup detector stall duration threshold in seconds. The soft lockup detector threshold is set to twice the value. A value of 0 disables both. Default is 10 seconds.\nHard lockup (nmi_watchdog) fires when a CPU stops taking timer interrupts altogether. Soft lockup fires when a task hogs a CPU for twice as long without scheduling.\nIn a guest, disabling the hard lockup detector is right. A descheduled vCPU can trip it through no fault of its own, and the detector\u0026rsquo;s perf-counter work causes VM exits for a signal that was nowt but noise.\nOn the host it is a judgement call, and it depends on what noise you are actually chasing. Over-provisioning starves guests, not the host kernel — the host\u0026rsquo;s physical CPUs keep taking interrupts however packed the VMs are. So the log noise a busy hypervisor throws is overwhelmingly soft lockup and RCU stall messages, not NMI hard-lockup reports. If that is the noise, nmi_watchdog=0 will not silence it, and softlockup_panic=0 will not either — that stops the panic, not the messages.\nThe targeted knobs are:\nwatchdog_thresh=30 # hard 30s, soft 60s — scale to taste nowatchdog # honest single flag: disables both detectors sysctl -w kernel.soft_watchdog=0 # runtime, keeps hard-lockup detection There is a separate and better reason to disable the hard lockup detector on a busy host, which has nothing to do with noise: it consumes a hardware performance counter per CPU. That is why the parameter accepts rNNN to configure a raw perf event. If you are doing PMU-based profiling, or running a deliberately over-provisioned box where you have already accepted latency variance as the price of density, giving that counter back is a reasonable trade — and small stalls you have consciously signed up for are not incidents.\nJust make the choice for that reason rather than for the noise reason, because only one of them is true.\nconsoleblank=0 does nothing. The kernel documents the console blank timeout as \u0026ldquo;A value of 0 disables the blank timer. Defaults to 0.\u0026rdquo; It is already off. Harmless, but it is the second parameter in common circulation that has no effect, and carrying it makes a line look considered when it is copied.\nQuick Reference: Where Each Flag Belongs Flag Host Guest Notes iommu=pt yes no Host owns the IOMMU. Only applies in a guest if you are running nested passthrough with a vIOMMU amd_iommu=pgtbl_v2 yes no Host DMA-API domains. Verify you got v2 rather than the v1 fallback amd_iommu=on — — Not a valid option. Kernel logs \u0026ldquo;Unknown option - \u0026lsquo;on\u0026rsquo;\u0026rdquo; pcie_acs_override=… yes no Proxmox patch, host topology only. Weakens isolation by design pcie_aspm=off yes no There are no real PCIe links in a guest; the host owns the physical link pci=pcie_bus_perf yes not reliably Host sets MPS on the wire. Use pcie_bus_safe instead if devices arrive over USB4 processor.max_cstate=1 yes no Real idle states are the host\u0026rsquo;s. Add intel_idle.max_cstate=1 on Intel amd_pstate=disable yes no Guests do not control CPU frequency cpuidle.off=1 with care yes Free in a guest. On the host it costs boost headroom and overlaps max_cstate nmi_watchdog=0 judgement yes Right in a guest. On the host, do it for the PMU counter, not for the noise softlockup_panic=0 no yes Guest stalls are often the host\u0026rsquo;s scheduler. Usually already the default consoleblank=0 — — No-op. Kernel default is already 0 mitigations=off both both Valid on either, different risk calculus: guest-to-host leaks on the host, process isolation in the guest default_hugepagesz + hugepages= both both Host: backing VM memory. Guest: a workload inside the VM that wants them. Different purposes, same flags watchdog_thresh= / nowatchdog both both Host: quieten stalls you accepted. Guest: the detector was never trustworthy Three genuinely belong on both sides, and it is worth being exact that \u0026ldquo;both\u0026rdquo; does not mean \u0026ldquo;for the same reason\u0026rdquo;:\nmitigations=off on the host is about guest-to-host and guest-to-guest leakage. Inside a guest it is about process isolation within that VM. You can reasonably disable it in one place and not the other. Hugepages on the host back guest RAM; in a guest they back an application. Reserving them twice for the same memory is waste, so decide which layer wants them. Watchdog tuning is a noise decision on the host and a correctness decision in the guest. Everything else is one side or the other, and two of them are not choices at all.\nThe Two Lines Host, on an AMD box doing passthrough, single-tenant:\nGRUB_CMDLINE_LINUX_DEFAULT=\u0026#34;quiet amd_iommu=pgtbl_v2 iommu=pt pcie_acs_override=downstream,multifunction pcie_aspm=off pci=pcie_bus_perf default_hugepagesz=1G hugepages=64 processor.max_cstate=1 amd_pstate=disable mitigations=off\u0026#34; Swap pci=pcie_bus_perf for pcie_bus_safe if anything arrives over USB4. Drop mitigations=off if the guests are not all yours. On Intel, intel_iommu=on replaces the AMD flags and intel_idle.max_cstate=1 joins the C-state cap.\nLinux guest:\nGRUB_CMDLINE_LINUX_DEFAULT=\u0026#34;quiet nmi_watchdog=0 softlockup_panic=0 cpuidle.off=1\u0026#34; And leave kvm-clock alone in the guest — do not force tsc or hpet. The paravirtual clock exists precisely because the counters are not yours. That is the whole argument in the VM timekeeping article.\nWhat To Delete If you inherited a line from a forum post, these two are the first things to remove, because they cost you nowt and prove the line was never tested:\namd_iommu=on — not a valid option; the kernel logs \u0026ldquo;Unknown option - \u0026lsquo;on\u0026rsquo;\u0026rdquo; and carries on consoleblank=0 — already the default And check the remainder against /proc/cmdline after a reboot. Every flag on that line should be one you can name a reason for.\nIf you cannot say what a flag does, it is not tuning. It is superstition. And it will be copied into the next build by someone who trusts you.\nReferences Linux kernel — the kernel\u0026rsquo;s command-line parameters — mitigations=, default_hugepagesz=, hugepages=, watchdog_thresh=, nowatchdog, consoleblank=, processor.max_cstate=, amd_pstate=, amd_iommu= and the pci=pcie_bus_* policies Linux kernel source — drivers/iommu/amd/init.c — parse_amd_iommu_options(), and the v2 page-table capability check that falls back to v1 Linux kernel source — arch/x86/kernel/pci-dma.c — iommu=pt calling iommu_set_default_passthrough() Linux kernel — Thunderbolt — security levels, and connected devices as DMA masters Proxmox VE — System Administration — proxmox-boot-tool status, editing the kernel command line for systemd-boot versus GRUB ","permalink":"https://blogs.damiendye.uk/en/proxmox/kernel-boot-flags-host-guest/","summary":"Most Proxmox tuning lines you find online are one blob of kernel flags. Half of them belong on the hypervisor, half belong inside the guest, two of the popular ones do nothing at all, and a few change meaning depending on which side of the boundary they land.","title":"Two Boot Lines — Which Kernel Flags Belong on a Proxmox Host, and Which Belong in the Guest"},{"content":"Why You Would Want One QEMU can emulate an actual NVMe controller — not a paravirtual device that needs a driver you supply, but a PCIe NVMe controller that a guest recognises as a normal SSD and drives with the NVMe support it already has.\nThat is the entire appeal, and it is worth more than it sounds.\nLinux has had an in-box nvme driver for years. Windows has shipped stornvme since Windows 8.1 and Server 2012 R2. So a guest boots, enumerates a PCIe NVMe controller, loads its own driver and finds a disk. No VirtIO ISO, no driver injection at install time, and no \u0026ldquo;no drives found\u0026rdquo; screen halfway through a Windows installer.\nAnyone who has sat looking at that screen, with the VirtIO ISO mounted and the installer still insisting there are no disks, will see the appeal immediately.\nThe second reason is that it behaves like NVMe all the way up. nvme-cli works. Namespaces are real. LBA formats, metadata bytes and protection information are all configurable. Which makes it a very good place to practise the operations you should not be practising on hardware that holds data.\nHow To Add One Proxmox has no GUI checkbox or config key for this. It is a raw QEMU device, so it goes in args: in /etc/pve/qemu-server/\u0026lt;vmid\u0026gt;.conf.\nThe QEMU documentation gives the minimal pair: a backing drive with no interface, and the controller that consumes it. Edited straight into /etc/pve/qemu-server/\u0026lt;vmid\u0026gt;.conf, unquoted:\nargs: -drive file=/var/lib/vz/images/100/nvm.img,if=none,id=nvmidentifier -device nvme,serial=LAB-NVME-01,drive=nvmidentifier if=none matters: it tells QEMU not to attach the drive to a default controller, because the -device nvme line is going to claim it. The id= on the drive and the drive= on the device have to match. That pairing is what joins the two halves.\nThe serial= is mandatory; QEMU refuses to start the VM without one. Choose something you will recognise, because it is exactly what the guest reports back in nvme list and smartctl, and \u0026ldquo;which of these four identical virtual drives is which\u0026rdquo; is a question you will eventually ask.\nQuoting: The Bit That Catches Everyone Whether you quote that string depends on where you are typing it, and getting it the wrong way round is the most common reason one of these fails on the first attempt.\nProxmox stores the args: value and later splits it with Text::ParseWords::shellwords. So in the config file, quotes are honoured and stripped. A fully quoted string becomes a single argument:\n# WRONG in the config file — collapses to one argv element QEMU cannot parse args: \u0026#34;-drive file=…,if=none,id=nvmidentifier -device nvme,serial=…,drive=nvmidentifier\u0026#34; Run through shellwords, that yields exactly one element. Unquoted, the same line yields the four QEMU actually needs: -drive, its parameter blob, -device, its parameter blob.\nOn the command line it is the opposite, because there you are quoting for your shell, not for Proxmox. Here the quotes are required, and what gets stored in the config is the unquoted value:\nqm set 100 --args \u0026#34;-drive file=/var/lib/vz/images/100/nvm.img,if=none,id=nvmidentifier -device nvme,serial=LAB-NVME-01,drive=nvmidentifier\u0026#34; Both of those are correct. They just are not interchangeable. If you build the line with qm set, check the result with qm config 100 afterwards and you will see it stored bare. That is the form the config file wants.\nCreate the backing image first if it does not exist:\nqemu-img create -f raw /var/lib/vz/images/100/nvm.img 32G More Than One Namespace For anything beyond a single disk, split the controller from its namespaces:\n-device nvme,id=nvme-ctrl-0,serial=deadbeef -drive file=nvm-1.img,if=none,id=nvm-1 -device nvme-ns,drive=nvm-1 Namespace identifiers are allocated from 1 upwards automatically. This is the configuration that makes the device genuinely useful for learning, because namespace management is the part of NVMe most people never get to touch.\nA 4Kn Virtual Namespace The namespace takes the usual block-size properties, and QEMU derives the LBA data size directly from them. hw/nvme/ns.c computes the format exponent as ds = 31 - clz32(ns-\u0026gt;blkconf.logical_block_size). So this gives you a proper 4K-native namespace:\n-device nvme-ns,drive=nvm-1,logical_block_size=4096,physical_block_size=4096 The namespace device also accepts ms for metadata bytes per LBA, mset for extended LBAs, and pi and pif for protection information type and guard format.\nThat is a complete laboratory for everything in the 4Kn and 512e article — 512-byte versus 4096-byte logical blocks, metadata-bearing formats, T10-PI — on a device you can destroy as often as you like.\nCheck It Landed From inside the guest:\nlsblk -o NAME,MODEL,SIZE,LOG-SEC,PHY-SEC nvme list nvme id-ns -H /dev/nvme0n1 | grep -i \u0026#34;lbaf\\|data size\u0026#34; You should see a real NVMe namespace, with the block sizes you asked for.\nWhat You Give Up Three things, and the first two are not performance trade-offs. They are capability removals. Know them before you put anything on the device.\nWhat Proxmox manages, and what an args-attached device sits outside of In the VM config, as a drive scsi0: local-zfs:vm-100-disk-0,iothread=1 Proxmox owns the volume and knows it exists Included in vzdump / PBS backups yes PVE snapshots yes Live migration yes Resize and Move Disk in the GUI yes Counted in the storage view yes Attached through args: -device nvme,drive=nvm1,serial=… a raw QEMU device — PVE knows nothing of it Included in backups no PVE snapshots no Live migration no Resize and Move Disk no Counted in the storage view no Only one of those \"no\"s announces itself. Live migration fails with an error, because QEMU marks the device unmigratable. The backup simply succeeds without the disk in it — which is why this belongs in your runbook, not just in your memory. Proxmox manages what is in the VM config as a drive. An args-attached controller is outside that, so every feature built on the storage layer simply does not apply to it. 1. Live Migration Is Off This is not a Proxmox limitation or an oversight. QEMU declares the device unmigratable in the device model itself. From hw/nvme/ctrl.c in QEMU 10.2:\nstatic const VMStateDescription nvme_vmstate = { .name = \u0026#34;nvme\u0026#34;, .unmigratable = 1, }; Three lines, and the middle one is the whole story. The controller has no migration state, so QEMU refuses the migration rather than attempting it. That is the right failure. You get an error, not a guest that resumes on another node with a confused disk.\nThere is a second, independent reason it cannot work: Proxmox does not know the disk exists. Even if QEMU could move the device state, nothing in PVE\u0026rsquo;s migration logic would arrange for the backing volume to be available on the target.\nWorth watching, though: QEMU\u0026rsquo;s development branch has replaced the blanket flag with a nvme_set_migration_blockers() function that permits migration and blocks it only for specific features. More than one namespace, for instance, where the comment notes \u0026ldquo;we don\u0026rsquo;t handle this in migration code yet\u0026rdquo;. That has not appeared in a release up to and including 10.2, so it does not help you today, but this restriction looks likely to soften. Check your own QEMU version rather than trusting an article.\n2. Proxmox Backups Will Not See It vzdump and Proxmox Backup Server back up the volumes that appear in the VM config as drives — scsi0, virtio0, and so on. A disk attached through args: is not one of those. It is a raw QEMU device that PVE knows nowt about.\nSo the backup runs, reports success, and does not contain the device.\nThat failure mode is worse than an error, because nothing tells you. The same applies across the board: no PVE snapshots, no disk resize from the GUI, no Move Disk, no accounting in the storage view. If you created the volume through PVE and then detached it, PVE may not clean it up either. An orphan waiting to confuse somebody later.\nIf data is going to live on one of these, back it up from inside the guest, and write down somewhere that the hypervisor is not covering it.\n3. It Is Not Faster Than VirtIO SCSI This one surprises people, because \u0026ldquo;NVMe\u0026rdquo; reads like a performance feature. Here it is not one.\nVirtIO SCSI and VirtIO block are paravirtual: the guest driver and the hypervisor share a ring buffer designed for exactly this job, and the guest knows it is talking to a hypervisor.\nThe emulated NVMe controller is the opposite by design. It presents real NVMe registers, so the guest programs it as though it were hardware. Every doorbell write is an MMIO access that traps into the hypervisor. Correct, and more expensive per IO than putting a descriptor on a ring.\nA shared ring versus emulated registers guest │ hypervisor VirtIO SCSI — paravirtual guest driver knows it is a VM shared ring both ends understand it host block layer 1 notify A descriptor goes on the ring and the host is notified. The ring straddles the boundary on purpose. Emulated NVMe — real registers guest driver thinks it is hardware NVMe registers doorbells, queues, MMIO host block layer every doorbell write traps trap + emulate, per IO The guest is doing exactly what it would do to a physical controller, which is the point — its own driver works unmodified. It is also why this is a compatibility feature and not a performance one. Interrupt coalescing is unsupported and off by default: correct hardware behaviour, at the cost of emulating hardware. The paravirtual path is a ring the guest and host both understand. The emulated path makes the guest drive registers, and each doorbell write is a trap — accurate hardware behaviour, at hardware-emulation cost. QEMU\u0026rsquo;s own documentation is candid about the device\u0026rsquo;s rough edges too: interrupt coalescing \u0026ldquo;is not supported and is disabled by default\u0026rdquo;, and accounting numbers in the SMART/Health log page \u0026ldquo;are reset when the device is power cycled\u0026rdquo;.\nNone of that makes it slow in absolute terms. It is perfectly usable. It just means you should never pick it hoping for more throughput than VirtIO SCSI gives you. Pick it for the driver, or for the NVMe semantics.\nWhere It Actually Earns Its Place Installing a guest with no VirtIO media. A Windows installer that cannot see a VirtIO SCSI disk will see an NVMe one, because the driver is already in the image. Install onto it, then decide whether to switch to VirtIO afterwards. Appliances and images you do not control. Anything shipped as a fixed image that lacks VirtIO drivers, and that you would rather not rebuild. Learning and lab work. nvme format --lbaf, namespace creation and attachment, metadata and protection information — the operations that are destructive and vendor-dependent on real hardware are free here. This is the safest way to build the muscle memory before touching a drive that matters. Reproducing somebody else\u0026rsquo;s topology. If you are debugging a customer\u0026rsquo;s NVMe layout, an emulated controller with matching namespaces and block sizes is a much faster loop than borrowing their hardware. What To Use Instead In Production For a VM that needs performance, PVE features, and a quiet life: VirtIO SCSI single, with iothread=1, discard=on and ssd=1, on cache=none. That is the arrangement that keeps live migration, backups, snapshots and the storage view all working.\nFor a VM that needs the last few percent and can give up those features on purpose, the answer is not an emulated NVMe device. It is real passthrough, with its own hard trade-offs, covered in the IOMMU tax article.\nThe emulated NVMe device sits in neither camp. As such, it is a compatibility and lab tool, and it is very good at being that.\nUse it for the job it is good at and it will not let you down. Ask it to be a performance feature and it will let you down very promptly.\nReferences QEMU — NVMe Emulation — the -drive/-device nvme syntax, nvme-ns for multiple namespaces, the ms/mset/pi/pif namespace parameters, and the stated limitations on interrupt coalescing and SMART accounting QEMU source — hw/nvme/ctrl.c — the nvme_vmstate declaration with .unmigratable = 1 in the 10.2 release QEMU source — hw/nvme/ns.c — the namespace deriving its LBA format from logical_block_size Proxmox VE — Backup and Restore — what vzdump covers, and the per-volume backup options that exist for drives PVE manages Proxmox VE — Qemu/KVM Virtual Machines — VirtIO SCSI, iothread, discard and the supported disk options ","permalink":"https://blogs.damiendye.uk/en/proxmox/virtual-nvme-proxmox/","summary":"QEMU can emulate a real NVMe controller, so the guest uses its own in-box NVMe driver with no VirtIO media required. It is also unmigratable by design, invisible to Proxmox backups, and no faster than VirtIO SCSI. Here is how to add one and when it is worth it.","title":"A Virtual NVMe Device in Proxmox — and the Three Things You Give Up"},{"content":"The Assumption Every Clock Makes A computer\u0026rsquo;s clock works by counting something regular and trusting that it keeps counting. A crystal oscillates, a counter increments, and software multiplies out how much time has passed.\nVirtualisation breaks the trusting part.\nThe kernel\u0026rsquo;s own KVM timekeeping documentation puts the problem in one sentence: \u0026ldquo;the virtual operating system does not run with 100% usage of the CPU, despite the fact that it may very well make that assumption.\u0026rdquo; Everything below follows from that.\nYour vCPU Is Not Always Running A vCPU is a thread on the host. It runs when the host scheduler says so.\nWhen it is not running, the guest is not merely idle — it is absent. It cannot count, it cannot service a timer interrupt, and it has no way to know how long it was gone. The host tracks that as steal time, which is the honest name for \u0026ldquo;time that happened to you rather than for you\u0026rdquo;.\nTimer interrupts are where this hurts most. A guest that asks for a periodic tick is asking the host to deliver interrupts at a fixed rate, and the host cannot always oblige. Again from the kernel documentation: \u0026ldquo;the host virtualization engine may not be able to deliver the proper number of interrupts per second, and so guest time may fall behind.\u0026rdquo;\nWhy a guest's timer ticks stop being evenly spaced Bare metal — the CPU is always yours running timer ticks, evenly spaced — counting them gives you the time In a VM — the vCPU is a thread on somebody else's scheduler running descheduled descheduled ticks that were due during the gaps arrive late, or bunched, or not at all The guest cannot see the dashed periods. From inside, the clock simply produced fewer ticks than it should have — which is why the kernel documentation says a guest's time \"may fall behind\" when the host cannot deliver the interrupts it asked for. The host calls the dashed regions steal time. The guest calls them nothing at all, because it was not there. The guest\u0026rsquo;s own view is the bottom row: ticks that arrive late, ticks that arrive in bursts, and gaps it cannot account for. It is measuring the host\u0026rsquo;s scheduler as much as the passage of time. The higher the tick rate the worse it gets, and an overcommitted host makes it worse again. This is also why a busy host degrades the timekeeping of quiet guests. They are all queuing for the same physical cores.\nThe Counters Are Not Yours Either If periodic ticks are unreliable, the obvious answer is to read a counter instead. That has its own problems.\nThe TSC is the fast one, and the kernel documentation is blunt about it: \u0026ldquo;The TSC is a CPU-local clock in most implementations… the TSCs of different CPUs may start at different times.\u0026rdquo; Its rate can vary with processor power states, and on older parts it stops entirely when the core idles. A vCPU that migrates between physical cores can therefore read a counter that disagrees with the one it read a microsecond ago.\nThe alternatives are worse in a different way. The HPET, the PIT and the ACPI PM timer are all emulated devices, so every read is a trap into the hypervisor. Correct, and expensive enough that a guest reading the clock in a hot loop will notice.\nThis is why paravirtual clocks exist. On KVM, kvm-clock lets the host publish its own timekeeping into a shared structure that the guest reads directly: no trap, no counting, no assumption that the guest was awake. It is also why you should leave the guest\u0026rsquo;s clocksource alone rather than forcing tsc or hpet because a forum post said it was faster.\nMigration, Snapshots and Suspend Live migration, snapshot restore and suspend/resume all do the same thing to a guest clock: they stop it, and then start it again somewhere else.\nWhat the guest sees is not drift, it is a step. The clock was one value, and now it is a different value, with nowt in between. Migrating to a host whose TSC runs at a different frequency compounds it.\nStep changes matter because the software that corrects clocks is built to correct drift, not teleportation.\nThis Is Not a KVM Problem It is tempting to read all of the above as a KVM shortcoming. It is not.\nEvery hypervisor ships a paravirtual clock, because every hypervisor has the same structural problem: KVM has kvm-clock, Hyper-V has its reference TSC page, VMware has a pseudo-performance counter plus Tools-based sync, Xen has its pvclock. Those are four independent implementations of one workaround.\nThe kernel documentation\u0026rsquo;s own conclusion is that there is no perfect solution here. Only trade-offs between accuracy, performance and complexity.\nIf you want it from a vendor rather than from kernel developers, Microsoft\u0026rsquo;s support boundary for high-accuracy time is remarkably candid. To claim 50 ms accuracy on a virtualised Windows system, one of the stated requirements is that \u0026ldquo;the one-day average CPU utilization of the host must not exceed 90%.\u0026rdquo; For 1 ms, the host must stay under 80%.\nRead that again: the accuracy of the guest\u0026rsquo;s clock is documented as conditional on how busy the host is. That is not a Windows quirk, it is the same physics as the kernel doc describes, written down as a support boundary.\nThe three layers of guest timekeeping, and which ones the guest actually owns Wall-clock synchronisation what the actual time is NTP over the network jitter, asymmetry, stratum ptp_kvm hypercall asks the host directly Paravirtual clock the host publishes its own timekeeping Every hypervisor ships one, because they all have this problem: kvm-clock Hyper-V TSC page VMware pseudo-perf Xen pvclock four independent implementations of one workaround Hardware counters not yours in a VM TSC HPET PIT ACPI PM timer CPU-local and rate-variable, or emulated — so every read is a trap into the hypervisor The guest owns its choice at the top. The middle layer is the hypervisor lending it a clock that was running the whole time. Three layers, and the guest only genuinely owns the top one. Every hypervisor\u0026rsquo;s paravirtual clock exists to work around the same missing guarantee in the layer below. Why NTP Inside the Guest Is the Wrong Tool NTP is good at what it was designed for: a machine with a real oscillator that runs slightly fast or slow, corrected by measuring the round trip to a remote server and gently slewing the local clock.\nEach of those assumptions is shaky in a VM.\nThe local oscillator is not slightly wrong, it is intermittently absent. The round-trip measurement is taken by a process that may be descheduled between reading the clock and sending the packet, which corrupts the measurement itself. And the corrections needed after a migration are steps, which a slewing algorithm handles badly or refuses outright.\nchrony copes far better than ntpd here. It slews faster, tolerates steps, and is honest about its own uncertainty. But it is still solving the wrong problem: pulling time across a network from a stratum-2 server 20 ms away, when the correct time is sitting in the hypervisor on the other side of a single memory boundary.\nWhat you get in practice is a guest that is mostly right, occasionally tens or hundreds of milliseconds out, and never quite able to tell you which.\nWhat Skew Actually Breaks Nobody cares about clocks for their own sake. They care when something stops working.\nKerberos and Active Directory Kerberos is time-dependent by design, because a ticket\u0026rsquo;s validity is expressed as a time window.\nMIT\u0026rsquo;s krb5.conf documentation defines clockskew as \u0026ldquo;the maximum allowable amount of clockskew in seconds that the library will tolerate before assuming that a Kerberos message is invalid\u0026rdquo;, and the default is 300 seconds — five minutes.\nCross that and authentication does not degrade, it fails. Since Active Directory authentication is Kerberos, that means domain logons, net use, SQL Server connections, Exchange, file shares — the lot. Five minutes sounds generous until a guest steps backwards after a snapshot restore.\nTLS A certificate carries a validity window: notBefore and notAfter. A guest whose clock is behind will reject a certificate that was issued this morning because, as far as it knows, the certificate is not valid yet. A guest whose clock is ahead will reject one that has not actually expired.\nThe same arithmetic governs OCSP and CRL freshness, JWT nbf/exp claims, and TOTP multi-factor codes, which live in 30-second windows. A clock 45 seconds out is an authentication outage with a very confusing error message.\nCeph Ceph monitors care about this more than almost anything else in the stack, because their consensus depends on it.\nThe MON_CLOCK_SKEW health check fires when \u0026ldquo;the clocks on hosts running Ceph Monitor daemons are not well-synchronized\u0026rdquo;. Specifically, when the skew exceeds mon_clock_drift_allowed. The documentation\u0026rsquo;s advice is to sync with ntpd or chrony against multiple sources, and it notes that monitor-to-monitor synchronisation particularly matters.\nYou can raise mon_clock_drift_allowed, but the docs are clear it has to stay \u0026ldquo;significantly below the mon_lease interval\u0026rdquo;. As such, it is a small budget, and spending it to paper over a hypervisor timekeeping problem is not a good trade.\nEverything Else Logs across hosts stop correlating, which turns an incident timeline into guesswork. Database replication and distributed consensus — etcd, Galera, anything doing leader leases — get unhappy. Backup and monitoring windows drift out of alignment with the thing they were meant to observe.\nWindows Guests Skew Differently Windows deserves its own note, because its time service was built with different goals and it shows.\nMicrosoft states plainly that versions before Windows 10 1607 / Server 2016 \u0026ldquo;can\u0026rsquo;t guarantee highly accurate time\u0026rdquo;. What the Windows Time service provided on those releases was \u0026ldquo;the necessary time accuracy to satisfy Kerberos version 5 authentication requirements\u0026rdquo; and \u0026ldquo;loosely accurate time\u0026rdquo; for machines in a common AD forest. Tighter than that was \u0026ldquo;outside of the design specification… and weren\u0026rsquo;t supported.\u0026rdquo;\nIn other words, older Windows aims to stay inside the five-minute Kerberos window, not to be right. Which is fine until something in your estate needs better. A virtualised Windows box that is 90 seconds out will authenticate happily while writing logs that cannot be correlated with anything.\nWindows 10 and Server 2016 onward can do 1 s, 50 ms or even 1 ms — but only under the conditions quoted earlier, including the host CPU utilisation limits. Microsoft also notes that \u0026ldquo;anything that introduces network asymmetry, such as a one-way satellite connection or high CPU load on the target system, will negatively influence accuracy\u0026rdquo;. A contended vCPU is high CPU load on the target system by another name.\nThere is no ptp_kvm for Windows. What you have instead:\nHyper-V clock enlightenments. Proxmox already exposes these to Windows guests. PVE::QemuServer::CPUConfig sets hv_time alongside hv_vapic, hv_spinlocks, hv_relaxed and hv_synic. hv_time is the paravirtual clock, and it is the Windows-side equivalent of kvm-clock. This is on by default for Windows-typed VMs; there is nothing to enable. The QEMU guest agent. With the agent installed, the host can push its time into the guest after a resume or snapshot restore, which handles the step-change case that NTP handles worst. Pick one authority. The classic Windows-in-a-VM failure is two time sources fighting: host-to-guest sync and domain hierarchy sync, both correcting the same clock in opposite directions. For a domain-joined guest, let the domain hierarchy win and stop the host pushing time at it. For a standalone guest, host sync is fine. Never both. If you want to know exactly what your hypervisor is telling a given VM about time, ask it rather than guessing:\n# Everything Proxmox actually passes to QEMU for this VM, including -rtc and CPU flags qm showcmd \u0026lt;vmid\u0026gt; --pretty The Fix on QEMU/KVM: ptp_kvm For Linux guests on KVM there is a proper answer, and it is not \u0026ldquo;more NTP servers\u0026rdquo;.\nptp_kvm lets the guest ask the host what time it is, directly, through a hypercall — KVM_HC_CLOCK_PAIRING on x86, and an equivalent firmware call on arm64. The kernel presents that as a PTP hardware clock device, so from userspace it looks like any other precision clock source, and chrony can use it as a reference clock.\nThe properties that matter:\nNo network. No jitter, no asymmetry, no stratum, no packets. The path is a memory boundary. Sub-microsecond. The accuracy is bounded by the hypercall, not by a round trip across a datacentre. It bypasses the broken part. The guest is not counting anything or estimating a round trip. It is reading a value the host computed with a clock that was running the whole time. NTP across the network versus ptp_kvm across a memory boundary chrony against network pools guest clock keeps stopping network jitter, asymmetry stratum 2 server tens of ms away the round trip is measured by the clock being corrected — and the process doing the measuring can be descheduled mid-measurement chrony against ptp_kvm guest reads /dev/ptp_kvm host clock never stopped running hypercall KVM_HC_CLOCK_PAIRING a memory boundary, not a network sub-µs No packets, no stratum, no asymmetry, and nothing being estimated. The guest is not working out what time it is — it is being told, by the only participant that was awake for the whole interval. Both paths end with the guest setting its clock. One of them measures a network round trip using a clock that keeps stopping; the other asks the hypervisor. Implementation Apply this to every Linux VM that has the driver. The host hypervisor needs working NTP or PTP of its own. ptp_kvm hands the guest the host\u0026rsquo;s time, so it inherits the host\u0026rsquo;s error.\n1. Load the kernel module at boot.\n# /etc/modules-load.d/ptp_kvm.conf ptp_kvm 2. Give it a stable name and let chrony read it.\n# /etc/udev/rules.d/90-ptp-kvm.rules ACTION==\u0026#34;add\u0026#34;, SUBSYSTEM==\u0026#34;ptp\u0026#34;, ATTR{clock_name}==\u0026#34;kvm\u0026#34;, SYMLINK+=\u0026#34;ptp_kvm\u0026#34;, GROUP=\u0026#34;chrony\u0026#34;, MODE=\u0026#34;0660\u0026#34; The symlink matters because PTP device numbering is not stable. /dev/ptp0 may be a NIC\u0026rsquo;s clock on one boot and the KVM clock on the next. Matching on clock_name gets the right one every time.\n3. Point chrony at it, and remove the pools.\nEdit /etc/chrony/chrony.conf on Debian and Ubuntu, or /etc/chrony.conf on RHEL family. Delete the pool lines and use:\nrefclock PHC /dev/ptp_kvm poll 2 stratum 1 delay 0.0004 Removing the pools is not optional tidiness. Leaving them in asks chrony to reconcile a sub-microsecond local reference against internet servers tens of milliseconds away, and the internet sources can only make the answer worse.\n4. Restart and check.\nsystemctl restart chronyd # or chrony, on Debian/Ubuntu Verify # The symlink exists and points at the KVM clock ls -l /dev/ptp_kvm cat /sys/class/ptp/ptp*/clock_name # chrony should be using PHC0 as its selected source chronyc sources -v # and the offset should be microseconds, not milliseconds chronyc tracking In chronyc sources, the PHC refclock appears as #* PHC0 once selected. The # marks a local hardware reference rather than a network peer, and the * marks it as the one in use. If you see it listed but not selected, chrony has not accepted it: check permissions on the device and that the chrony group in the udev rule matches the user chrony actually runs as on your distribution.\nCaveats The host must be right. This makes the guest agree with the host, which is only useful if the host agrees with reality. Give the hypervisors real NTP or PTP. Every guest needs it. A fleet where half the VMs use ptp_kvm and half use internet pools is a fleet with two time authorities. Live migration is fine, and is the point. After migrating, the guest reads its new host\u0026rsquo;s clock. Provided the hosts agree with each other, the guest never sees a step. KVM only. It is a KVM hypercall. Nested or foreign hypervisors will not present the device, and the udev rule simply will not fire. That is a clean failure rather than a silent wrong answer. What Not To Do Don\u0026rsquo;t force the clocksource. Leave kvm-clock alone. Forcing tsc or hpet on the guest kernel command line trades a paravirtual clock designed for this situation for a counter that was never yours. Don\u0026rsquo;t run ntpdate or hwclock from cron. That is a step change on a schedule, which is exactly what databases and Kerberos hate. Don\u0026rsquo;t keep the pools \u0026ldquo;as a fallback\u0026rdquo;. With a working refclock they are not a fallback, they are a second opinion from a worse source. Don\u0026rsquo;t raise mon_clock_drift_allowed and call it fixed. You have spent part of a budget that exists for network reality, in order to tolerate a problem with a known solution. That last one is worth saying plainly: widening the threshold is not fixing the clock. It is moving the alarm so it stops going off.\nReferences Linux kernel — KVM timekeeping — why guest time falls behind, the TSC\u0026rsquo;s local-clock problem, and the conclusion that only trade-offs exist Linux kernel — PTP_KVM — the hypercall interface behind the PTP clock device Linux kernel — KVM x86 hypercalls — KVM_HC_CLOCK_PAIRING, the x86 side of it chrony — chrony.conf — the refclock directive and PHC driver options MIT Kerberos — krb5.conf — clockskew and its 300-second default Microsoft — support boundary for high accuracy time — the accuracy targets, and the host CPU utilisation conditions for virtualised systems Ceph — health checks — MON_CLOCK_SKEW, mon_clock_drift_allowed and its relationship to mon_lease ","permalink":"https://blogs.damiendye.uk/en/proxmox/vm-time-ptp-kvm/","summary":"A virtual machine\u0026rsquo;s clock is built on an assumption virtualisation breaks: that the CPU keeps counting. Why guest time drifts on every hypervisor, what the skew actually breaks — Kerberos, TLS, Ceph, Windows — and how ptp_kvm fixes it properly on QEMU/KVM.","title":"A VM's Clock Cannot Be Trusted — and the ptp_kvm Fix for QEMU/KVM"},{"content":"Three Ways a Drive Can Present Itself A drive has two block sizes, and the difference between them is the whole story.\nThe physical sector is what the medium actually works in — the smallest unit the drive can read or write without doing extra work. The logical block is what the drive tells the host it works in — the unit the host addresses.\nThere are three combinations in the wild.\n512n — native 512. Both sizes are 512 bytes. This is the old world: pre-2010 hard drives, and nowt you would buy new today at any size worth having.\n512e — 512-byte emulation. The physical sector is 4096 bytes, but the drive reports 512-byte logical blocks and translates in firmware. This exists for one reason: compatibility with operating systems, bootloaders and applications that assumed 512 forever. Sensible enough as engineering, and the root of very nearly all the trouble that follows.\n4Kn — native 4K. Both sizes are 4096 bytes. The drive tells the truth, the host addresses it in the unit the medium uses, and no translation layer sits in between.\n512n, 512e and 4Kn: what the host addresses versus what the medium uses upper row = logical, what the host addresses · lower row = physical, what the medium works in 512n 512512 512512 512512 512512 1:1, and honest — but obsolete. Nowt new ships like this. 512e 8 × 512 B logical — what the host is told one 4096 B physical sector — what actually exists firmware translates, every access The host addresses something that is not there. Writes that do not fill a sector cost extra. 4Kn one 4096 B logical block one 4096 B physical sector nothing to translate The drive tells the truth, so nothing downstream can make a decision on a wrong number. 512e is the interesting case: the host is addressing something that does not exist, and the firmware maintains the fiction on every access. What 512e Does On Every Unaligned Write A read is cheap in all three cases. The drive reads the 4K sector and hands back whichever 512-byte slice you asked for.\nWrites are where the fiction gets expensive.\nIf the host writes a single 512-byte logical block, the drive cannot write 512 bytes. The medium has no such unit. So it does this instead:\nRead the whole 4096-byte physical sector. Merge the 512 bytes of new data into it. Write the whole 4096-byte sector back. That is a read-modify-write, and it turns one small write into a read plus a write. On a spinning disk that means waiting for the platter to come round again. A full rotation at 7,200rpm is something like 8ms you did not budget for. On flash the cost is different in kind, and it does not go away when the write finishes. That one has its own section below.\nThe read-modify-write penalty, and what misalignment does to it 4Kn — aligned 4 KB write 4 KB of new data fills the sector 1 write nothing to read first 512e — writing one 512 B logical block 1. read the whole physical sector 4096 B read back 2. merge the 512 B only this much actually changed 3. write the whole sector back 4096 B written 1 read + 1 write a rotation on a disk; a page programmed on flash 512e — misaligned 4 KB write (the expensive one) one 4 KB write, offset by half a sector physical sector n physical sector n+1 2 read-modify-writes on every write, forever Modern partitioning tools align to the physical sector by default, so this is usually inherited from an old install or a cloned image rather than freshly created. It does not fix itself. A 4K-aligned write of 4K needs no read at all. Everything else does — and a write that straddles a sector boundary needs two. Misalignment is the version of this that bites hardest. If a partition starts at an odd 512-byte offset — the classic being the old 63-sector convention — then every 4K filesystem write straddles two physical sectors. That is not one read-modify-write, it is two, on every single write, forever, until somebody repartitions the drive.\nWrite Amplification on Flash On a hard drive the read-modify-write costs you a rotation and then it is over. On flash it costs you drive life, and that is a bill you pay once and keep paying.\nWrite amplification is the ratio of what actually got written to the NAND against what the host asked to write. A ratio of 1.0 would mean the drive wrote exactly what it was given. It is never 1.0, because flash cannot overwrite in place: the drive programs a fresh page, marks the old one stale, and garbage collection later relocates the surviving pages so a whole block can be erased. Those relocations are writes too.\n512e adds an avoidable layer on top of that, and the reason is a detail worth stating clearly: the drive\u0026rsquo;s translation layer maps in units of around 4 KB no matter what logical block size it advertises. A drive reporting 512-byte blocks is still tracking 4 KB units internally.\nSo a 512-byte write from the host becomes, inside the drive: read the 4 KB mapping unit, merge in the 512 bytes, program a new 4 KB unit. 4096 bytes reach the NAND so that 512 bytes could change. Eight times the write, for that write.\nMisalignment is less dramatic per write and worse in total. A 4 KB write that straddles two mapping units dirties both of them, so 8 KB is programmed for 4 KB of data. That is 2× amplification on every single write, permanently, until the partition table is fixed.\nWhy a misaligned write doubles what reaches the NAND reading downwards: what the host wrote → the drive's 4 KB mapping units → what reached the NAND 512e — 4 KB write, misaligned 4 KB from the host unit n unit n+1 4 KB rewritten 4 KB rewritten 8 KB written for 4 KB — WAF 2× on every write, until the partition table is fixed 4Kn — 4 KB write, aligned 4 KB from the host unit n 4 KB written 4 KB for 4 KB — WAF ≈ 1× the floor, before garbage collection Neither side escapes garbage collection: the drive still has to relocate whatever is live in a block before it can erase it, and that work scales with how much was written in the first place. 4Kn does not make amplification go away — it removes the part of it you were paying for nothing. The host asked for the same 4 KB in both cases. On the left it lands across two mapping units, so both are rewritten — and garbage collection will later move whatever is still live in them. The knock-on effects are what make this matter rather than just being untidy:\nMore NAND writes means garbage collection runs more often, and its relocations are themselves amplification. An SLC write cache fills sooner, so the drive drops to its slower steady state earlier and sustained write throughput falls. Program/erase cycles are consumed in proportion to what reached the NAND, not to what the host sent. Double the amplification and you have halved the drive\u0026rsquo;s life for the same workload. Endurance ratings — TBW, DWPD — are quoted against host writes. Amplification eats that margin quietly, and the drive wears out ahead of the warranty arithmetic you did when you bought it. Direct and Synchronous Writes Are the Worst Case Most of the time the page cache hides all of this. The kernel accumulates small writes, merges them, and issues 4 KB or larger IO to the drive, so the read-modify-write never happens.\nTwo flags remove that protection, and applications that care about durability set both.\nO_DIRECT bypasses the page cache. There is now nothing between the application and the drive to coalesce a 512-byte write into something sector-sized.\nO_SYNC or O_DSYNC requires the write to be on stable media before the call returns. That stops the drive absorbing the write into a volatile buffer and combining it with its neighbours later.\nPut them together and each 512-byte write is a complete read-modify-write cycle that has to finish now, on its own, with nothing to amortise it against. That is where amplification on a 512e drive approaches its theoretical 8×, and it is exactly the pattern a database commit log, a ZFS ZIL or a Ceph BlueStore WAL produces.\nPower-loss protection is what rescues this on enterprise hardware. A drive with PLP can acknowledge a synchronous write once it is in a capacitor-backed DRAM buffer, which is durable, and still coalesce internally before programming NAND. A consumer drive without PLP has to reach the flash before it may answer, so it pays the full cost per write. One more reason PLP is not optional for this kind of workload.\nIn Proxmox the cache mode decides which of these you get:\ncache=none is O_DIRECT — the PVE default, and the right choice for all-flash HA. The host page cache is out of the path, so guest write patterns reach the drive as the guest issued them. cache=directsync is O_DIRECT plus O_DSYNC — every write synchronous. It is a niche setting for dedicated database-log drives and unusable for general workloads. On a 512e backing device, cache=directsync with a guest committing in 512-byte records is about the least efficient arrangement available: no host coalescing, no drive coalescing, and a read-modify-write per commit.\nAnd here is the part that makes 4Kn structural rather than merely preferable. O_DIRECT requires the offset, length and buffer to be aligned to the drive\u0026rsquo;s logical block size. On a 512e drive that is 512 bytes, so a 512-byte direct write is legal and the drive quietly pays for it. On 4Kn the logical block is 4096, so the smallest direct write the kernel will accept is 4 KB. The pathological pattern stops being something you have to avoid and becomes something the stack cannot express.\nOverhead on the Medium The second cost is structural, and it is the reason Advanced Format exists at all.\nA sector is not just its data. On a hard drive each one carries a sync mark so the head knows where the sector begins, a gap so consecutive sectors do not run into each other, an address marker, and an ECC field to correct read errors.\nWith 512-byte sectors you pay all of that eight times for every 4K of data. With one 4K sector you pay it once.\nThat has two consequences, and the second matters more than the first:\nSome of the platter that was overhead becomes usable capacity. This was the industry\u0026rsquo;s stated motivation during the Advanced Format transition, in the low single-digit percentages. The ECC field can be much larger for the same total overhead. One strong code protecting 4096 bytes corrects far more than eight weak codes protecting 512 bytes each. As areal density climbed, that stopped being a nicety and became the only way to keep error rates acceptable. Per-sector overhead is paid eight times at 512 bytes and once at 4K 512-byte sectors — the same 4 KB of data 8 sectors × (sync + data + ECC + gap) 8 × the per-sector overhead, and eight small ECC fields One 4096-byte sector — same data 1 × the overhead, and one much larger ECC field sync mark, address marker, inter-sector gap ECC your data Not to scale — the overhead fields are exaggerated so they are visible at all. The recovered capacity was worth low single-digit percent. The stronger ECC is why the industry actually moved. Eight sets of sync marks, gaps and ECC, or one. The recovered space is the small win; the stronger error correction is the reason the industry moved. Flash has no sync marks or rotational gaps, but the same logic applies a level up: NAND is programmed in pages, pages are far larger than 512 bytes, and the drive\u0026rsquo;s mapping tables have an entry per addressable unit. Smaller logical blocks mean more metadata to track the same capacity.\nOverhead in the Host: Commands and Interrupts The third cost is the one people miss, because it is not on the drive at all.\nThe logical block size sets the floor on how small an IO can be. On a 512-byte logical drive, a filesystem or an application is free to issue a 512-byte write, and each one is a complete IO: a command built and submitted, a doorbell write, a completion queue entry, and an interrupt to say it finished.\nEvery one of those has a fixed cost that does not care how much data was involved. Move 4KB as eight 512-byte commands and you pay that cost eight times. Move it as one 4K command and you pay it once.\nThe logical block size sets the floor on IO size, and every command has a fixed cost 512 B logical blocks — an application may issue 512 B writes 4 KB of data becomes 8 commands 512 B512 B 512 B512 B 512 B512 B 512 B512 B 8 × submission + doorbell 8 × completion entry 8 × interrupt opportunity 8 × the fixed per-command cost for exactly the same 4 KB of data 4 KB logical blocks — 4 KB is the floor one 4 KB command 1 × submission, 1 × completion, 1 × interrupt 1 × the fixed cost This only helps where small IO is actually being issued: one command can describe many blocks, so a 1 MB write is one command either way. The same 4 KB of data. The drive is not the bottleneck here — the per-command and per-completion cost in the host is. Two honest qualifications, because this is where the argument is usually overstated.\nFor large IO, the logical block size changes nothing about the command count. A single command can describe many blocks, so a 1MB sequential write is one command whether the blocks are 512 bytes or 4K. The saving is real only where small IO is being issued.\nWhere small IO is being issued, though, the effect is not subtle: each of those eight requests is one the kernel has to build, schedule and complete, each with its own MSI-X interrupt and user-to-kernel transition, and at high queue depths that is how you get an interrupt storm. Inside a VM it is worse again, because every one of those interrupts is also a guest context switch.\nAnd modern NVMe controllers coalesce interrupts, so the interrupt count is not simply the command count. The submission and completion work per command remains, though, and at a few hundred thousand IOPS the per-command CPU cost is a measurable fraction of a core. This is the same arithmetic as posted interrupts in the passthrough series — small fixed costs multiplied by a very large number.\n4Kn, Sector Metadata and Hardware RAID Going to a 4096 + 0 format has a consequence that catches people out, and it lands squarely on hardware RAID.\nSome RAID controllers and storage arrays do not use plain 512 or 4096-byte sectors at all. They format drives to an extended size — 520 or 528 bytes, or the 4K equivalents like 4104, 4160 and 4224 — because those extra bytes per sector are where the controller keeps its own metadata. That is T10-PI/DIF protection information, or vendor integrity data, stored inline with the very data it describes.\nA 4096 + 0 format has nowhere to put it. The sector is data, end to end, and that is the entire point of choosing it.\nSo a controller that wants inline metadata has three options, and none of them are free:\nRefuse the drive. Reformat it back to an extended format, undoing the 4Kn work you just did. Keep its metadata somewhere else on the drive. The third is where flash punishes you. Metadata written separately from the data it describes is a second write, at a different offset, landing in a different mapping unit. Another NAND page programmed for every one you actually meant to write. That is the amplification from the section above, reintroduced deliberately, in order to carry integrity metadata that the drive could have held inline for nowt if you had left it in an extended format.\nYou cannot have both. Either the sector carries the controller\u0026rsquo;s metadata, or it carries only your data.\nWhich Is Why Flash Has Rather Undermined the Case for Hardware RAID The rest of this is judgement rather than mechanism, so take it as that.\nA hardware RAID controller is firmware RAID running on a dedicated processor. The \u0026ldquo;hardware\u0026rdquo; is a CPU, some DRAM and a battery. Not a fundamentally different way of computing parity. What it historically bought you was a battery-backed write cache and parity offload, and on flash both arguments have weakened badly. Enterprise NVMe already has a power-loss-protected cache of its own, and the controller becomes a bandwidth ceiling in front of devices that can each saturate several gigabytes per second.\nIt also costs you things you now actively want:\nDevice state disappears. SMART detail, wear indicators and the vendor logs that let you compute write amplification are all behind an opaque abstraction. No end-to-end checksums. A controller verifies parity, which detects a missing drive, not a wrong answer from a present one. ZFS and Ceph checksum the data itself and can tell you which copy is wrong — silent corruption a controller passes straight through. As such, the controller is solving the wrong problem. Vendor metadata on the drives ties the array to a controller family, which is its own kind of unreliable when the controller is what fails. Parity RAID does its own read-modify-write on partial-stripe writes, stacking on top of everything in the sections above. For Ceph this is not even a preference. Proxmox\u0026rsquo;s own hyper-converged guidance is that disks must be presented in HBA or pass-through mode, not behind a RAID controller, and ZFS wants exactly the same thing for the same reasons.\nSo the arrangement that follows from all of this is: an HBA rather than a RAID controller, drives formatted 4Kn with zero metadata, and redundancy plus checksums done by ZFS or Ceph, which can actually tell you when a drive lied. If something in your estate genuinely needs an extended sector format, that is a deliberate either/or to settle while the drives are still empty. Not something to discover after the OSDs are built.\nWhere 512e Actually Bites It would be dishonest to claim 512e ruins a modern system, because usually it does not.\nA current Linux stack reads the physical sector size, aligns partitions to it — parted and sfdisk both do this by default now — and uses 4K filesystem blocks. In that configuration the host issues 4K-aligned 4K IO, the drive never needs a read-modify-write, and 512e costs you close to nowt.\nThe problems are specific:\nMisaligned partitions, usually inherited from an old install or a cloned image. Two read-modify-writes on every write. ZFS with ashift=9 on a 512e drive, because ZFS believed the reported 512. Every record write becomes a read-modify-write, and you cannot change ashift after the fact — the pool has to be rebuilt. Applications that write 512-byte records with O_DIRECT, bypassing the page cache\u0026rsquo;s coalescing. Some databases and a lot of bespoke software do this. Anything that trusts the logical size to be the real one. That is the actual harm in the emulation: it hands out a number that is wrong, and things downstream make decisions with it. 4Kn removes the whole category. The drive cannot lie about a sector size it does not have.\nIt is worth being straight about the size of the prize, though. Seagate\u0026rsquo;s own guidance is that 4Kn is clearly worth chasing when the stack is fully optimised for 4K and you are counting every IOPS — a tuned all-flash tier, say. Below that, on a correctly aligned modern Linux, the performance difference for aligned IO is often small. The other argument is fleet consistency: a uniformly 4Kn estate has no mixed-format surprises in it, and nobody has to remember which drives lie.\nMost Drives Can Be Converted — If the Vendor Allows It This is the part that gets missed: 512e is often a format setting, not a property of the hardware. A great many enterprise SAS and SATA drives, and most enterprise NVMe, ship reporting 512 bytes and will happily reformat to 4Kn.\nAll of the following destroy every byte on the device. There is no in-place conversion.\nNVMe — nvme-cli Look at what the namespace supports first:\n# Lists each LBA format and marks which one is in use nvme id-ns -H /dev/nvme0n1 | grep -i \u0026#34;lbaf\\|data size\u0026#34; You want a format with Data Size 4096 and Metadata Size 0, marked as best and not currently in use. Then apply it:\n# -l/--lbaf selects the LBA format index from the list above nvme format /dev/nvme0n1 --lbaf=1 --force The metadata size matters as much as the data size. Some factory formats reserve extra bytes per sector — 520, or 4160 — to carry end-to-end T10-PI/DIF protection metadata. If nothing in your stack consumes that, it is padding on every sector, so pick the zero-metadata format and be rid of it. Choosing a metadata-bearing format by accident also gets you a drive that behaves differently from the one you meant to create, and combining a sector-size change with a PI change can force a slow full format rather than a fast one.\nFor a whole host\u0026rsquo;s worth of drives, loop it. This needs shopt -s extglob for the extended globs, and it selects only formats that are 4096/0 and not in use:\nshopt -s extglob for dev in /dev/nvme+([0-9])n+([0-9]); do # Skip anything that is not actually there [ -e \u0026#34;$dev\u0026#34; ] || continue # An LBA format with 4096-byte data, 0-byte metadata, marked Best, not in use lbaf=$(nvme id-ns -H \u0026#34;$dev\u0026#34; \\ | grep -P \u0026#39;(?=.*Metadata Size: 0)(?=.*Data Size: 4096)(?=.*Best)(?!.*in use)\u0026#39; \\ | awk \u0026#39;{found=$3} END {print (found != \u0026#34;\u0026#34; ? found : -1)}\u0026#39;) if [ \u0026#34;$lbaf\u0026#34; != \u0026#34;-1\u0026#34; ]; then echo \u0026#34;Formatting $dev using LBA Format: $lbaf\u0026#34; nvme format --force --lbaf=\u0026#34;$lbaf\u0026#34; \u0026#34;$dev\u0026#34; else echo \u0026#34;Skipping $dev: no matching LBA format found.\u0026#34; fi done Two things in there are doing more work than they look.\nThe glob matches namespaces — nvme0n1, nvme12n3 — and deliberately does not match partitions like nvme0n1p1, because the pattern ends after the digits following the n. That is the difference between reformatting a namespace and doing something unrecoverable to a running system.\nAnd the selection only ever converts a drive that actually offers what you asked for. Anything else falls through to -1 and is skipped, which covers three separate cases:\nThe drive only offers 512. No 4096-byte format exists, so there is nothing to convert to and the loop leaves it alone. It does not try, and it does not fail halfway. The drive is already 4Kn. The 4096/0 format is the one in use, and (?!.*in use) excludes it — so a second run over the same host is a no-op. No needless reformat of every drive. The only 4096 formats carry metadata. A 4096 + 8 format does not satisfy Metadata Size: 0, so the loop will not quietly hand you a T10-PI drive you did not ask for. In other words it fails closed. When it is unsure, it skips.\nOne portability note: those lookaheads need GNU grep\u0026rsquo;s -P (PCRE) mode. On a system where grep is something else, the pattern matches nothing and every drive is skipped. Annoying, but at least it errs in the safe direction.\nRead the loop before you run it, all the same. Where it does match, it reformats without further prompting. nvme format --force does not ask twice. It belongs in provisioning, on a machine whose drives hold nothing, never on a host with a live OSD, pool or VM disk.\nOn FreeBSD the equivalent is nvmecontrol, where -f is the format index:\nnvmecontrol format -f 1 nvme0ns1 SAS and SATA — openSeaChest Seagate\u0026rsquo;s openSeaChest is cross-platform, open source, and works on other vendors\u0026rsquo; drives too.\n# Find the handle openSeaChest_Format --scan # Ask the drive which sector sizes it will accept openSeaChest_Format -d /dev/sg1 --showSupportedFormats # Convert. The confirmation string is deliberately hard to type by accident. openSeaChest_Format -d /dev/sg1 --setSectorSize 4096 \\ --confirm this-will-erase-data-and-may-render-the-drive-inoperable That confirmation phrase is not me being dramatic. It is the literal string the tool requires, and the \u0026ldquo;may render the drive inoperable\u0026rdquo; part is real. A low-level format interrupted by a power cut can leave a drive needing another format before it will work at all.\nUnderneath, the operation differs by transport: SAS and SCSI use Format Unit, SATA uses Set Sector Configuration Ext — the fast-format path — and NVMe uses NVM Format. For a SAS drive you can drive Format Unit directly, and note that this option takes the shorter confirmation string:\nopenSeaChest_Format -d /dev/sg1 --formatUnit 4096 --poll \\ --confirm this-will-erase-data Two different options, two different confirmation strings — get them the wrong way round and the tool refuses.\nWhere the drive supports a fast format, the sector size changes in seconds rather than hours; the drive then does its integrity and background work afterwards, and writing your real data onto it reduces that background time. A full format writes zeroes end to end and can take many hours to days on a large spinning disk.\nopenSeaChest_Format -d /dev/sg1 --setSectorSize 4096 --fastFormat \\ --confirm this-will-erase-data-and-may-render-the-drive-inoperable SCSI — sg_format For anything that speaks SCSI, sg3_utils will do the same job:\n# --size requires --format; expect hours on a large spinning disk sg_format --format --size=4096 /dev/sdb # Fast format where the drive supports it — seconds instead of hours sg_format --format --size=4096 --ffmt=1 /dev/sdb sg_format gives you a 15-second countdown before it commits, which --quick skips. Its documentation also warns of a specific failure worth knowing: if the block-size change succeeds but the format then fails, the drive can end up in a \u0026ldquo;format corrupt\u0026rdquo; state and needs another format to recover.\nBefore You Convert Anything Check the boot path. A 4Kn drive as a boot device needs UEFI and an OS that supports it. Modern Linux is fine. Older Windows is not, and some hardware RAID controllers still refuse 4Kn entirely. Do it before the drive holds anything. Retrofitting means evacuate, convert, restore. Do one, then check. Convert a single drive, confirm the reported sizes, then do the rest. Expect hours on a spinning disk without fast format. Do not start a low-level format on a machine you need back soon. Checking What You Have # LOG-SEC is what the host addresses, PHY-SEC is what the medium uses lsblk -o NAME,MODEL,SIZE,LOG-SEC,PHY-SEC # The same, from sysfs cat /sys/block/sda/queue/logical_block_size cat /sys/block/sda/queue/physical_block_size # SMART states both, and this is the clearest 512e signature there is smartctl -a /dev/sda | grep -i \u0026#34;sector size\u0026#34; 512 bytes logical, 4096 bytes physical is a 512e drive. Matching numbers mean native — 512n if both are 512, 4Kn if both are 4096.\nAnd check the partitions actually line up:\nparted /dev/sda align-check optimal 1 ZFS, Ceph and Virtual Disks ZFS — set ashift=12 explicitly when creating a pool, and do not rely on the drive\u0026rsquo;s reported size, because on 512e it will tell you 9 and be wrong. It cannot be changed later.\nCeph — BlueStore\u0026rsquo;s minimum allocation size should be 4 KB on flash. Modern Ceph defaults to 4096, but older builds defaulted higher — around 16 KB on SSD and 64 KB on HDD — and that suits RBD VM disks badly, because they issue many small random 4 KB writes and a 4 KB write landing in a 16 or 64 KB allocation unit both amplifies the write and wastes the remainder on padding. It is fixed when the OSD is created, so it has to be set before creating or rebuilding:\nceph config set global bluestore_min_alloc_size_ssd 4096 # new or rebuilt OSDs only There is a matching bluestore_min_alloc_size_hdd. Existing OSDs keep whatever they were built with, so changing it means rebuilding them.\nVirtual disks — a guest sees whatever the hypervisor presents, not the underlying drive, so a 4Kn drive under a VM still hands the guest 512-byte blocks unless you say otherwise. Keeping the stack 4K end to end means telling QEMU to present 4K, which in Proxmox is a raw argument line in /etc/pve/qemu-server/\u0026lt;vmid\u0026gt;.conf:\nargs: -global scsi-hd.physical_block_size=4k -global scsi-hd.logical_block_size=4096 That line is what actually forces QEMU into 4Kn for those disks — -global applies it to every scsi-hd device on the VM, so the guest is told 4096 for both logical and physical block size and partitions and aligns accordingly.\nDo not quote the whole string. Proxmox parses args: with Text::ParseWords::shellwords, so this:\nargs: \u0026#34;-global scsi-hd.physical_block_size=4k -global scsi-hd.logical_block_size=4096\u0026#34; collapses into a single argument — the quotes are honoured and stripped, and QEMU is handed one long unparseable option rather than four. Unquoted, the same line splits into -global, scsi-hd.physical_block_size=4k, -global, scsi-hd.logical_block_size=4096, which is what you want. It is an easy mistake to make because quoting is the right instinct on a command line, and the failure — a VM that will not start, complaining about the option — does not point back at the quotes.\nDo this before installing the guest OS. Changing a disk\u0026rsquo;s block size under an already-installed system can leave it unbootable, because the partition layout and bootloader were written for 512-byte sectors. Note also that args: is an expert escape hatch outside the GUI\u0026rsquo;s management, and the interaction with live migration and snapshots is worth re-checking against current Proxmox documentation.\nWhen the Guest Says 512 and the Host Says 4K This is the case worth understanding properly, because it is what you get by default after doing all the work above.\nQEMU presents 512-byte logical blocks to the guest unless told otherwise, whatever the backing device is. So you can convert every drive in the host to 4Kn, and the VMs on top will still be told 512 — and they will believe it.\nThe guest then partitions on 512-byte boundaries because it may, and issues 512-byte IO because it may. But the host device now genuinely has a 4096-byte logical block, and it will not accept a 512-byte write. Something has to reconcile the two, and that something is the host: QEMU reads the surrounding 4 KB, merges the guest\u0026rsquo;s 512 bytes into it, and writes the whole thing back.\nYou have not removed the emulation. You have moved it off the drive\u0026rsquo;s firmware and into your hypervisor, where it costs host CPU and a bounce buffer instead of drive cycles.\nWith cache=none this is not a soft penalty either. O_DIRECT against a 4Kn device requires 4 KB-aligned offsets and lengths, so a sub-4K guest write cannot simply be passed through — the alignment has to be fixed up in QEMU before the IO is issued at all.\nAnd the misalignment case comes back one layer up. A guest partitioned on 512-byte granularity puts its 4 KB filesystem writes at offsets that straddle two host 4 KB blocks, so each one becomes two read-modify-writes on the host. Same failure as a misaligned partition on a bare 512e drive, except now it is happening inside a VM where nobody is looking for it.\nSo the rule is: if the host is 4Kn, present 4K to the guest as well, and do it before the OS goes on.\nThe exception is guest support, which is the whole reason 512e exists:\nLinux guests handle 4Kn without fuss. Windows supports 4Kn for data volumes from Windows 8 and Server 2012 onward, and booting from 4Kn wants UEFI. Anything older — Windows 7 and back — cannot do 4Kn at all. For those, a 512-presenting virtual disk is the price of running them, and the host will do the reconciling. If that matters, keep those guests on storage where it costs you least rather than on your fastest tier. Check what the guest actually ended up with, from inside the guest:\nlsblk -o NAME,LOG-SEC,PHY-SEC 512 there on a 4Kn host means the reconciliation above is happening on every unaligned write.\nReferences OpenZFS — Workload Tuning — ashift, why 2^ashift is the smallest possible IO on a vdev, and the observation that \u0026ldquo;many devices misreport their sector sizes\u0026rdquo; nvme-format man page — --lbaf, --namespace-id, --ses and --force sg_format man page — --size with --format, the fast-format option, and the format-corrupt warning openSeaChest — Seagate\u0026rsquo;s open-source drive utilities, including --setSectorSize and --showSupportedFormats openSeaChest wiki — Format, Fast Format, And Sector Sizes — which transport uses which command, and when fast format applies nvmecontrol(8) — the FreeBSD equivalent Linux block layer documentation — how the kernel models logical and physical block sizes ","permalink":"https://blogs.damiendye.uk/en/proxmox/block-sizes-4kn-512e/","summary":"512e drives present 512-byte sectors they do not have, and the firmware makes up the difference on every unaligned write. What that costs on the medium, in the host and in write amplification — why direct synchronous writes are the worst case — and how to convert a fleet to 4Kn.","title":"4Kn, 512e and 512n — Why Native 4K Wins, and What the Emulation Costs"},{"content":"What a BAR Is, and Why GPUs Outgrew It Every PCIe device exposes one or more Base Address Registers. A BAR tells the system how much address space the device wants, and the firmware maps it into the host\u0026rsquo;s memory map. Once it is mapped, the CPU reaches the device\u0026rsquo;s memory with ordinary loads and stores.\nThe size of a BAR used to be fixed in silicon. The firmware read it at power-on and that was the end of the conversation.\nThat was fine while BARs were small. A network card wants a few tens of kilobytes for its registers. An NVMe controller wants 16KB to 64KB.\nGraphics cards broke the arrangement. A modern card has 8, 12, 16 or 24GB of video memory, and the obvious thing to do is map all of it so the CPU can write anywhere in it. Fixed BARs could not express that, so the convention became a small window — usually 256MB — that the driver repoints over the framebuffer, copying through it in pieces. It works. It is just a daft amount of bookkeeping to reach memory you have already bought and fitted.\nA fixed aperture exposes one slice at a time; Resizable BAR maps the lot Fixed BAR — a 256 MB window 24 GB of VRAM, drawn as slices the window The driver repoints the window and copies through it, one slice at a time. Slices are illustrative, not to scale. Resizable BAR — all of it the same 24 GB, mapped once one BAR covers the lot The CPU writes where it means to. No window to move. The 256MB aperture is less a bandwidth limit than a bookkeeping tax: the driver spends its time moving the window instead of moving data. What Resizable BAR Changes Resizable BAR is a PCIe capability that lets a BAR\u0026rsquo;s size be negotiated instead of fixed. The device advertises the sizes it can support, and the firmware or the operating system programs one of them.\nFor a GPU that means the whole framebuffer can be mapped at once. The driver stops paging through a window and writes where it means to write. The kernel reports the capability as Physical Resizable BAR, and it is vendor-neutral. An AMD card works with it on an Intel platform and vice versa.\nWhat the Three Vendors Actually Do The interesting part is that the vendors do not agree on how much it matters.\nThree vendors, one capability, three different postures Intel Arc / Arc Pro REQUIRED “Required to get a good experience with Arc” Off, the frametime spikes that were already there get bigger. Platforms: 10th Gen Core and newer, most Ryzen 3000, all Ryzen 5000. NVIDIA PER-GAME PROFILE RTX 30 Series onward, March 2021 “A few percent, up to 12%” — and some titles get slower, so the driver enables it only where it tested faster. Needed a VBIOS + SBIOS update. AMD Radeon SMART ACCESS MEMORY Same capability, sold under a brand name Launched with RX 6000 and Ryzen 5000 as a paired feature, but the capability is standard PCIe and works cross-vendor. Patchier on Vega and Polaris. The capability is identical in all three cases — what differs is whether the vendor treats it as a prerequisite, an optimisation to be applied selectively, or a product feature. Same capability, three different postures. Intel treats it as a prerequisite; NVIDIA ships it switched on only for games it has tested; AMD sells it as a feature. Intel Arc and Arc Pro — Intel Calls It Required Intel is the emphatic one. Their guidance says Resizable BAR is \u0026ldquo;required to get a good experience with Intel® Arc™ hardware.\u0026rdquo; That is unusually strong language for a platform feature, and it shows how the Arc driver is built rather than a marketing choice.\nOn the symptom side, Intel describes the effect of turning it off in terms of consistency rather than averages: with \u0026ldquo;ReBAR off\u0026rdquo; you \u0026ldquo;will generally result in spikes that were already there getting bigger.\u0026rdquo; That is a frametime-stability argument, and it is the same shape of problem as ASPM\u0026rsquo;s effect on latency tails. The average hides it.\nIntel lists 10th Generation and newer Intel Core processors, most Ryzen 3000 series and all Ryzen 5000 CPUs as supported platforms, and treats it as a motherboard BIOS feature you have to go and enable.\nThe practical upshot for Arc and Arc Pro is simple. If you have put an Arc card in a machine and performance is disappointing or uneven, check this before you check anything else. As such, it is not a tuning knob on those cards so much as a prerequisite.\nNVIDIA — Ampere Onward, and Only Where It Helps NVIDIA added support for GeForce RTX 30 Series cards and laptops in March 2021. Getting it working needed a supported VBIOS, a compatible CPU and motherboard, a motherboard firmware update, and a current driver. The RTX 3060 shipped with it; the 3060 Ti, 3070, 3080 and 3090 could need a firmware update to get it.\nNVIDIA\u0026rsquo;s own performance framing is refreshingly unglamorous. They found \u0026ldquo;some titles benefit from a few percent, up to 12%\u0026rdquo; while \u0026ldquo;there are also titles that see a decrease in performance.\u0026rdquo;\nTheir response to that is the detail worth knowing: rather than leave it on globally, NVIDIA pre-tests titles and uses per-game profiles to enable Resizable BAR only where it measured as a gain. So on an NVIDIA card, \u0026ldquo;is ReBAR on?\u0026rdquo; has a per-application answer, and a benchmark that shows nowt may simply be a title the driver decided to leave alone.\nAMD — Smart Access Memory Is the Same Thing AMD\u0026rsquo;s branding for it is Smart Access Memory, introduced alongside the Radeon RX 6000 series and Ryzen 5000 CPUs. The marketing implies an AMD-CPU-plus-AMD-GPU pairing, and that is how it was launched, but the capability underneath is the standard PCIe one. Enable Resizable BAR in firmware with a Radeon card in an Intel machine and you get the same feature without the branding.\nRadeon Pro W-series cards support it too. On older architectures — Vega and Polaris — support is patchier, and that is also where virtualisation trouble lives.\nResizable BAR and AI Work This is where the feature gets misunderstood, so it is worth being clear about what it touches.\nResizable BAR changes how the CPU reaches GPU memory. That is the transfer path — staging model weights, pushing batches, reading results back. It does not touch the GPU\u0026rsquo;s own memory bandwidth, and it does not make a matrix multiply faster.\nSo the honest summary is that it affects the loading and feeding of a model, not the arithmetic. For a large model, the host-to-device copy at load time is a genuine cost, and a full-size BAR lets the CPU write straight into device memory instead of shuttling through a 256MB porthole. For the steady state of an inference run — where the weights are already resident and the work is compute-bound. Expect nowt.\nThree specifics are worth knowing.\nArc Pro for AI inherits Intel\u0026rsquo;s verdict. If Intel says the card needs Resizable BAR for a good experience, that applies to a card running oneAPI or PyTorch just as much as one running a game. Arc and Arc Pro cards doing inference should have it enabled, full stop.\nOn NVIDIA, the number you want is BAR1. BAR1 is the host-visible window onto device memory, and it is what host-mapped allocations and GPUDirect RDMA go through:\n# How much device memory is actually host-visible nvidia-smi -q | grep -A3 \u0026#34;BAR1 Memory Usage\u0026#34; A data-centre card is built with a large BAR1 already. On a desktop card, Resizable BAR is what makes that window big rather than tiny. If you are doing RDMA straight from a NIC into GPU memory, this is not a nicety.\nDo not confuse this with ROCm\u0026rsquo;s requirement. AMD\u0026rsquo;s ROCm system requirements call for CPUs supporting PCIe atomics — \u0026ldquo;modern CPUs after the release of 1st generation AMD Zen CPU and Intel™ Haswell\u0026rdquo; — and say nothing about BAR sizing. Those are two different platform requirements that both live in the same BIOS menu, and people conflate them constantly. Check the one you actually need.\nThe failure mode that bites AI builds hardest is not performance at all, and it is covered below: several large-BAR GPUs in one machine can run the address space out.\nWhat It Needs to Work Four things, and they are all firmware-level:\nResizable BAR enabled in the motherboard firmware. Often off by default. Above 4G Decoding enabled. A 24GB BAR cannot fit below the 4GB line, so the platform has to be willing to allocate address space above it. Turn this on even with no GPU present. It costs nothing. UEFI boot, with CSM off. Legacy compatibility mode and large BARs do not mix. Current firmware and drivers, particularly on the boards and cards from the 2020–2021 transition, where support arrived by update rather than at launch. A large BAR only fits above 4GB, and the guest aperture has to be big enough to hold it Where a resized BAR can actually live system RAM 0 32-bit MMIO hole 4 GB 64-bit MMIO high GPU 0 24 GB BAR GPU 1 — 24 GB GPU 2 — 24 GB GPU 3 — 24 GB Crowded, and only 4 GB wide in total A 24 GB BAR cannot go here. A 256 MB one barely could. Usable only with Above 4G Decoding A firmware setting. Off by default on some boards, and without it a large BAR is never allocated at all. Every card needs its own room up here Four 24 GB cards want 96 GB of 64-bit MMIO allocated, plus everything else on the bus. Not every consumer board's firmware will do it. The tell: fine with two cards, one refuses to initialise with four, and dmesg reports no space for the BAR. Not to scale: the 64-bit region is vastly larger than the 4 GB below it, which is rather the point. Above 4G Decoding is what makes the upper region usable at all — and every card needs its own room up there, which is where multi-GPU AI builds hit the wall. How to Check on Linux Whether the card has the capability, and which sizes it offers:\n# Substitute your card\u0026#39;s address from lspci lspci -vvs 0000:XX:00.0 | grep -A6 \u0026#34;Physical Resizable BAR\u0026#34; The kernel also exposes this in sysfs, one file per resizable BAR:\ncat /sys/bus/pci/devices/0000:XX:00.0/resource1_resize That value is a bitmap of supported sizes, not a size. Bit 0 means 1MB, bit 1 means 2MB, bit 2 means 4MB, and the size for a given bit is 2 ^ (bit + 20). So 00000000000001c0 has bits 6, 7 and 8 set, meaning the BAR can be 64MB, 128MB or 256MB.\nTo see what is actually in force, read the assigned regions:\nlspci -vvs 0000:XX:00.0 | grep -i Region A card running with its full framebuffer mapped shows a region matching its VRAM size rather than a 256MB one.\nResizing by Hand You can write the bit position yourself:\n# bit 7 -\u0026gt; 2 ^ (7 + 20) = 128MB echo 7 \u0026gt; /sys/bus/pci/devices/0000:XX:00.0/resource1_resize The conditions attached are strict, and worth reading before trying it on a machine you care about. Every driver must be unbound from the device first. Peer devices under the same parent bridge may need soft-removing. On a VGA device, writing a resize value tears down the low-level console drivers. Anything holding the resourceN sysfs files open has to let go.\nThe kernel documentation is also blunt about the result: success is not guaranteed. The resize fails if there is no address space to place the larger BAR. Which takes you straight back to Above 4G Decoding.\nWhen It Goes Wrong dmesg says it cannot assign the BAR. Messages in the shape of BAR 0: no space for [mem size ...] mean allocation failed, not that the card is faulty. Above 4G Decoding is the first thing to check.\nSeveral GPUs and one of them will not initialise. This is the multi-GPU and AI-rig failure. Four cards with 24GB BARs need 96GB of 64-bit MMIO space allocated, plus everything else, and not every consumer board\u0026rsquo;s firmware will do it. The symptom is that the machine is fine with two cards and falls over with four.\nPerformance went down. On NVIDIA that may be the driver\u0026rsquo;s own conclusion, given they enable it per title exactly because some workloads regress. On the others, measure both ways rather than assuming.\nNothing changed at all. The most likely outcome for a workload that was never limited by the CPU\u0026rsquo;s window into VRAM.\nPassing a Large-BAR GPU Through to a VM Worth flagging because it surprises people: a guest does not inherit the host\u0026rsquo;s memory map. The VM\u0026rsquo;s firmware builds its own 64-bit MMIO aperture, and OVMF\u0026rsquo;s default is far smaller than a modern GPU needs, so the card either fails to initialise or falls back to a small BAR. That is a Proxmox and QEMU topic rather than a GPU one, and it lives in the IOMMU tax article alongside the machine-type requirements from Always Use Q35, Not i440fx.\nReferences Intel — Resizable BAR and Intel Arc Graphics — Intel\u0026rsquo;s own statement that it is required for a good experience on Arc, and the supported platform list NVIDIA — Resizable BAR support for GeForce RTX 30 Series — the requirements, the measured range, and the per-game profile approach AMD — Smart Access Memory — AMD\u0026rsquo;s branding and platform pairing Linux kernel sysfs-bus-pci ABI — resourceN_resize — the bitmap, the 2 ^ (bit + 20) sizing, and the unbind conditions ROCm system requirements — the PCIe atomics requirement that gets confused with this one ","permalink":"https://blogs.damiendye.uk/en/proxmox/pcie-resizable-bar/","summary":"Resizable BAR lets the CPU map a GPU\u0026rsquo;s whole framebuffer instead of peering at it through a 256MB window. Intel calls it required for Arc, NVIDIA enables it per game, AMD sells it as Smart Access Memory — and for AI work it changes the transfer, not the maths.","title":"PCIe Resizable BAR and Modern GPUs — Intel Arc, NVIDIA and AMD"},{"content":"The Problem This came up during a customer engagement where we were designing a Proxmox VE deployment with NVMe passthrough for a latency-sensitive workload. The customer had done their own benchmarking before the call. On the host, fio against the NVMe drive reported 700K random read IOPS with sub-10µs completion latency. Inside the VM, using the same drive with the same test, they were getting roughly half that.\nThey\u0026rsquo;d already checked the obvious things. The drive hadn\u0026rsquo;t changed. The firmware hadn\u0026rsquo;t changed. The PCIe slot hadn\u0026rsquo;t moved. They were starting to wonder whether passthrough was the wrong approach entirely.\nIt wasn\u0026rsquo;t. What they were seeing is the IOMMU tax. It catches folk out because nobody tells you about it before you\u0026rsquo;ve committed to the passthrough design. The good news is that most of the overhead is recoverable once you understand where it comes from.\nDeep Dive The Problems That Had to Be Solved Giving a VM direct access to a physical PCIe device sounds straightforward. In practice, it\u0026rsquo;s one of the harder problems in systems virtualisation. A few things that work automatically on bare metal become dangerous when a device is shared between a host and a guest.\nDMA Isolation This is the fundamental problem.\nPCIe devices don\u0026rsquo;t go through the CPU to read and write memory. They use Direct Memory Access. They write straight to physical RAM addresses. On bare metal, that\u0026rsquo;s fine. The device and the OS trust each other.\nUnder virtualisation, the guest VM has its own view of physical memory. The addresses the guest driver gives to the NVMe controller are guest physical addresses. They don\u0026rsquo;t map to the same locations in host RAM. If the device uses them directly, it reads and writes the wrong memory. That corrupts the host, other VMs, or both.\nWorse, a malicious or buggy guest driver could deliberately program the device to DMA into any part of host memory. That\u0026rsquo;s effectively root access to the entire machine without ever exploiting a hypervisor bug.\nThe solution is the IOMMU — a hardware translation unit (Intel VT-d, AMD-Vi) that sits between every PCIe device and main memory. It maintains its own page tables, separate from the CPU\u0026rsquo;s. Every DMA request from the device passes through the IOMMU, which translates guest physical addresses to host physical addresses and blocks any access outside the guest\u0026rsquo;s allocated memory regions.\nWithout the IOMMU, safe passthrough is impossible. With it, the device is contained.\nDevice Grouping The IOMMU doesn\u0026rsquo;t isolate individual devices. It isolates groups.\nThe PCIe specification defines Access Control Services (ACS) which govern whether devices on the same bus can talk to each other directly — peer-to-peer DMA — without going through the root complex where the IOMMU sits. If two devices share a PCIe switch that doesn\u0026rsquo;t enforce ACS, one device can DMA into the other\u0026rsquo;s memory space, bypassing the IOMMU entirely.\nThe kernel groups devices that can potentially reach each other without IOMMU enforcement into a single IOMMU group. If your NVMe controller shares a group with another device, passing through just the NVMe breaks the isolation model. The other device in the group could still be used as a side channel around the IOMMU.\nServer-grade hardware with proper ACS support on every bridge and switch typically gives each device its own group. Consumer and workstation boards often lump multiple devices together because the PCIe root complex doesn\u0026rsquo;t implement ACS on every port.\nProxmox carries a kernel patch — pcie_acs_override — that tells the kernel to treat every device as isolated regardless of hardware ACS support. It works in practice, but it\u0026rsquo;s lying to the kernel about the hardware topology. On a production system, clean groups backed by actual hardware ACS are always preferable.\nInterrupt Delivery On bare metal, when an NVMe controller completes an IO operation, it fires an MSI-X interrupt directly to the CPU. The CPU handles it in a few hundred nanoseconds.\nUnder virtualisation, that interrupt has to reach the guest, not the host. The naive approach is to trap every interrupt in the hypervisor, trigger a VM exit, inject the interrupt into the guest, and resume. That works, but each VM exit costs 5–20µs. At high IOPS — hundreds of thousands of interrupts per second — the overhead is substantial.\nThe hardware solution is posted interrupts. Intel\u0026rsquo;s APICv and AMD\u0026rsquo;s AVIC allow the IOMMU to write the interrupt directly into the guest\u0026rsquo;s virtual APIC page without causing a VM exit at all. The guest sees the interrupt as if it came from bare-metal hardware. The overhead drops to a few hundred nanoseconds.\nNot all platforms support posted interrupts. Older CPUs, some workstation chipsets, and some BIOS versions don\u0026rsquo;t expose the capability. When they\u0026rsquo;re absent, every interrupt goes through the slow path, and there\u0026rsquo;s no software workaround.\nDevice Reset When a VM shuts down or crashes, the passed-through device needs to return to a clean, known state. Otherwise it can\u0026rsquo;t be re-assigned to another VM or reclaimed by the host.\nOn bare metal, the OS does an orderly shutdown of the device driver. Under passthrough, the guest might crash, the user might force-stop the VM, or the hypervisor might kill the process. The device could be mid-transfer with DMA operations in flight.\nPCIe defines Function Level Reset (FLR) for this — a way to reset a single device function without affecting the rest of the bus. NVMe controllers generally support FLR and handle it well. GPUs are notoriously bad at it, but that\u0026rsquo;s a different article.\nIf FLR isn\u0026rsquo;t supported, the fallback is a secondary bus reset, which resets everything behind that PCIe bridge. If the bridge has other devices on it, they all get reset too. In the worst case, a full host reboot is the only way to reclaim the device.\nAddress Translation Overhead The IOMMU solves the safety problem. As such, it is not optional. But it introduces a performance one.\nEvery DMA operation now passes through an extra level of address translation. The IOMMU has its own TLB — the IOTLB — and when it hits, the overhead is small. When it misses, the IOMMU has to walk its page tables, and that adds real latency to every affected IO operation.\nThis is the IOMMU tax. The rest of this article is about understanding where it comes from and how to minimise it.\nEvery DMA goes through the IOMMU: a hit is cheap, a miss walks the page tables Guest VM NVMe driver hands out guest physical addresses NVMe controller writes memory directly — no CPU DMA IOMMU IOTLB its own page tables translate, then permit or refuse Host memory this VM's pages translation lands here host and other VMs unreachable by design Without the IOMMU, the controller would write guest addresses straight into host RAM and corrupt whatever lives there. IOTLB hit — the translation is already cached and costs almost nothing. IOTLB miss — the IOMMU walks its page tables, and that latency lands on this IO. This is the tax. Nothing reaches memory without passing through here. That is the safety guarantee, and the translation it performs is the cost. BAR Mapping and Address Space Every PCIe device exposes one or more Base Address Registers (BARs) that map the device\u0026rsquo;s internal registers and memory into the host\u0026rsquo;s MMIO address space. The host CPU accesses the device through these mappings. For passthrough to work, the hypervisor has to present these mappings correctly to the guest.\nTraditionally, BAR sizes were fixed at boot by the BIOS and fitted within the legacy 32-bit MMIO window below 4GB. That worked when BARs were small. Modern GPUs have changed the picture. A 24GB framebuffer needs a 24GB BAR, which doesn\u0026rsquo;t fit in a 32-bit address space.\nResizable BAR (ReBAR) — also marketed as AMD Smart Access Memory (SAM) — is a PCIe capability that allows the BAR size to be renegotiated after boot. For GPUs, this is a significant feature. Instead of accessing the framebuffer through a 256MB window and paging through it in chunks, the host maps the entire VRAM at once.\nFor NVMe, the direct impact is smaller. NVMe controller BARs are typically 16KB to 64KB for the controller register set (BAR0). The NVMe specification defines a Controller Memory Buffer (CMB) that can expose a larger BAR for host-resident submission queues, but most drives don\u0026rsquo;t implement it. ReBAR doesn\u0026rsquo;t change NVMe throughput the way it does for GPUs.\nThe reason it matters in an NVMe passthrough context is the shared PCIe environment. If you\u0026rsquo;re passing through an NVMe drive alongside a GPU on the same host, the GPU\u0026rsquo;s resized BAR needs address space above the 4GB boundary. The BIOS, the IOMMU, and the virtual PCIe topology all need to accommodate that. Getting the address space allocation wrong means devices don\u0026rsquo;t initialise, and the NVMe passthrough fails alongside everything else.\nPower Loss Exposure On a virtual disk, the hypervisor and storage layer handle write ordering and crash consistency. With passthrough, the guest is talking directly to flash. If the host loses power mid-write, whatever the NVMe controller\u0026rsquo;s firmware does — or doesn\u0026rsquo;t do — with its write cache determines whether you lose data.\nEnterprise NVMe drives carry power loss protection (PLP) capacitors that flush the write cache safely during a power failure. Consumer drives without PLP may not. With passthrough, there\u0026rsquo;s no hypervisor safety net between the guest and the hardware.\nFor a production workload, an enterprise drive with PLP isn\u0026rsquo;t optional.\nOperational Trade-offs Passthrough also removes capabilities that virtual disks provide.\nA passed-through device is physically bolted to a specific host. The VM cannot be live-migrated while the device is attached. In a Proxmox cluster with HA, a node failure means the VM goes down and cold-starts on another node. There\u0026rsquo;s no seamless failover.\nThe drive is also invisible to vzdump and Proxmox Backup Server. It won\u0026rsquo;t be included in VM snapshots or scheduled backups. A separate backup strategy — guest-level, filesystem-level, or application-level — needs to be in place before the workload goes live.\nHow VFIO Passthrough Actually Works When you pass a PCIe device through to a VM, the hypervisor hands the guest direct control of the device\u0026rsquo;s MMIO registers. The guest driver talks to the NVMe controller as if it were running on bare metal. That part is near-native. MMIO register access goes through Extended Page Tables (EPT on Intel, NPT on AMD) and typically completes without a VM exit.\nThe DMA path is where the cost appears. Every DMA operation goes through the IOMMU for address translation, and as described above, that translation has a price. Especially on IOTLB misses.\nWhy the Benchmarks Look Worse Than Reality Here\u0026rsquo;s where most people go wrong with their testing.\nA Proxmox forum thread that prompted this article had users running fio with iodepth=1. At that queue depth, fio submits one IO, waits for it to complete, then submits the next. The test is just measuring per-IO latency. Every microsecond of IOMMU overhead shows up in full.\nThe numbers from that thread tell the story clearly. Bare metal completion latency averaged around 10µs. Inside the VM, it averaged around 28µs. That extra ~18µs per IO is the IOMMU translation overhead. At iodepth=1, it directly halves throughput because throughput equals 1 / latency when there\u0026rsquo;s only one IO in flight.\nBump the queue depth to 32 or 64 — which is how NVMe drives are designed to operate — and the picture changes. With multiple IOs in flight, the IOMMU overhead is amortised across all of them. The controller processes completions while new translations are happening. Throughput recovers to within a few percent of bare metal.\nThe practical takeaway is this. If your workload runs at queue depths above 4, the IOMMU throughput penalty is likely negligible. That covers most database, virtualisation, and storage workloads. If your workload is latency-sensitive at low queue depths — certain real-time applications, synchronous metadata operations — you\u0026rsquo;ll feel it.\nThe IOMMU penalty is a queue-depth artefact more than a throughput ceiling VM throughput as a share of bare metal realistic workloads live in here 0% 25% 50% 75% 100% the benchmark everyone runs iodepth=1 measures pure per-IO latency, so the 10 µs against 28 µs shows up in full within a few percent 1 2 4 8 16 32 64 queue depth (fio iodepth) Illustrative, from the article's own 10 µs / 28 µs figures. Same hardware, same test, same overhead — only the queue depth changes. The overhead is the same at every point on this curve. All that changes is how many IOs are in flight to amortise it over. Run your benchmarks at realistic queue depths before concluding passthrough is too slow:\n# Bare metal baseline — run on the host before binding to vfio-pci fio --name=randread --ioengine=libaio --direct=1 --bs=4k \\ --iodepth=32 --numjobs=4 --rw=randread --size=1G \\ --filename=/dev/nvme0n1 --runtime=30 --time_based \\ --group_reporting # Same test inside the VM after passthrough fio --name=randread --ioengine=libaio --direct=1 --bs=4k \\ --iodepth=32 --numjobs=4 --rw=randread --size=1G \\ --filename=/dev/nvme0n1 --runtime=30 --time_based \\ --group_reporting Compare the clat (completion latency) percentiles and the IOPS figures. At iodepth=32 with four jobs, the gap should be single-digit percentage points, not 50%.\nNUMA Alignment This has its own article. See NUMA Alignment on Proxmox VE — Why It Matters and How to Get It Right.\nThe short version: on multi-socket systems, every PCIe device is wired to a specific socket. If the NVMe drive is on NUMA node 1 and the VM\u0026rsquo;s vCPUs are pinned to node 0, every DMA completion crosses the inter-socket link. That adds 50–100ns per operation. At high IOPS, the throughput difference between aligned and misaligned NUMA is 20–30%.\nCheck with cat /sys/bus/pci/devices/0000:XX:00.0/numa_node, then pin the VM\u0026rsquo;s vCPUs to cores on the same node with the affinity parameter in the VM config. On single-socket systems, this isn\u0026rsquo;t a concern.\nPCIe Active State Power Management (ASPM) This has its own article. See PCIe ASPM and Why You Should Disable It for Passthrough.\nThe short version: ASPM allows PCIe links to enter low-power states when idle. Under passthrough, the host still controls the physical link but the guest owns the device. When the guest submits IO and the link is asleep, the wake-up time adds latency. The symptom is a wide spread in your clat percentiles. The p99 might be 5–10x higher than the average while the mean looks fine.\nDisable it on the host with pcie_aspm=off in the kernel command line. Also add disable_idle_d3=1 to the vfio-pci module options if your NVMe controller has power state recovery issues. The Samsung 990 EVO Plus is a known offender.\nMaxPayloadSize (MPS) This has its own article. See PCIe MaxPayloadSize — A Free Performance Win for Passthrough.\nThe short version: PCIe devices transfer data in Transaction Layer Packets. QEMU\u0026rsquo;s virtual root complex defaults to a 128-byte maximum payload. Most devices support 256 or 512 bytes. Adding pci=pcie_bus_perf to the host kernel command line sets MPS to the maximum each device\u0026rsquo;s parent bus allows. It\u0026rsquo;s a small throughput improvement — low single-digit percentage — but it\u0026rsquo;s free with no downside.\nResizable BAR and MMIO Aperture The BAR mapping problem described earlier has practical steps on both the BIOS and VM side.\nFirst, enable Above 4G Decoding in the BIOS. This allows BARs to be mapped into address space above the 4GB boundary, which is required for any device with large BARs. Enable it even for NVMe-only passthrough. It has no downside and avoids problems if you add a GPU or other large-BAR device later.\nIf ReBAR is available in the BIOS, enable that too. It won\u0026rsquo;t affect NVMe performance directly, but it allows GPUs on the same host to use their full framebuffer mapping.\nOn the VM side, QEMU\u0026rsquo;s virtual Q35 root complex needs a large enough 64-bit MMIO window for the guest to see resized BARs. By default, OVMF allocates a relatively small window. For NVMe-only passthrough this is fine. NVMe BARs fit comfortably. But if the VM has both an NVMe and a GPU passed through, increase the MMIO aperture:\nargs: -fw_cfg name=opt/ovmf/X-PciMmio64Mb,string=65536 This tells OVMF to allocate 64GB of 64-bit MMIO space, enough for most GPU framebuffers alongside the NVMe controller\u0026rsquo;s small BAR.\nQEMU\u0026rsquo;s ReBAR support has been improving but is still not seamless. Some AMD GPUs (Vega and newer) trigger driver errors (Code 43 in Windows) with ReBAR enabled under QEMU. If you hit that, disable ReBAR in the BIOS as a first step. NVMe passthrough won\u0026rsquo;t be affected either way.\nFor what Resizable BAR does on the GPU side of the fence, and why Intel treats it as mandatory on Arc while NVIDIA enables it per game, see PCIe Resizable BAR and Modern GPUs.\nInterrupt Handling The interrupt delivery problem described earlier has a practical tuning step. Posted interrupts (APICv on Intel, AVIC on AMD) may not be enabled by default.\nCheck whether they\u0026rsquo;re active:\n# Intel — look for \u0026#34;Posted-Interrupts\u0026#34; in dmesg dmesg | grep -i \u0026#34;posted\u0026#34; # AMD — check AVIC support dmesg | grep -i \u0026#34;avic\u0026#34; Trapped interrupt delivery versus posted interrupts No posted interrupts — the completion takes the long way round NVMeraises MSI-X hypervisortraps it injectinto the guest guest resumeshandles it VM exit VM entry 5–20 µs per interrupt At a few hundred thousand IOPS that is not a rounding error — it is the dominant cost. Posted interrupts (Intel APICv / AMD AVIC) — straight in NVMeraises MSI-X IOMMUwrites it directly guest's virtual APIC page the guest just sees an interrupt no exit no exit ~100s of ns There is no software workaround: posted interrupts are a hardware capability. But they are not always enabled by default, so check for them before concluding the slow path is unavoidable. Four steps and two VM exits, or one write into the guest\u0026rsquo;s APIC page. At high IOPS the difference stops being academic. On AMD EPYC systems, enable AVIC in the KVM module if it isn\u0026rsquo;t on by default:\n# /etc/modprobe.d/kvm.conf options kvm_amd avic=1 On Intel systems, APICv with posted interrupts is typically enabled automatically when VT-d is active.\nIf your platform doesn\u0026rsquo;t support posted interrupts, there\u0026rsquo;s no software workaround. It\u0026rsquo;s a hardware capability. But it\u0026rsquo;s worth verifying it\u0026rsquo;s actually turned on before assuming the slow path is unavoidable.\nInterrupt Affinity and Queue Alignment NVMe controllers use multiple submission and completion queue pairs. Typically one per CPU core. When the VM\u0026rsquo;s vCPUs don\u0026rsquo;t align with the physical cores handling the NVMe interrupts, completions have to cross cores via inter-processor interrupts. That adds latency.\nInside the guest, check how many IO queues the NVMe driver has created and how they\u0026rsquo;re mapped:\n# List NVMe IO queues cat /proc/interrupts | grep nvme # Check affinity for irq in $(grep nvme /proc/interrupts | awk \u0026#39;{print $1}\u0026#39; | tr -d \u0026#39;:\u0026#39;); do echo \u0026#34;IRQ $irq: $(cat /proc/irq/$irq/smp_affinity_list)\u0026#34; done Ideally, each NVMe IO queue\u0026rsquo;s interrupt should be affinitised to the vCPU that submits to that queue. Most modern NVMe drivers handle this automatically. But it\u0026rsquo;s worth verifying, especially if you\u0026rsquo;ve manually pinned vCPUs or reduced the vCPU count below the controller\u0026rsquo;s queue count.\nAlways Use Q35, Not i440fx This deserves its own article. See Always Use Q35, Not i440fx.\nThe short version: i440fx presents a flat legacy PCI bus. Q35 presents a proper PCIe root complex. Passthrough devices on i440fx appear as legacy PCI, which breaks MSI-X multi-queue interrupt delivery. NVMe controllers need MSI-X for their queue-per-core architecture. Without it, all IO completions funnel through a single interrupt and you get a bottleneck at high IOPS that no amount of kernel tuning will fix.\nIn Proxmox 8.x and newer, Q35 is the default for new VMs. If you\u0026rsquo;re doing passthrough on an older VM that\u0026rsquo;s still i440fx, switch it. RHEL 10 has formally deprecated i440fx, and the wider KVM ecosystem is following.\nPutting It All Together Here\u0026rsquo;s a summary of the tuning steps in order of impact.\nThe tuning steps in order of impact, and what each one actually changes in order of impact what it changes 1 NUMA alignment 20–30% of throughput on a multi-socket box. Nowt else on this list comes close. THROUGHPUT 2 ASPM off Removes wake-up latency from the tail. The average barely moves; p99 does. LATENCY JITTER 3 Realistic queue depths Changes nothing on the machine. Stops you drawing the wrong conclusion from iodepth=1. THE MEASUREMENT 4 Posted interrupts (APICv / AVIC) Microseconds to nanoseconds per interrupt — but only if the platform has it. Verify, don't assume. PER-INTERRUPT COST 5 MaxPayloadSize — pci=pcie_bus_perf Low single-digit percent. Free, no downside, so set it — but do not expect to see it. WIRE OVERHEAD 6 vfio-pci disable_idle_d3 Keeps a fussy controller from dying in D3. Buys reliability, not speed. RELIABILITY Ranked, deliberately not drawn as bars: these do not share a unit, so a bar chart would invite a comparison that is not there. Work down this list, not across it. The first entry is worth more than the rest combined on a multi-socket machine. NUMA alignment — make sure the NVMe drive and the VM\u0026rsquo;s vCPUs are on the same NUMA node. This alone can account for a 20–30% throughput difference on multi-socket systems.\nASPM off — add pcie_aspm=off to the host kernel command line. Eliminates latency jitter from PCIe link power state transitions.\nRealistic queue depths — test at iodepth=32 or higher, not iodepth=1. The IOMMU overhead that dominates at low queue depths is amortised at realistic depths.\nPosted interrupts — verify APICv (Intel) or AVIC (AMD) is active. Reduces per-interrupt overhead from microseconds to nanoseconds.\nMPS optimisation — add pci=pcie_bus_perf to the host kernel command line. Sets MaxPayloadSize to the maximum the topology supports.\nvfio-pci power management — add disable_idle_d3=1 if your NVMe controller has power state issues under passthrough.\nA combined host kernel command line for a Proxmox node doing NVMe passthrough on an AMD EPYC system would look summat like:\nGRUB_CMDLINE_LINUX_DEFAULT=\u0026#34;quiet amd_iommu=on iommu=pt pcie_aspm=off pci=pcie_bus_perf\u0026#34; For Intel:\nGRUB_CMDLINE_LINUX_DEFAULT=\u0026#34;quiet intel_iommu=on iommu=pt pcie_aspm=off pci=pcie_bus_perf\u0026#34; When Passthrough Isn\u0026rsquo;t Worth It Before going down this path, it\u0026rsquo;s worth asking whether you actually need NVMe passthrough at all.\nVirtIO-SCSI and VirtIO-BLK with an NVMe-backed virtual disk are already very efficient. The overhead compared to passthrough is typically 5–10% on latency. The throughput difference is negligible for most workloads.\nPassthrough makes sense when you need the guest OS to manage the device directly. That includes SMART monitoring, firmware updates, TRIM/discard control, and specific NVMe features like reservations. It also makes sense for latency-sensitive workloads where even a few microseconds matter. Certain database engines and real-time data ingest fall into that category.\nFor everything else, the operational trade-offs covered earlier — loss of live migration, loss of snapshots and backup integration — usually outweigh the small performance gain.\nThere\u0026rsquo;s nowt clever about choosing the harder path when the easier one does the job.\nVerifying Your Changes After applying the tuning steps, verify everything is working as expected:\n# Host side — confirm IOMMU is in passthrough mode dmesg | grep -i iommu # Confirm ASPM is disabled lspci -vv | grep -i \u0026#34;ASPM Disabled\u0026#34; # Check MPS on the NVMe controller lspci -vv -s XX:00.0 | grep MaxPayload # Inside the VM — run the fio comparison fio --name=randread --ioengine=libaio --direct=1 --bs=4k \\ --iodepth=32 --numjobs=4 --rw=randread --size=1G \\ --filename=/dev/nvme0n1 --runtime=30 --time_based \\ --group_reporting Compare the VM results against your earlier bare-metal baseline. At iodepth=32, you should see throughput within 5% of bare metal. Completion latency averages should be within 10–15µs of the host figures. If the gap is still large, check NUMA alignment first. It\u0026rsquo;s the most commonly overlooked factor. It also costs nothing but a config change, which makes it the best sort of problem to be left with.\nReferences Linux kernel PCI documentation — MPS and MRRS tuning options — the authoritative source for pcie_bus_perf, pcie_bus_safe, and related parameters Proxmox VE Administration Guide — PCI(e) Passthrough — official Proxmox documentation on VFIO device passthrough configuration Proxmox Forum — NVMe Passthrough Performance — the community discussion that prompted this article ","permalink":"https://blogs.damiendye.uk/en/proxmox/pcie-passthrough-performance-the-iommu-tax/","summary":"Why PCIe devices lose throughput when passed through to a VM via VFIO, and the practical tuning steps that claw most of it back.","title":"PCIe Passthrough Performance on Proxmox VE — The IOMMU Tax and How to Minimise It"},{"content":"What NUMA Is NUMA stands for Non-Uniform Memory Access. On a single-socket system, every CPU core accesses all of the system\u0026rsquo;s RAM through the same memory controller. Access time is the same no matter which core makes the request or where the data sits in physical memory.\nOn a multi-socket system, each CPU socket has its own memory controller and its own bank of RAM. A core on socket 0 can access the RAM attached to socket 0 quickly — that\u0026rsquo;s local memory. It can also access the RAM attached to socket 1, but that request has to cross the inter-socket link (Intel UPI, AMD Infinity Fabric). That\u0026rsquo;s remote memory, and it\u0026rsquo;s slower.\nThe kernel calls each socket-and-its-local-memory a NUMA node. A dual-socket AMD EPYC system has at least two NUMA nodes. Some EPYC processors expose four NUMA nodes per socket (one per CCD), giving eight nodes on a dual-socket board.\nThe performance difference between local and remote memory access isn\u0026rsquo;t subtle. Local access is around 80–100ns. Remote access is around 130–200ns. That\u0026rsquo;s a 50–100% penalty per memory operation. On its own that is a very small number. Paid on every memory access for the life of the VM, it stops being small.\nWhy It Matters for Virtualisation When Proxmox creates a VM, it allocates vCPUs and RAM. By default, those vCPUs can be scheduled on any physical core on any socket. The VM\u0026rsquo;s RAM can be allocated from any NUMA node\u0026rsquo;s memory pool.\nIf the scheduler puts a vCPU on socket 0 and the VM\u0026rsquo;s RAM is on socket 1, every memory access that vCPU makes crosses the inter-socket link. If the vCPUs bounce between sockets — which they will if they\u0026rsquo;re not pinned — the memory access pattern becomes a mess. Some accesses are local, some are remote, and the VM\u0026rsquo;s performance jitters to match.\nFor general workloads — a web server, a file server, a desktop VM — this is often liveable. The overhead is there but it\u0026rsquo;s spread across many operations and doesn\u0026rsquo;t dominate.\nFor IO-intensive workloads — databases, storage servers, anything doing heavy disk or network IO — the penalty compounds. Every DMA completion, every interrupt delivery, every buffer copy involves memory access. If those accesses are crossing sockets, the overhead adds up fast.\nWhy It Matters for Passthrough PCIe devices are physically wired to a specific CPU socket. Each socket has its own PCIe root complex. The NVMe drive in slot 3 might be on socket 0\u0026rsquo;s PCIe lanes. The GPU in slot 5 might be on socket 1\u0026rsquo;s.\nWhen a device performs DMA, the data goes into the memory attached to whatever NUMA node the IOMMU maps it to. If the VM\u0026rsquo;s RAM is allocated from the device\u0026rsquo;s local node, the DMA write goes straight to local memory. If the RAM is on the other node, every DMA operation crosses the inter-socket link.\nFor an NVMe drive doing hundreds of thousands of IOPS, that 50–100ns penalty per operation adds up. At iodepth=32 with 4KB random reads, the throughput difference between aligned and misaligned NUMA can be 20–30%. That\u0026rsquo;s before you\u0026rsquo;ve looked at IOMMU overhead, ASPM, MPS, or owt else.\nMisaligned NUMA sends every DMA across the inter-socket link; aligned keeps it local Misaligned — the VM is on node 0, the drive is on node 1 NUMA node 0 cores 0–15 RAM 128 GB VM: vCPUs pinned 0–15, RAM allocated here NUMA node 1 cores 16–31 RAM 128 GB NVMe — root complex 1 UPI / IF every DMA crosses the link — 130–200 ns Aligned — vCPUs, RAM and the drive all on node 1 NUMA node 0 cores 0–15 RAM 128 GB free for other VMs NUMA node 1 cores 16–31 RAM 128 GB NVMe — root complex 1 VM: affinity 16-31, numa0 hostnodes=1, policy=bind idle 80–100 ns At iodepth=32 with 4 KB random reads the gap between these two is 20–30% of throughput — before IOMMU overhead, ASPM or MPS enter the picture. On a single-socket system there is no second node and no inter-socket link, so none of this applies. The same hardware either way. The only difference is which node the VM was pinned to — and whether the drive\u0026rsquo;s DMA has to cross the link to reach the VM\u0026rsquo;s memory. The same applies to network cards, GPUs, and any other passed-through device. The device\u0026rsquo;s DMA traffic should land in local memory, and the vCPUs processing that traffic should be on the same node.\nHow to Check Your Topology Find Which NUMA Node a Device Is On # Replace 0000:XX:00.0 with your device\u0026#39;s PCI address from lspci cat /sys/bus/pci/devices/0000:XX:00.0/numa_node This returns the NUMA node number. If it returns -1, the kernel couldn\u0026rsquo;t determine the node. That sometimes happens with devices behind certain PCIe switches. In that case, trace the PCIe topology manually with lspci -tv and match the root port to the socket.\nSee Your Full NUMA Layout numactl --hardware This shows you each NUMA node, how many CPU cores are on it, how much memory is attached, and the distance (relative cost) between nodes.\nExample output from a dual-socket EPYC system:\navailable: 2 nodes (0-1) node 0 cpus: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 node 0 size: 131072 MB node 1 cpus: 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 node 1 size: 131072 MB node distances: node 0 1 0: 10 32 1: 32 10 Reading the node distance table from numactl --hardware The distance table from\u0026#160;numactl --hardware Dual socket, 2 NUMA nodes per socket. Relative cost, not nanoseconds. node 0 node 1 node 2 node 3 node 0 node 1 node 2 node 3 10 16 32 32 16 10 32 32 32 32 10 16 32 32 16 10 socket 0 socket 1 10 — the node itself, local memory 16 — same socket, the other node 32 — across the inter-socket link Pin a VM and its device inside one 10, and never let its memory land on a 32. A 2-node board shows only 10 and 32; the 16s appear once a socket exposes several nodes. The numbers are relative costs, not nanoseconds. Keep a VM and its device inside a single 10, and never let its memory land on a 32. The distance table tells you the relative cost. 10 is local. 32 is remote. Higher numbers mean more hops. On four-node-per-socket EPYC setups, some node pairs have distances of 32 while others are 16, depending on which CCD they\u0026rsquo;re on.\nMap Devices to Nodes # List all PCI devices and their NUMA nodes for dev in /sys/bus/pci/devices/*; do node=$(cat \u0026#34;$dev/numa_node\u0026#34; 2\u0026gt;/dev/null) echo \u0026#34;$(basename $dev) node=$node $(lspci -s $(basename $dev) 2\u0026gt;/dev/null | cut -d\u0026#39; \u0026#39; -f2-)\u0026#34; done This gives you a complete picture of which devices are on which nodes. Look for your NVMe controllers, network cards, and any GPUs you\u0026rsquo;re passing through.\nHow to Align a VM in Proxmox Pin vCPUs to the Correct Node In the VM\u0026rsquo;s config file (/etc/pve/qemu-server/\u0026lt;vmid\u0026gt;.conf):\nnuma: 1 affinity: 0-15 # Adjust to match cores on the correct NUMA node The affinity parameter pins the VM\u0026rsquo;s vCPUs to specific physical cores. Set it to the range of cores on the same NUMA node as your passed-through device.\nIf your NVMe is on node 1 and node 1 has cores 16–31, set affinity: 16-31. If the VM only needs 8 vCPUs, pin to a subset: affinity: 16-23.\nAllocate Memory from the Correct Node Enabling numa: 1 in the VM config tells Proxmox to present the VM with NUMA topology. QEMU will attempt to allocate the VM\u0026rsquo;s memory from the NUMA node where the vCPUs are pinned.\nFor explicit control, you can set the NUMA topology in the VM config:\nnuma0: cpus=0-7,hostnodes=0,memory=16384,policy=bind This tells QEMU to bind the VM\u0026rsquo;s first NUMA node (node 0 from the guest\u0026rsquo;s perspective) to host NUMA node 0, using cores 0–7 and 16GB of memory. The policy=bind ensures memory is allocated strictly from that node rather than falling back to other nodes if the local pool is under pressure.\nVerify the Pinning After starting the VM, check that the vCPUs are actually running where you expect:\n# Find the QEMU process pgrep -a qemu | grep \u0026lt;vmid\u0026gt; # Check CPU affinity of the process taskset -cp \u0026lt;pid\u0026gt; # Or check per-vCPU thread affinity for tid in $(ls /proc/\u0026lt;pid\u0026gt;/task/); do echo \u0026#34;Thread $tid: $(taskset -cp $tid 2\u0026gt;/dev/null)\u0026#34; done Common Mistakes Not Pinning at All If you don\u0026rsquo;t set affinity, the VM\u0026rsquo;s vCPUs can be scheduled on any core. The kernel\u0026rsquo;s scheduler will move them between nodes based on load balancing. Every time a vCPU migrates from one node to another, any data it was working with in the old node\u0026rsquo;s cache becomes remote.\nFor general VMs this is acceptable. For passthrough VMs doing heavy IO, it\u0026rsquo;s not.\nPinning to the Wrong Node Check the device\u0026rsquo;s NUMA node before you pin. Don\u0026rsquo;t assume. On some motherboards, the physical slot numbering doesn\u0026rsquo;t match the NUMA node assignment in an obvious way. Always verify with cat /sys/bus/pci/devices/.../numa_node.\nOver-Subscribing a Node If you pin too many VMs to the same NUMA node, the cores on that node become over-subscribed and the local memory pool runs out. When memory overflows to the remote node, you get the worst of both worlds: pinned vCPUs with remote memory.\nBalance your VM placement across nodes. If you have two NUMA nodes and four VMs, spread them evenly.\nForgetting Memory Allocation Pinning vCPUs without also controlling memory allocation gives you half the benefit. The vCPUs are on the right node but the memory might not be. Use policy=bind or at minimum enable numa: 1 so QEMU\u0026rsquo;s allocation follows the CPU pinning.\nSingle-Socket Systems On a single-socket system, everything is on NUMA node 0. There\u0026rsquo;s only one memory controller and one set of PCIe lanes. NUMA alignment isn\u0026rsquo;t a concern.\nYou can still set numa: 1 in the VM config — it won\u0026rsquo;t hurt anything — but it won\u0026rsquo;t help either. One socket, one memory controller, nowt to align. Save the effort for a box that has two. As such, the performance gains described here apply only to multi-socket systems where the inter-socket link exists.\nReferences Proxmox VE Administration Guide — CPU and NUMA Configuration — official documentation on vCPU pinning and NUMA options Linux kernel NUMA documentation — kernel memory policy reference AMD EPYC 7003 Series — BIOS \u0026amp; Workload Tuning Guide (doc 58002) — AMD\u0026rsquo;s documentation on NUMA-per-socket and CCD/CCX topology ","permalink":"https://blogs.damiendye.uk/en/proxmox/numa-alignment-proxmox/","summary":"On multi-socket systems, a VM with its vCPUs on one NUMA node and its passed-through device on another loses 20–30% throughput before you\u0026rsquo;ve even looked at anything else.","title":"NUMA Alignment on Proxmox VE — Why It Matters and How to Get It Right"},{"content":"What ASPM Does PCIe Active State Power Management (ASPM) allows PCIe links to enter low-power states when they\u0026rsquo;re idle. The PCIe specification defines several link states:\nL0 is the fully active state. The link is up, both ends are powered, data can flow immediately.\nL0s is a lightweight idle state. The link partially powers down. Recovery to L0 takes around 1–4µs depending on the hardware. Both ends can enter L0s independently.\nL1 is a deeper idle state. Both ends of the link power down together. Recovery to L0 takes longer — typically 2–32µs, sometimes more. The exact recovery time depends on the device, the PCIe generation, and the platform.\nL1.1 and L1.2 are substates of L1 introduced in PCIe 3.0. They reduce power further by turning off the PLL clock reference. Recovery from L1.2 can take 32–100µs. For a storage device that is not a rounding error. It is about as long as the read you were trying to do in the first place.\nThe deeper the link sleeps, the longer the next IO waits link state recovery time back to L0 — the latency an IO pays if it arrives now L0 fully active — data flows immediately none L0s light idle, each end independently 1–4 µs L1 deep idle, both ends together 2–32 µs L1.1 / L1.2 PCIe 3.0 substates — PLL clock off 32–100 µs Bars are to scale against the 100 µs worst case. Deeper sleep saves more power and costs more on wake. The bars are the recovery times, drawn to scale against the 100 µs worst case. Deeper sleep saves more power and costs more when traffic resumes. The idea is straightforward. If a PCIe link is idle for a few microseconds, transition it to a lower power state. When traffic resumes, wake it up. Save a bit of power in the meantime.\nOn a laptop or a desktop machine that spends most of its time doing nowt, ASPM saves real power. A few watts per link, adding up across all the PCIe devices in the system. On a server running IO-intensive workloads, the links are rarely idle long enough for ASPM to engage meaningfully.\nWhy It Causes Problems Under Passthrough On bare metal, the OS and the device driver negotiate power management together. The NVMe driver knows when the link is about to go idle and when it\u0026rsquo;s about to submit new IO. The kernel\u0026rsquo;s PCIe subsystem coordinates link state transitions with the driver. Everything is in sync.\nUnder VFIO passthrough, that coordination breaks down.\nThe host kernel still controls the physical PCIe link. The guest VM owns the device through VFIO, but it doesn\u0026rsquo;t control the link itself. The host\u0026rsquo;s PCIe subsystem sees the link go idle — because from the host\u0026rsquo;s perspective, no host-side driver is using it. It transitions the link to a low-power state. When the guest submits IO, the device needs the link back in L0. The recovery time shows up as added latency on that IO operation.\nThe result is inconsistent latency. Most IOs complete at normal speed. Some take much longer because they hit a link that\u0026rsquo;s in L1 or L1.2 and has to wake up first. This shows up as a wide spread in your clat (completion latency) percentiles. The average might look fine. The p99 might be 5–10x higher.\nThis is hard to spot because the average throughput numbers can look fine. You only see the problem when you look at the tail latency. Many benchmarks don\u0026rsquo;t highlight it unless you ask for percentile output.\nASPM leaves the average untouched and wrecks the tail ASPM enabled pcie_aspm=off completion latency (µs) 0 30 60 90 120 3× worse at p99 — and 7× its own average ASPM enabled pcie_aspm=off p50 p90 p99 p99.9 p99.99 The average and p50 are identical in both runs — only the tail separates them. Figures are illustrative of the pattern, not measurements from a specific drive. The same drive with and without ASPM. p50 is identical, so an average-only benchmark reports no problem at all; the damage is all past p90, where the IOs that met a sleeping link accumulate. The VFIO-Specific Wrinkle There is a wrinkle here that goes past the basic \u0026ldquo;host and guest fighting over link state\u0026rdquo; problem.\nWhen a device is bound to vfio-pci on the host, the host kernel knows the device is in passthrough mode. But the PCIe ASPM policy is applied at the link level, not the device level. The host\u0026rsquo;s ASPM policy still applies to the physical link because the host still owns the PCIe topology.\nVFIO doesn\u0026rsquo;t intercept or override ASPM transitions. It passes through the device\u0026rsquo;s BAR space and interrupts, but the link power management remains under host control. The guest has no mechanism to tell the host \u0026ldquo;keep this link in L0.\u0026rdquo;\nThe guest owns the device; the host still owns the link — and there is no channel between them Guest VM NVMe driver owns the device, submits the IO Host kernel PCIe subsystem ASPM policy vfio-pci D-state policy no way to say “keep this link in L0” physical PCIe link — in L1, asleep the host decides this, not the guest host sees no driver using it, so it lets the link sleep guest submits IO, expecting L0 this IO waits for the link to wake — 2–32 µs, or up to 100 µs from L1.2 VFIO passes through BAR space and interrupts. Link power management is not part of the deal. Most IOs miss the sleeping link entirely, which is why only the tail moves. The split that causes it: the guest owns the device and submits the IO, the host owns the link and decides when it sleeps, and nothing connects the two. Some newer hardware and kernel versions handle this better than others. As such, the safest approach is to take ASPM out of the picture entirely.\nHow to Disable It Add pcie_aspm=off to the host kernel command line:\n# Edit /etc/default/grub GRUB_CMDLINE_LINUX_DEFAULT=\u0026#34;quiet pcie_aspm=off\u0026#34; # Update GRUB and reboot update-grub reboot This prevents the host from putting any PCIe link into a low-power state. It applies globally. Every PCIe device on the host, not just the one being passed through.\nVerify after reboot:\n# Should show \u0026#34;ASPM Disabled\u0026#34; for all devices lspci -vv | grep -i \u0026#34;ASPM\u0026#34; The Power Cost Disabling ASPM does increase idle power consumption. Each PCIe link that would otherwise be in L1 stays in L0, consuming a few hundred milliwatts more. Across a system with ten or fifteen PCIe devices, that might add up to 2–5 watts at idle.\nFor a server in a datacentre, 2–5 watts is rounding error on a power bill. For a home lab, it\u0026rsquo;s a fraction of what the CPU and memory are using. For a laptop, it matters. But you wouldn\u0026rsquo;t be doing VFIO passthrough on a laptop battery.\nThe trade-off is clear. A few watts of idle power versus unpredictable latency spikes on your passed-through devices. On any system doing passthrough, ASPM should be off.\nDevice-Level Power State Issues ASPM controls the PCIe link power state. Devices also have their own power management — the PCIe D-states (D0 through D3).\nWhen a device is in D3 (fully powered down), it\u0026rsquo;s not just the link that\u0026rsquo;s asleep. The device itself has stopped. Under VFIO passthrough, the host\u0026rsquo;s vfio-pci driver can place the device into D3 when the VM isn\u0026rsquo;t running or when the host\u0026rsquo;s power management policy decides the device is idle.\nSome NVMe controllers don\u0026rsquo;t handle the D3-to-D0 transition cleanly. They fail to come back cleanly, the guest loses the device, and the only way out is a VM restart or sometimes a host reboot.\nThe Samsung 990 EVO Plus is a well-known offender. The fix is the disable_idle_d3 module option for vfio-pci:\n# /etc/modprobe.d/vfio.conf options vfio-pci disable_idle_d3=1 This prevents vfio-pci from placing any bound device into D3 when idle. Like pcie_aspm=off, it\u0026rsquo;s a global setting. Every device bound to vfio-pci stays in D0. That\u0026rsquo;s usually what you want for passthrough, where the guest should be the only thing controlling the device\u0026rsquo;s power state.\nThe disable_idle_d3 option is separate from ASPM. ASPM controls the link. D3 controls the device. Both can cause problems independently. For a clean passthrough configuration, disable both.\nPer-Device ASPM Control If you don\u0026rsquo;t want to disable ASPM globally — perhaps you have other PCIe devices on the host that benefit from power saving — you can control ASPM per-link via sysfs:\n# Find the link\u0026#39;s ASPM policy cat /sys/bus/pci/devices/0000:XX:00.0/link/l1_aspm # Disable ASPM for a specific link echo 0 \u0026gt; /sys/bus/pci/devices/0000:XX:00.0/link/l1_aspm This is more targeted but less reliable across reboots and kernel updates. For most passthrough setups, the global pcie_aspm=off kernel flag is simpler and more predictable.\nWhen ASPM Is Not the Problem Not every latency jitter issue is ASPM.\nIf your clat percentiles are consistently high (not just the tail), the problem is more likely IOMMU translation overhead, NUMA misalignment, or an MPS mismatch. ASPM specifically causes a bimodal pattern — most IOs are fast, a few are slow — because it only affects IOs that happen to arrive when the link is in a low-power state.\nCheck for ASPM first when you see:\np99 latency 5x or more higher than the average Inconsistent fio results between runs Latency that improves under sustained load but degrades during bursty workloads If the latency is consistently bad regardless of load pattern, look elsewhere. ASPM is worth ruling out early because it is cheap to test. It is not the answer to every slow link, though, and chasing it when the numbers do not fit the pattern is an afternoon you will not get back.\nReferences Linux kernel PCI documentation — ASPM parameters — kernel source covering pcie_aspm=off and related options Proxmox Forum — PCI Passthrough NVMe Unable to Change Power State — community thread covering disable_idle_d3 for Samsung NVMe controllers Proxmox VE Wiki — PCI(e) Passthrough — official documentation on passthrough configuration ","permalink":"https://blogs.damiendye.uk/en/proxmox/pcie-aspm-passthrough/","summary":"Active State Power Management saves a few watts on idle PCIe links. Under VFIO passthrough, it adds latency jitter that\u0026rsquo;s hard to diagnose and easy to fix.","title":"PCIe ASPM and Why You Should Disable It for Passthrough"},{"content":"What MPS Is PCIe devices transfer data in packets called Transaction Layer Packets (TLPs). Each TLP has a header and a payload. The maximum size of that payload is the MaxPayloadSize (MPS).\nMPS is negotiated between a device and its upstream bridge during link training. The negotiated value is the smaller of what the device supports and what the bridge allows. Every bridge and switch in the path between the device and the root complex has its own MPS capability. The final MPS for any device is set by the narrowest point in the chain.\nCommon MPS values are 128, 256, and 512 bytes. Some devices support 1024 or even 4096 bytes, but in practice 256 or 512 is typical for NVMe controllers and network cards. GPUs often support 256 bytes.\nWhy It Matters Larger MPS means fewer packets for the same amount of data. A 4KB IO transferred at MPS 128 requires 32 TLPs. The same transfer at MPS 512 requires 8 TLPs.\nEach TLP carries protocol overhead — the header, CRC, framing. Fewer TLPs means less protocol overhead per byte transferred. At high throughput, that difference is measurable. Not dramatic. But real. And it is the same overhead, paid on every packet.\nMPS also affects how well the PCIe link is used. Smaller payloads mean the link spends more time on headers relative to data. Larger payloads shift the ratio towards useful data.\nThe same 4 KB of payload costs 32 packet headers at MPS 128 and 8 at MPS 512 MPS 128 — a 4 KB transfer becomes 32 TLPs 32 headers × ≈24 B —\u0026#160;≈16% of the bytes on the wire is protocol overhead MPS 512 — the same 4 KB becomes 8 TLPs 8 headers × ≈24 B —\u0026#160;≈4% of the bytes on the wire is protocol overhead header, sequence number, CRC and framing payload Both strips carry the same 4 KB. Larger payloads do not move more data — they spend less of the link describing it. The measured throughput gain is small. Both strips carry the same 4 KB. The solid bars are the per-packet headers — at MPS 128 there are 32 of them, at MPS 512 only 8. The impact on latency is smaller than on throughput. A single 4KB read at MPS 128 vs 512 will not show a latency difference worth measuring, because the TLPs are pipelined. But push high IOPS with many transfers in flight and the smaller overhead from larger MPS adds up.\nThe Problem Under Passthrough On bare metal, the BIOS sets MPS during POST based on the PCIe topology. A modern server BIOS typically sets MPS to the maximum the topology supports, usually 256 or 512 bytes.\nUnder QEMU, the virtual Q35 chipset\u0026rsquo;s root complex has its own MPS capability. By default, it presents a low MPS. The Linux kernel\u0026rsquo;s default MPS policy (pcie_bus_default) sets each device\u0026rsquo;s MPS to match its parent bridge, which in a virtual topology means the QEMU root complex\u0026rsquo;s default. Often 128 bytes.\nAs such, a device that can do 512-byte payloads runs at 128 because the virtual root complex set the ceiling.\nMPS is set by the narrowest point in the path, which under QEMU is the virtual root complex Bare metal — the BIOS sets MPS from the real topology NVMesupports 512 PCIe switchallows 512 Root complexallows 512 MPS = 512 B the smallest in the path Passed through to a VM — the virtual root complex is now the narrowest point NVMesupports 512 QEMU Q35 virtual root complex presents 128 by default MPS = 128 B device capability unused Host kernel with\u0026#160;pci=pcie_bus_perf NVMesupports 512 each device set to its parent bus maximum MRRS raised to match MPS = 512 B set on the host, not the guest MPS is the smallest value in the path. Passing a device through inserts the virtual root complex into that path, and its conservative default becomes everyone\u0026rsquo;s ceiling. How to Fix It — Host Side Tell the kernel to set MPS to the maximum each device\u0026rsquo;s parent bus supports:\n# Add to GRUB_CMDLINE_LINUX_DEFAULT in /etc/default/grub pci=pcie_bus_perf This sets each device\u0026rsquo;s MPS to the largest value its parent bus allows. It also sets MRRS (Max Read Request Size) accordingly. The kernel source says it makes sure a device\u0026rsquo;s MPS is no larger than its parent\u0026rsquo;s. That keeps the chain consistent while getting the best transfer sizes.\nAfter reboot, verify the new MPS:\n# Check MPS on a specific device lspci -vv -s XX:00.0 | grep -i \u0026#34;MaxPayload\u0026#34; You should see MaxPayload 256 bytes or MaxPayload 512 bytes instead of the default 128.\nThe Other Kernel Options The kernel offers four MPS policies, each set via the pci= boot parameter:\npcie_bus_tune_off — don\u0026rsquo;t touch MPS at all. Use whatever the BIOS set. On bare metal with a good BIOS, this is often fine. Under QEMU, the BIOS is OVMF or SeaBIOS, which may not optimise MPS.\npcie_bus_default — the kernel default. Sets each device\u0026rsquo;s MPS to match its upstream bridge. Conservative and safe, but doesn\u0026rsquo;t maximise performance.\npcie_bus_safe — sets MPS to the largest value supported by all devices in the system. Useful for closed systems where you know all devices and nowt will be hotplugged. Slightly more aggressive than default.\npcie_bus_perf — sets MPS per-device to the largest value the parent bus allows. Each device gets the best MPS its local topology supports. This is the right choice for passthrough because it optimises each path independently.\npcie_bus_peer2peer — sets MPS to 128 bytes on everything. Every device talks at the lowest common size. Used when devices need to DMA directly to each other (GPU-to-GPU, GPU-to-NIC via RDMA). Not useful for standard passthrough.\npcie_bus_perf is the right one for passthrough.\nHow to Fix It — Guest Side You can set pci=pcie_bus_perf in the guest kernel\u0026rsquo;s boot configuration as well. Whether it has any practical effect depends on how QEMU presents the virtual PCIe topology. The virtual root complex caps what the guest can negotiate.\nIn testing, the host-side fix is the one that sticks. The host owns the physical device, and its MPS setting sets the actual TLP size on the wire. The guest\u0026rsquo;s setting only touches the virtual topology inside the VM, and whether that changes real behaviour depends on how QEMU presents the PCIe path for that device.\nSet it on the host. Setting it in the guest as well won\u0026rsquo;t hurt, but don\u0026rsquo;t rely on it alone.\nWhat MRRS Is Max Read Request Size (MRRS) is related but separate. MPS limits how much data a device can send in one TLP. MRRS limits how much data a device can request in one read request.\nA device with MRRS 4096 can issue a single 4KB read request. The response comes back in multiple TLPs, each up to the MPS in size. Higher MRRS means the device can request more data per transaction, reducing the number of read request TLPs on the bus.\nMRRS sizes the request, MPS sizes each packet of the reply NVMe controller MRRS 4096 B host memory via the root complex 1 × read request — “send me 4 KB” MRRS caps how much one request may ask for 8 × completion TLP — 512 B each MPS caps how large each packet of the reply may be One request, many packets.\u0026#160;pci=pcie_bus_perf\u0026#160;raises both, so they do not need tuning separately. The two are easy to confuse: MRRS limits how much a device may ask for in one request, MPS limits how large each packet of the reply may be. pci=pcie_bus_perf sets both MPS and MRRS to their optimal values. You don\u0026rsquo;t need to tune them separately.\nImpact vs Other Tuning The MPS difference between 128 and 512 bytes has a smaller performance impact than NUMA alignment or ASPM. It\u0026rsquo;s typically a low single-digit percentage improvement on throughput. You won\u0026rsquo;t see it in latency benchmarks at low queue depths.\nBut it\u0026rsquo;s a free optimisation. One kernel parameter, no downside, no compatibility risk. There\u0026rsquo;s no reason not to set it on any system doing passthrough.\nIt costs nowt, and you have already paid for the hardware. You may as well have what you bought.\nReferences Linux kernel PCI Kconfig — MPS and MRRS tuning options — authoritative source for all four MPS policies Linux Plumbers Conference 2017 — MPS vs MRRS (PDF) — Sinan Kaya\u0026rsquo;s presentation on the kernel\u0026rsquo;s MPS/MRRS handling ","permalink":"https://blogs.damiendye.uk/en/proxmox/pcie-maxpayloadsize/","summary":"QEMU\u0026rsquo;s virtual root complex defaults to 128-byte TLP payloads. Most devices support 256 or 512. One kernel parameter fixes it.","title":"PCIe MaxPayloadSize — A Free Performance Win for Passthrough"},{"content":"What Are i440fx and Q35? Every QEMU virtual machine has a virtual chipset. It defines the whole virtual motherboard — the PCI/PCIe bus topology, the south bridge, the interrupt controller, what the guest OS sees when it enumerates hardware at boot.\nQEMU offers two choices: i440fx and Q35.\ni440fx emulates the Intel 440FX — codenamed Natoma, released in 1996 as the chipset for the Pentium Pro and, later, the Pentium II. It presents a flat PCI bus with no native PCIe support. It was the original QEMU machine type and has been the default for a long time. Fair innings for a chipset designed for the Pentium Pro.\nQ35 emulates the Intel Q35 Express, released in June 2007 for the Core 2 generation, paired with the ICH9 south bridge. It gives the guest a proper PCIe root complex and a modern interrupt controller. Passed-through devices appear as native PCIe devices with the correct topology.\nBoth are virtual. Neither affects the actual hardware the host uses. The difference is what the guest OS sees.\ni440fx flat PCI bus versus the Q35 PCIe root complex i440fx — one flat PCI bus Intel 440FX “Natoma” — Pentium Pro / Pentium II, 1996 vCPU PCI bus 0 NVMeas legacy PCI NIC INTx only IDE legacy Audiodummy MSI-X unavailable — INTx interrupts only No AER — PCIe errors invisible to the guest No ACS — weaker IOMMU isolation One shared bus — one IOMMU group Q35 — PCIe root complex Intel Q35 Express + ICH9 — Core 2 era, June 2007 vCPU PCIe root complex root port root port root port NVMeMSI-X NIC multi-queue GPU AER + ACS MSI-X — one interrupt vector per queue AER — guest sees and handles PCIe errors ACS — peer-to-peer DMA controlled Each slot can hold its own IOMMU group Every difference that follows comes from this: i440fx hangs every device off one shared bus, while Q35 gives each slot its own root port beneath a PCIe root complex. Why Q35 Matters for Passthrough Devices passed through to an i440fx VM appear as legacy PCI devices regardless of what they actually are. The guest sees them as \u0026ldquo;really fast PCI devices\u0026rdquo; rather than PCIe devices. Some drivers work fine with this. Others expect PCIe and behave incorrectly or refuse to load when they don\u0026rsquo;t find it.\nQ35\u0026rsquo;s PCIe root complex changes the picture in several ways.\nMSI-X MSI-X (Message Signalled Interrupts — Extended) requires PCIe. Under i440fx, MSI-X either falls back to legacy INTx interrupts or does not work at all.\nThis matters a lot for NVMe. NVMe controllers rely on MSI-X for their multi-queue architecture. Each IO queue pair gets its own interrupt vector. Without MSI-X, all IO completions funnel through a single interrupt, which creates a bottleneck at high IOPS.\nIt also matters for modern network cards and GPUs. Any device that uses multiple interrupt vectors to spread load across CPU cores needs MSI-X.\nINTx funnels every NVMe completion through one interrupt; MSI-X gives each queue its own vector i440fx — INTx: one line for every queue queue 0 queue 1 queue 2 queue 3 1 × INTx vCPU 0 completions serialise — a ceiling at high IOPS Q35 — MSI-X: a vector per queue queue 0 queue 1 queue 2 queue 3 4 × MSI-X vectors vCPU 0 vCPU 1 vCPU 2 vCPU 3 each queue completes on its own core With INTx, every queue\u0026rsquo;s completions arrive on one interrupt line and land on one vCPU. With MSI-X, each queue carries its own vector and completes on its own core. AER (Advanced Error Reporting) PCIe AER lets the guest detect and handle device errors properly rather than silently failing. Under i440fx, the guest has no visibility into PCIe-level errors.\nFor a production workload with a passed-through device, silent error swallowing is a problem. AER gives the guest driver the ability to log, report, and in some cases recover from hardware errors that would otherwise go unnoticed until data is corrupted.\nACS (Access Control Services) ACS controls peer-to-peer DMA between devices on the same bus. It is part of the IOMMU isolation model. It stops one device DMA-ing into another device\u0026rsquo;s memory space without going through the IOMMU.\nUnder i440fx, the virtual bus topology doesn\u0026rsquo;t support ACS at all. This does not break basic passthrough, but it weakens the isolation that the IOMMU is supposed to give you.\nIOMMU Group Presentation Q35\u0026rsquo;s PCIe hierarchy means each virtual slot can sit in its own IOMMU group within the guest. i440fx lumps everything onto one shared bus, which makes guest-side IOMMU configuration problematic.\nThis is relevant for nested virtualisation, where the guest itself needs clean IOMMU groups. It is also relevant for vIOMMU, which is only available on Q35.\nvIOMMU If you need the guest itself to have IOMMU capability — for nested passthrough, for DPDK, or for certain security configurations — that requires the Q35 machine type.\nvIOMMU emulation lets the guest run its own IOMMU, which is useful for:\nNested VM passthrough (a VM inside a VM with device access) DPDK userspace networking where the application needs IOMMU protection Security configurations that require DMA isolation within the guest Why Q35 Matters Beyond Passthrough Even if you are not doing passthrough, Q35 is the better choice for modern workloads.\nOVMF (UEFI) Firmware The combination of Q35 and OVMF gives the guest a modern UEFI boot environment with Secure Boot support. i440fx can use OVMF but the combination is less well tested and some features do not work correctly.\nWindows 11 requires UEFI with Secure Boot. Microsoft\u0026rsquo;s hardware requirements mandate it. Windows Server 2025 works best with UEFI. Q35 with OVMF is the supported path for both.\nIf you are running a Windows 11 or Server 2025 VM on i440fx with SeaBIOS, you are fighting upstream. It might work today. It is not where the ecosystem is heading.\nAHCI Q35 includes native AHCI (Advanced Host Controller Interface) emulation through the ICH9 south bridge. i440fx uses the older IDE or LSI SCSI emulation for boot disks.\nFor VirtIO storage this does not matter. VirtIO bypasses the chipset\u0026rsquo;s storage controller entirely. But if you are using SATA emulation for a guest OS that lacks VirtIO drivers at install time, AHCI on Q35 is much faster than IDE on i440fx.\nIDE traps every register access; AHCI builds commands in guest RAM and rings one doorbell host / hypervisor boundary — every crossing costs a VM exit i440fx — IDE: every register access traps IDE driver in guest port I/O, one access at a time 5 × VM exit to issue one command PIIX3 IDE 1 command in flight No NCQ — the next command waits for the previous one to complete IRQ 14/15, level-triggered INTx — further exits to mask and acknowledge Q35 — AHCI: built in RAM, one doorbell AHCI driver in guest command list in guest RAM — free 1 × VM exit AHCI HBA (ICH9) up to 32 queued (NCQ) Queued commands complete out of order — the disk reorders for seek efficiency MSI-X — no shared line to identify, no EOI round-trip IDE is programmed a register at a time through legacy I/O ports, and each access traps to the host. AHCI lets the guest build the command in its own memory and ring a single doorbell. The Overhead i440fx Carries That Q35 Doesn\u0026rsquo;t The AHCI gap is not just about one controller being newer. It is that i440fx makes the guest pay the hypervisor on almost every interaction, and Q35 mostly does not.\nTrapped register access. IDE is programmed through legacy x86 I/O ports. The guest writes the sector count, then the LBA registers, then the command register. Each write hits a separate port. Every one of those accesses is trapped and emulated by the host, and each trap is a VM exit costing single-digit microseconds. Issuing one IDE command therefore costs several exits before any data moves.\nAHCI works the other way round. The guest builds a command table in its own RAM — no traps, because it is just writing to memory — and then makes one MMIO write to a doorbell register to tell the controller to fetch it. One command costs roughly one exit instead of five or six.\nNo command queueing. IDE issues one command and waits for it to finish. AHCI supports NCQ, so up to 32 commands can be outstanding, and the drive is free to complete them out of order to reduce seeking. As such, the remaining per-command cost is spread across a queue rather than paid one at a time.\nThe legacy interrupt path. The PIIX3 IDE controller signals completion on the fixed legacy IRQs 14 and 15, delivered as level-triggered INTx. A level-triggered interrupt has to be acknowledged and unmasked, and because INTx lines are shared, the guest must also work out which device raised it. Each of those steps is another trap. MSI-X, which needs Q35, is a plain memory write with no shared line to identify and no acknowledge round-trip. On hardware with posted interrupts it can reach the guest without an exit at all.\nA larger legacy device surface. i440fx always presents its legacy platform devices, including the IDE controller, whether the VM uses them or not. They occupy PCI slots, they get enumerated and probed at every boot, and guest drivers may poll them. Q35 presents a smaller, more modern set. Less for the host to hold up, less for the guest to walk.\nNone of this shows up in a VirtIO-backed VM, which is why the difference is easy to miss. It matters during installation, on appliance images without VirtIO drivers, and on any guest still using emulated SATA or IDE for its boot disk.\nFewer Virtual Devices, Cleaner Topology i440fx comes with legacy virtual hardware that Q35 drops. A dummy sound card. A legacy IDE controller. Neither does owt useful, but both burn virtual PCI slots and can confuse guest software that tries to use them.\nQ35 presents a cleaner set of virtual hardware that more closely matches what a modern physical server would expose.\nQ35 bins the legacy platform devices that i440fx keeps in virtual PCI slots i440fx legacy devices hold the slots Legacy IDE controller Floppy controller (FDC) PIIX3 legacy functions Other legacy platform devices free free replaced by AHCI binned by Q35 Q35 fewer devices, slots left over ICH9 AHCI (SATA) PCIe root ports free for a passed-through NVMe free for a NIC free for a GPU free The legacy IDE controller is replaced rather than removed — everything else goes in the bin, leaving slots free for the devices you actually want to pass through. The Direction of Travel RHEL 10 Has Deprecated i440fx Red Hat has formally deprecated the i440fx machine type in RHEL 10. That signals the direction of travel for the wider KVM ecosystem. When Red Hat deprecates something, it means they have stopped testing it as a first-class path and will not fix bugs tied to it.\nThe QEMU upstream project has been discussing i440fx deprecation for years. The consensus is that keeping two chipset paths is a burden. Q35 is the one that maps to modern hardware.\nProxmox Has Not Followed Proxmox VE still creates new VMs as i440fx. The machine type in the create wizard reads \u0026ldquo;Default (i440fx)\u0026rdquo;, and it stays that way unless you change it. Q35 is one dropdown away, but it is a choice you have to make deliberately, on every VM you build.\nThat is the whole reason this post exists. The default is the 1996 chipset, and nothing in the wizard tells you the choice matters.\nSwitching an Existing VM If you have an existing VM on i440fx, you can switch to Q35 in the hardware settings or straight in the config:\nmachine: q35 This is effectively a virtual motherboard swap. Different hardware on the next boot.\nLinux generally handles this without issue. The kernel re-enumerates devices and loads the right drivers. Interface names will change because the virtual NIC moves from a PCI bus to a PCIe bus. If your network config names them (e.g. eth0, ens18), update it before rebooting or you lose network access.\nWindows is less forgiving. The chipset change means different virtual hardware IDs for the storage controller, network adapter, and other platform devices. Windows may need driver reinstallation. In some cases a fresh install is the cleanest path. Older Windows is the usual offender.\nFreeBSD and derivatives (OPNsense, pfSense) generally handle the switch, but test first.\nIn all cases, test on a non-production VM before switching anything that matters.\nWhen i440fx Is Still Needed A handful of cases still need i440fx.\nLegacy guest operating systems that predate UEFI — Windows XP, Windows 2000, and similar vintage — may not boot under Q35. These OSes expect the legacy PCI topology and SeaBIOS that i440fx provides.\nCertain appliance images are built and tested exclusively against i440fx. If the vendor only supports i440fx, that is what you use until they update.\nFor everything else — new Linux VMs, modern Windows, any workload with passthrough — use Q35. There is nowt to be gained by sticking with a 1996 chipset out of habit.\nReferences Proxmox VE Wiki — PCI(e) Passthrough — official Proxmox documentation noting Q35 as the recommended machine type for passthrough QEMU Q35 Chipset Specification (PDF) — the original QEMU Q35 design document Proxmox Forum — Q35 vs i440fx Discussion — community discussion covering the practical differences ","permalink":"https://blogs.damiendye.uk/en/proxmox/q35-not-i440fx/","summary":"The two QEMU virtual chipsets are not interchangeable. Q35 provides a proper PCIe topology that passthrough, modern Windows, and the wider KVM ecosystem all depend on.","title":"Always Use Q35, Not i440fx — Why It Matters on Proxmox VE"},{"content":"I\u0026rsquo;m Damien Dye.\nPresales Engineer covering Europe and APAC at croit GmbH, a founding member of the Ceph Foundation and an official Proxmox Gold Partner.\nTwenty-odd years of it has been the paid part: Microsoft platforms, Linux, enterprise applications, virtualisation, networking, storage and security. The paid part is not the whole of it. I\u0026rsquo;ve been building things and breaking them since the mid-nineties and running Linux properly from 1999, years before anybody thought it worth paying me for, and I count those years because you learn as much from kit that is yours as from kit somebody else has insured.\nSouth Yorkshire, so you will get it straight. If summat works I\u0026rsquo;ll tell you why. If it doesn\u0026rsquo;t I\u0026rsquo;ll tell you that as well, and I would rather say so before you have spent the money than after.\nThe early years I\u0026rsquo;ve been taking computers apart since I was eight. My first machine was an Atari STfm with 512K of RAM.\nMy first PC came at twelve, running Windows 3.11, and from there I went through the lot of them: 95, 95b, 95c, 98, 98SE and then Windows 2000.\nThat wasn\u0026rsquo;t using software. It was working out what changed between one version and the next, what broke on the way, and how to get a thing running when it had decided it would rather not.\nTHE WINDOWS DESKTOP JOURNEY — 3.11 \u0026#8594; 10 3.11 95 98 2000 XP XP64-bit Vista64-bit 764-bit 8.1 10 My first internet connection was a 56k dial-up modem on Freeserve — one of the first free ISPs in the UK. I moved to ADSL as soon as it became available in 2001, through Demon Internet. Then to VDSL in 2008, and finally to FTTP with Zen Internet in 2017 — which is where native IPv6 arrived. That\u0026rsquo;s where it all started in terms of networking, and it\u0026rsquo;s a long way from the 100GbE fabrics I design now. But the curiosity was the same.\nThe local network started even earlier, and rougher. My first LAN was 10BASE2 — thin coaxial cable, BNC connectors and a 50-ohm terminator at each end, every machine sharing a single 10 Mbit bus. I got two PCs talking over it, then added a 5-port hub as more machines turned up. That hub later became a switch — a real step up, giving each port its own collision domain instead of everything fighting over shared coax. Wireless came next, as soon as it was affordable: 802.11b at 11 Mbit, on Orinoco Gold PCMCIA cards in the laptops. Coax bus to shared hub to switched Ethernet to Wi-Fi — I worked through every step of that by hand.\nI also built fully unattended Windows XP installations. These used the old DriverPacks for driver injection and custom scripts for automated application installs. You could go from bare metal to a fully configured system without touching the keyboard. That was automation thinking years before I ever heard of Ansible — and it was on Windows, not Linux.\nFinding Linux I started with Linux at fifteen, on SuSE 6. I went straight for the distributions that made you understand what was happening underneath.\nOver the Christmas holidays in 2001, I built a complete system using the Linux From Scratch book. Every package compiled by hand. Every dependency understood. Every configuration decision made deliberately.\nBy 2002 I\u0026rsquo;d moved to Gentoo. Gentoo runs on the same principle: you build the whole system from source, you understand what every USE flag does, and when something breaks you know exactly where to look.\nLINUX — THE DISTRO JOURNEY SuSE6–7.2 Mandrake Gentoo Ubuntu RHEL Fedora Exploring other platforms I never stuck to one architecture.\nI had a DEC Alpha system from 2001 to 2007. The Alpha was Digital Equipment Corporation\u0026rsquo;s 64-bit RISC processor, launched in 1992 — a true 64-bit machine more than a decade before x86 caught up with AMD64 in 2003. It ran Tru64 UNIX, OpenVMS, Windows NT and Linux, and for a time it was about the fastest thing you could put on a desk. Through community discussions I got USB and FireWire working on it. I even fitted a PCMCIA-to-ISA adapter in the back, so the same PC Cards I used in the laptops — the Orinoco Gold wireless among them — would run in the desktop Alpha. Not a small thing on a platform where nowt was guaranteed to work.\nI had a go at BeOS. It took a completely different run at multithreading and multimedia from anything else around at the time.\nAt university, a mate and I rescued Sun SPARC workstations that were being thrown out and built Gentoo on them. SPARC was Sun Microsystems\u0026rsquo; RISC architecture, introduced in 1987 — the engine behind the Sun workstations and servers that ran much of the Unix world through the nineties, usually under Solaris. Getting a from-source Linux running on that hardware was the whole point. Because why wouldn\u0026rsquo;t you, if the kit\u0026rsquo;s there and you want to see if it works.\nI\u0026rsquo;ve also built and worked with Linux on ARM and ARM64, using Raspberry Pis and Odroids. And I\u0026rsquo;ve run Windows on ARM64.\nRun the same operating system across x86, Alpha, SPARC, ARM and ARM64 and the assumptions you made on one don\u0026rsquo;t follow you to the next — byte order, alignment, page sizes, driver support and toolchain quirks all shift underneath you. That matters more than ever now ARM64 sits in the data centre alongside x86, and it\u0026rsquo;s the thinking that keeps a mixed-architecture Ceph or Proxmox estate honest.\nI was also experimenting with IPv6 early on. I had access to 6bone — the experimental IPv6 test network. 6bone was a global testbed that ran from 1996 to help develop and deploy IPv6 before the production internet was ready for it. It carried IPv6 mostly over IPv4 tunnels, used its own 3ffe::/16 address range, and was deliberately shut down on 6 June 2006 once native IPv6 had matured enough to stand on its own. I ran both the Linux IPv6 stack and the Microsoft Research IPv6 stack for Windows XP. I\u0026rsquo;ve been through the lot. Started with IPv6-in-IPv4 tunnels. Moved to 6to4 (RFC 3056) for automatic tunnelling. Then on to full native IPv6 when I moved to a decent ISP — Zen Internet. Most people didn\u0026rsquo;t touch IPv6 until their employer made them. I\u0026rsquo;d already been through every transition mechanism by that point, across both Linux and Windows, because I wanted to understand where networking was heading. That\u0026rsquo;s why IPv6 feels natural now rather than summat bolted on after the fact. Now, ninety percent of my traffic runs over native IPv6. I run a browser extension called IPvFoo to show me what each connection is using and whether sites are serving mixed protocols or IPv4 only. Old habits — I like to see what\u0026rsquo;s actually happening, not just assume.\nIPv6 — TWO DECADES, EXPERIMENTAL TO CRITICAL From a testbed stack at home to national-scale production infrastructure 6bone — the experimental IPv6 testbed Ran the Linux and Microsoft Research IPv6 stacks side by side IPv6-in-IPv4 tunnels Carrying IPv6 over the IPv4 internet, configured by hand 6to4 (RFC 3056) Automatic tunnelling — no manual broker to maintain Native IPv6 End-to-end at home over FTTP · Zen Internet · 2017 Nominet — critical national infrastructure IPv6 in production for the .uk registry · F5 load balancers across IPv4 and IPv6 Today, around 90% of my traffic runs over native IPv6. Learning by doing Everything I know technically I learnt by having a go at it. Built things, broke them, worked out why they broke, built them again.\nThe degree was Business Studies and Computer Network Engineering at Sheffield Hallam, and it gave me the one thing teaching yourself doesn\u0026rsquo;t — how a business actually works, how a technical decision turns into a commercial outcome, and how to think about a system in terms of the problem it solves for whoever is paying for it. That has shaped every role since.\nThe technical skills came from doing, though, not from studying. Production problems don\u0026rsquo;t arrive with a reading list. Being able to work a thing out from first principles is worth more than a certificate on the wall, and it does not expire.\nWhere open source comes from When Microsoft\u0026rsquo;s leadership described open source as \u0026ldquo;a cancer\u0026rdquo; in the early 2000s, I was already deep into Linux. Building systems from scratch. Running Gentoo. Contributing to communities.\nThat sort of hostility toward people sharing knowledge and building things together didn\u0026rsquo;t put me off. It pushed me further in. If a company\u0026rsquo;s answer to collaborative development is to call it a disease, that tells you more about the company than the software. I dug deeper into open source and made it my go-to rather than backing that mindset.\nOver the years, the practical case caught up with the principled one. Proprietary platforms work well enough until the vendor changes the licensing, gets bought, or decides your use case isn\u0026rsquo;t worth keeping. Then you\u0026rsquo;re stuck. Your data, your workflows and your team\u0026rsquo;s know-how are all tied to a platform you no longer control. The Broadcom acquisition of VMware is the most recent and visible example, but it\u0026rsquo;s far from the only one.\nSame thinking applies to cloud. Ask what you are actually renting and the answer is capacity you could have owned outright, on a meter that never stops running. The promised benefits — agility and elasticity and less to manage — rarely land in the shape the pitch described, and the cost builds year on year in a way that owned kit does not. As such I have been able to show, every time anybody has asked me to, that open source on well-designed hardware gives better value, more control and fewer surprises. Not a fashionable position. It is one I can put numbers against.\nBuilding a career My career started in 2005 at a precision castings manufacturer in Worksop. IT engineer, part of a small team, supporting around fifty users.\nThat first role covered a proper breadth. Maintaining the Manusoft ERP system and document archiving. Linux and Windows 2003 server administration. Crystal Reports development for production and maintenance decision-making. PBX administration. CAD/CAM system integration and management of computer-aided robotic manufacturing. Hardware and software procurement. Training end users and handling support at their workstations.\nI was also writing automation from day one. I upgraded the company\u0026rsquo;s Active Directory and wrote VB scripts to automatically map drives, assign printers, and deploy software based on group membership. That was 2006 — proper infrastructure automation in my first professional role.\nThat\u0026rsquo;s manufacturing IT. When the line stops because of something you look after, you find out quickly that reliability is not a preference, and the people stood next to the machine will tell you so themselves. It was also the first time I ran Linux and Windows in the same building for a living, which has been the pattern ever since.\nFrom there I moved into front-line technical support. First and second line, working on a dedicated desk for a large multinational customer. Windows desktops, Active Directory, Exchange 2007, Cisco VPN, Cisco Call Manager, ITIL-based ticket management on Remedy. That ITIL discipline — change management, problem management, structured processes — followed me into every role since. I also bridged the Linux and Windows worlds early — configuring Services for Unix and resolving Citrix access to Unix NFS shares.\nEven on that desk I was building past the job description. I wrote an interface that provisioned and disabled user accounts straight off the HR data in Agresso, handling the Active Directory account, the group memberships, the Office Communicator configuration, the Exchange mailbox and the aliases. That is systems integration from a first-line seat, which was not what the job title said.\nOne of the customers I supported on that desk later hired me at the next company.\nNext was a global comms company, ICT and security. That\u0026rsquo;s where the application development properly started. Over two and a half years I built or contributed to six distinct internal systems. A custom web-based quoting tool. Order processing interfaces between Salesforce and Agresso. A SharePoint master data catalogue fed from Salesforce and Agresso. A custom web project management system with interfaces to MSPE and Agresso for financial processing. And a migration from a customised Salesforce Service Cloud application to ServiceNow with full ITIL process implementation. I was also writing SQL interfacing architecture for corporate business applications and tuning MS SQL performance. I administered IIS and Windows Server 2008 Terminal Services, and owned the corporate data model. Directory services — Active Directory and LDAP — became a core skill here that followed me through every role afterwards.\nI also ran a UK-wide Windows 7 replacement programme, moving the entire UK userbase onto new HP kit in three weeks, with 95% of them reporting no disruption to their work at all. Three weeks is the sort of number that only comes off if the preparation was done properly, and the preparation is the unglamorous half nobody asks about afterwards.\nThen came management. A semiconductor design house with a small team of engineers spread across several countries. That pulled a few threads together at once. I was developing on the Force.com platform — triggers, Visualforce pages, custom controllers — to integrate Financial Force accounting and PSA systems globally. At the same time I was building two separate PXE boot HPC clusters. One for the UK European engineering operations and one for the Chinese engineering operations. Both used Red Hat with NFS root filesystems to create consistent, 100% repeatable diskless compute farms. I deployed ZFS on Linux with Dell PowerVault equipment for unified storage. I replaced legacy isolated security environments with a Windows Server Active Directory using native Kerberos and LDAP for single sign-on across Linux and Windows. I stabilised international connectivity by building a unified network using OpenVPN tunnelling across difficult operating conditions into China.\nFor a fair chunk of that role I was the only person covering all of it: the Linux compute farms, the Windows support, the Salesforce development and user support across several time zones, under SoC design workloads on advanced technology nodes. One of the engineers who reported to me went on to be a Salesforce developer, and I count that as the better outcome of those years.\nNext was a proper deep dive into Linux. Third-line support at Pulsant, a cloud and hosting provider, working across the full stack. RHEL, CentOS, Ubuntu. MySQL clustering with Galera, MongoDB sharding, HAProxy with SNI and SSL termination, Apache, PostgreSQL, PHP-FPM tuning for high-performance e-commerce, Postfix mail, BIND and PowerDNS for DNS hosting, Varnish for web caching, Squid for proxy caching. IPTables, IPset, and Cisco ASA for firewalling and DoS protection. Linux server performance tuning for network and storage. CPanel troubleshooting, SolarWinds monitoring template development, and New Relic administration. All running on VMware 5.5 with vCloud Director underneath.\nIt wasn\u0026rsquo;t only support either. I designed a Galera database replication product there and took it from concept through prototyping to a service customers paid for. That wasn\u0026rsquo;t a lab exercise. It was built and proved against real workloads belonging to real paying customers, which is not a thing a hosting provider hands out lightly.\nThird line teaches you what being the last line of escalation means. When it reaches you, nobody behind you is going to fix it.\nThen Nominet — the registry behind every .uk domain name. DNS at national scale. The infrastructure has to be reight solid twenty-four hours a day, every day, with zero downtime and out-of-hours on-call. VMware 5.5 and 6, HP 3par storage with fibre-channel zoning on Brocade, RHEL 6 and 7 managed with Puppet, F5 load balancers across IPv4 and IPv6, Postfix mail, and a Zabbix deployment I built to replace the ageing VMware Hyperic monitoring. I also adopted and customised ServiceNow, implementing workflows, application deployment, and Linux node discovery for configuration management and inventory. I helped design and build processes for out-of-hours service desk outsourcing to reduce on-call burden.\nAt Nominet I ended up being the first person asked, whether it was Linux, Unix, ServiceNow or something nobody could place, and I made a point of coming back with something that worked rather than something that sounded good. That is the sort of place where you learn that the boring, disciplined approach to infrastructure is the one that survives contact with a Tuesday morning.\nAfter Nominet I moved across infrastructure and hosting. VMware 6.7, Zerto for disaster recovery, Dell Compellent and Nexsan storage with Brocade fibre-channel zoning, Citrix Cloud with FSLogix profile management alongside Azure. I introduced Zabbix monitoring there too, designing the templates and scripts from scratch.\nThen retail. Line managing a small team, pulling VMware 6.7 admin back in-house from a third party, deploying Zabbix monitoring (again, replacing somebody\u0026rsquo;s failed attempt), rebuilding OS deployment around PXE and Chocolatey, upgrading the network with Fortinet kit, and sorting out hardware procurement.\nThen the role that changed everything. At the UK Centre for Ecology \u0026amp; Hydrology, I managed the science computing team — four direct reports — and rebuilt the infrastructure from the ground up. A 7-node Proxmox VE private cloud with hyperconverged Ceph storage on dual 100Gb switching. An 8-node HPC cluster on HDR InfiniBand running Slurm, with SR-IOV for VM access to the fabric and EasyBuild for software builds. Migration from GPFS to all-NVMe Ceph storage. Ansible with Netbox as the source of truth for everything — patch management, configuration drift control, integrations into Cloudflare, PowerDNS, and Active Directory. OSPF dynamic routing for cloud networks. Local repository mirrors for maximum deployment speed and consistency. CentOS 7 to Rocky 9 redeployment of the HPC environment. Cloudflare for DNS (including DNSSEC), DDoS protection, and Zero Trust remote access. This included migrating 25 authoritative DNS zones from self-hosted BIND 9 to Cloudflare in two working days. Multiple registries were updated, and the whole thing was integrated with Ansible and Let\u0026rsquo;s Encrypt.\nAll of it on open-source tools, tight budget, built deliberately to stay clear of the Broadcom and VMware licensing trap before it shut. That settled the argument. Open-source infrastructure at that scale is not just workable — it is better, and I have the cluster and the invoices to say so.\nThe other half of that job was the four people in the team. Scientists do not care what the storage is called. They care whether the job runs tonight, and whether the person they ask can explain the answer without making them feel daft for asking.\nHow it all comes together The move into presales at croit wasn\u0026rsquo;t a change of direction. It was everything landing in one place.\nManufacturing IT Reliability isn't optional when production depends on you Front-line support Learning to listen, communicate and stay patient Application development Building systems that solve real business problems Global IT management Carrying the technical load while developing people Deep Linux engineering The last line of escalation — nobody else to call Critical DNS at Nominet National-scale infrastructure and discipline Infrastructure \u0026amp; hosting VMware, storage and disaster recovery at scale Science computing (UKCEH) Rebuilt on Proxmox VE + Ceph + HPC, open source throughout Presales at croit Design the system, defend the design, be honest about it Manufacturing IT taught me that reliability stops being a preference the moment production depends on it. Front-line support taught me to listen before doing anything else. Application development taught me to build systems that solve a business problem rather than an interesting one. Managing a global team taught me to carry the technical load and still develop the people around me, which is harder than either half on its own. Third-line Linux taught me what the last line of escalation feels like. Running Windows and Linux side by side for twenty years taught me how platforms actually behave under load, as against how the datasheet says they will. VMware, which I ran in nearly every role from 2008 on, taught me what enterprise virtualisation looks like at scale — and then what happens when the commercial ground shifts underneath a platform the whole estate is sat on. Storage and networking taught me where the hard problems live. Security has run through the lot, from firewalls and VPNs early on to DNSSEC, Zero Trust and hardening now. Nominet taught me discipline.\nPut all that together and you get presales. You design the system, then you defend the design, and you stay honest about what it will not do, because somebody is about to spend real money on your say-so.\nAt croit, that means working with outfits across Europe and Asia-Pacific who are rethinking their virtualisation and storage. The conversations are usually about leaving VMware for Proxmox VE with Ceph. The workloads coming across are mostly Windows, so the cross-platform experience isn\u0026rsquo;t old news. It\u0026rsquo;s what I do now.\nThe part I care about is the honesty. What I put in front of somebody has to be the thing that works in their building, at their scale, with the constraints they actually have, and if their existing kit does most of the job already then that is what I will tell them.\nTHE STACK I DESIGN \u0026amp; BUILD TODAY Automation \u0026amp; source of truth Ansible · NetBox · Let's Encrypt Compute Proxmox VE — KVM virtual machines + LXC containers Storage Ceph — RBD block · CephFS · all-NVMe Network 25/100GbE fabric · BGP / OSPF · native IPv6 Foundation Open source, on hardware you own Storage Storage kept being where the hardest problems sat. It wasn\u0026rsquo;t a deliberate career plan — it just worked out that way.\nThe fascination started young. In the 90s I dreamed of Iomega Zip and Jaz disks — 100MB, and then a whole gigabyte, on a single removable cartridge, when the floppies everyone else swapped around held just 1.44MB. They were expensive kit, so for years they stayed a want rather than a have. By the time I finally got a USB Zip drive, USB flash drives had just started to appear — and the format I\u0026rsquo;d wanted for so long was already on its way out. An early lesson in how fast storage moves, and how quickly today\u0026rsquo;s must-have becomes tomorrow\u0026rsquo;s drawer clutter.\nOptical was part of the same story. In 1998 I had a quad-speed HP CD-RW drive, and it was a cracker. It could read and rewrite discs that later, faster drives simply refused — so it earned its keep as a rescue drive long after it should have been retired.\nIt started at the interconnect. I\u0026rsquo;ve worked with storage hardware across the lot. IDE, multiple generations of SCSI, SATA, SAS, and NVMe on the direct-attached side. ATA over Ethernet and iSCSI on the network storage side. HP 3par, Dell Compellent, Dell PowerVault, Nexsan — each with its own quirks and failure modes.\nFrom there it went up through the stack. Clustered LVM taught me how shared storage behaves when multiple nodes need concurrent access — and what happens when locking and fencing aren\u0026rsquo;t right. ZFS taught me what happens when you actually think about data integrity at the filesystem level. GPFS showed me parallel filesystems at scale. Ceph taught me how distributed systems behave differently under failure.\nOver the years, what started as \u0026ldquo;the person who also does the storage\u0026rdquo; became \u0026ldquo;the person you call when the storage needs to be designed properly.\u0026rdquo;\nNetworking Networking has been there from the start. It\u0026rsquo;s not a supporting skill — it\u0026rsquo;s a core one. TCP/IP, DHCP, and DNS span every single role I\u0026rsquo;ve held from McKenna onwards.\nWorking at Nominet meant working on the infrastructure behind the UK\u0026rsquo;s domain name registry. That\u0026rsquo;s DNS at national scale. Naming, resolution, delegation, and the expectation that it works every single time.\nI ran authoritative zones on BIND 9, configured DNSSEC, set up F5 load balancers for IPv4 and IPv6, and later migrated DNS estates to Cloudflare with API automation and Let\u0026rsquo;s Encrypt integration. I\u0026rsquo;ve built PXE boot DNS infrastructure for cluster deployments in both the UK and China. I understand DNS from both sides. Self-hosted, where you own every failure. And managed, where you\u0026rsquo;re trusting a provider and need to verify that trust through monitoring.\nBeyond DNS, networking runs through every role I\u0026rsquo;ve held. VLANs, bonding, LACP, fibre-channel zoning with Brocade, Fortinet and Cisco ASA firewalling, IPTables and IPset. 25GbE and 100GbE fabric design, BGP, OSPF dynamic routing, MTU management, IPv6 architecture. VPN tunnelling — from OpenVPN across difficult international links to WireGuard and cloudflared for modern secure access. Samba has been a professional skill across multiple roles too, bridging Linux file sharing and Windows domain integration long before I was testing the AD implementation in alpha.\nA Ceph cluster is a network application. The storage performance ceiling is set by the network underneath it. The failure modes are network failure modes. That\u0026rsquo;s not summat you learn from a textbook. It\u0026rsquo;s summat you learn from troubleshooting at two in the morning.\nLinux I\u0026rsquo;ve worked across RHEL, Debian, SUSE, and Fedora families for over two decades. Professionally, that\u0026rsquo;s meant RHEL and CentOS for live services, Ubuntu for application hosting, Rocky for HPC, and Gentoo and Fedora for personal use. I\u0026rsquo;ve passed LinkedIn\u0026rsquo;s Linux skill assessment.\nThe depth came from problems that have no Stack Overflow answer, which is where you end up learning the PCI subsystem properly rather than by rumour: IOMMU groups, ACS, VFIO, SR-IOV, Resizable BAR and what DMA translation is quietly costing you, and kernel boot parameters as controls that change how the machine behaves rather than a list to copy off a wiki because somebody\u0026rsquo;s blog said it fixed their thing.\nStorage and virtualisation performance problems nearly always trace back to a layer nobody looked at. That is where my Linux knowledge lives.\nWindows Linux is where I spend my time now, but the Windows side is just as real and it has been there in every job I have had. Windows Server, Active Directory, LDAP, Exchange, IIS, Microsoft SQL Server. None of that is a legacy skill I have put down.\nIt matters at the guest level, which is where most people stop looking. When a Windows workload is running on KVM, the person who can tell you how that guest behaves under a given CPU topology, and can trace a performance complaint back to a Microsoft patch instead of blaming the hypervisor, is the person who has spent years on both sides of the fence. Most VM performance arguments I have watched were won by somebody knowing the guest, not the host.\nThe people I have supported run from professors and board directors to the team that looks after the buildings. The job is the same in every case. Find out what they actually need, get it working, and explain it in a way that doesn\u0026rsquo;t leave them feeling daft for having asked.\nVMware VMware is the platform I grew up on professionally, across seven roles from 2008 onwards, and I ran the full stack of it: ESXi, vCenter, vSAN, vSphere clustering, vMotion. I know what it does well. I know where it doesn\u0026rsquo;t. And I watched the Broadcom acquisition rewrite the commercial reality underneath outfits that had built the whole estate on it, which is a different thing from reading about it.\nYou cannot help somebody leave a platform you never properly worked on. As such I can tell them what they are giving up, what they are getting, and which part of the migration will be worse than they have been led to expect.\nAutomation Automation thinking started in my first professional role in 2006. VB scripts at McKenna to automate AD drive mapping, printer assignment, and software deployment. Then HR-to-AD provisioning tools at BT Engage IT. Then PXE boot systems for diskless compute clusters at Sondrel. Then Puppet at Nominet. Then Chocolatey packaging and PXE rebuild at a retail company. Then ServiceNow workflow customisation across multiple roles.\nNow it\u0026rsquo;s Ansible with Netbox as the source of truth. Not because they\u0026rsquo;re trendy, but because infrastructure that can\u0026rsquo;t be rebuilt from code isn\u0026rsquo;t infrastructure you can trust.\nDocumentation is the same argument. Knowledge that lives only in somebody\u0026rsquo;s head is a single point of failure, and it walks out of the building at five o\u0026rsquo;clock with the rest of them. Same as a disk with no redundancy behind it, and treated with the same seriousness.\nMonitoring My monitoring background started with SolarWinds at Pulsant, running scale monitoring across the hosting estate. Zabbix came later, and it has earned its own mention because I\u0026rsquo;ve put it in at nearly every place since. It began at Nominet, where I ran the tender and the testing and then moved us onto Zabbix to replace the ageing VMware Hyperic setup, and it won on being genuinely intuitive rather than another thing wearing Nagios underneath. After that I put it in from scratch at a hosting and infrastructure platform, replaced somebody\u0026rsquo;s failed attempt at it in retail, and used it to watch an 8-node HPC cluster at UKCEH.\nEach time I designed the monitoring templates and wrote the scripts myself. By the time I sat the Zabbix Specialist and Professional exams in 2018, I\u0026rsquo;d already been deploying it for years.\nCommunication Working in critical infrastructure teaches you to be precise. Working in presales teaches you to be clear. The two are related but not the same.\nBeing precise means nowt if the person listening cannot follow the reasoning, so I have had to learn to do the same explanation twice over: once for the engineer who wants to see the CRUSH map, and once for whoever has to sign off why the budget is the number it is. I don\u0026rsquo;t dumb owt down. I just make every step in the reasoning visible, and let them stop me where they want to.\nYears of end users taught me a different sort of patience. The people who reckon themselves the smartest in the room are usually the ones who behave the most helplessly. The ones who open with \u0026ldquo;I\u0026rsquo;m not very technical\u0026rdquo; tend to listen, follow the steps and get it sorted in ten minutes. And the ones who tell you they know exactly what they are doing have generally made it worse before they picked the phone up.\nThe business systems background helps here more than people expect, because I have always had to translate in both directions. SQL architecture for an operations manager, a storage design for a CTO, or sitting next to somebody at their own desk showing them the thing they were never shown. Same skill every time.\nCommunity Sharing knowledge has run through the whole career.\nI was active on the Gentoo forums when I was building systems from source and needed to understand why a USE flag combination was breaking a compile. I contributed on the Samba forums when I was working through file sharing and domain integration problems across Linux and Windows. That went beyond just asking questions. I built a Samba-based Active Directory domain controller while the AD support was still in alpha. I tested it against Windows 2000, XP, and Vista clients and fed back into the community.\nNow I\u0026rsquo;m an active contributor on the Proxmox community forums as DamienDye. I help with NVMe passthrough performance, Windows VM tuning, Ceph troubleshooting, and cluster networking.\nThe technologies change. The principle doesn\u0026rsquo;t. If I have worked a problem out, there is nowt to be gained by sitting on it, and somebody at two in the morning six months from now will be glad it is written down.\nI also have a habit of digging into problems beyond what the task strictly needs. I\u0026rsquo;ve designed a DIY NVMe power loss protection board because I wanted to understand exactly why FTL corruption happens at the hardware level. I\u0026rsquo;ve researched self-hosted email protocols because I wanted to understand JMAP from the RFC up rather than just trusting a provider. I built this blog because writing things down properly is how you find the gaps in your own understanding.\nCommunity off the keyboard Not all of it has been forums.\nIn 2018 I was one of the founders of the Longford Park Community Association in Banbury, on the estate where I live. Four phases of new build, a community centre written into the plans, and nowt in existence to run it. So a handful of residents set summat up.\nI did the IT end first, because that is what I had to give. The lpca.org.uk domain went live in January 2018, and behind it I built the website, the mail systems and the committee lists, and wrote the privacy policy.\nI was on the committee from the start, and I was its chairman from December 2018 to February 2020. Fourteen months, and the things that actually mattered all landed inside them. We got it registered as a charity on 11 February 2019, number 1181953. I sorted the commercial lease for the building. Then we opened the centre.\nThat is what the chair was for. Not agendas and minutes. Getting a building open that a few hundred households could walk into. Resident volunteers still run it across phases 1 to 4 of the development — three rooms, a kitchen and a car park, hired out by the hour to whoever on the estate wants them.\nVolunteer committees run on goodwill, and goodwill is not a governance model. So I held people to the constitution, myself included. I asked why we were recruiting from outside when the constitution said engage the residents, and I pulled up a mailout that was going to land in every spam folder on the estate. Not to be awkward. A residents\u0026rsquo; association that cannot reach the residents has already failed at the only job it has, and I sent the links to fix it along with the complaint.\nThe centre is still open and I am not on the committee, and that is the right way round. Anything that only works while you are stood there holding it up was never built properly.\nThis site This blog is a Hugo static site using the PaperMod theme. It\u0026rsquo;s hosted on Cloudflare Workers, serving the built site as static assets.\nI write here about the infrastructure I work with day to day: Proxmox VE, Ceph, Ansible, Netbox, certificates and whatever else I have been digging into that week. It is written to be used, with the commands and the numbers in it, because a post you cannot follow at the keyboard is decoration. No marketing fluff. If summat has rough edges, the post says so.\nGet in touch You can find me on LinkedIn or on the Proxmox forums.\nIf you are looking at Ceph or Proxmox for your place and want a proper conversation about it rather than a pitch, drop me a line. Bring the workload, the constraints and the budget you actually have. I will tell you what it will do, what it won\u0026rsquo;t, and if the honest answer is that you should keep what you have got and configure it properly, you will get that answer as well.\n","permalink":"https://blogs.damiendye.uk/en/about/","summary":"Damien Dye — Presales Engineer, infrastructure specialist, and open-source advocate.","title":"About"}]