IPsec was a good idea. Put the encryption at the network layer, underneath everything, and every protocol running over IP inherits confidentiality and integrity without being told. No library to link. No per-application certificate. No rewriting the thing you already shipped. The packet leaves protected and arrives protected, and the routers in between carry it without knowing or caring what is inside.

That design rests on one assumption, and it is stated in the standard rather than implied: a packet’s address identifies the machine it came from. A security association is looked up on the destination address, the protocol number and the SPI1. The Authentication Header goes further and signs the IP header itself, source and destination included2. The address is not routing metadata to IPsec. It is part of the identity and part of the integrity check.

Then this industry spent thirty years taking the address away.

First NAT, to make an office fit behind one line. Then carrier-grade NAT, to make several hundred houses fit behind one address, because turning on IPv6 was work and buying a box was procurement — which is the whole subject of We Never Ran Out of Addresses. We Ran Out of Effort. and I will not go back over it here. What matters for this post is the consequence. The one thing IPsec was built on is the one thing the modern access network no longer provides.

So IPsec got patched. Wrap the encrypted packet in UDP so a translator has a port to rewrite. Zero the checksum so nothing revalidates it. Send a one-byte packet every twenty seconds, for ever, so a table in somebody else’s kit does not forget you exist. Move the identity off the address and onto a name. Abandon the Authentication Header, because it cannot survive a rewritten header by construction. When a hotel blocks UDP, wrap the whole lot in TCP as well3.

Every one of those is a real, standardised, vendor-supported fix. Together they are a protocol held in place by its own scaffolding. And the scaffolding is the argument: you do not spend a quarter of a century propping something up because it is fundamentally sound.

The conclusion I have come to is that IPsec should be retired in full. Not tuned, not re-proposed with better ciphers, not kept for site-to-site because that bit still works. Retired, with dates, the way PPTP should have been retired a decade before anyone got round to it. What follows is the evidence, the diagrams, the vendor documentation that says all of this in the vendors’ own words, and — because most people reading this still have to keep the things running on Monday — a working method for diagnosing IPsec faults in the meantime.

What It Actually Costs You, In Tickets

Before the standards, here is the bill, in the order you will meet it.

The tunnel drops on a clock. Every hour, or every eight, or after twenty minutes of no traffic. It comes back when somebody opens a file, so half the users never report it and the other half report it as “the VPN is slow”. Nobody has changed anything.

Two people in the same house cannot both connect. The second one comes up, the first one goes down. They ring the service desk separately, so the tickets never meet and nobody spots the pattern for a fortnight.

Small things work and big things hang. Login works. Teams works. Ping works. Copying a file stalls at the same place every time, and opening a large page in an internal application sits there until it times out.

The tunnel is up and no traffic passes. Both ends say established. Both ends are pleased with themselves. Nothing moves.

Nothing can reach in. Site-to-site to the branch that moved onto a fibre altnet no longer establishes in the direction it used to, and nobody can say why, only that “their IP changed”.

Turning on QoS broke encryption. Somebody prioritised voice, and now the far end is discarding packets as replays.

Not one of those is a misconfiguration in the ordinary sense. Every one of them is IPsec meeting the network as it now is. The rest of this post is why, in order, and how to prove which one you have got.

IPsec Is Not Carried By The Network. It Is The Network.

Start with what it was, because the design is genuinely good and the failures only make sense against it.

ESP is not a protocol running over TCP or UDP. It is a transport protocol, IP protocol number 50, sat directly on IP in the same slot TCP and UDP sit in. AH is protocol 51. Neither has a port field, because neither needs one: on the internet IPsec was designed for, the destination address already names exactly one machine, and the SPI in the ESP header names which security association on that machine. Address plus protocol plus SPI. That triple is the lookup1.

The design as specified: the address names the machine, so no ports are needed and either end can startIPsec as specified: the address is the identityHost A203.0.113.10Routersforward onlyHost B198.51.100.7either endcan startWhat leaves the machineouter IP headersrc 203.0.113.10 to 198.51.100.7protocol50 = ESPSPI4 Bsequence4 Bencrypted payloadyour packet, unreadable in transitICV16 BAssociation lookup = destination address + protocol + SPI. No port field anywhere, because none is needed.Three things this design assumes, all of them true in 1995:The address on the packet is the machine. Nothing in the path rewrites a header. Either end may be dialled.The Authentication Header goes further still and signs the outer header itself, source and destination included,so a receiver can prove the addresses were not touched on the way.Not to scale. ESP overhead varies with the cipher; the figures here are AES-GCM in tunnel mode over IPv4.
The design as specified. ESP sits directly on IP as protocol 50, with no ports, because it does not need them — the destination address names the host and the SPI names the association on it. AH signs the IP header itself. Both ends hold a real, reachable address, either end can start the conversation, and nothing in the middle needs to understand the payload.

Look at what that buys. There is no handshake port to expose, no session layer to get wrong, no application that has to opt in. Either end can start. The middle of the network is dumb, which is exactly what the middle of a network should be. A router forwards protocol 50 the same way it forwards protocol 6, and the fact that it cannot read the payload is the point rather than a limitation.

It is a clean design. It is also, in 2026, a description of an internet that most people reading this cannot buy.

Then Somebody Put A Translator In Every Path

A NAT rewrites the source address, and often the source port, so several machines can share one address. That is all it does. Against IPsec it is close to a complete demolition, and the IETF was straight enough about it to publish an entire document listing the pieces: RFC 3715, IPsec-Network Address Translation (NAT) Compatibility Requirements4. Sixteen separate incompatibilities. Here are the ones that matter.

AH is finished, by construction. In RFC 3715’s own words: “Since the AH header incorporates the IP source and destination addresses in the keyed message integrity check, NAT or reverse NAT devices making changes to address fields will invalidate the message integrity check.”4 There is no fix for that and there was never going to be. A protocol that signs the header cannot cross a box whose whole job is rewriting the header. AH did not get worked around. It got abandoned.

There are no ports to translate. A NAT doing port translation needs a port. ESP has not got one. Cisco put it plainly in the Catalyst configuration guide: “If PAT found a legislative IP address and port, it would drop the Encapsulating Security Payload (ESP) packet.”5 The packet is not rejected on policy. It is dropped because the box has nowhere to write the thing it needs to write.

The identity stops matching the packet. Again from RFC 3715: “Where IP addresses are used as identifiers in Internet Key Exchange Protocol (IKE) Phase 1 or Phase 2, modification of the IP source or destination addresses by NATs or reverse NATs will result in a mismatch between the identifiers and the addresses in the IP header.”4 An identity scheme that names machines by address cannot survive a device that renames machines for a living.

Two machines can pick the same SPI. The SPI is chosen by the receiver and only has to be unique to it. Put two hosts behind one address and the translator has two associations with nothing to tell them apart.

Nothing can come in first. A NAT builds its table from outbound packets. There is no outbound packet until someone starts, and on a NATted link only the inside can start. Half the protocol’s symmetry is gone.

One translator in the path breaks five things at once, and every fix for them costs somethingOne translator in the path, and five things break at onceHost192.0.2.20Translatorrewrites source to 203.0.113.5Gateway198.51.100.7nothing can start from this sideWhat the rewrite destroys1The signed header no longer matches.AH covers the source and destination addresses, so theintegrity check fails by construction. There is no fix. AH is simply not usable across a translator.2There is no port to rewrite.ESP is IP protocol 50 and carries no port field, so a translatordoing port translation has nothing to work with and drops the packet.3The identity stops matching the packet.The key exchange named the peer by its address.The address in the header is now somebody else's.4Two hosts can choose the same SPI.The receiver picks it and only guarantees it is uniqueto itself, so the translator has two associations it cannot tell apart.5Half the symmetry is gone.The table is built from outbound packets, so only the inside can start.The compromise, and what each part of it costsWrap the whole ESP packet in UDP on port 4500, so the translator has something it understands.Abandon the Authentication Header entirely, because nothing can save it.Zero the UDP checksum, because a checksum over rewritten addresses would only fail.Move the identity off the address and onto a name the peer asserts instead.Send one byte every twenty seconds, for ever, so a table in somebody else's kit does not forget you.It works. That is not in dispute. The argument is that nothing sound needs five concessions to cross one box.
The same tunnel with one translator in the path. Five things break at once: the signed header no longer matches, there is no port for the translator to rewrite, the identity no longer matches the source address, two hosts can choose the same SPI, and nothing outside can start a conversation. The row underneath is what the industry did about it — and each fix is something given up.

The Fix Was Real, And Every Part Of It Cost Something

NAT traversal works. That is not in dispute and I am not going to pretend otherwise. I have run plenty of tunnels through plenty of NATs. What I want on the record is the bill, because it is paid every day by everyone and almost nobody itemises it.

The mechanism is RFC 3948: detect a translator during the key exchange, then put the whole ESP packet inside a UDP datagram on port 4500 so the translator has something it understands6. Cisco’s description of the wire format is exact: after encryption, “a UDP header and a non-IKE marker (which is 8 bytes in length) are inserted between the original IP header and ESP header”5. Juniper’s is shorter and says the same thing: “NAT-T encapsulates both IKE and ESP traffic within UDP with port 4500 used as both the source and destination port.”7

Now the itemised bill.

The Authentication Header is gone. Not deprecated politely. Unusable. RFC 8221 now says plainly that using ESP and AH together is NOT RECOMMENDED8, and the honest reason is that the one thing AH did that ESP does not is the one thing NAT destroys.

The UDP checksum is deliberately zeroed. RFC 3948 requires it: with the addresses rewritten, a checksum computed over them would fail, so the standard’s answer is to stop computing it. “If the protocol header after the ESP header is a UDP header, set the checksum field to zero in the UDP header.”6 A layer of error detection removed to keep a layer of address rewriting working.

You now send packets to keep a table warm. RFC 3948 defines a keepalive as one byte, 0xFF, sent when nothing else has gone out for a configurable interval whose default is twenty seconds6. Juniper states the reason without decoration: “Because NAT devices age out stale UDP translations, keepalive messages are required between the peers.”7 So a laptop on a train, doing nothing at all, transmits every twenty seconds, for ever, because a table in a box neither end owns will otherwise forget it exists. Multiply by a fleet. That is the radio waking up, the battery going down, and the mobile network carrying traffic whose entire purpose is to prevent forgetting.

You now firewall two ports instead of one protocol, and Cisco’s own restriction list requires static translation rules for both 500 and 4500 to be in place before any of it works5.

And read the rest of that restriction list, because it is the vendor telling you the shape of the thing. Dynamic NAT policies: unsupported. IPv6 traffic: incompatible with the feature. IPsec and NAT on the same device: cannot both operate5. That is not a configuration guide. That is a list of places the scaffolding does not reach.

Nowt about that is elegant, and none of it is anybody’s fault in particular. It is what happens when you keep a layer-3 protocol alive on a network that stopped honouring layer 3.

Carrier-Grade NAT Removed What Was Left

Ordinary NAT took the address off the machine and gave it to the site. You still owned the translator, so you could forward a port, pin a mapping, lengthen a timer, or put the concentrator in front of it.

Carrier-grade NAT takes the address off the site and gives it to several hundred strangers, and the translator belongs to your ISP. Everything you used to be able to do about it, you now cannot.

And be clear about what is actually in the path, because this is the bit people get wrong. The router in the house is still doing NAT. It has not been switched off. It still translates your laptop on 192.168.1.20 to whatever address the line holds — except the line now holds 100.64.12.7, which is shared space, not a public address. The carrier then translates that again. So the packet crosses two translators before it reaches the internet, and that is the ordinary case, not an unusual one.

Carrier-grade NAT means two translations, two timers, and no reachable address at either of themCarrier-grade NAT is two translations, not oneLaptop192.168.1.20Router in the housetranslation oneThe line100.64.12.7Carrier translatortranslation two, to 203.0.113.9yours, and the only table you can seetheirs, shared with several hundred homes, invisible to younothing from the internet can reach either of those addressesTwo tables, two timers, and the shorter one wins.Your keepalive has to beat whichever ofthe two ages out first, and you can only read one of them. The one that matters is the one you cannot see.Port forwarding still works and achieves nothing.The router happily forwards a port from anaddress the internet cannot reach. UPnP and PCP report success and open a door onto a corridor. This is why"I have forwarded 500 and 4500 and it still will not come up" is such a common and such a misleading ticket.The responder role is gone.Two branches on consumer fibre cannot dial each other at all.Something in the middle has to introduce them, and now you depend on a company you never chose.Identity cannot be the address.Several subscribers arrive at the far end as one address, sowhatever tells them apart, it is not the field IPsec was designed to use. And there is a published ceiling:on the larger SRX platforms Juniper state no more than 1,000 tunnels from any one translated address.The quickest confirmation costs nothing: read the WAN address in the router. If it starts 100.64, that is theshared space set aside in RFC 6598, you are behind carrier-grade NAT, and half the settings page is decorative.A business line with a real static address has one translation and a reachable endpoint. That is what the extramonthly charge is actually selling you: the thing every machine used to have for nowt.
The packet is translated twice: once by the router in the house, once by the carrier. Neither address it holds is reachable from outside. That means two mapping tables with two independent timers where only one is visible to you, port forwarding that succeeds and achieves nothing, no responder role at all, and a peer identity that can no longer be an address.

You are double-NATted, and only one of the tables is yours. Two translations means two mapping tables, two ageing timers, and two chances of the mapping going away. Your keepalive has to beat whichever expires first, and you can read exactly one of them. Worse, the two interact: the router in the house may have its own idea about IPsec and try to help with an application gateway, so the source port your client thinks it is using is not the one leaving the house, and is not the one leaving the carrier either.

Port forwarding still works, and does nothing at all. This is the ticket that eats the most time. Somebody forwards UDP 500 and 4500 on the home router, the router accepts it, the settings page says the rule is active — and nothing can use it, because it forwards from an address the internet cannot reach. UPnP and PCP behave the same way: the client asks for a mapping, the router grants it, and the port opens onto a corridor. Everything reports success and nothing works. The box will tell you owt you want to hear.

The quickest way to settle it costs nothing. Read the WAN address in the router. If it begins 100.64, that is the shared space set aside in RFC 6598, you are behind carrier-grade NAT, and half that settings page is decorative.

The responder role no longer exists. A device behind CGNAT cannot be the end anybody dials. Site-to-site between two branches on consumer fibre — normal, cheap, and exactly what a small business wants — requires at least one end to hold a real address, or a third party in the middle to introduce them. That third party is a company you now depend on because your ISP would not give you an address.

The timer belongs to someone else. RFC 4787 tells NAT operators a UDP mapping “MUST NOT expire in less than two minutes” and recommends five or more9. That is the floor and the advice, not a promise, and you cannot inspect what your carrier actually does. As such the keepalive stops being a tuning option. It is load-bearing, on a timer you are guessing at.

Your peer’s identity cannot be its address. Several subscribers reach the far end as one address. Whatever the concentrator uses to tell them apart, it is not the IP header — which is the thing IPsec was designed to use.

There is a hard ceiling, and the vendors publish it. Juniper document that on SRX5400, SRX5600 and SRX5800, “the total number of tunnels from a given public translated IP cannot exceed 1000 tunnels”7. Read that as an operator rather than as a spec line. Your VPN concentrator has a per-shared-address limit, the sharing is done by a carrier you have no contract with, and how close you are to that limit depends on how many of your users happen to sit behind the same one. There is no counter you can look at. There is only the day it starts failing for some people and not others.

Everything in this section is downstream of one decision this country made and kept making. I have written that case out in full elsewhere and I am not repeating it. The point here is narrower: IPsec’s core assumption was deleted by the access network, and IPsec has been living on workarounds ever since.

And The IPv6-Only Network Does Not Save It Either

Here is where I have to be honest against my own argument, because the obvious reply to everything above is: fine, so give every machine a real IPv6 address and IPsec works as specified again.

It does, between two ends that both have one. That is not the network most people are on.

The IPv6-only access networks that actually exist, mobile in particular, reach the IPv4 internet through NAT64, which is not a NAT at all in the ordinary sense. It is a protocol translator, rewriting an IPv6 packet into an IPv4 one. And RFC 6146 names what it will carry, and names what it will not, without any ambiguity at all:

“The current specification only defines how stateful NAT64 translates unicast packets carrying TCP, UDP, and ICMP traffic. Multicast packets and other protocols, including the Stream Control Transmission Protocol (SCTP), the Datagram Congestion Control Protocol (DCCP), and IPsec, are out of the scope of this specification.”10

IPsec, by name, out of scope. And packets carrying anything outside that list “SHOULD be discarded”10.

So on an IPv6-only phone or laptop trying to reach an IPv4 concentrator, and that is most concentrators, native ESP does not get dropped by a firewall or mangled by a translator. It is never carried in the first place. The translator is doing precisely what its standard tells it to do.

The industry’s answer to that is, inevitably, another layer: 464XLAT, which gives the device a local IPv4 stack and translates twice, out of IPv4 and back, so that things NAT64 cannot carry work anyway11. It is already on the list of bolt-ons in the IPv6 post and I am not going to re-argue it. What is worth saying here is the shape: a protocol that broke on NAT in 2004 also breaks on the translation built for the IPv6 transition in 2011, and both are answered by wrapping it in something else.

That is the test a protocol has to pass to belong in 2026. Does it work on the network people actually have — behind a carrier’s shared address, on an IPv6-only mobile network, through a hotel that only passes TCP 443? Anything built on a UDP port passes all three without being told. IPsec needs a different workaround for each, and on the third one it needs TCP encapsulation as well3.

Here is a failure that has nothing to do with NAT, is purely modern, and gets almost no airtime.

Routers and switches spread traffic over parallel paths. Link aggregation, equal-cost multipath: both work the same way, by hashing the five-tuple. Source address, destination address, protocol, source port, destination port. ESP has no ports. So every packet of a tunnel between the same two addresses hashes identically, and the whole tunnel lands on one member link, regardless of how many you bought.

Two 10G links and one IPsec tunnel gives you 10G. Four gives you 10G. The kit is working exactly as designed.

The workaround is the usual shape. Some silicon can hash on the SPI instead, since each SPI names one association and therefore one flow, but that is a feature you have to have bought rather than something you can assume of a path you do not own. There is an active IETF draft whose entire purpose is to wrap ESP in yet another UDP header so that ordinary routers can hash it, and it says why in one sentence: “Although the ESP SPI field within the IPsec packets can be used as the load-balancing key, but it cannot be used by legacy switches and routers.”12 Its problem statement is equally blunt about what people do instead: “Many cloud service providers allow customers to establish multiple IPsec VPN tunnels in parallel to enable ECMP and increase aggregate bandwidth. However, this approach is not ideal, as each tunnel typically requires its own public IP address, leading to higher public IP consumption and increased operational overhead.”12

Sit with that for a second. The recommended way to make an encrypted link go faster is to build several of them, each burning a public IPv4 address — during an address shortage — because the protocol has no port number to hash on. Meanwhile a UDP-based tunnel gets multipath for free, from hardware that shipped fifteen years ago, because it has a port like everything else on the modern internet.

That is not a legacy problem waiting to age out. It is a live limit on new builds, today, at exactly the speeds people are now buying.

The MTU Tax, And Who Pays It

Every tunnel costs bytes. IPsec costs more than most, and the way it fails when it runs out is the worst kind of fault: intermittent, size-dependent, and invisible to every test anybody runs first.

What each tunnel spends per packet before any of your data goes inBytes gone before your data starts, to scaleheaders, per packettotalNative ESPtunnel, AES-GCMIP20ESP8IV8trailer + tag1854 BESP in UDPport 4500, for NATIP20UDP8ESP8IV8trailer + tag1862 BL2TP/IPsecAES-CBC, NAT-TIP20UDP8ESP8IV16UDP8L2TP6PPP4trailer + tag1888 BPPTPGRE, protocol 47IP20GRE16PPP4no integrity tag at all40 BWireGuardone UDP portIP20UDP8header16tag1660 BThe shaded blocks are the price of somebody else's translator.In the L2TP row, three ofthem sit inside the encryption: a second UDP header, a session layer, and the framing from a dial-up modem.PPTP is cheapest because it protects nothing.The 40 bytes buy no integrity check, and MS-CHAPv2was reduced to a single DES operation in 2012. Cheap is not the measure. What the bytes buy is the measure.One worked configuration each, over IPv4. Exact figures move with the cipher, the mode and the address family.
Bytes spent on every packet before any of your data goes in, for one worked configuration each. Native ESP is lean. Wrapping it for NAT costs eight more. L2TP/IPsec carries a UDP header, an L2TP header and a PPP header inside the encryption — dial-up framing, encrypted, in 2026. PPTP looks cheap because it does not carry an authentication tag at all, which is the whole problem with it.

Work it through for one configuration rather than hand-waving. ESP in tunnel mode over IPv4 with AES-GCM: 20 bytes of outer IP header, 8 of ESP header, 8 of nonce, 2 minimum of trailer, 16 of integrity tag. 54 bytes before any of your packet goes in. Wrap it for NAT traversal and the UDP header makes it 62. On a 1500-byte path that leaves 1438, and the moment anything upstream is running PPPoE at 1492 you are down again.

Then the bit that makes it a fault rather than an arithmetic problem. A sender finds out a packet was too big only by receiving an ICMP error back — Fragmentation Needed on IPv4, Packet Too Big on IPv6. If anything on the path drops those errors, the sender never learns, and it keeps sending packets that keep dying. That is the classic black hole: the handshake completes because handshakes are small, and the transfer hangs because transfers are not.

Cisco have maintained a whole document about this since the GRE and IPsec era, and it is still one of the better explanations of the interaction in anyone’s library13. The reason it needs maintaining is that people keep blocking ICMP wholesale and then wondering why tunnels behave oddly.

I have written the two halves of that argument out already and they hold here without restating: which ICMP messages are load-bearing and which one is not, in Ping: The Diagnostic Tool That Opens a Whole Lot More, and how to find the exact hop that is eating your traffic, in The Firewall Is Eleven Hops Away. The short version for this post: the errors are the mechanism, echo is not, and a border policy that drops all of ICMP has broken your VPN in a way that will be blamed on the VPN.

And Your Own Quality Of Service Can Break It

One more, because it catches good engineers doing the right thing.

ESP carries a sequence number and the receiver keeps a replay window of 64 packets by default on Cisco platforms14. Now prioritise voice on the sending router. Low-latency queueing does what you asked and reorders packets relative to the sequence they were encrypted in. If a packet falls outside the window by the time it arrives, the far end discards it as a replay, and the counter that goes up is a security counter.

Cisco’s own words: “Certain QoS features, such as Low Latency Queueing (LLQ), could cause IPsec packet delivery to become out-of-order and dropped by the receiving endpoint due to a replay check failure.”14 Their answer is to widen the window to 1024 where the platform supports it, or to adopt a multiple-sequence-number-space extension that maps QoS classes to separate sequence spaces within one association14.

So: turning on a standard feature of your own network breaks your own tunnel, and the remedy is another protocol extension. It is the same shape as everything above.

L2TP: Dial-Up Framing, Encrypted, In 2026

L2TP is not a security protocol and never claimed to be. RFC 2661 is a tunnelling protocol for carrying PPP sessions, published in 1999, with no confidentiality of its own — its own security section points you at IPsec for packet-level protection15. That pairing is RFC 3193, and the result is the stack in the diagram above: an outer IP header, a UDP header for NAT traversal, ESP, then inside the encryption another UDP header on port 1701, an L2TP header, and a PPP header.

PPP. The framing from dial-up modems, carried inside an encrypted tunnel, over the internet, in 2026, because that is what L2TP was written to carry.

Count what that costs on a worked example — outer IP 20, UDP 8, ESP header and IV 24, inner UDP 8, L2TP 6, PPP 4, trailer and tag 18 — and you are spending about 88 bytes per packet to move data that native ESP moves for 54. Thirty-four bytes, every packet, for a session layer that adds nothing you wanted and a link layer designed for a telephone line.

The overhead is the least of it.

It multiplies the NAT problem rather than dividing it. L2TP/IPsec conventionally uses ESP in transport mode, which is the mode NAT hurts most, and the well-known result is that many implementations cannot support two clients behind one address at all. Two people in one house, or forty in one office, or several hundred behind a carrier’s shared address. The far end sees one address and cannot tell the sessions apart, so the second connection replaces the first. That is the ticket from the top of this post, and it is not a bug in anybody’s product — it is the identity-by-address problem arriving in the place where the most people meet it.

And in the field it is usually deployed with one shared secret for everybody. Because the secret is configured on the client profile and handed out with the setup instructions, the pre-shared key is in the onboarding document, the wiki page, the email to new starters, and every laptop that ever left. It is not a second factor. It is a password that authenticates the gateway to nobody in particular and has never once been rotated.

L2TP is not a protocol that has aged badly. It is a protocol that was carrying the wrong thing on day one, and got bolted to IPsec to make up for what it could not do at all.

PPTP Was Never Secure, And Is Still On Sale

PPTP deserves two paragraphs, not a section, and the reason it gets them is that people still ship it.

It was never a standard. RFC 2637 is Informational. A vendor protocol written up, not something the IETF ever recommended. It carries PPP inside GRE, IP protocol 47, which like ESP has no ports, so it needs its own special-case handling in every NAT on the path — the “PPTP passthrough” checkbox, which on a great many home routers supports exactly one session at a time.

The security ended publicly in 2012. Marlinspike and Hulton showed that MS-CHAPv2’s security reduces to a single DES operation regardless of password length, built chapcrack to extract the handshake, and wired it to a cracking service that returned the key within a day for twenty dollars — a 100% success rate, not a probability16. Their conclusion was that PPTP traffic should be considered unencrypted. Apple agreed with their feet and removed PPTP from the built-in client in macOS Sierra and iOS 10 in 2016, and still publish the warning17.

Ten years on from that, PPTP is still a menu item on routers being sold this year, still in vendor how-to pages, still the thing somebody turns on because it is the one that works first time. It works first time because it is not doing the job.

Everything That Calls Itself An IPsec VPN

“IPsec VPN” is not a protocol. It is a family, and the length of the list below is the argument, because no two products implement the same subset of it, and the gaps between those subsets are where every interop job you have ever hated actually lives.

PieceWhat it addsWhere it stands
ESP, IP protocol 50The encryption and integrity itselfRFC 4303 — current18
AH, IP protocol 51Integrity over the IP header tooRFC 4302 — unusable through NAT2
IKEv1The original key exchangeDeprecated, RFCs moved to Historic19
IKEv2The current key exchangeRFC 729620
IPComp, IP protocol 108Compresses before encrypting, with its own associationsRFC 317321
PF_KEY v2A kernel API so a daemon can load the keysRFC 236722
NAT traversalWraps ESP in UDP 4500 so a translator can copeRFC 3947 / 39486
TCP encapsulationFor networks that block UDP as wellRFC 9329, replacing RFC 82293
IKEv2 fragmentationBecause the key exchange itself outgrew the MTURFC 738323
MOBIKESo the tunnel survives the address changingRFC 455524
Dead peer detectionA heartbeat, because nothing else tells youRFC 370625
XAUTHUser authentication — password, token, RADIUSNever an RFC. Expired draft, 200126
Mode-ConfigHands the client an address, DNS and routesNever an RFC either26
L2TP/IPsecCarries PPP inside the tunnelRFC 2661 + RFC 319327
GRE or VTI over IPsecGives you a routable interface to run a protocol onVendor architecture on top of ESP
DMVPNmGRE plus NHRP plus IPsec, so spokes find each otherVendor architecture, NHRP RFC 233228
GETVPNGroup keys with no pairwise tunnel at allGDOI, RFC 640729
PPTPThe thing IPsec was supposed to replaceRFC 2637 — Informational, never a standard30

Now look at the two rows in bold, because they are the ones that should stop you.

For the best part of two decades, the standard way to log a user into a corporate IPsec VPN was XAUTH — your username and password, your token, your RADIUS server — with Mode-Config handing the client its address, DNS servers and routes. Between them they are the entire remote-access experience. Every “Cisco IPsec” client, every VPN icon in a tray, every set of joining instructions.

Neither of them is a standard. XAUTH was an individual internet draft that expired in 2001 and was archived without ever becoming an RFC26. The stated reason is worth reading, because it is a committee explaining why it would not do its job: the draft records that the IPSRA working group would not accept any protocol extending ISAKMP or IKE, and the IPsec working group refused anything dealing with remote access26. So the most widely deployed part of the most widely deployed VPN protocol was left homeless, implemented anyway by every vendor from their own reading of an expired draft, and shipped to millions of users.

IKEv2 eventually fixed both, and it is worth saying so: authentication moved to EAP and the client configuration payloads went into the core specification20. But read the dates. The most-used feature of the most-used VPN protocol ran on an expired draft for about a decade before it had a standard at all, and the installed base carried on running the draft version for years after that. A protocol family does not get credit for eventually standardising the part everybody was already using.

That is the family you are running. Some of it is standards-track and current. Some of it is Historic. Some of it never got as far as being anything. And a product that says “IPsec VPN” on the datasheet has told you approximately nothing about which of these eighteen things it does, which is why connecting two of them together is a fortnight, a spreadsheet of proposals, and a phone call to somebody who has done it before.

The Fisher-Price OS (Windows) Has Never Really Interoperated

This is the part where somebody says the problem is really Linux, so let us do it with sources.

By default, the Windows client will not connect to an IPsec server that sits behind NAT at all. Not “will struggle”. Will refuse. The fix is a registry value called AssumeUDPEncapsulationContextOnSendRule under HKEY_LOCAL_MACHINE\SYSTEM\CurrentControlSet\Services\PolicyAgent, set to 1 if the server is behind a translator or 2 if both ends are, on every client and on the server, followed by a reboot31. Microsoft’s own words: “By default, Windows Vista and Windows Server 2008 don’t support Internet Protocol security (IPsec) network address translation (NAT) Traversal (NAT-T) security associations to servers that are located behind a NAT device.”31

Two things about that. The first is that Microsoft co-authored the NAT traversal standard it is refusing to use. Their name is on RFC 3947 and RFC 39486. The second is that the page telling you to edit the registry was last revised in February 202631. Twenty years on, a registry hack on every endpoint is still the answer, and it is still being maintained as the answer.

And read the sentence Microsoft puts just above it: “If you must use IPsec for communication, use public IP addresses for all servers that you can connect to from the Internet.”31

That is this entire post, in the vendor’s own documentation. Do not put IPsec behind NAT. Give every machine a real address. They wrote it down, and then the industry spent two decades doing the opposite and billing for the difference.

It does not stop at NAT. Go and read what an open-source gateway has to document to accept a Windows client.

  • The gateway certificate needs an extended key usage that exists for this and nothing else. serverAuth, OID 1.3.6.1.5.5.7.3.1, plus IP Security IKE Intermediate, OID 1.3.6.1.5.5.8.2.232. Your CA has to be told to emit an OID most tooling has never heard of, or the connection fails with a policy error and no useful message.
  • A second registry value is needed before the client will offer decent cryptography. strongSwan documents adding NegotiateDH2048_AES256 under Rasman\Parameters to get AES-256-CBC and a 2048-bit group33. Read that the other way round, which is the way that matters: without a registry edit, the default proposal is weaker than that.
  • Server-initiated rekeying is rejected by clients behind NAT, and the documented workaround is to disable rekeying on the gateway and let the client start it33.
  • Standard IKEv2 extensions are simply absent — no IKE redirection, no multiple authentication rounds33.

None of those is a Linux bug. Every one is an open-source project writing down what it has to do to accommodate one vendor’s reading of a standard that vendor helped write.

That is the pattern, and it is thirty years old. PPTP was Microsoft’s protocol, written up as Informational and never standardised30. MS-CHAPv2 and MPPE were Microsoft’s authentication and encryption, and both were broken in public16. SSTP is a Microsoft tunnel nobody else terminates. DirectAccess was IPsec, and it was Windows at both ends by design — and it has now been deprecated and is being removed, with customers pushed onto Always On VPN instead34. Not one of those was ever a protocol the rest of us could meet in the middle. They were a protocol you joined, and if you did not run the right operating system on both ends you got the registry hack, the odd OID and the workaround page.

So let me be plain, since it is my blog. The Fisher-Price OS (Windows) has never been a genuine peer in an open protocol stack, because being a peer was never what it was for. It hides the machine from the person using it as a design goal, and a stack you cannot see is a stack you cannot make interoperate. If it is the only operating system you have ever administered, the sections above about reading kernel counters and watching the wire will have read like another language, and that is the gap — not a preference, a gap.

None of which need cost the reader anything. Every diagnostic in this post runs from any Unix on the network, pointed at whatever is broken, and it does not care in the slightest what the far end is running. And if that operating system is the only one you have, the diagnostic section below carries its own tools for it as well — the capture, the two cmdlets, and the error codes with what each one is really telling you. A broken tunnel still has to be fixed on Monday. Then the replacement I argue for at the end has one client, behaving the same way, on every platform including that one, which is the first time in thirty years that has been true of a VPN.

The Deprecation Record Reads Like An Obituary

Set the workarounds aside and just read what the standards bodies have done to this family over the years. Not opinion. Requirement levels, in published RFCs.

WhatWhere it stands nowSource
IKEv1Deprecated; RFCs 2407, 2408 and 2409 moved to HistoricRFC 9395, 202319
DES in ESPMUST NOTRFC 82218
3DES in ESPSHOULD NOTRFC 82218
HMAC-MD5-96MUST NOTRFC 82218
ESP together with AHNOT RECOMMENDEDRFC 82218
Encryption-only ESPShown insecure, and broken in practice in 2007Degabriele and Paterson35
IPsec on an IPv6 nodeDowngraded from MUST to SHOULDRFC 6434, 201136
PPTPNever a standard; Informational onlyRFC 263730

The last-but-one row is the one I would put in front of anyone who tells me IPsec is fine and the network is the problem. IPv6 originally mandated IPsec — it was the security story, written into the node requirements. In 2011 the IETF changed its mind: “Previously, IPv6 mandated implementation of IPsec and recommended the key management approach of IKE. This document updates that recommendation by making support of the IPsec Architecture a SHOULD for all IPv6 nodes.”36

Even the address family that would have given IPsec back everything NAT took away stopped requiring it fifteen years ago. That is not the network failing IPsec. That is the people who designed the network deciding it had not earned the mandate.

The Complexity Was Flagged In 1999, In Writing

None of this is hindsight, and that is what makes it worth writing down.

In 1999 Niels Ferguson and Bruce Schneier were commissioned to evaluate IPsec. Their report is short, plain and worth reading whole. It opens with “IPsec was a great disappointment to us. Given the quality of the people that worked on it and the time that was spent on it, we expected a much better result.”37 It names the cause: “Our main criticism of IPsec is its complexity. IPsec contains too many options and too much flexibility; there are often several ways of doing the same or similar things. This is a typical committee effect.”37

And it made three recommendations that read now like a list of things that happened anyway, twenty years late and the hard way:

  • Eliminate transport mode. “We therefore recommend that transport mode be eliminated.”37 Transport mode is the mode L2TP/IPsec uses, and it is the mode NAT hurts worst.
  • Eliminate AH. “We conclude that eliminating transport mode allows the elimination of the AH protocol as well, without loss of functionality.”37 NAT eliminated it instead, for a worse reason.
  • Never allow encryption without authentication. They warned that administrators “will be quite likely to configure ESP for only encryption, believing that it provides security.”37

Eight years later that last one stopped being a warning. Degabriele and Paterson published attacks that “break any RFC-compliant implementation of IPsec making use of encryption-only ESP” — ciphertext-only, needing nothing more than the ability to watch traffic and inject packets35. The prediction was in the public record for eight years and the standard still permitted the configuration.

The verdict of that 1999 report is the sentence I keep coming back to: “We have found serious security weaknesses in all major components of IPsec. As always in security, there is no prize for getting 90% right; you have to get everything right.”37

Two more data points, and then I will leave it.

Logjam, 2015. The team behind it scanned a 1% sample of IPv4 for IKE and found that 86.1% of IKEv1 and 91.0% of IKEv2 servers supported the 1024-bit Oakley Group 2, and that 66.1% of profiled IKEv1 servers preferred it. Their conclusion: precomputation against a second 1024-bit group “would allow decryption of traffic to 66% of IPsec VPNs”, and the published intelligence documents on VPN exploitation are “consistent with having achieved such a break”38. Cryptographic agility, the thing IPsec has most of, is what let almost everybody sit on the same weak group for fifteen years.

CVE-2016-1287. A buffer overflow in the IKEv1 and IKEv2 code on Cisco ASA, reachable by sending crafted UDP packets, giving remote code execution before authentication39. Think about where that box sits. It is the device you deliberately exposed to the entire internet, running the most option-laden protocol in the estate, with a fragment reassembler in front of the parser, holding the keys to everything inside. The complexity Ferguson and Schneier warned about is not an abstraction. It is attack surface, on the one machine you cannot put behind anything.

Diagnosing It While You Still Run It

You cannot turn all this off this afternoon, so here is how to work on it. This is the method, in the order that costs least, with what each result actually means.

Six ways it gives up, placed on the path where each one happensWhere each failure actually livesClientpolicy and routesHome routertranslation oneCarriertranslation twoThe internetfilters and MTUGatewayselectors and identity1234561Nothing on the wire at all.The traffic never reached IPsec. Stop looking at crypto.ip xfrm policyand the route, and whatever the host firewall is doing.2Port 500 both ways, 4500 never appears.NAT traversal was not negotiated.tcpdump -ni eth0 'udp port 500 or udp port 4500 or ip proto 50'Bare ESP will not survive it.3Dies after idle, revives on first use.A mapping aged out at one of the two translators.Your keepalive is losing to a timer you cannot read. Shorten it, and stop trusting the default.4Outbound counter rising, inbound flat.ESP is dying in one direction, in transit.ip -s xfrm stateat both ends, then find the hop that is eating it.5Login fine, large transfers hang.Path MTU, and the ICMP errors are not coming back.ping -M do -s 1400stepping down, then clamp the segment size and fix the ICMP rule.6Both ends say up, nothing passes.Traffic selectors, policy or routing — not the keys.And if a second user knocked the first one off the tunnel, it is peer identity, not capacity.One capture at the border for thirty seconds answers the first three. Do that before you open either console.
The six failures in the order to test for them, with the symptom, the check and what the answer means. Each row is a different layer of the stack giving up, and the first three are answered by watching the wire for thirty seconds — which is why that is the first thing to do, not the last.

Watch The Wire First, Not The Console

Both consoles will tell you what they believe. The wire tells you what happened. One capture at the border, thirty seconds, answers the first three questions at once:

tcpdump -ni eth0 'udp port 500 or udp port 4500 or ip proto 50 or ip6 proto 50'
  • Nothing outbound at all — the problem is in front of IPsec: routing, policy, or a host firewall. Stop looking at crypto.
  • Outbound only, nothing back — your packets are leaving and their answers are not arriving. Filtering in transit, a dead peer, or the far end rejecting silently.
  • UDP 500 both ways but 4500 never appears — NAT traversal was not negotiated. Either one end has it disabled or the detection failed.
  • Protocol 50 on the wire while one end is behind NAT — the negotiation decided there was no translator when there is one. It will never come back.

Reading The Key Exchange Failure

IKEv2 tells you why it refused, and the notify names are specific enough to diagnose from alone. It helps to have the shape of the whole exchange in front of you first, because every notify below belongs to a particular rung of it.

The IKEv2 exchange step by step, and which fault lives at each stepEvery step of the exchange, and the fault that lives on itClientbehind a translatorGatewayreal addressWhat fails hereIKE_SA_INIT request — UDP 500nothing back — 809 ERROR_VPN_TIMEOUTIKE_SA_INIT response — UDP 500NO_PROPOSAL_CHOSEN, INVALID_KE_PAYLOADNAT detected, both ends float to 4500no float — proto 50 dies at the translatorIKE_AUTH — UDP 4500, encryptedtoo big, fragment lost, retransmit, timeoutIKE_AUTH response — CHILD_SA createdAUTHENTICATION_FAILED — 13801, 13806ESP in UDP 4500 — your trafficTS_UNACCEPTABLE — up, and nothing moveskeepalive — 1 byte every 20 s, for everno keepalive — mapping expires on idleThe float is the hinge.Above it everything is UDP 500. Below it everything is UDP 4500. If the float never happens you are puttingprotocol 50 into a translator that has nothing to rewrite, and it will never come back.The size fault lives in one rung.IKE_AUTH carries the certificate chain, so it is the only large message on this ladder. A pre-shared key that connectswhere a certificate does not is a dropped fragment, not a bad certificate.Created is not the same as working.The child association can exist and still carry nothing, if the two ends disagree about which traffic it covers.Every fault in the right-hand column is reported at the client as a timeout, whatever it actually was.
The exchange from the first packet to steady state, with the fault that lives on each step. The float from 500 to 4500 is the hinge: everything above it is one port, everything below it is another, and if the float never happens you are putting protocol 50 into a translator that cannot carry it. IKE_AUTH is the only large message here, which is why a pre-shared key can connect where a certificate does not — that is a dropped fragment, not a bad certificate.
NotifyWhat it actually meansWhere to look
NO_PROPOSAL_CHOSENNot one of the offered cipher/integrity/DH/PRF combinations is acceptable to the far endBoth proposal lists; expect a deprecated algorithm on one side
INVALID_KE_PAYLOADDiffie-Hellman group mismatch — you offered one group, it wants anotherThe DH group, first thing in the proposal
AUTHENTICATION_FAILEDKey or certificate or identity wrong — a pre-shared key mismatch, an expired certificate, or an ID that is not what the peer expectsThe identity, not just the secret
TS_UNACCEPTABLEThe traffic selectors do not overlap — you asked to protect subnets the peer will not protectBoth ends’ selector configuration
INVALID_SPIA packet arrived for an association that no longer exists, usually after a one-sided restartWhether one end has rekeyed or rebooted

Juniper’s Phase 2 guidance says the same thing about the commonest of them: “no proposal chosen” means “the device did not accept any of the IKE Phase 2 proposals that the peer sent”, and the fix is a mutually acceptable proposal rather than a repeated restart40.

On strongSwan, the state of everything in one command:

swanctl --list-sas                 # what is established, and what it negotiated
swanctl --log                      # the negotiation as it happens

On Cisco, show crypto ikev2 sa and show crypto ipsec sa, with debug crypto ikev2 when it will not come up41. On Junos, show security ike security-associations and show security ipsec security-associations, with the negotiation in show log kmd-logs42.

Was NAT Traversal Actually Negotiated?

This is the check people skip, and it explains a large share of “it works from the office and not from home”.

Each end sends hashes of the addresses and ports it believes are in play. If the hash the far end computes from the packet it received does not match the one you sent, there is a translator between you, and both ends move to UDP 4500. If detection fails — one side has traversal disabled, or something in the middle is mangling the exchange — both ends carry on with bare ESP, which will not survive the translator.

So: see port 4500 in the capture, from both directions, or there is no traversal happening. Do not take the console’s word for it.

Then check the keepalive is actually running and its interval is below whatever your carrier ages mappings at. The default is twenty seconds6; the standard’s floor for NAT operators is two minutes9; what your particular carrier does, you cannot see. If the tunnel dies after idle and revives on traffic, this is your fault every time.

The Kernel Counters Almost Nobody Reads

On Linux the transform layer keeps a full error breakdown, and it is the fastest way to turn “it does not work” into a specific cause. The counters are documented by the kernel itself43.

cat /proc/net/xfrm_stat        # error counters, by cause
ip -s xfrm state               # per-SA packet and byte counters
ip xfrm policy                 # what should be protected, and in which direction

It reads better as a path than as a list. The packet passes through five stages, and each one has its own counter:

Five stages, and the counter that names the one your packet died atFollow one packet, and let the counter name the stage it died atYour policydoes this traffic get protected?XfrmOutPolBlockyou are dropping it yourself, on purpose, in policyYour associationencrypt, seal, number itXfrmOutNoStatespolicy matched and there is no association to carry itThe pathtranslator, MTU, filtersnothing increments, at either endthis is where NAT, MTU and filtering live, and no counter can see any of itTheir associationdest + protocol + SPIXfrmInNoStates · XfrmInStateProtoError · XfrmInStateSeqErrorit arrived and the crypto did not work out: wrong SPI, wrong key, out of windowTheir policywas this meant to be protected?XfrmInTmplMismatch · XfrmInNoPolsthe crypto was fine and the policy disagreed about whether it should have beenStage three is the one with no counter.Everything a quarter of a century of translation did to this protocol happens there, and the transform layer atneither end can see a single packet of it. That is the whole reason a capture comes before a console.Stages four and five are opposite faults.Four means the crypto did not work out. Five means it did, and something disagreed about whether it shouldhave. They get looked at in the wrong order almost every time, because five looks like a crypto fault and is not.Counter names are Linux. On Junos the same stages read out of show security ipsec statistics; on theFisher-Price OS (Windows), out of Get-NetIPsecQuickModeSA and the Windows Firewall with Advanced Security log.
Five stages, and the counter that names the one your packet died at. Stage three is the only one with no counter at all — the translator, the MTU and every filter in between live there, and the transform layer at neither end can see a single packet of it. That is the whole argument for reaching for a capture before a console. Stages four and five are opposite faults and get looked at in the wrong order almost every time.

Watch which one moves while the fault happens:

CounterKernel’s descriptionWhat it means on the day
XfrmInNoStates“No state is found i.e. Either inbound SPI, address, or IPsec protocol at SA is wrong”Their packets are arriving for an association you do not have — usually a one-sided rekey or restart
XfrmInStateSeqError“Sequence error i.e. Sequence number is out of window”Reordering or replay-window trouble; see the QoS interaction above
XfrmInStateProtoError“Transformation protocol specific error e.g. SA key is wrong”Keys disagree — the association survived a rekey on one side only
XfrmInTmplMismatch“No matching template for states e.g. Inbound SAs are correct but SP rule is wrong”The association is right and the policy is wrong
XfrmInNoPols“No policy is found for states e.g. Inbound SAs are correct but no SP is found”Protected traffic arriving that nothing asked to protect
XfrmOutPolBlock“Policy discards”You are dropping it yourself, on purpose, in policy
XfrmOutNoStates“No state is found”Traffic matched a policy with no association to carry it — the tunnel never came up

XfrmInTmplMismatch and XfrmInNoPols are the two worth knowing by sight, because both mean the crypto is fine and the policy is not, which is the opposite of where everybody looks first.

It Says Up And Nothing Moves

Both ends established, no traffic. Read the per-association counters in both directions:

ip -s xfrm state
  • Outbound bytes rising, inbound flat — you are encrypting and sending, and nothing is coming back. Either your ESP is not reaching them or theirs is not reaching you. Ask the far end for their outbound counter; if it is rising too, the packets are dying in transit and the next question is where, which is a TTL question rather than a crypto one.
  • Both flat — nothing is being offered to the tunnel. Routing or policy, not IPsec. On a route-based setup check the route actually points at the tunnel interface; on a policy-based one check the selectors.
  • Both rising, applications still broken — it is not the tunnel. Go and look at what is on the far side.

Juniper’s guidance for this case is the same instinct in their idiom: if only the outbound packet counter on the session is incrementing, confirm with the peer whether the traffic is being received at all44.

Small Things Work, Big Things Hang

MTU. It is always MTU. The test takes ten seconds:

ping -M do -s 1400 10.0.0.1      # inside the tunnel, do-not-fragment set
ping -M do -s 1300 10.0.0.1      # step down until it succeeds

Where it starts succeeding tells you the real usable size. Then make TCP find out for itself, by clamping the advertised segment size to the path rather than hoping every ICMP error survives the journey:

nft add rule inet filter forward tcp flags syn tcp option maxseg size set rt mtu

And fix the underlying cause too, which is nearly always an over-broad ICMP rule at a border somewhere. Drop echo if you want. I have argued elsewhere that it deserves dropping. But keep Fragmentation Needed and Packet Too Big. They are the mechanism, not a nicety.

It Drops On A Clock

Time the failures. The interval names the cause on its own.

  • A fixed period matching a configured lifetime — rekey. The association expires and the replacement negotiation is failing or racing. Check both ends’ lifetimes; mismatched values are normal and fine, but a hard lifetime on one end shorter than the other’s soft lifetime produces exactly this.
  • After a period of no traffic, back on first use — a NAT mapping expired. Keepalive interval, or the absence of one.
  • Dead-peer detection tearing it down while the link is fine — the probes are being lost rather than the peer being dead, often because the probes are the only traffic and the mapping has already gone.

The Second User Kills The First

Two peers reaching the concentrator from one address, authenticating as the same identity. The gateway has a choice between keeping the old association and replacing it, and a common default is to replace. So the second connection wins and the first one silently dies.

Give every peer a genuinely unique identity rather than an address or a shared name, and set the gateway to keep multiple associations from one address rather than assuming one per peer. Then test it the only way that counts: two clients, one address, at the same time. If your acceptance test has never had two users behind one NAT, you have not tested the case that most of your users are in.

The Same Faults, On The Fisher-Price OS (Windows)

Every check above runs from a Unix box, and if you have one on that network then use it, because it will tell you the truth faster and it does not care what the far end runs. But a lot of readers have a client that will not connect, an operating system that hides the machine from them on purpose, and nothing else to look at. So here is the same method again, in the same order, with the tools that operating system actually ships.

Watch the wire. There is no tcpdump, but there is a capture. Run it elevated, reproduce the failure, stop it:

netsh wfp capture start cab=on file=ipsec
netsh wfp capture stop

That writes a .cab. Inside it is the trace of what the filtering platform and the key exchange actually did while the fault was happening, which is more than either console will admit to. For a live look rather than an archive, the same command set prints to the console with file=-: netsh wfp show state “Displays the current state of WFP and IPsec”, and netsh wfp show ikeevents “Displays recent Internet Key Exchange (IKE) epoch events matching the specified parameters”, filtered to one peer45.

netsh wfp show state file=-
netsh wfp show ikeevents remoteaddr=203.0.113.5 file=-

show ikeevents is the nearest thing here to the exchange log every other stack writes without being asked. Worth knowing it is there. Otherwise the interface hands you a three-digit number and nothing else, and you are diagnosing a key exchange by guesswork.

Read the associations, and read both of them. The equivalents of ip -s xfrm state and swanctl --list-sas are two cmdlets, and the split between them is the whole diagnosis:

Get-NetIPsecMainModeSA
Get-NetIPsecQuickModeSA

Main mode is the key exchange. Quick mode carries the packets. Microsoft puts the relationship plainly: “There is only one main mode SA between a pair of computers, but there can be many quick mode SAs”46. So main mode present with quick mode empty is the same fault as an IKE_SA up with no CHILD_SA under it, and it means the same thing: the two ends agreed on how to talk and then failed to agree on what to protect. Look at the traffic selectors, not the ciphers.

Read the logs, in the two places they hide. Connection-level failures land in the Application log against the RasClient source, and Microsoft’s own note on reading them is the useful bit: “All error messages return the error code at the end of the message”47. That number is the diagnosis, and the next section is what the numbers mean. Policy and filtering decisions land somewhere else entirely, under Applications and Services Logs, in the Windows Firewall with Advanced Security channels. Two logs, two teams, one fault.

Collect it properly when you have to escalate. The supported bundle is TSS — TSS.ps1 -Scenario NET_VPN on the client, TSS.ps1 -Scenario NET_RAS on the server, started before you reproduce the fault and stopped after48. Learn it before somebody asks you for it.

When It Sits On Connecting, And Then Times Out

This is the fault that fills the tickets, and the word in the error is a lie. Start by naming the code. The code is specific even when the message is useless.

A connect timeout is one of three things, and none of them is a clockWhat the client calls a timeout, and what it actually wasThe client says: timed out809, 718, 828, 930, 638 — five codes, one word, and not one of them is a clockIt is one of three things, and one test tells you whichNothing arrivedthe silence caseTest: capture at the client — does UDP 500 leave and nothing come back?Filtering in the path, or NAT traversal never negotiated, so 4500 was never sent at all.It arrived, too bigthe fragment caseTest: does a pre-shared key connect where a certificate does not?IKE_AUTH carries the chain, so it fragments, and the fragment is dropped in the path.Never the tunnelthe wrong-layer caseTest: does the gateway log an authentication failure at the same second?930 is RADIUS. 812 is the authentication method. 13801 and 13806 are certificates.None of the three is a timer.Lengthening the timeout is the one thing the interface invites you to do, and the one thing that has never oncefixed any of them. The word is there because the layer reporting the failure cannot see far enough to say more.Test in that order. The first costs thirty seconds of capture, the second costs one connection attempt, the third costs a log.
Five error codes carry the word timeout and not one of them is a clock. It is one of three things and a single test separates them: a capture says whether anything came back at all, one connection attempt with a pre-shared key says whether the certificate exchange was simply too big to arrive, and the gateway’s own log says whether the tunnel was ever the problem. Test them in that order, because that is also the order of what they cost.
CodeName in raserror.hWhat actually happened
809ERROR_VPN_TIMEOUTNothing came back at all. Microsoft’s stated cause: “the UDP 500 or 4500 ports on the VPN server or firewall are blocked”47 — but blocked covers three different things and only one is a deny rule. Usually 4500 was never sent, because NAT traversal was never negotiated, or the helper that used to carry it has been switched off. Read on before you ask for a firewall change
789ERROR_OAKLEY_GENERAL_PROCESSING“The L2TP connection attempt failed because the security layer encountered a processing error during initial negotiations” — credentials or certificates, not the network
718ERROR_PPP_TIMEOUTThe IPsec part worked. PPP inside the L2TP tunnel got no answer
828ERROR_IDLE_TIMEOUT“The connection was terminated because of idle timeout” — a server-side setting, deliberately
930ERROR_AUTH_SERVER_TIMEOUTRADIUS did not answer in time. Nothing to do with IPsec whatsoever
638ERROR_REQUEST_TIMEOUTThe generic one. Treat it as no information and go to the wire

Every one of those is sourced from Microsoft’s own error list49. Now read 809 again. It is called ERROR_VPN_TIMEOUT, and the documented cause is a blocked port. The client is not timing out because the far end is slow. It is timing out because the far end is silent, and silence is the only failure mode a protocol with no ports and no handshake visibility can report.

Be careful with that word blocked, though, because it is carrying more than it can hold. Three different things wear it, and only one of them is a deny rule.

Nothing ever sent 4500 in the first place. NAT traversal was not negotiated, so the client carried on speaking ESP, and ESP gives a translator nothing to rewrite. The default on that platform is reason enough on its own: without the registry value, it will not form a NAT-T association to a server behind a translator at all31. So 4500 is not blocked. It was never tried.

The helper that used to paper over it has been switched off. Firewalls and home routers carry per-protocol helpers — IPsec passthrough on consumer kit, and the whole application-layer gateway family behind it — that read a protocol the NAT cannot handle and open the return path for it. On Linux the automatic form of that was turned off by default at kernel 4.7 “for security reasons”, with the guidance since being to attach a helper deliberately with a rule or not at all; the helpers it covers include the one for PPTP50.

And switching them off was the right call. A helper is a piece of your firewall that parses a payload and then punches a hole based on what it read. NAT Slipstreaming is the bill for that: a browser visiting a page, traffic shaped so the router’s SIP or H.323 helper reads it as a call, and a pinhole opened through the NAT — in the 2021 version, to any internal address, not just the machine that loaded the page51. Turning helpers off closes that. It also stops your IPsec working. Both are true at once, and the second is not an argument for undoing the first.

Which is the post in miniature, again. The protocol only worked because boxes in the middle were reading traffic that was not theirs and opening holes on its behalf, and the industry has spent the last decade correctly deciding to stop doing that.

So work the causes in this order, cheapest first.

One: nothing is coming back. Capture at the client, or ask the gateway. If UDP 500 leaves and nothing returns, it is filtering or reachability and no timer will fix it. If 500 completes both ways and 4500 never appears, NAT traversal did not negotiate. And if either end is behind a translator, you are back at the registry value from earlier in this post: AssumeUDPEncapsulationContextOnSendRule, 1 or 2, on both machines, then a reboot31. Without it the client refuses by design and reports it as a timeout.

Two: the answer is too big to arrive. This one wastes whole afternoons. Certificate authentication makes the second exchange large, because it carries a chain, and a large exchange fragments. Fragments get dropped by the same middleboxes as everything else in this post, the client retransmits into the same hole, and then it gives up and says timeout. The tell is diagnostic on its own: a pre-shared key connects and a certificate does not. That is not a certificate fault. That is a size fault, because the pre-shared key exchange is small enough to fit. Standard IKEv2 fragmentation exists precisely for this23, and strongSwan records when this platform got it: “IKEv2 fragmentation is supported since the v1803 release of Windows 10 and Windows Server”33. Anything older, or a gateway with fragmentation disabled, and you are relying on a path that will carry a fragmented UDP datagram. Plenty will not.

Three: it connects, then drops on a clock. Time it. If it dies after a fixed idle period, that is 828 and it is configuration, not a fault — -IdleDisconnectSeconds on the server, with -SALifeTimeSeconds, -MMSALifeTimeSeconds and -SADataSizeForRenegotiationKilobytes as the other three clocks that can end a session52. That last one ends it on volume rather than time, which is why a session can die reliably during a large file copy and never during a day of email. And if it dies at rekey with the client behind NAT, that is the documented interop fault from earlier: the client rejects a server-initiated rekey with Microsoft error 13863, and the gateway-side answer is to stop initiating and let the client do it33.

Four: it is not the tunnel at all. 930 is RADIUS. 812 is an authentication method the server would not accept47. 13801 and 13806 are certificates — wrong extended key usage, expired, missing root, or a server name that does not match the certificate’s subject47. The tunnel negotiation was fine in all four cases, and if you spend the afternoon on ciphers you will not find any of them.

Here is the thing worth taking away from that table. Six error codes, five of them with the word timeout in the name, and not one of them is a timeout in fact. They are a blocked port, a lost fragment, a policy decision and a RADIUS server, all wearing the same word, because the layer reporting the failure cannot see enough of what happened to say anything more useful. Lengthening the timer fixes none of them, and lengthening the timer is what the interface invites you to do.

What To Run Instead

I am not going to pretend the replacement is exotic. It is in the kernel and it has been for years.

WireGuard is one UDP port, one key per peer, no cipher negotiation and no protocol agility at all. The author’s position on that is deliberate and stated: “It intentionally lacks cipher and protocol agility. If holes are found in the underlying primitives, all endpoints will be required to update. As shown by the continuing torrent of SSL/TLS vulnerabilities, cipher agility increases complexity monumentally.”53 That single decision deletes NO_PROPOSAL_CHOSEN, INVALID_KE_PAYLOAD, downgrade attacks and the Logjam finding in one go, because there is nothing to negotiate and nothing to downgrade.

It is honest about the trade too, and so am I: no agility means that when a primitive does fall, you update the entire fleet rather than flipping a config line. That is a real operational cost and it is the right one to pay.

The rest lines up against the list above almost item for item. Having a UDP port means NAT and CGNAT treat it like any other flow, and ECMP and LAG hash it like any other flow. Roaming is built in rather than bolted on — an authenticated packet from a new address moves the peer’s endpoint, so a phone going from WiFi to mobile does not renegotiate anything. It answers nothing to an unauthenticated packet, so a scanner finds a closed port where IPsec would find a concentrator to talk to. And it is under 4,000 lines of code53 against a stack that needs a document listing its sixteen incompatibilities with one middlebox.

On performance it simply measured faster than both IPsec configurations it was benchmarked against, at 1,011 Mbit/s against 881 and 825, with lower latency53. I would not retire a protocol over a benchmark. I mention it because the last argument standing for IPsec is usually performance, and it is not true either.

For getting a person to an application — rather than a network to a network — the answer is not a tunnel at all. Identity at the front door, the application published through it, nothing routed. I have built that out with Proxmox and Cloudflare Access in Zero Trust VDI Without the Cloud Bill, and the relevant point for this post is that a user who needs three internal applications does not need a route to your entire estate.

And say the quiet part about hardware. Yes, there are NICs and ASICs with ESP offload, and that is a genuine argument for IPsec on specific kit at specific speeds. It is an argument about silicon someone already sold you, not about the protocol being right. As such it dates, and it dates fast. PPTP survived in exactly the same way, for exactly as long as the checkbox existed.

Retiring It Properly

Retirement is a plan with dates, not a feeling. Here is the one I would put my name to.

Stop new IPsec deployments now. Not “prefer alternatives”. Stop. Every new tunnel is a tunnel someone has to migrate later, and the ones being built today will still be running in 2035 unless somebody says no this week.

Do remote access first, because that is where every failure in this post lands hardest: the CGNAT, the shared address, the keepalives, the second user in the same house, the MTU black hole on a random hotel network. It is also the easiest to move, because the endpoints are managed and the change is a client.

Then site-to-site over the public internet, which is the same set of problems with fewer endpoints and a maintenance window.

Leave for last the tunnels you do not own both ends of: a partner, a regulator, a carrier’s managed service. Those move when the contract moves, and the way to make them move is the next point.

Stop buying kit whose only tunnel is IPsec. Put it in the tender. A line asking for a modern, port-based, roaming-capable tunnel is a line a vendor either answers or does not, and it is how the installed-base argument finally loses. That argument is the only one keeping any of this alive, and it is only ever defeated by purchasing.

Write down which tunnels are left and why, and put a date on each. A protocol nobody has owned for ten years is how PPTP got to 2026. An inventory with dates is the difference between retiring something and merely disliking it.

If You Are Still Deploying It, Can You Call Yourself A Professional?

That is a genuine question and I am going to answer it honestly, because it is the one the rest of this post has been walking towards.

It depends on whether you know. And knowing is not something that happens to you. Making sure you know is the job.

If you are standing up a new IPsec remote-access service this year and you cannot say, without looking any of it up, why ESP has no ports, what a residential line under carrier-grade NAT does to it, why a NAT64 network will not carry it at all, or what that one-byte packet every twenty seconds is actually for — then no. Not on this, not yet. You are not choosing a protocol. You are repeating a shape, because the last one looked like that and nobody in the room asked why — including you. The failing is not the gap. Everybody has gaps. The failing is building across one you never went and closed. As such the person who inherits it in 2035 gets a decade of tickets that were all avoidable on the day you drew it.

Let me put the right name on that, because the polite version has been in circulation for twenty years and it has changed nothing. Somebody who deploys a protocol they cannot explain is not an engineer. They are a follower. They read the cue cards — the vendor’s reference architecture, the last change request, a diagram somebody drew in 2014 and nobody has opened since — and they read them out with real conviction, and there is no understanding anywhere behind the performance. From the outside it looks exactly like competence. It goes on looking exactly like competence right up to the first failure the cards do not cover, and from that moment it is the only thing in the room that matters.

The cue cards are also the reason this protocol is still here. Nobody stood at a whiteboard in 2026 and argued for IPsec on the merits. It got deployed again because it was on the card, and the card was written by a vendor whose interest is that you keep buying the box that terminates it. That is how something outlives its own obituary by twenty years — not by being defended, but by never once being made to justify itself in a room where somebody could tell the difference.

And this gets called out. Out loud, in the room, at the time — not muttered in the corridor afterwards. If somebody puts the word professional next to their name, the word arrives with an invitation to be asked, and asking is not rude. Being asked and having an answer is the entire difference between the word and a business card.

It matters most when you are paying for it. A consultancy, an MSP, a vendor’s professional services arm, the integrator on the framework — what you are buying is judgement, and judgement is the one thing you cannot inspect on delivery. So inspect it beforehand. Ask why this protocol and not another one. Ask what happens to it on a carrier-grade NAT line, on an IPv6-only mobile network, on a path that quietly drops fragments. Ask which of those they have personally hit and what they did about it. You will know inside two minutes whether you are being told something or being read to, and two minutes is a considerably cheaper test than four years of tickets. And if the answer is cue cards and you sign anyway, that is a decision as well — it has just become yours rather than theirs.

If you can say all of that, and you are deploying it anyway because a regulator names it by protocol, because the partner’s kit terminates nothing else, or because the replacement is in next year’s budget and this has to work in March — then yes, obviously, and you are doing the job properly. Constraints are real and I have built round worse. What separates the two is not the protocol on the diagram. It is whether you wrote down why, and whether there is a date next to it.

The indefensible position is the middle one. Knowing enough to be uneasy, and building it anyway because nobody made you justify it. Not engineering. Habit, with a change number attached — and it is exactly how L2TP ended up on a datasheet printed this year.

So ask it of yourself before the tender closes rather than after. Professional is not a word about which protocols you know. It is a word about whether you can defend the one you picked, out loud, to somebody who knows the failure modes. If you can, deploy whatever the constraints demand and sleep fine. If you cannot, you have just found what to go and read tonight, and that is not an insult. Everybody was a follower once, without exception, me included. The part that is not defensible is choosing to stay one and calling it a career.

A Good Idea Is Allowed To Be Over

I want to be fair to IPsec, because it deserves it.

It was the right instinct. Security belongs low in the stack, where everything inherits it and no application has to be trusted to get it right. Binding the association to the address was not a mistake in 1995 — the address was the machine, and building on that was correct. The people who wrote it were serious people solving a real problem, and the WireGuard paper, of all documents, is the one that puts it best: the layering of IPsec is sound, everything is in the right place, to academic perfection53.

And then the ground moved. Not because IPsec did anything wrong, but because this industry decided that addresses were a cost to be managed rather than a thing every machine gets, and built twenty-five years of translation to avoid the alternative. IPsec’s foundation was quietly withdrawn, and rather than admit it, we shimmed it. UDP encapsulation. Then keepalives. Then TCP encapsulation for the networks that block UDP. Then another UDP header so a router can find something to hash on. Each fix reasonable on its own; the stack of them is a protocol being kept upright by people who are paid to keep it upright.

That is the part worth being angry about, and it is not really about IPsec at all. Our industry is very good at maintaining things and very bad at ending them. Maintenance is billable, budgeted, staffed and safe. Retirement is a decision somebody has to sign, with their name on it, and no immediate reward. So PPTP lived fourteen years past the proof it was worthless, L2TP still ships on boxes sold this year, and IKEv1 kept negotiating tunnels for a decade after it went Historic. Not because anyone defended them. Because nobody was ever required to kill them.

There is no prize for getting 90% right. Ferguson and Schneier wrote that about IPsec in 1999, and what has happened since is thirty years of the industry getting the last 10% wrong in a new way each time and calling the patch a solution.

A good idea is still allowed to be over. Knowing when to stop maintaining something is a skill, and it is the one this trade is worst at. Somebody has to be the one who says a protocol has had its life, writes the date down, and takes the consequences of being the one who said it. Otherwise we will still be sending a one-byte packet every twenty seconds in 2040, to keep a table warm in a box we do not own, on a network that took our addresses off us and charged us for the privilege.


  1. RFC 4301 — Security Architecture for the Internet Protocol, December 2005. Defines the security association and the triple it is looked up on: destination address, security protocol and SPI. ↩︎ ↩︎

  2. RFC 4302 — IP Authentication Header, December 2005. AH’s integrity check covers the immutable fields of the IP header, source and destination addresses included. ↩︎ ↩︎

  3. RFC 9329 — TCP Encapsulation of Internet Key Exchange Protocol (IKE) and IPsec Packets, November 2022, replacing RFC 8229. Exists because middleboxes on public networks block UDP. ↩︎ ↩︎ ↩︎

  4. RFC 3715 — IPsec-Network Address Translation (NAT) Compatibility Requirements, March 2004. Sixteen enumerated incompatibilities, including: “Since the AH header incorporates the IP source and destination addresses in the keyed message integrity check, NAT or reverse NAT devices making changes to address fields will invalidate the message integrity check,” and “Where IP addresses are used as identifiers in Internet Key Exchange Protocol (IKE) Phase 1 or Phase 2, modification of the IP source or destination addresses by NATs or reverse NATs will result in a mismatch between the identifiers and the addresses in the IP header.” ↩︎ ↩︎ ↩︎

  5. Cisco — Configuring IPsec NAT-Traversal, Security Configuration Guide, Cisco IOS XE 17.15.x (Catalyst 9300 Switches). “If PAT found a legislative IP address and port, it would drop the Encapsulating Security Payload (ESP) packet”; the UDP header and an 8-byte non-IKE marker are inserted between the outer IP header and the ESP header; the restriction list includes static rules for both port 500 and 4500, no support for dynamic NAT policies, no IPv6, and that IPsec and NAT cannot both operate on the same device. ↩︎ ↩︎ ↩︎ ↩︎

  6. RFC 3948 — UDP Encapsulation of IPsec ESP Packets, January 2005, with the negotiation in RFC 3947. Defines the keepalive as “a one-octet-long payload with the value 0xFF”, sent “if no other packet to the peer has been sent in M seconds. M is a locally configurable parameter with a default value of 20 seconds,” and requires the UDP checksum be zeroed: “If the protocol header after the ESP header is a UDP header, set the checksum field to zero in the UDP header.” Microsoft, Cisco, F-Secure, Nortel and SafeNet are all on the author lists of RFC 3947 and RFC 3948. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  7. Juniper — Route-Based and Policy-Based VPNs with NAT-T, Junos OS IPsec VPN User Guide. “NAT-T encapsulates both IKE and ESP traffic within UDP with port 4500 used as both the source and destination port”; “Because NAT devices age out stale UDP translations, keepalive messages are required between the peers”; and on SRX5400, SRX5600 and SRX5800, “the total number of tunnels from a given public translated IP cannot exceed 1000 tunnels.” ↩︎ ↩︎ ↩︎

  8. RFC 8221 — Cryptographic Algorithm Implementation Requirements and Usage Guidance for ESP and AH, October 2017. ENCR_DES MUST NOT, ENCR_3DES SHOULD NOT, AUTH_HMAC_MD5_96 MUST NOT; ESP combined with AH is NOT RECOMMENDED. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  9. RFC 4787 — Network Address Translation (NAT) Behavioral Requirements for Unicast UDP, January 2007. REQ-5: “A NAT UDP mapping timer MUST NOT expire in less than two minutes,” with “a default value of five minutes or more for the NAT UDP mapping timer is RECOMMENDED”. ↩︎ ↩︎

  10. RFC 6146 — Stateful NAT64: Network Address and Protocol Translation from IPv6 Clients to IPv4 Servers, April 2011. “The current specification only defines how stateful NAT64 translates unicast packets carrying TCP, UDP, and ICMP traffic. Multicast packets and other protocols, including the Stream Control Transmission Protocol (SCTP), the Datagram Congestion Control Protocol (DCCP), and IPsec, are out of the scope of this specification.” Packets carrying anything else “SHOULD be discarded”. ↩︎ ↩︎

  11. RFC 6877 — 464XLAT: Combination of Stateful and Stateless Translation, April 2013. Gives an IPv6-only device a local IPv4 stack so that traffic NAT64 cannot carry works anyway. ↩︎

  12. IETF — draft-xu-ipsecme-esp-in-udp-lb, Encapsulating IPsec ESP in UDP for Load-balancing. “Although the ESP SPI field within the IPsec packets can be used as the load-balancing key, but it cannot be used by legacy switches and routers”; “Many cloud service providers allow customers to establish multiple IPsec VPN tunnels in parallel to enable ECMP and increase aggregate bandwidth. However, this approach is not ideal, as each tunnel typically requires its own public IP address, leading to higher public IP consumption and increased operational overhead.” ↩︎ ↩︎

  13. Cisco — Resolve IPv4 Fragmentation, MTU, MSS, and PMTUD Issues with GRE and IPsec. The long-standing reference on tunnel overhead, path MTU discovery and what breaks when the ICMP errors do not come back. ↩︎

  14. Cisco — Troubleshoot IPsec Anti-Replay Check Failures. “Certain QoS features, such as Low Latency Queueing (LLQ), could cause IPsec packet delivery to become out-of-order and dropped by the receiving endpoint due to a replay check failure.” Default window 64 packets; 1024 on later platforms; the alternative remedy is multiple sequence number spaces per association. ↩︎ ↩︎ ↩︎

  15. RFC 2661 — Layer Two Tunneling Protocol “L2TP”, August 1999. A tunnelling protocol for PPP with no packet-level confidentiality of its own; its security section points at IPsec. ↩︎

  16. Moxie Marlinspike and David Hulton — Divide and Conquer: Cracking MS-CHAPv2 with a 100% Success Rate, 2012; tool at github.com/moxie0/chapcrack. MS-CHAPv2’s security reduces to a single DES operation regardless of password length; contemporary write-up at The Register. Their conclusion was that PPTP traffic should be considered unencrypted. ↩︎ ↩︎

  17. Apple — If you see a “VPN Using PPTP May Not Be Secure” alert. PPTP was removed from the built-in client in macOS Sierra and iOS 10, 2016. ↩︎

  18. RFC 4303 — IP Encapsulating Security Payload (ESP), December 2005. IP protocol 50; the ESP header carries the SPI and sequence number and has no port field. ↩︎

  19. RFC 9395 — Deprecation of the Internet Key Exchange Version 1 (IKEv1) Protocol and Obsoleted Algorithms, April 2023. “Internet Key Exchange Version 1 (IKEv1) has been deprecated, and RFCs 2407, 2408, and 2409 have been moved to Historic status.” ↩︎ ↩︎

  20. RFC 7296 — Internet Key Exchange Protocol Version 2 (IKEv2), October 2014. The current key exchange, and the source of the notify names in the diagnostics table. ↩︎ ↩︎

  21. RFC 3173 — IP Payload Compression Protocol (IPComp), September 2001. Its own IP protocol number and its own associations, negotiated alongside ESP. ↩︎

  22. RFC 2367 — PF_KEY Key Management API, Version 2, July 1998. The kernel interface a key daemon uses to install associations. ↩︎

  23. RFC 7383 — IKEv2 Message Fragmentation, November 2014. “This document describes a way to avoid IP fragmentation of large Internet Key Exchange Protocol version 2 (IKEv2) messages.” ↩︎ ↩︎

  24. RFC 4555 — IKEv2 Mobility and Multihoming Protocol (MOBIKE), June 2006. Lets an established tunnel survive a change of address. ↩︎

  25. RFC 3706 — A Traffic-Based Method of Detecting Dead Internet Key Exchange (IKE) Peers, February 2004. ↩︎

  26. draft-beaulieu-ike-xauth — Extended Authentication within IKE (XAUTH). Last revision 02, October 2001, status Expired, “Expired & archived”, never published as an RFC. The draft records that it was offered as Informational because “the IPSRA working group will not accept any protocol which extends ISAKMP or IKE, and the IPsec working group refuses to accept any protocols that deal with remote access.” Mode-Config shared the same fate. ↩︎ ↩︎ ↩︎ ↩︎

  27. RFC 3193 — Securing L2TP using IPsec, November 2001. Co-authored at Microsoft. ↩︎

  28. RFC 2332 — NBMA Next Hop Resolution Protocol (NHRP), April 1998. The piece that lets DMVPN spokes find each other. ↩︎

  29. RFC 6407 — The Group Domain of Interpretation, October 2011. Group keying, as used by GETVPN. ↩︎

  30. RFC 2637 — Point-to-Point Tunneling Protocol (PPTP), July 1999. Category: Informational. A vendor protocol written up, never a standard. ↩︎ ↩︎ ↩︎

  31. Microsoft — Configure L2TP/IPsec server behind NAT-T device, KB 926179, last revised 12 February 2026. “By default, Windows Vista and Windows Server 2008 don’t support Internet Protocol security (IPsec) network address translation (NAT) Traversal (NAT-T) security associations to servers that are located behind a NAT device”; the AssumeUDPEncapsulationContextOnSendRule DWORD under HKEY_LOCAL_MACHINE\SYSTEM\CurrentControlSet\Services\PolicyAgent takes 0 (default, cannot), 1 (server behind NAT) or 2 (both ends behind NAT), and the machine must be restarted. Also: “If you must use IPsec for communication, use public IP addresses for all servers that you can connect to from the Internet.” ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  32. strongSwan — Windows Certificate Requirements. The gateway certificate needs the serverAuth EKU, OID 1.3.6.1.5.5.7.3.1, and the IP Security IKE Intermediate EKU, OID 1.3.6.1.5.5.8.2.2. ↩︎

  33. strongSwan — Windows Clients. Documents adding the NegotiateDH2048_AES256 DWORD under Rasman\Parameters to get AES-256-CBC and MODP-2048; the rekeying workaround for clients behind NAT (rekey_time = 0 on the gateway, letting the client initiate); and that the Windows client “does not currently support IKE redirection (RFC 5685) and multiple authentication rounds (RFC 4739).” Also: “IKEv2 fragmentation is supported since the v1803 release of Windows 10 and Windows Server”, and clients behind NAT refuse a server-initiated CHILD_SA rekey with Microsoft error 13863. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  34. Microsoft — DirectAccess and Remote Access Always On VPN migration overview. DirectAccess is deprecated and will be removed in a future release of Windows Server; customers are directed to Always On VPN. ↩︎

  35. Jean Paul Degabriele and Kenneth G. Paterson — Attacking the IPsec Standards in Encryption-only Configurations, IEEE Symposium on Security and Privacy, 2007. Attacks that “break any RFC-compliant implementation of IPsec making use of encryption-only ESP”, ciphertext-only, needing only the ability to eavesdrop and to inject traffic. ↩︎ ↩︎

  36. RFC 6434 — IPv6 Node Requirements, December 2011. “Previously, IPv6 mandated implementation of IPsec and recommended the key management approach of IKE. This document updates that recommendation by making support of the IPsec Architecture [RFC4301] a SHOULD for all IPv6 nodes.” ↩︎ ↩︎

  37. Niels Ferguson and Bruce Schneier — A Cryptographic Evaluation of IPsec, Counterpane Internet Security, 1999. “IPsec was a great disappointment to us”; “Our main criticism of IPsec is its complexity”; “We therefore recommend that transport mode be eliminated”; “We conclude that eliminating transport mode allows the elimination of the AH protocol as well, without loss of functionality”; and “We have found serious security weaknesses in all major components of IPsec. As always in security, there is no prize for getting 90% right; you have to get everything right.” ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  38. David Adrian et al. — Imperfect Forward Secrecy: How Diffie-Hellman Fails in Practice, ACM CCS 2015. Precomputation for a second 1024-bit group “would allow decryption of traffic to 66% of IPsec VPNs”; 86.1% of scanned IKEv1 and 91.0% of IKEv2 servers supported Oakley Group 2, and 66.1% of profiled IKEv1 servers preferred it. ↩︎

  39. CVE-2016-1287 and Cisco’s advisory, Cisco ASA Software IKEv1 and IKEv2 Buffer Overflow Vulnerability. Remote code execution before authentication, reached by crafted UDP packets to the IKE service. ↩︎

  40. Juniper — How to Analyze IKE Phase 2 VPN Status Messages. “No proposal chosen” means the device “did not accept any of the IKE Phase 2 proposals that the peer sent”. ↩︎

  41. Cisco — Understand and Use Debug Commands to Troubleshoot IPsec, and Troubleshoot Common L2L and Remote Access IPsec VPN Issues. ↩︎

  42. Juniper — Troubleshoot a VPN Tunnel That is Down. ↩︎

  43. Linux kernel — XFRM proc counters. The descriptions in the table are the kernel documentation’s own. ↩︎

  44. Juniper — Troubleshoot a VPN That Is Up But Not Passing Traffic. “If only the pkts counter in the out direction of the session is incrementing, then validate with the VPN peer that the traffic is being received.” ↩︎

  45. Microsoft — netsh wfp, Windows Commands. netsh wfp capture start “Starts a capture session for network events processed by WFP” and defaults to writing wfpdiag.cab; netsh wfp show state “Displays the current state of WFP and IPsec”; netsh wfp show ikeevents “Displays recent Internet Key Exchange (IKE) epoch events matching the specified parameters” and takes a remoteaddr= filter. Setting file=- on either show prints to the console instead of writing XML. ↩︎

  46. Microsoft — Get-NetIPsecQuickModeSA, NetSecurity module. “There is only one main mode SA between a pair of computers, but there can be many quick mode SAs,” and monitoring them “can provide information about which peers are currently connected to this computer, and which protection suite is protecting the data exchanged between them.” ↩︎

  47. Microsoft — Troubleshoot Always On VPN, last revised 12 February 2026. On reading client logs: “look for events labeled RasClient. All error messages return the error code at the end of the message.” Error 809’s cause: “You can encounter this issue when the UDP 500 or 4500 ports on the VPN server or firewall are blocked.” Error 812 is an authentication method mismatch between the server’s policy and the client’s profile. Error 13801’s four listed causes are a machine certificate without Server Authentication under Enhanced Key Usage, an expired RAS machine certificate, a client missing the root certificate, and a client whose “VPN server name doesn’t match the subjectName value on the server certificate”; 13806 is “IKE can’t find a valid machine certificate”. ↩︎ ↩︎ ↩︎ ↩︎

  48. Microsoft — Guidance for troubleshooting Remote Access (VPN and AOVPN). The supported collection is TSS, run elevated: TSS.ps1 -Scenario NET_VPN on the client and TSS.ps1 -Scenario NET_RAS on the server, reproducing the fault between start and stop, with the traces written to C:\MS_DATA. ↩︎

  49. Microsoft — Routing and Remote Access Error Codes, the codes defined in raserror.h. 638 ERROR_REQUEST_TIMEOUT; 718 ERROR_PPP_TIMEOUT; 789 ERROR_OAKLEY_GENERAL_PROCESSING, “The L2TP connection attempt failed because the security layer encountered a processing error during initial negotiations with the remote computer”; 809 ERROR_VPN_TIMEOUT, “The network connection between your computer and the VPN server could not be established because the remote server is not responding”; 828 ERROR_IDLE_TIMEOUT, “The connection was terminated because of idle timeout”; 930 ERROR_AUTH_SERVER_TIMEOUT, “The authentication server did not respond to authentication requests in a timely fashion”. ↩︎

  50. firewalld — Automatic Helper Assignment. “With kernel 4.7 and up the automatic helper assignment in kernel has been turned off by default”, controlled by the sysctl at /proc/sys/net/netfilter/nf_conntrack_helper, and “for the secure use of iptables and connection tracking helpers it is recommended to turn AutomaticHelpers off”. The kernel’s own message on the change names the reason and the replacement: default automatic helper assignment “has been turned off for security reasons”, use the CT target to attach helpers instead. The helpers covered include ftp, irc, sip, h323, tftp, snmp and pptp. ↩︎

  51. Samy Kamkar — NAT Slipstreaming, v1 31 October 2020, v2 26 January 2021 with Ben Seri and Gregory Vishnipolsky of Armis. The attack abuses “the Application Level Gateway (ALG) connection tracking mechanism built into NATs, routers, and firewalls” to “bypass victim NAT and connect directly back to any port on any machine on the network, exposing previously protected/hidden services and systems”. v1 used the SIP gateway on port 5060; v2 used H.323 on 1720, which let the pinhole be aimed at any internal host rather than only the machine that loaded the page. ↩︎

  52. Microsoft — Set-VpnServerConfiguration, RemoteAccess module. -IdleDisconnectSeconds “Specifies the time, in seconds, after which an idle connection is terminated”; -SALifeTimeSeconds and -MMSALifeTimeSeconds set the quick mode and main mode lifetimes; -SADataSizeForRenegotiationKilobytes “Specifies the number of kilobytes that are allowed to transfer using a security association (SA), after which the SA will be renegotiated.” ↩︎

  53. Jason A. Donenfeld — WireGuard: Next Generation Kernel Network Tunnel, NDSS 2017. “It intentionally lacks cipher and protocol agility. If holes are found in the underlying primitives, all endpoints will be required to update. As shown by the continuing torrent of SSL/TLS vulnerabilities, cipher agility increases complexity monumentally”; “implemented for Linux in less than 4,000 lines of code”; “it is important to stress, however, that the layering of IPsec is correct and sound; everything is in the right place with IPsec, to academic perfection.” Benchmarks: 1,011 Mbit/s against 881 and 825 for two IPsec cipher suites, and 0.403 ms ping against 0.501 and 0.508. ↩︎ ↩︎ ↩︎ ↩︎