A connection that times out tells you almost nothing. The far end might have nothing listening, a firewall three hops away might be dropping your SYN in silence without generating so much as a log line, or the route might simply not exist — and from where you are sitting every one of those looks the same. You wait. Nothing happens.

So the ticket gets written as “port 445 is blocked somewhere”, and “somewhere” is what makes it bounce. Your provider checks their edge, finds it clean, and hands it back. You check your host firewall, find it clean, and hand it back. A week goes by. Nothing is fixed.

The word doing the damage is “somewhere”. It does not have to be there. Every router between you and the destination is obliged to tell you when it is the one, and that obligation has been written into the standards since 1995. You just have to ask in the right way.

Why This Needs Spelling Out

Because too many people in this industry cannot do the basics, and it costs their employers real money every week.

I do not mean juniors. I mean people with years behind them, certifications on the wall and senior in the job title, whose diagnosis of a timed-out port stops at “it’s blocked” and goes no further. They run ping. It fails, or it works, and either way they have learned nothing about the port they were asked about. Then the ticket goes to the provider, the provider bounces it, and a fortnight of somebody’s salary goes into a thread that could have been one command.

None of this is hard. Telling a drop from a reject, reading the TTL on a reply, knowing that a star in a traceroute means nothing on its own — that is an afternoon to learn and it lasts a career. The reason people do not know it is not that they are thick. It is that nobody teaches it. Vendor training teaches you a vendor’s console. Certifications teach you the exam. The fundamentals underneath get assumed on day one and never actually covered, so people arrive at senior roles having never been shown, and by then it is embarrassing to ask.

So rather than moan about it, here it is written down. This is the basics, spelled out, with the commands and a script you can run today.

And a word on why I bothered: I ran this against my own line while writing it and found three faults I did not know I had. An SMB drop eleven hops out. Forged SMTP resets one hop away. A firewall rule that exists on IPv4 and not on IPv6. An afternoon, no root, on a line I look after myself and pay attention to. Have a think about what is sat unnoticed on the networks somebody is being paid to run.

Why the Fisher-Price OS (Windows) Is Not In Here

Two reasons it is not here. One of them is technical and one of them is not, and I would rather give you both than pretend it is all engineering.

The technical one is that it cannot do this. tracert sends ICMP echo requests and nothing else — Microsoft’s own reference describes it as “sending Internet Control Message Protocol (ICMP) echo Request or ICMPv6 messages to the destination with incrementally increasing time to live (TTL) field values”, and there is no port parameter anywhere in its syntax. pathping is that same tool with statistics bolted on. Test-NetConnection will tell you a TCP port is open or shut and nothing whatsoever about how far away the answer came from. Not one of them can probe the port you actually care about at a chosen distance, which is the entire method in this post.

Nor can you script your way round it. Setting the TTL on a socket is easy enough in .NET, but reading the ICMP error back is the hard half, and there is no equivalent of the trick this post leans on — no way to have the kernel report the error on the ordinary socket that caused it. That leaves a raw socket, and Microsoft’s own documentation says “only members of the Administrators group can create sockets of type SOCK_RAW”. So the unprivileged route does not exist and the privileged one wants a local admin token. You are into third-party downloads before you have started.

The other reason is that I do not care, and I would rather say that than dress it up. The name is not a cheap shot either, it is a description. It is the OS you get handed when you have never had another, it hides the machine from you as a design goal rather than an accident, and the moment you want to ask the network a precise question it turns out the tool was never built — because the people it is built for were never expected to ask. Thirty years on and tracert still cannot aim a packet at a port.

It is not a proper operating system for real IT people, and if it is the only one you have ever used then you are not doing this job at the level this post is written for. I am aware that lands badly. I am not writing to be liked, I am writing down what I measure for people who measure things, and I do not much care that somebody who has only ever run it would rather I put it another way.

The tool list is the argument, not the opinion. Every Unix in the table further down will let you set a hop limit and choose a protocol, four of them from a single command, and Linux will do the whole measurement without so much as sudo. That is not because they are harder to use. It is because they were built by people who expected whoever was at the keyboard to want to know things. An operating system whose diagnostic ends at ping has told you plainly what it thinks of the person using it, and thirty years of people accepting that is how we ended up with an industry that cannot locate a dropped packet.

So I do not run it, I have not run it in anger for years, and I am not writing it a section of its own to look even-handed about a gap that is real. Everything I build on and everything worth measuring from is Unix, and that is where this post lives.

If the broken box happens to be running the Fisher-Price OS (Windows), that changes nothing about the method. Get a shell on anything else — a Linux VM, a Mac, a Raspberry Pi on the same switch — and walk the hop limit towards it. The measurement does not care in the slightest what the far end runs. It only cares what is in between.

TTL Is a Hop Budget, and Every Router Owes You a Receipt

The IPv4 header has an 8-bit Time to Live field. The name is a leftover: it was specified in seconds, and nothing has treated it as seconds for decades. RFC 1812 settled the argument in 1995 and made the hop-count reading normative:

Each router (or other module) that handles a packet MUST decrement the TTL by at least one, even if the elapsed time was much less than a second. Since this is very often the case, the TTL is effectively a hop count limit on how far a datagram can propagate through the Internet.

Then the part that matters here, from the same section:

If the TTL is reduced to zero (or less), the packet MUST be discarded, and if the destination is not a multicast address the router MUST send an ICMP Time Exceeded message, Code 0 (TTL Exceeded in Transit) message to the source.

That is a MUST. Not a nicety, not a suggestion. RFC 792 in 1981 only said a gateway “may also notify the source host”; thirteen years later the requirement was tightened, and RFC 1812 says why in as many words:

ICMP Time Exceeded messages are required because the traceroute diagnostic tool depends on them.

IPv6 dropped the pretence and renamed the field. RFC 8200 calls it Hop Limit — “8-bit unsigned integer. Decremented by 1 by each node that forwards the packet” — and the expiry message became ICMPv6 type 3, code 0, “hop limit exceeded in transit”.

Read that as an instrument rather than a rule and it says something useful. Every router on the path is a beacon you can address by distance. Set the hop limit to 4 and the fourth router identifies itself. You do not need to know the topology, you do not need access to anything, and you do not need the operator’s cooperation. You need one packet per hop.

Traceroute has done exactly this since the late 1980s. What it does badly is the thing you actually care about, because by default it probes UDP ports up in the 33434 range, which is a port nobody filters and nobody serves, so it tells you about a path that nothing real ever uses. A clean traceroute to a host you cannot reach on TCP/445 proves only that UDP/33434 gets there. Which was never the question.

So probe the port you care about.

Read What Came Back, Not Whether Something Came Back

Before counting hops, look at what the far end does when you reach it with a normal hop limit. There are five distinct answers and people routinely collapse them into one.

What comes backWhat it meansWho sent it
SYN-ACK, the connection opensthe port is openthe host, or something answering for it
TCP RSTactively refusedthe host with nothing listening, or a device configured to reject
ICMP 3/13, communication administratively prohibiteda device is refusing on policy and saying sothat device — its source address is your answer
ICMP 3/1, 3/2, 3/3host, protocol or port unreachablethe last router, or the host
nothing at allsomebody is dropping in silenceunknown, so go and measure it

The third row is the one worth changing your habits over. When a firewall is configured to reject rather than drop, it puts its own address in the source field of the ICMP and hands you the culprit for free. On Linux that is what nft ... reject with icmpx admin-prohibited produces, and what iptables -j REJECT --reject-with icmp-admin-prohibited has always produced. Most tools throw it away and print “filtered”. nmap --reason will show it. So will the script further down.

Rows two and five are the interesting failure. A silent drop is a policy decision to tell you nothing, and because it is the default on nearly every commercial firewall it is the one you will actually meet. Expect silence.

Row two deserves suspicion too. A reset is not proof the host sent it, and I will come back to that with a live example, because I found one on my own line while writing this.

Walk the Port You Care About, Twice

The method is two runs and a diff. That is the whole of it.

  1. Walk the hop limit up from 1 to 20 using the exact protocol and port that fails, and write down which router answers at each hop.
  2. Do the same with something that works — ideally the same host and a port that opens.
  3. The hop where the answers stop on run one, and keep going on run two, is the device dropping you. Run two gives you its address.

Why the answers stop rather than change: a router applies its inbound access list before it does anything else with the packet, so if the policy says drop then the packet is gone before the forwarding path ever looks at the hop limit, no Time Exceeded is generated, and the device never puts its name to what it did. As such, the silence starts at the offending hop, not after it.

I wrote a small tool for the walking because none of the native ones do it portably. It sets IP_TTL (or IPV6_UNICAST_HOPS) on an ordinary socket, connects, and reads back the ICMP error. On Linux IP_RECVERR reports that error on the very socket that provoked it, which means the whole thing runs unprivileged — no root, no raw sockets, no capabilities. On macOS, the BSDs and Solaris the kernel will not hand you the ICMP that way, so it falls back to a raw socket and needs root.

hopfind.py — the TTL walker hopfind.py · 11 kB

It is 279 lines of standard library and nothing else, and the whole of it is printed at the end of this post if you would rather read it than download it.

python3 hopfind.py example.net 445      # the port under suspicion
python3 hopfind.py example.net 443      # the reference run
python3 hopfind.py example.net 53 --proto udp
python3 hopfind.py 2001:db8::1 443 -6

Use the same protocol for both runs. Comparing a TCP trail against an ICMP trail is comparing two paths, because load balancing hashes on the five-tuple and ICMP has no ports to hash. Same host, same protocol, different port is the honest comparison.

Everything below is real output from my own line on 28 August 2026, run as an unprivileged user on Fedora. Every address in it has been rewritten into the documentation ranges — RFC 5737 for IPv4, RFC 3849 for IPv6, with the one interface identifier altered as well. So 198.51.100.x is my own router and my ISP, 203.0.113.x is transit and peering, 192.0.2.x is the far network, and 2001:db8::/32 is the whole IPv6 path. Structure is preserved throughout: same prefix boundaries, same host-part shapes, same number of distinct networks. The hop numbers, the timings, the ICMP types and which hop went quiet are exactly as measured.

One Host, Three Ports, Three Different Faults

Same destination throughout: a host on the public internet, fourteen hops away, with 443 open. First the reference run.

$ python3 hopfind.py 192.0.2.4 443 --max 15
walking to 192.0.2.4  TCP/443  hop limit 1-15
  1  198.51.100.254                              0.3 ms  ICMP 11/0 time exceeded in-transit
  2  198.51.100.133                             28.8 ms  ICMP 11/0 time exceeded in-transit
  3  *
  4  198.51.100.153                              6.3 ms  ICMP 11/0 time exceeded in-transit
  5  *
  6  203.0.113.240                              16.5 ms  ICMP 11/0 time exceeded in-transit
  7  203.0.113.188                              16.5 ms  ICMP 11/0 time exceeded in-transit
  8  203.0.113.185                              16.5 ms  ICMP 11/0 time exceeded in-transit
  9  203.0.113.15                               15.6 ms  ICMP 11/0 time exceeded in-transit
 10  203.0.113.125                              16.1 ms  ICMP 11/0 time exceeded in-transit
 11  192.0.2.31                                 26.6 ms  ICMP 11/0 time exceeded in-transit
 12  *
 13  *
 14  192.0.2.4                                  19.5 ms  connected

Verdict: TCP/443 is open. It answered at hop 14.

Note hops 3, 5, 12 and 13. Four routers on a path that plainly works said nothing, because plenty of kit is configured not to generate ICMP for itself, or rate-limits it hard. A star is not evidence of a firewall. Take that one away if nowt else. The signal is never the presence of stars in a single run; it is where two runs stop agreeing.

Now the same host on 445, which times out from here.

$ python3 hopfind.py 192.0.2.4 445 --max 15
walking to 192.0.2.4  TCP/445  hop limit 1-15
  1  198.51.100.254                              0.3 ms  ICMP 11/0 time exceeded in-transit
  2  198.51.100.133                              8.0 ms  ICMP 11/0 time exceeded in-transit
  3  *
  4  198.51.100.153                              6.4 ms  ICMP 11/0 time exceeded in-transit
  5  203.0.113.76                                5.6 ms  ICMP 11/0 time exceeded in-transit
  6  203.0.113.240                              15.9 ms  ICMP 11/0 time exceeded in-transit
  7  203.0.113.188                              15.8 ms  ICMP 11/0 time exceeded in-transit
  8  203.0.113.185                              16.6 ms  ICMP 11/0 time exceeded in-transit
  9  203.0.113.15                               15.9 ms  ICMP 11/0 time exceeded in-transit
 10  203.0.113.125                              15.0 ms  ICMP 11/0 time exceeded in-transit
 11  *
 12  *
 13  *
 14  *
 15  *

Verdict: answers stop after hop 10 (203.0.113.125).
         Whatever swallows TCP/445 is hop 11.
         Walk a port that works and read off the address at hop 11.

Hop 11 answered the 443 run in 26.6 ms and said nothing at all to the 445 run. Same box, same path, same ten routers in front of it. Cross-reference the run that works and hop 11 has a name: 192.0.2.31. That is the device dropping SMB, three hops short of the destination and eight hops past my provider’s edge. Not mine, and not my ISP’s.

Hop 5 makes the point about stars from the other direction. It was a star on the 443 run and answered on the 445 run — the opposite way round from the fault. That could be ICMP rate limiting, or it could be the two runs taking different paths through a load balancer. I do not know which, and neither will you. Repeat both runs before you believe either.

Two hop-limit walks to the same host, and the hop where they stop agreeingTwo walks to the same host, and the hop where they stop agreeingOne probe per hop limit. A shaded cell means that router sent back ICMP Time Exceeded and named itself.router answerednothing came backsilent from here on1234567891011121314hop limit set on the probeTCP/443reference, it opens••*•*••••••**✓opensTCP/445under test, it times out••*•••••••****hop 11 answered one walk and not the otherHops 3, 5, 12 and 13 said nothing on a path that plainly works, so a star on its own means nowt.Hop 11 replied to the reference walk in 26.6 ms and never replied to the test walk at all.That is the boundary, and the reference walk is what gives it an address: 192.0.2.31.
The same two runs, side by side. The only cell that matters is hop 11, and what matters about it is the disagreement: it answered one run and not the other.

Then port 25. This one stopped me.

$ python3 hopfind.py 192.0.2.4 25 --max 15
walking to 192.0.2.4  TCP/25  hop limit 1-15
  1  192.0.2.4                                   0.4 ms  TCP reset

Verdict: a reset came back to a probe with a hop limit of 1.
         Nothing more than one hop away can have sent it, so check the reply
         TTL before you believe the host did.

A hop limit of 1 means the packet died at my own router. It went one hop. It cannot have travelled fourteen. Yet a TCP reset came back in 0.4 ms with the destination’s address on it, and my kernel dutifully reported connection refused. Without the hop limit set I would have read that as “the far end has no mail server” and closed the ticket.

Something one hop away is forging resets for outbound SMTP and signing them with the destination’s address. Blocking outbound 25 is an entirely ordinary thing for a consumer router or an ISP to do, and doing it with a reset rather than a drop is arguably the polite version, but the reset carries somebody else’s address and I had no idea mine did it. The hop limit is what caught it, and nothing else in the reply would have.

Off By One: Where the Access List Sits in the Pipeline

The verdict above says the dropper is hop 11 because the answers stopped after hop 10. Be careful with that arithmetic, because it depends on the order the offending device does two jobs.

Most kit applies the inbound policy first and the hop-limit check second. Deny hits, the packet is discarded, and no Time Exceeded is ever generated, so the device never appears and the silence begins at its own hop number. That is the case above. It is also the common one.

Some platforms handle hop-limit expiry in the fast path before policy is evaluated. There, the device answers the probe addressed to it and only drops probes aimed past it, so it appears normally and the silence begins one hop later.

Why the offending hop usually does not appear: policy is evaluated before the hop-limit checkInside the hop that is dropping you, the order of two checks decides what you seeYour probe arrives with one hop left on its budget, and it matches a rule that says deny.Policy first — nearly all firewalls, and every case measured in this postthe router at hop Nprobe ininbound policyhop-limit checkforwarddiscarded before anything looks at the hop limitNo ICMP is generated, so this router never puts its name to the drop.Your walk goes quiet at hop N.Expiry first — some platforms handle it in the fast paththe router at hop Nprobe inhop-limit checkinbound policyforwardexpired, so ICMP 11/0 goes back with this router's addressIt answers for itself and only swallows probes aimed past it.Your walk goes quiet from hop N+1.
Why the hop doing the dropping usually stays invisible. On nearly all firewalls the deny is evaluated first, so the probe is gone before the forwarding path notices the hop limit expired and no ICMP is ever generated. On kit that handles expiry in the fast path, the same device answers the probe addressed to it and only swallows the ones aimed further on.

So the honest reading of “answers stop after hop N” is: the dropper is hop N+1 if it binned your probe before noticing the budget had run out, or hop N itself if it answered the probe addressed to it and swallowed everything aimed past. Two adjacent devices, and the run that works names them both. Quote the addresses, not the hop count — a hop count means nothing to the person reading your ticket, who is counting from somewhere else.

The Reply’s TTL Tells You Who Really Answered

The SMTP reset above was caught by the hop limit going out. There is a second, independent check available in every reply that comes back, and it costs nothing.

Initial TTL values are not standardised, but in practice there are three:

Starts atTypical sender
64Linux, macOS, the BSDs, illumos, most hosts
255Cisco IOS, Junos, Solaris, most network kit’s own traffic
128the Fisher-Price OS (Windows), which you still need for reading a reply off one

Subtract the TTL you received from the next value up and you have the hop count back. A reply arriving with TTL 50 started at 64 and came 14 hops. One arriving with TTL 250 started at 255 and came 5. ping prints it without being asked:

ping -c1 192.0.2.4          # ttl=50 → 14 hops away
tcpdump -n -v 'icmp'        # -v prints the ttl of every packet it shows

Two things fall out of that. Both are free.

A reply whose reverse hop count does not match the host’s other replies was not sent by the host. If an echo reply from a server comes back 14 hops out and the RST on port 25 comes back 1 hop out, a middlebox wrote the RST. Same trick as the section above, from the other end, and it works even when you cannot set the outbound TTL.

A reply that started at 255 came from network kit, not from a server. Useful when you are trying to work out whether the thing rejecting you is the host or the router in front of it.

To see the field on TCP rather than ICMP you need a capture, and the filter is worth memorising:

# every hop-limit expiry coming back to you, IPv4 and IPv6
tcpdump -n -v 'icmp[icmptype] == 11 or icmp6[icmp6type] == 3'

# who is resetting you, and from how far
tcpdump -n -v 'tcp[tcpflags] & tcp-rst != 0'

tcpdump prints the expiry as ICMP time exceeded in-transit — the same phrase the standards use, and the same event that shows up as “TTL expired in transit” on platforms that word it that way.

The Same Trick With Native Tools, on Five Unixes

If you would rather not run a script, the native tools will do most of it. They just disagree with each other more than you would expect. -P in particular means three different things depending on whose traceroute you are holding, and one of them will quietly ruin your test.

Linux (traceroute 2.1.x)macOS / FreeBSDOpenBSDNetBSDSolaris 11
TCP probes-T-P tcpnot usablenono
ICMP probes-I-I-I-I-I
UDP to a fixed port-U -p N-e -p Nnonono
destination port-p N (constant for TCP)-p N (increments without -e)-p N (increments)-p N (increments)-p N (increments)
what -P meansnot usedprobe protocolnumeric protocol, “will not work reliably for most protocols”set DF and probe path MTUpause between probes, in seconds
needs privilegesyes, for -T and -Iyesyesyesyes

Every one of those needs raw sockets, so every one needs privileges — though macOS and the BSDs usually ship traceroute setuid root, so you may not have to type sudo in front of it. Linux does not, and Fedora does not, which is half the reason the script above exists.

Three traps in that table, and I have watched all three waste an afternoon.

On the BSDs and macOS, -p is a base port that increments with every probe. So traceroute -P tcp -p 445 host tests 445, then 446, then 447, and by hop 10 you are asking about a port nobody has ever heard of. You want -e as well, which the man page calls firewall evasion mode and which actually just means “keep the port still”:

sudo traceroute -P tcp -e -p 445 example.net     # macOS, FreeBSD

On Linux, plain -p does the same thing for the default UDP method and you need -U -p for a constant UDP port. For TCP, -T -p is already constant — the man page is explicit that “for TCP and others specifies just the (constant) destination port to connect”.

sudo traceroute -T -p 445 example.net            # Linux
sudo traceroute -U -p 53  example.net            # Linux, UDP/53 specifically

On Solaris and NetBSD there is no way to pin the port at all, and on Solaris -P is a pause in seconds, so a copied Linux command line will run without error and measure nothing you asked about. Solaris also has no TCP probe mode. This is the case where you want the script.

Redox is the odd one out and worth a sentence because I expect the question. Its whole network toolkit is netutils — dns, ifconfig, nc, ping, telnetd, wget. No traceroute, no tcpdump, nowt to capture with. If a Redox box is one end of the problem, measure from the other end and point the walk at it.

mtr deserves a mention too, because it does the repeat-and-average part that the tables above make you do by hand:

sudo mtr -T -P 445 --report --report-cycles 20 example.net

Run that against the failing port and again against a working one, side by side. Same method, prettier output.

None Of This Works If Somebody Blocks ICMP

Every measurement in this post is made of ICMP errors travelling back to me. Block those and the entire diagnostic goes dark — and so does a good deal else.

Blocking ICMP wholesale is still treated as a security posture in places. It has not been a defensible one since the 1990s. The attacks it is imagined to stop were ping of death and smurf, both of which were fixed in the stacks rather than at the border, and both of which were fixed before some of the engineers still repeating the advice were born. What blanket blocking stops now is diagnosis. Nothing else.

RFC 1812 does not leave room for interpretation on the point: Time Exceeded is a MUST, and the standard states its reason is that traceroute depends on it. Drop it and you have broken a tool the internet’s own router requirements document names as the justification for the message existing.

Path MTU discovery is the expensive one. It needs ICMP 3/4, fragmentation needed, to come back to the sender. Filter that and you get the fault every network engineer has chased at least once, the one where the handshake completes and small transfers work and anything carrying a full-size packet hangs forever. SSH connects and scp stalls. The page loads and the image never arrives. Nothing in the logs. Nothing to grep for.

On IPv6 it stops being a matter of taste. RFC 4890 §4.3.1 lists the messages a firewall must not drop:

  • Destination Unreachable (Type 1) - All codes
  • Packet Too Big (Type 2)
  • Time Exceeded (Type 3) - Code 0 only
  • Parameter Problem (Type 4) - Codes 1 and 2 only

and on Packet Too Big it is blunt about the consequence: “Effectively, parts of the Internet will become inaccessible.” IPv6 routers do not fragment. If Packet Too Big cannot reach the sender, there is no recovery path.

The control is a rate limit, not a drop. Permit type 3 and type 11 inbound, count them, cap them at something like a hundred a second, log whatever exceeds the cap, and you have kept the diagnostics, kept path MTU discovery working, and kept every bit of the protection the blanket rule was imagined to be providing in the first place. Restricting how much of something you accept is a control. Refusing all of it and calling that hardening is just refusing to be measured.

For anyone running an MSP: if your customer’s line drops ICMP errors, you have removed their ability to prove which network a fault is in — and your own. The next time a fault sits between two providers who both say it is clean, that is the bill for the policy.

Except Echo. Drop That.

Everything above is about ICMP errors. Echo is a different animal, and it is the one part of the protocol I would take off the wire at the border.

Look at what the standard requires of it. In IPv4, RFC 792 says of an echo request that “the data received in the echo message must be returned in the echo reply message”. IPv6 is blunter still — RFC 4443 defines the field as “zero or more octets of arbitrary data” and then requires that it “MUST be returned entirely and unmodified in the ICMPv6 Echo Reply message”.

Read that as an attacker rather than as an operator. The standard obliges every host on earth to accept a block of bytes you choose and hand it straight back to you. That is not a side effect. It is a mandated, bidirectional, arbitrary-payload channel, running over a protocol most firewalls pass without inspecting and most logging records as a packet count rather than as content.

People have been building tunnels on it for thirty years. Loki did it in Phrack 49 in 1996. Ptunnel will carry a full TCP session inside ping and has been a package away for two decades. If your egress policy is “block everything, allow ICMP because the network team need it”, you do not have an egress policy. You have a VPN with extra steps, and the traffic leaves looking like somebody testing whether the internet is up.

RFC 4890 disagrees with me, and it is worth saying so plainly rather than quoting only the half that suits. §4.3.1 puts Echo Request and Echo Response in the same must-not-drop list as the errors. Then read the justification it gives:

For Teredo tunneling [RFC4380] to IPv6 nodes on the site to be possible, it is essential that the connectivity checking messages are allowed through the firewall.

The stated reason to keep echo open is that somebody needs to build a tunnel through your firewall with it. That is my argument, written down by the people making the opposite case.

So the policy is narrow, not blanket:

# transit rules. permit the errors, drop the ping
ip protocol icmp icmp type { destination-unreachable, time-exceeded, parameter-problem } \
    limit rate 100/second accept
ip protocol icmp icmp type { echo-request, echo-reply } drop

ip6 nexthdr ipv6-icmp icmpv6 type { destination-unreachable, packet-too-big, \
    time-exceeded, parameter-problem } limit rate 100/second accept
ip6 nexthdr ipv6-icmp icmpv6 type { echo-request, echo-reply } drop

On IPv6, do not carry that pattern onto a link-local or host chain without keeping neighbour discovery. Types 133 to 137 — nd-router-solicit through nd-redirect — are how IPv6 does the job ARP does in IPv4. Drop those and the segment stops working within minutes, and it will not look like a firewall fault. Filter echo at the border, not on the wire between a host and its own router.

What does that cost you? ping across the boundary, and nothing else. Everything in this post keeps working, because not one measurement here sends an echo request. hopfind.py walks TCP and UDP and reads the errors that come back; traceroute -T and -U do the same. Path MTU discovery needs Packet Too Big, which is an error. The reverse-TTL trick works on any reply, and a TCP handshake will give you one. The only thing you lose is the least informative tool in the box, and this whole post is an argument about why stopping at ping is the problem in the first place.

My own line already does exactly this, though I doubt it was deliberate. ping to my default gateway gets 100% loss, and ICMP Time Exceeded from that same gateway comes back in 0.3 ms — as every trace in this post shows. Echo shut, errors open. Whoever shipped that firmware got the right answer, and I only found out they had by going looking.

The Same Rule, Written Once

Here is the fault I did not expect to find in my own house. A public DNS resolver, walked on TCP/443 over both families, minutes apart.

$ python3 hopfind.py 192.0.2.53 443 --max 10
walking to 192.0.2.53  TCP/443  hop limit 1-10
  1  *
  2  *
  3  *
  4  *
  5  *
  6  *
  7  *
  8  *
  9  *
 10  *

Verdict: nothing answered at all, not even the first hop.

Nothing at all. Not one hop. My own router did not even report the expiry it must have generated, the same expiry it reported in 0.3 ms for every other walk in this post — so the drop happens at hop 1, before the hop-limit check ever runs, and hop 1 is mine.

The path is fine, which the same destination proves on UDP:

$ python3 hopfind.py 192.0.2.53 53 --proto udp --max 10
  1  198.51.100.254                              0.3 ms  ICMP 11/0 time exceeded in-transit
  2  198.51.100.133                              5.4 ms  ICMP 11/0 time exceeded in-transit
  3  *
  4  198.51.100.167                             14.5 ms  ICMP 11/0 time exceeded in-transit
  5  203.0.113.50                                6.3 ms  ICMP 11/0 time exceeded in-transit
  6  203.0.113.174                               5.7 ms  ICMP 11/0 time exceeded in-transit
  7  203.0.113.201                               6.4 ms  ICMP 11/0 time exceeded in-transit
  8  *

Seven hops of clean answers to the same address. So it is not routing and it is not the destination — something on my line drops TCP to that host and passes UDP to it.

Then the same resolver, same port, over IPv6:

$ python3 hopfind.py 2001:db8:53::53 443 -6 --max 10
walking to 2001:db8:53::53  TCP/443  hop limit 1-10
  1  2001:db8:1:ee:beef:abcd:fec0:2a30           0.5 ms  ICMP 3/0 hop limit exceeded in-transit
  2  2001:db8:1::15c                            24.6 ms  ICMP 3/0 hop limit exceeded in-transit
  3  *
  4  2001:db8:2:200::50                          5.5 ms  ICMP 3/0 hop limit exceeded in-transit
  5  2001:db8:2:2::4                             5.7 ms  ICMP 3/0 hop limit exceeded in-transit
  6  2001:db8:2:991::                           16.6 ms  ICMP 3/0 hop limit exceeded in-transit
  7  2001:db8:53::53                             5.7 ms  connected

Verdict: TCP/443 is open. It answered at hop 7.

Straight through. Seven hops, no drama. Same service, same port, same intent, and the rule exists on one address family only.

Whoever wrote that rule wrote it for IPv4 and never wrote the twin. I have no idea what it was meant to achieve — my guess is something about keeping DNS local — but whatever it was, it has been achieving it on half the traffic for as long as this line has had IPv6. If it was there for a reason, it does not work. If it was not, it should not be there.

That is the everyday version of the dual-stack problem, and it is far more common than the arguments about whether to deploy IPv6 at all. Two rulebooks. One maintained.

The Script, In Full

No dependencies, no install, nothing but the standard library. Python 3.6 or later, and on Linux no privileges at all.

The two classes are the whole of the portability story. ErrorQueue is the Linux path: arm IP_RECVERR on the socket, and after the probe fails, read MSG_ERRQUEUE and pull the router’s address out of the sock_extended_err structure the kernel appends it to. RawIcmp is everywhere else: open a raw ICMP socket, read whatever arrives, take the source address off the packet. The first needs nothing, the second needs root, and the rest of the program does not care which one it was handed.

One detail is worth pointing out because it is the difference between a right answer and a plausible one. In probe() the error queue is drained before SO_ERROR is consulted. An ICMP error reaches a TCP socket as a plain errno — ICMP 3/3 port unreachable arrives as ECONNREFUSED, exactly like a real reset — so checking SO_ERROR first would have reported the forged SMTP reset further up this post as an honest refusal from the far end. Read the queue first and ee_origin tells you a router spoke.

#!/usr/bin/env python3
"""hopfind - work out how many hops away the thing blocking your port is.

Walks the IPv4 TTL, or the IPv6 hop limit, up from 1 and records which router
answers at each step - using the protocol and port you actually care about
instead of traceroute's default UDP high ports.

Run it twice. Once against something that works, once against the port that
does not. The hop where the answers stop is the device dropping you, and the
run that works gives you its address.

    python3 hopfind.py example.net 445            # the port under suspicion
    python3 hopfind.py example.net 443            # the reference run
    python3 hopfind.py example.net 53 --proto udp
    python3 hopfind.py 2001:db8::1 443 -6

On Linux this needs no privileges at all: IP_RECVERR and IPV6_RECVERR hand the
ICMP errors back on the ordinary socket that caused them. On macOS, the BSDs
and Solaris the errors have to be read off a raw ICMP socket, which means root.

Written for https://blogs.damiendye.uk/networking/how-far-away-is-the-firewall/
Public domain. Do what you like with it.
"""

import argparse
import errno
import os
import select
import socket
import struct
import sys
import time

# Linux socket options. Absent from the socket module on some builds, so they
# are spelled out rather than looked up.
IP_RECVERR = 11
IPV6_RECVERR = 25

# ee_origin values from linux/errqueue.h. Anything else means the errno came
# from the local stack rather than from a router.
SO_EE_ORIGIN_ICMP = 2
SO_EE_ORIGIN_ICMP6 = 3

ICMP_V4 = {
    (11, 0): "time exceeded in-transit",
    (11, 1): "fragment reassembly time exceeded",
    (3, 0): "net unreachable",
    (3, 1): "host unreachable",
    (3, 2): "protocol unreachable",
    (3, 3): "port unreachable",
    (3, 4): "fragmentation needed",
    (3, 9): "net administratively prohibited",
    (3, 10): "host administratively prohibited",
    (3, 13): "communication administratively prohibited",
    (5, 0): "redirect",
}

ICMP_V6 = {
    (3, 0): "hop limit exceeded in-transit",
    (3, 1): "fragment reassembly time exceeded",
    (1, 0): "no route to destination",
    (1, 1): "communication administratively prohibited",
    (1, 3): "address unreachable",
    (1, 4): "port unreachable",
    (2, 0): "packet too big",
}


def describe(family, icmp_type, icmp_code):
    table = ICMP_V4 if family == socket.AF_INET else ICMP_V6
    return table.get((icmp_type, icmp_code), "unrecognised")


def is_expiry(family, icmp_type):
    """Was this the router saying 'your hop budget ran out here'?"""
    return icmp_type == (11 if family == socket.AF_INET else 3)


class ErrorQueue:
    """Linux. The kernel reports the ICMP error on the socket that provoked it."""

    def arm(self, sock, family):
        if family == socket.AF_INET:
            sock.setsockopt(socket.IPPROTO_IP, IP_RECVERR, 1)
        else:
            sock.setsockopt(socket.IPPROTO_IPV6, IPV6_RECVERR, 1)

    def extra_readers(self):
        return []

    def collect(self, sock, family):
        try:
            _, ancillary, _, _ = sock.recvmsg(0, 1024, socket.MSG_ERRQUEUE)
        except OSError:
            return None
        wanted = (socket.IPPROTO_IP, IP_RECVERR) if family == socket.AF_INET \
            else (socket.IPPROTO_IPV6, IPV6_RECVERR)
        for level, kind, data in ancillary:
            if (level, kind) != wanted or len(data) < 16:
                continue
            # struct sock_extended_err, then the sockaddr of the router that
            # sent the error - SO_EE_OFFENDER in the kernel headers.
            _, origin, icmp_type, icmp_code = struct.unpack_from("=IBBB", data, 0)
            if origin not in (SO_EE_ORIGIN_ICMP, SO_EE_ORIGIN_ICMP6):
                return None
            addr = None
            if len(data) >= 24:
                offender_family, = struct.unpack_from("=H", data, 16)
                if offender_family == socket.AF_INET:
                    addr = socket.inet_ntoa(data[20:24])
                elif offender_family == socket.AF_INET6 and len(data) >= 40:
                    addr = socket.inet_ntop(socket.AF_INET6, data[24:40])
            return addr, icmp_type, icmp_code
        return None


class RawIcmp:
    """macOS, the BSDs, illumos, Solaris. Read the ICMP off a raw socket, as root."""

    def __init__(self, family):
        proto = socket.IPPROTO_ICMP if family == socket.AF_INET else socket.IPPROTO_ICMPV6
        self.sock = socket.socket(family, socket.SOCK_RAW, proto)
        self.sock.setblocking(False)

    def arm(self, sock, family):
        pass

    def extra_readers(self):
        return [self.sock]

    def collect(self, sock, family):
        try:
            packet, peer = self.sock.recvfrom(1500)
        except OSError:
            return None
        if family == socket.AF_INET:
            # BSD raw sockets hand back the IP header too.
            header_len = (packet[0] & 0x0F) * 4
            packet = packet[header_len:]
        if len(packet) < 2:
            return None
        return peer[0], packet[0], packet[1]


def probe(dest, port, proto, family, hop_limit, timeout, listener):
    """One probe at one hop limit. Returns (icmp, socket_state, note)."""
    kind = socket.SOCK_STREAM if proto == "tcp" else socket.SOCK_DGRAM
    sock = socket.socket(family, kind)
    if family == socket.AF_INET:
        sock.setsockopt(socket.IPPROTO_IP, socket.IP_TTL, hop_limit)
    else:
        sock.setsockopt(socket.IPPROTO_IPV6, socket.IPV6_UNICAST_HOPS, hop_limit)
    listener.arm(sock, family)
    sock.setblocking(False)

    try:
        if kind == socket.SOCK_DGRAM:
            sock.connect((dest, port))
            sock.send(b"\x00" * 32)
        else:
            try:
                sock.connect((dest, port))
            except BlockingIOError:
                pass
    except OSError as exc:
        sock.close()
        return None, None, "local error: %s" % exc.strerror

    readers = [sock] + listener.extra_readers()
    writers = [] if kind == socket.SOCK_DGRAM else [sock]
    deadline = time.time() + timeout
    icmp = state = None

    while time.time() < deadline:
        ready_r, ready_w, ready_x = select.select(
            readers, writers, [sock], max(0.01, deadline - time.time()))
        if not (ready_r or ready_w or ready_x):
            continue
        # Drain the error queue first, always. An ICMP error reaches a TCP
        # socket as a plain errno, so SO_ERROR on its own cannot tell you
        # whether a router spoke or the far end did.
        icmp = listener.collect(sock, family)
        if icmp:
            break
        if ready_w:
            err = sock.getsockopt(socket.SOL_SOCKET, socket.SO_ERROR)
            if err == 0:
                state = "connected"
            elif err == errno.ECONNREFUSED:
                state = "TCP reset"
            else:
                state = os.strerror(err)
            break

    sock.close()
    if icmp or state:
        return icmp, state, None
    return None, None, "no reply"


def walk(dest, port, proto, family, first, last, timeout, listener):
    print("walking to %s  %s/%d  hop limit %d-%d" % (dest, proto.upper(), port, first, last))
    answered = []
    for hop in range(first, last + 1):
        started = time.time()
        icmp, state, _ = probe(dest, port, proto, family, hop, timeout, listener)
        rtt = (time.time() - started) * 1000
        if icmp:
            addr, icmp_type, icmp_code = icmp
            print(" %2d  %-39s %7.1f ms  ICMP %d/%d %s"
                  % (hop, addr or "?", rtt, icmp_type, icmp_code,
                     describe(family, icmp_type, icmp_code)))
            if is_expiry(family, icmp_type):
                answered.append((hop, addr))
            else:
                return answered, hop, "icmp-reject", addr
        elif state:
            print(" %2d  %-39s %7.1f ms  %s" % (hop, dest, rtt, state))
            return answered, hop, state, dest
        else:
            print(" %2d  *" % hop)
    return answered, None, "silent", None


def main():
    parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
    parser.add_argument("host")
    parser.add_argument("port", nargs="?", type=int, default=443)
    parser.add_argument("--proto", choices=("tcp", "udp"), default="tcp")
    parser.add_argument("--first", type=int, default=1, help="hop limit to start at")
    parser.add_argument("--max", type=int, default=20, help="hop limit to stop at")
    parser.add_argument("--wait", type=float, default=2.0, help="seconds to wait per hop")
    parser.add_argument("-6", dest="v6", action="store_true", help="force IPv6")
    parser.add_argument("-4", dest="v4", action="store_true", help="force IPv4")
    args = parser.parse_args()

    family = socket.AF_INET6 if args.v6 else socket.AF_INET
    kind = socket.SOCK_STREAM if args.proto == "tcp" else socket.SOCK_DGRAM
    dest = socket.getaddrinfo(args.host, args.port, family, kind)[0][4][0]

    if sys.platform.startswith("linux"):
        listener = ErrorQueue()
    else:
        try:
            listener = RawIcmp(family)
        except PermissionError:
            sys.exit("%s cannot report ICMP errors on a normal socket, so this "
                     "needs a raw one. Run it as root." % sys.platform)

    answered, stop, why, who = walk(dest, args.port, args.proto, family,
                                    args.first, args.max, args.wait, listener)
    what = "%s/%d" % (args.proto.upper(), args.port)
    print()

    if why == "connected":
        print("Verdict: %s is open. It answered at hop %d." % (what, stop))
    elif why == "icmp-reject":
        print("Verdict: %s at hop %d is refusing %s on policy, and is honest "
              "enough to say so." % (who, stop, what))
    elif why == "TCP reset":
        print("Verdict: a reset came back to a probe with a hop limit of %d." % stop)
        print("         Nothing more than %s away can have sent it, so check the reply"
              % ("one hop" if stop == 1 else "%d hops" % stop))
        print("         TTL before you believe the host did.")
    elif answered:
        last_hop, last_addr = answered[-1]
        print("Verdict: answers stop after hop %d (%s)." % (last_hop, last_addr))
        print("         Whatever swallows %s is hop %d." % (what, last_hop + 1))
        print("         Walk a port that works and read off the address at hop %d."
              % (last_hop + 1))
    else:
        print("Verdict: nothing answered at all, not even the first hop. Either the")
        print("         first hop is the one dropping you, or the ICMP errors are being")
        print("         filtered on the way back. Walk a port that works to tell those")
        print("         two apart.")


if __name__ == "__main__":
    main()

What This Cannot Tell You

The method is cheap and it is honest about most things, but it is not a topology scanner. Be straight in the ticket about what you actually measured.

Different five-tuples can take different paths. ECMP hashes the source and destination ports into the choice of next hop, so two runs on two different ports are not guaranteed to traverse the same routers at all, which is one of the two things that could explain hop 5 above answering one run and staying quiet on the other. Repeat both runs. A boundary that moves is unproven.

MPLS hides hops. A label-switched core can present as one hop, or as none at all. Any count across somebody else’s backbone is a lower bound.

ICMP generation is rate limited nearly everywhere. Probe faster than the router will answer and you manufacture your own stars. hopfind.py sends one probe per hop and waits; that is deliberate.

The return path need not match the outbound one. The hop count out is not the hop count back, and reverse-TTL arithmetic measures the return leg only.

Anycast means the host at hop N may not be the same box twice. Public resolvers and CDNs, in particular.

A completed handshake does not mean the session survives. A stateful firewall can permit the SYN and drop what follows on inspection. If the connection opens and then dies, this is the wrong instrument — go and capture.

You have found the first device that drops, not the one anybody will admit to. In a CGN or a carrier network the address at hop N+1 may be one of several boxes behind one address. It is still the right thing to quote, because it is a fact about the path.

What It Is Actually For

Ending the bouncing. That is the whole return on the exercise.

“Port 445 is blocked somewhere” is an invitation to hand the ticket back. This is not:

TCP/445 to 192.0.2.4 is dropped silently at hop 11, address 192.0.2.31. Hop 11 answers ICMP Time Exceeded to TCP/443 on the same path in 26 ms and answers nothing at all on 445, so the drop is a policy decision on that device, not a routing fault. Ten hops in front of it are clean. Reproduced four times over twenty minutes, from an unprivileged shell, script attached.

Nobody hands that back. It names a device, states what it did, states what it did not do, and shows the working. Whether they choose to change it is still their call. But the week of ping-pong is over, and it took four commands.

Everything in it came out of an 8-bit field that was specified as a timer in 1981, has never once been used as one, and quietly turns “somewhere” into an address.

Worth learning to read.