Traceroute Lies: A Typical Misinterpretation Of Output
movingpackets.net
movingpackets.net
John alluded to this in his "side note" but often (the majority of cases, IME) the TTL will not be decremented inside of the MPLS core so you, as an end user, will have no visibility into it. You'll just see, for example, hop #4 (Los Angeles) then hop #5 (New York City), completely unaware that the IP packet actually passed through (MPLS "P") routers in Phoenix, Denver, Kansas City, Chicago, and Washington, DC, in between.
You'll see the same thing -- although at a different layer -- when Q-in-Q tunneling is in use (assuming L2 protocols are also being tunneled).
N.B.: On common Linux distributions, there are usually several traceroute variants available. They are not all created equal. If a "regular" (UDP) traceroute won't work, you can try using ICMP or even TCP.
[0]: https://youtu.be/a1IaRAVGPEE
[1]: https://www.nanog.org/sites/default/files/tuesday_steenberge...
What's needed is some kind of L2 Traceroute. It's maddening that an increasingly common piece of routing infrastructure is invisible to almost all common troubleshooting tools.
Edit: Cisco has one for their routers [0]. Since each hop would have to be traced locally, an L2 Trace would require orchestrating against all hardware devices in the route which introduces legal complications. A new discovery protocol might be needed.
[0] https://www.cisco.com/c/en/us/td/docs/switches/lan/catalyst6...
2ms 3ms 38ms 5ms 8ms 9ms (target)
That 38ms showing packet losss too. I suspect it's the control plane that doesn't prioritise responses.
As I have no access to intermediate hops, nor even a business relationship with the owner, traceroutes, and things like bgp.he.net, help me work out where problems may be occurring and sometimes let me reroute.
Even on private wires it's hard to explain to network providers what an outage means. Sometimes a 10ms outage can be the difference between your application working and failing. Had a mikrotik at one event this year that I think was bouncing the queuing process from core to core. Looking at the rtp before and after showed it often throwing 30 packets out of order - about 15ms. Occasionally it would simply drop them, and that causes a major problem with broadcast video as normal FEC can cope with 20 losses on the trot at most.
Had to change ISP in kiev a couple of years ago because of a frequent 136ms drop. I'd isolated it to the next upstream provider from our existing isp, but nothing more could be done as it seemed that was their only real transit provider, and 130ms doesn't really stop Facebook from working.
Replies may not be travelling the same path as the request, but traceroute offers no visibility of the return path.
Traceroute can missrepresent where latency is due to how it uses the TTL field, but increasing the sampling size usually averages out that issue.