Wireshark Is 25: The email that started it all and lessons learned along the way
blog.wireshark.org
blog.wireshark.org
It continues to blow my mind how many people there are in the world who consider themselves networking professionals but have never used or do not understand Wireshark. It is possibly the most important tool for actually understanding what's really happening in your network, without it you're effectively blind to so many things.
Just yesterday I used it to troubleshoot a weird behavior in a recently upgraded Asterisk/FreePBX system which would have probably taken me days to guess my way through without packet captures, but with them I was able to see clearly what was happening on the network and then track that back from there.
Congrats to everyone on the Wireshark team on 25 years of making network troubleshooting infinitely easier! I would 100% not be where I am today without it.
You might look for some pcap based CTFs with walkthroughs to get exposure to some of the more unique things you can do.
Just letting it run for a few min on your router and then powering a device up can also yield some interesting captures…
Most of my learning was just "capture the problem happening, capture what happens when it works right if possible, open up the relevant RFCs, then try to understand what's different and why.
I work in the VoIP industry so I'm dealing with a lot of NAT problems (insert rant here about lazy ISPs that still haven't enabled IPv6 on their networks) and my main protocol (SIP) is heavily inspired by HTTP and as a result is more or less human readable plaintext, so it was a relatively easy learning curve to just have Wireshark open on one side of the screen and the relevant RFCs on the other side.
All I can really say is have a problem you want to solve and start from there.
It is also a great way to figure out why programs without useful debug output die. Ie. after a program opens and reads a config file it doesn't like, it starts cleaning up and exits.
Lots of little things like that. Why is this program acting slow at startup when it should be fast? Oh, because it's opening and timing out on a socket connection with an unusually long timeout. Et cetera...
It's probably most useful to sysadmins working with binaries, or even if you do do have the source, it's usually a shorter path to the solution for any app/os interaction problem.
It's useful for certain classes of optimisation and tuning, because it will give timings and aggregate timings.
I'll use it for things as simple as "where is this program reading it's config files" - often useful when doco is poor and/or there are multiple config locations selected by conditional logic.
There's an "ltrace" as well, for share library tracing, although I've personally found that less useful - bugs that that shows are more likely to be code/logic problems rather that os/infrastructure interaction - which is to say, usually outside my job scope.
On commercial unix, the equivalent to strace is truss, and it's been around forever.
Like many, wirewhark, strace/truss are my go-to tools for a huge amount of troubleshooting.
For example if the open call fails with ENOENT, the file it's looking for doesn't exist, and strace will (also) tell you what file it's trying to open.
There is also ltrace, for library calls, although I find it less useful.
Do you mind sharing with us what was the problem and how you solved it with packet captures, if you have time? A blog post would be very interesting too.
In this case, the problem was the client wasn't actually sending the request, and with a sizable request that's visible even without decoding the https; although to be totally clear on what was happening, decoding was needed.
I've also debugged issued in remote networks where iirc, connections were being reset by some equipment local to the user. Seq/ack sequencing showed the resets were in response to a specific client sent packet and the timestamps showed it was impossible for that to have come from anywhere but equipment near the user.
For this bug [2], it took a lot of luck and patience to get a good capture, but once I did, the immediate problem became obvious: the machine I controlled was getting an icmp needs frag but DF set at the same mtu it was already using, and responding by sending the whole sendqueue at once, packetized to the new MTU that was the same as the old one. There's actually three problems here: a) there's no reason for the other side to send this packet (I found this is an already fixed linux bug with forwarding and large receive offload, but no way to contact the administrator of that router), b) our side shouldn't resend the whole sendqueue when the mtu changes, c) if the mtu didn't change, then there's no need to take any action. We only fixed c, but that solved the major problem: these resends would trigger more resends and we'd have periods of unavailability as the network was really busy.
This is pretty common when looking at wireshark; unless you work somewhere with full control of all clients and servers and a very network aware developer team, you're going to find lots of non-optimal or semi-broken stuff, and you've got to ignore it and focus on the majorly broken bit.
I had just upgraded and migrated one of my clients from an on premise FreePBX system that was a few years out of date and running on a repurposed desktop computer with a failing fan to a brand new instance running on a VPS. Everything was working fine with basic phone functionality, but their main ring group was taking a few seconds to stop ringing when answered. Calls would ring in to all phones effectively simultaneously as expected, but when someone answered the call certain phones kept ringing for almost four full seconds after that point.
In the past I had seen similar behaviors on AT&T DSL caused by their mandatory modem/router device having an anti-flood filter enabled by default which saw a bunch of nearly identical UDP packets hitting at once and dropped them after the first few. This site has cable internet through a dumb modem so I knew it wasn't that, but they had recently had their IT side taken over by a new company who put in a new firewall so that was a plausible answer.
Their IT however had been taken over from us so I wasn't about to go accusing them of getting it wrong without strong evidence. I'm also just that kind of person, I hate when someone blames me or my gear for problems we're not causing so I do my best to never be that guy either. I'll waste an extra few hours of mine any day of the week to be sure I'm not accusing someone else of getting it wrong without a reason.
I fired up sngrep on the server, waited for a call to come in, and saved all the SIP sessions that resulted. Download that file, load it up in Wireshark, and I see that while the INVITE messages to start ringing all went out more or less simultaneously (27 phones in ~5ms) the CANCEL messages that stop them from ringing once one answered were sent out sequentially, with the PBX waiting for the first one to respond and confirm it had stopped ringing before sending the next. Clearly this wasn't right, and it obviously wasn't a problem with the firewall either.
At that point I started looking at the Asterisk logs and saw that an AGI script was being run for each line that was ringing which wasn't there previously. That script was associated with a new FreePBX module for missed call notifications which was installed but unconfigured on the new server. It didn't indicate it was doing anything in the UI, but it sure seemed to be doing something in the logs.
I uninstalled that module and the next call all the CANCEL messages went out in ~5ms just like the INVITEs. I then filed a bug with FreePBX documenting what happened because I'm pretty sure it's not expected or desired for simply having that module installed to cause massive delays in ring groups.
---
In this case the packet captures demonstrated conclusively that the problem was on the server itself and not in the network. If the capture at the server had looked reasonable my next step would have been to have the IT vendor capture traffic on their firewall at the same time as I was capturing at the server so we could compare and see if it's getting messed with along the way, but here it was not necessary.
Like toast0 mentioned, captures help you narrow down where the problem is.
A network capture was the only good clue.
As part of the code, I wrote a packet dumper that put the Ethernet card (a Multibus card from a company called Exelan) into promiscuous mode. It didn't have a dissector like Wireshark, but just being able to dump raw packets in hex to a terminal was a huge advantage for debugging networks.
I love Wireshark and it's one of the first things I install on a new system.
I really need to dive deeper into networking…
Trouble is, the femtocell wanted a network connection, and our RF chamber didn't have an RJ45 passthrough. Some emails got sent, the chamber vendor could sell us a new passthrough module but it was on backorder, ETA two months or something.
So the following evening, I swung by the e-waste recycler where I used to volunteer years prior, which meant I could just give the proprietor a wave and then let myself into the back room and pick the pile. And sure enough, I found a couple of 8-port 10base-T ethernet hubs, with 10base-2 connections on the back for connection to a coax segment. I talked him up to twenty bucks so I'd have an expense to submit; the company did not deserve to get this for free.
Back in the RF lab the following day, it was a trivial matter to convert the BNC connector on the hubs to the N connector in the chamber wall, locate one of the hubs inside the chamber, and connect the femtocell to it. The one outside got the internet connection, which had been running at gigabit speeds but now found itself negotiating at 10/half! (I wonder if the campus networking folks get alerts when that happens. Because it's almost surely not what's intended, unless I'm around.)
The younger techs in the lab mere MYSTIFIED at this exotic hardware that could send Ethernet signals over coaxial cable! That must be expensive! How did you come up with it so fast! Whoever made that must've had this application in mind, but what a niche application! Amazing!
Layers. The physical signaling is vastly different but the content that rides on it can remain the same. If you study the OSI model (https://en.wikipedia.org/wiki/OSI_model) you will know more about it than me.
I don't know how faithfully modern (or ancient) Ethernet follows this model - it might predate this work. Some layers might be blended for the sake of efficiency, but there are definitely layers.
L2 switching also requires the Spanning Tree Protocol, which is definitely not a well-liked part of the stack.
More recently, we also run ethernet over optical fibers. The varience within the fiber family of cables is probably greater than the variance in either copper or coax.
Ethernet has also been run over power cables.
At least, at university: students like me that got hired cheaply and rewired everything. :)
That was not a fun summer, but I learned a lot.
> Is it still exactly the same protocol (is that even the right word?) as it was back then?
I would be surprised, given that coax is equivalent to 3 conductors, and catX cables have 8. And that's before we get into fibre. I would expect they have sime high-level protocol (frames etc.) that gets mapped onto the physical signaling, but I don't know much about that (resource suggestions welcome!).
I do know that going from a broadcast medium to switched point-to-point is a lot more efficient etc.
Plus the taps were notoriously unreliable (variable connection quality). And would cause reflections in the cable as well, which is fun.
The cables form the physical connection, on that Ethernet defines a way to determine who may send a message (essentially anybody can send while quiet, and if a conflict is detected everybody retries after a random time)
The big thing which changed is that we are often using switched networks, instead of all nodes attaching to the same cable, but that's a change in a higher layer.
Ethernet Designers where smart not tontine the spec to properties of a specific material for transport, but abstract ether where signals travel.
100base-tx increased the symbol rate, added speed and duplex negotiation (layered into the existing link pulse signaling), but otherwise kept things the same; you can even run a 100base-tx hub.
1000Base-T is a wide departure at the signalling level; all 4 pairs are used simultaneously, bidirectionally, the symbol rate is the same as 100base-tx, but each symbol carries more bits. But the ethernet frames are pretty much the same. (Larger frames started appearing around the same time as gigE, as I recall, but that might not be accurate)
Sadly, I've largely stopped using it because it appears to be unable to keep up with the data rates typically seen from servers these days. I believe the analysis is single-threaded, and doesn't seem to cache anything either. It struggles with captures "mere" gigabytes in size, which is just seconds on a 10 Gbps link.
Also the fact ethereal/wire shark could read files saved by Tcpdump meant I could ssh onto a remote server, fire Tcpdump, run wire shark in a client and when something failed I was able to look at the network stream "from both ends". It saved me hours and hours, from dodgy ISP Nat being evident at first glance, to misconfigured MPLS networks being provable (no more the routing team could just say : it looks good for us). No, there was proof... I bet countless people continue having the same experience with this software :-)
However, I have to correct one statement made in the article. Ethereal wasn't the first free gui network packet analyzer. There was a Microsoft tool I forgot the name of that was available even in Windows NT days, perhaps "netmon"? It was a long time ago. It was free and it predates ethereal. It only worked on Windows and it used it's own file format.
so...was it really free? sounds like it came with the OS that you paid for
Network Monitor, also called netmon (or Bloodhound internally), which actually had a documented (maybe unsupported IIRC, but still easy to tap into) API. I wrote a tcpdump wrapper around it, before Ethereal was a thing. The API, and hence netmon, became invalid with the "next-gen" TCP stack of Longhorn/Vista.
Eventually, MSNA (Microsoft Network Analyzer) came along, which worked on ETW and was able to analyze network and other ETW traces. You could write handlers for any protocol in a supported DSL. You could even make it parse log files and filter/analyze the data.
The New Microsoft being what they are, they killed MSNA because it was too powerful and useful to Windows developers. It probably wasn't used by a lot of people, but if you knew how to use it it was one of the most powerful analysis tools of its time.
Edit: Microsoft Message Analyzer, not Network Analyzer.
I still don't get why MS stopped its public distribution, although I do know it was pretty buggy as released...
And yeah, netmon is great. I still use it when I want to filer Windows captures on PID, since Wireshark won't do that. (Even though netsh or pktmon -- built in Windows tools for recording captures -- have it in the header...)
Linux has a lot of friction for normal users, but easy for sysadmins.
Windows has a lot of friction for sysadmins, but easy for users.
Thankfully the tide seems to be shifting finally with simple succinct tools like curl being installed by default.
I lost half of my hair on that job. Without Ethereal, I am sure I would have lost all of it and a lot more of my sanity too.
Also used it to learn about WiFi connection setup with acces point. Can see all the beacon packets and WiFi packets
At the beginning of my career, I once spent a week in a secure facility trying to understand an annoying network bug using tcpdump because we weren’t allowed to install wireshark. The whole thing turned out to be a combination of the worst bug I have ever seen in a standard library in our decade old version of GNAT (Ada lib - admittedly it had been corrected seven years before) and an ARP misconfiguration.
The whole week was awful and largely responsible for me moving on to greener pastures. It takes a special kind of character to enjoy these things.
Just like lots of things, you can collect all the data you want... But getting an actionable result is the trick. Wireshark does that.
25 years of effort has produced a really useful tool.
Wanted to learn Go so recently started working on a CLI packet capture tool like tcpdump that parses packets received on a raw socket. Got support for ethernet, ipv4, icmp, arp and udp so far.