Blip: A tool for seeing your internet latency
github.com
github.com
We had ridiculous bandwidth for the time, which users knew, and they would often come and ask "why can I stream youtube in HD but accessing my document/file takes ages" (ages being anything more than 2-3 seconds). A perfectly valid question. The answer was usually "illegal fishing", or "sharks". If the primary sub-sea cable was cut, we'd have to route through another which would add at least 400ms round trip extra.
Luckily this was a well known risk/issue for our parent company who had a team dedicated to negotiating peering and routing between office locations and their nearest regional DCs (these were not azure, aws, etc. usually local IX or Teclco). But they couldn't do anything about the sharks :(.
I had to double take when I saw illegal fishing, I thought you meant the digital kind and my mind went to ddos/similar :P
Average users will barely even notice a start time of up to 1 second. 500ms will seem fast to them, for starting a video.
Typing a document in Google Docs and having a 1000ms round-trip latency for every letter they type, however, will quickly become annoying.
Yeah, regardless of throughput, you're going to have a bad time with interactive workloads with high latency.
About a decade ago, one of our customers had a 25Gbit (IIRC) line that went cross country and had something like 150ms of latency. Then after hours they would basically cross replicate their entire data center through a couple machines they purchased from us (running some of our software). While their packet loss was fairly trivial, it was enough to cause some serious pain. Explaining why, if they insisted on TCP they were going to need an enormous amount of ram (for the time) in order to keep the thing running at line rate took about 6 months... Eventually someone ponied up the $$$$ for the ram, and the problem went away, and it sat there happily doing its thing, with only the occasional problems caused by other departments trying their own experiments and killing its throughput.
Changing to UDP would get rid of the congestion control but UDP on its own would have congestion issues killing its throughput.
Bittorrent created their own congestion control called ledbat and then added that onto UDP to create UTP which could help. There is also UDT which has its own congestion control algorithm designed for 10gbps+ connections, which would significantly fix the issue.
Here in Perth, Western Australia we are very far from everything. In the 90s/early 2000's our routing and latency was pretty terrible, especially over dialup and slow connections and that's how I thought online gaming was (300ms+++).
When there started to be "local" servers in Adelaide (thanks Internode!) and Sydney, this went down to 25-30ms for adelaide and 50ms for Sydney.
Now we have undersea cables to Singapore from Perth, so we actually get 2ms lower latency from PER->SNG than PER->SYD! Sometimes I get weird routing through Asia if a service determines I'm better suited for that. There is apparently also a cable from PER->Europe now, where we now get less latency to EU (it's now less than US, crazy).
Old Node was one of those rare companies where I actually believed they put their customers and employees before their bottom line. It's just TPG with different branding now.
(by the way, has anyone seen my orange watch?)
We'll see how being public keeps this up..
I particularly like that they've exposed some of the NBN troubleshooting tools to customers.
They also skipped over that "we'll provision more CVC capacity when we see it flatline for at least four hours a day" phase that most other retailers seemed to go through.
They don't quite have the same old school nerd appeal that Node had though. Things like free Usenet, games.on.net, unmetered downloads from the fairly well stocked mirror..
It's a mix of traffic, peering agreements between providers* and QoS rules on the routers you hop through along the way. After your request leaves your ISPs hardware often you have no choice where it goes next.
Something that wasn't mentioned was that this is why traders will pay millions to set up servers in DCs/IXs all the way between them and the market so they have more control over the route their traffic takes. Sometimes even running their own cables.
SEA-ME-WE-3, the cable you were talking about was the only game in town for that side of the country until about 4 years ago when ASC came online. ASC has over 10 times the capacity (60Tbit/s) of SEA-ME-WE-3 which was below 1 Tbit/s until 2009 and is not only 4.6 Tbit/s according to wiki. It's worth noting that SEA-ME-WE-3 is administered by Singtel there are 90+ other investors (who were all getting preference or had capacity agreements at launch).
* this is complicated, someone feel free to correct/expand but essentially carriers will lease capacity over particular connections and sometimes the agreements between two carriers will mean particular customers traffic may get preference, so on busy lines, businesses paying for higher SLAs get prefered over regular consumers who might shunted off to a lower grade connection where there's space.
There's a great example with NTT's partnership with Formula 1 where NTT engineers go in before race day to location making sure all that incredible Digital Twin data is getting fed back-to-base to manufacturer's HQs in the UK or Germany etc. as fast as possible. Sure there's the pit crew, but there's a whole team back home who are watching and analysing data and relaying advice. F1 is really an incredible technical feat in more ways than just the cars.
strangely I found the same was true for reliable video streaming from NZ to UK
instead of Auckland > Sydney > California > New York > London, the optimal path undoubtedly proved to be Auckland > Sydney > Singapore > NoMansLand > London
traversing America introduces all sorts of problems, I'm guessing due to deep packet inspection rather than distance
You make it sound like every other day some UNDERSEA CABLE gets cut - that's not a once-in-a-blue-moon thing?
Even rarer is something like the recent Shetland outage where virg cables were lost.
Most of the time I'd find out officially via network operators bulletin boards (i.e. those looking after AS and IXs) there wouldn't be anything reported in the local news because, for most people, as this article alludes to, the content edge servers they actually connect to are pretty close, both physically (distance) and virtually (hops).
"Don't these cables ever break? Yes! Cable faults are common. On average, there are over 100 each year.
You rarely hear about these cable faults because most companies that use cables follow a “safety in numbers” approach to usage, spreading their networks' capacity over multiple cables so that if one breaks, their network will run smoothly over other cables while service is restored on the damaged one.
Accidents like fishing vessels and ships dragging anchors account for two-thirds of all cable faults. Environmental factors like earthquakes also contribute to damage. Less commonly, underwater components can fail. Deliberate sabotage and shark bites are exceedingly rare. "
I'd expect cables through the Atlantic or Pacific ocean to get damaged much more rarely than cables through your local lake. Shallow waters and sports/amateur boating can damage the cables in shallower waters more easily than in the humongous depths of the greater oceans in the world.
Requests get blocked for mixed content if the site is loaded with HTTPS. That's what causes the constant stream of red blips after a few seconds.
So you might be wondering, hey, how did you make a javascript applet ping these arbitrary servers? What about cross-domain request protection?
Answer: I did it by just making the queries anyway, and seeing how long it takes to get the error message back that my request was refused because of cross-domain request protection. Yes, this results in an infinite number of error messages to your javascript console. Don't look at your javascript console and you'll be fine. Trust me on this.
the connection has to go through HTTP (and not HTTPS) as GP states, if browser blocks requests
> Answer: I did it by just making the queries anyway, and seeing how long it takes to get the error message back that my request was refused because of cross-domain request protection. Yes, this results in an infinite number of error messages to your javascript console. Don't look at your javascript console and you'll be fine. Trust me on this.
I'm hoping to use his WiFi debugging guidance from the readme at my in-laws: I was just experiencing remote desktop problems despite having a strong network connection. I suspect the culprit is Comcast's supplied router.
This also means if you rely on CORS to prevent XSRF attacks (which is maybe not the best idea), you must be ensuring that any request that came in would have been preflighted (for example, reject requests unless they have a special header)
"Preflight check" is such a wrong analogy, since with CORS you fly all the way to the destination to check if you're allowed to fly to the destination.
*cargo isn’t really the right analogy… but close enough
Nah, the whole thing is just a "Can I ask you a question?" implemented in code.
addendum: one can rarely state anything entirely accurate about CORS briefly.
Rather, it seems to be pinging my router (192.168.1.1).
> Blue blips are your ping time to apenwarr.ca
Rather, it seems to be pinging something in "Mexico City, MX" or "Moncton, CA":
What does the DNS checkbox do?
There is a hard-coded list of RFC 1918 addresses that are commonly used as routers that the program tries first to see if they are faster than gstatic.com.
https://github.com/apenwarr/blip/commit/b284f922b047e9032112...
Instead of apenwarr.ca, the code tries a bunch of sites from measurementlab.net:
https://github.com/apenwarr/blip/commit/20f99c1d641e8cc607b6...
The DNS checkbox does the following:
Generate a pseudorandom hostname for each test and looks it
up in DNS. The hostname happens to be in a domain that returns a valid
IP (incidentally always the same one) for every hostname you ask for.
This triggers a new HTTP connection, but also validates that your DNS
is working reliably without weird dropouts. I made it optional for
two reasons: it hits your DNS server pretty hard, which is more rude
than blip is usually; and DNS is surprisingly crappy, so you might want
to turn it off while looking to see if you have *non* DNS related
problems.
https://github.com/apenwarr/blip/commit/0678d0668c14e2c7a7f0...https://github.com/apenwarr/blip/commit/4a1640977303f6dc58db...
I use OpenWRT's Smart Queue Management package to avoid bufferbloat. Without it, sites like dslreports.net and fast.com show that my latency goes way up under load.
However, when I have SQM enabled, I see regular red blips with this tool. Disabling SQM makes the blips go away.
Strangely, even without SQM, latency as measured by this tool does not seem to change when I'm e.g. streaming a Youtube video in the background, as I would have expected given bufferbloat. (Does bufferbloat only apply to the specific packets being streamed?)
It's not too surprising that you don't see latency increase under light load like Youtube. "Under load" really means uploading (especially) or downloading (sometimes) at maximum speed. But Youtube usually doesn't download at maximum speed, because that would imply you can't keep up with its streaming video rate, and you'd get glitches. So your connection is likely not really "under load" at that time. Try uploading a large file somewhere and you should see an immediate change in blips.
I don't think I understand this enough to file a good bug report...
Chances are changing the SQM backend (eg. between fq-codel and cake) will at least change the behaviour and likely make the bug go away.
fq_codel/simplest.qos made the blips worse instead of better. fq_codel/simplest_tbf.qos reduced but did not eliminate the blips.
What hardware do you have? SQM is pretty CPU intensive. Mine is just powerful enough to be able to use SQM. Try checking if your cpu usage spikes from softirq interrupts.
I have cake configured on my border router.
In retrospect the industry really got shot in the foot on this ISP bandwidth race. ISP marketing kept going for bigger and bigger numbers but marketing only needs one big number (download speed) to sell subscriptions and so upload speed was de-prioritized leaving us with these wildly asymmetric connections (standard around here is 300/10), but if the upload is saturated (easy to do) the TCP Acknowledge packets can’t make it out and the whole connection grinds to a halt.
Those red marks are kind of a “slap on the wrist” that keeps the sender in line.
P.s. Try downloading something that can send to you at 200MBps (and therefore saturate your connection) or something that can receive file uploads at 10Mbps, or do both at the same time and see what that does to the test results.
And a good SQM algorithm shouldn't be shooting down packets from a tool like this unless the configured bandwidth limits are really low, because this tool isn't really generating all that much traffic.
A few hours earlier and I was seeing regular red blips, but the service is definitely highly variable. I'm also bouncing through two Unifi AP's before I hit the uplink.
Use of 'ping' was a bit confusing, as it's obviously not ICMP. I'd hoped Animats might have popped in with some observations around how his algorithm may help (or hinder) this kind of test. : )
Someone mentioned Riverbed (Steelheads) which have lots of knobs you can turn to deal with high latency and/or non-reliable connections - around window sizes, backoff / congestion settings, SCPS, and of course Nagle (or Neural Framing, as they called it, IIRC).
It's extremely hard (I've found it) to identify if packet loss is a problem in your house, or with the service you're using. Also, if you don't have a baseline for how your network normally operates, everything looks like a nail when you go digging.
I was having packet loss problems every 5 seconds in a particular game where it would spike up to around 15% then recover. I ran a ping plot to the router at 0.1x per second and had absolutely no packet loss. So is it me or is it the game server.
I was having exactly the problem GP described a few years ago, and this was the fix.
It just about drove me nuts. It got to the point where I replaced my AP, started dragging a 10 meter ethernet cable around my house for my laptop, and started to suspect esoteric things like the local airport weather radar triggering DFS[^1].
In the end, it just turned out to be the damn Location Services.
[1]: https://en.wikipedia.org/wiki/Dynamic_frequency_selection
I even went down the path of buying new powerline adapters only to have the same problem. I was unplugging things in the house and factory resetting everything. I've felt like I was going to crazy lengths to troubleshoot.
As a temporary measure, am I able to simply switch off the WiFi on those Mac devices before running those commands, to test?
I should have conditioned my previous comment with "if this is happening on an Apple device". In case it helps, I described what's happening under-the-hood in another comment: https://news.ycombinator.com/item?id=33451879
Location Services works by asking Apple if it knows the location of any of WiFi access points near you.† Apple knows these locations by using services like Skyhook who wardrive around mapping locations of BSSIDs (AP mac addresses).
So, when Location Services is on (which it is by default), macOS will periodically switch your wireless card to monitor mode to find those nearby WiFi access points. Doing so stops normal network traffic for around a few 100 ms on the device.
† On devices with a GPS (i.e. iPhones and iPads) it will also sometimes use the GPS, but does so sparingly because the GPS uses much more power.
> Apple ditched both Skyhook and Google location services and began relying on its own databases starting with the release of iOS 3.2 for the first-generation iPad in April of 2010
[1] https://appleinsider.com/articles/13/07/03/skyhook-accuses-g...
It's surprising to me how so many programs seem to check location on an interval timer, including programs that seem at a glance to have no reason to need to know the user's location. I wonder if there's some common SDK library that makes it easy to accidentally enable.
[1]: https://github.com/texstudio-org/texstudio/issues/62
Others that play the same game don't seem to experience the lag spikes as badly as I do, but I'm not sure whether they just get routed differently or if the problem is in my house or close to my house :(
The silly trick blip uses involves pinging non-encrypted HTTP web servers, which is not allowed from an encrypted web page. So you really have to load blip from a non-encrypted server.
This code is from 10 years ago. Most likely we could find some way to work with HTTPS nowadays, but alas, I don't have time to maintain it.
The blue ticks actually don't work, though. They're always accompanied by red and seem to be timing out instead of the usual exceptions.
I recently upgraded to a Mac Studio with 10GBase-T and was experiencing regular SSH disconnections. Blip allowed me to isolate this to some sort of sporadic incompatibility with the 10GBase-T SFP+ modules I used in my Ubiquiti switch (despite having tried two brands of SFP+). Switching to a switch with native 10GBase-T ports (Zyxel XGS1250-12, only £225) fixed this.
The multiple ring buffer timescales is genius!
When you connect to your favorite MMO, etc., it's doubtful that you're using the HTTP protocol to do so.
Still useful, but for one of many "Internet"-based scenarios.
TCP probably has higher overhead, but even that is going to be next to nothing unless you've got packet loss
traceroute on UN*X uses UDP by default. There is no overhead when using UDP.
ICMP Echo, I'm not sure about, but I don't believe it has overhead, but as I said, it is a depriortized and possibly rate-limited by routers.
HTTP 1.1 and 2.0 do not require TCP, although that is by far the most common way they are used.
What overhead would TCP bring after the three-way handshake on a uncontested stable network without packet loss? I would expect it to be no more than UDP.
> traceroute on UN*X uses UDP by default. There is no overhead when using UDP.
You are correct. UDP will get you lower numbers, I was more of saying that TCP will not be much worse (maybe within 1ms) under good network conditions.
If I ping apenwarr.co I get consistent results:
Pinging apenwarr.ca [74.207.252.179] with 32 bytes of data: Reply from 74.207.252.179: bytes=32 time=203ms TTL=51 Reply from 74.207.252.179: bytes=32 time=202ms TTL=51 Reply from 74.207.252.179: bytes=32 time=205ms TTL=51 Reply from 74.207.252.179: bytes=32 time=202ms TTL=51
And on my mobile, on the same network, my blue blips look normal enough. Just not on my wired desktop.
ping -c 4 gstatic.com | tail -1| awk '{print $4}' | cut -d '/' -f 2 | ts '%F %T' >> pinglog.txt
and tail the log file or use a tool like LiveGraph (https://live-graph.sourceforge.net/) to get a live graph of it.Very useful for diagnosing ISP service interruptions
Edit: Nvm it is because of https: https://news.ycombinator.com/item?id=33446511
So, on my Pixel 6, I see pretty bad latency: https://ibb.co/ft4ZmJr On my Mac, it looks decent: https://ibb.co/XJMXSzm
What can be done to improve it?
`iw dev $interfacename set power_save off`
Linux on various laptops/kernel's has been known to be a little too aggressive with the power savings and it absolutely kills first ping latency.
(Although at first glance this thing might be fast enough not to notice).
(I don't think it's the connection being flaky. Running `ping gstatic.com` in the terminal at the same time as the webpage is in the foreground shows a maximum time of 23ms after 180 pings.)
Ping/tracert will give you a _general_ idea of what is going on, but it is not something to definitively answer a question (i.e., a tracert may have an asymmetrical path -- travel one route to the destination, but the return packets travel another route, which you'll never see).
When I used ping to test my iPhone's wifi latency (1), I frequently saw 200ms results, even when the phone screen was on and I was using an app. Maybe there's aggressive power saving stuff causing latency on the network stack. I was very surprised as well. Friends with Google Pixel phones tried a similar test on their Wi-Fi networks and got similar results.
Edited to add: When I test with the Blip website, it never goes above 80ms for gstatic. Go figure.
1. Tested by a wired Mac pinging the default gateway, a Google Nest Wi-Fi Router, which was <1ms consistently, then I used the wired Mac to ping the phone
I have quite some wifi issues at home but when I check things like ping everything seems normal.
I'm going to try this out when I have issues in a video meeting and see if there is a correlation.
> Answer: I did it by just making the queries anyway, and seeing how long it takes to get the error message back that my request was refused because of cross-domain request protection. Yes, this results in an infinite number of error messages to your javascript console. Don't look at your javascript console and you'll be fine. Trust me on this.
The author mentioned the tool sends out requests, and just measures response time regardless of response status. (If they fail, they fail!)
And they measure response times from gstatic.com, which should have edge nodes near most locations.
Shot in the dark guess - perhaps you're hitting some kind of rate limit somewhere in that chain?
Requests get blocked for mixed content after a few seconds if the site is loaded with HTTPS.
Also https://ping.pe/ is good.
Is there something in particular that could be causing red dots at 3 second intervals at or above 2000 milliseconds?
Is this a mac ?
Requests get blocked for mixed content after a few seconds if the site is loaded with HTTPS.
I tried with HTTP and the red dots or bars disappeared. With HTTPS they are there.
Question - What are good ways to deal with dead zones?