Speedbump – a TCP proxy to simulate variable network latency
github.com
github.com
Turns out, I already had everything installed on my machine to do it, via `tc` (Explained a bit here: https://wiki.archlinux.org/title/advanced_traffic_control), which apparently came with the iproute2 package on my distribution.
With that, you'd be able to run something like this to add latency on a specific interface:
tc qdisc add dev eth0 root netem delay 100ms
Really easy to use, works well in docker container too, comes with a bunch of different conditions you can apply (delay, packet loss, duplication) and you might just have it installed already.That could be really useful for maybe simulating weather effects on satellite/RF over time?
I'm actually quite surprised that a similar frontend doesn't seem to exist as a commercial hardware product, unless we missed one in our search.
I think a programmable proxy with simple hooks is what I want
I had to write my own emulator once to simulate a specific commercial satellite terminal. The terminal had the behavior of queuing up packets until they hit a certain threshold or exceeded a time limit and then bursting them out in a big blob. It also "helpfully" reordered small packets to the front of the queue for better latency, which made TCP stacks very cross.
It turns out determining that a downstream service is slow is a lot harder than determining if it's unavailable, so it was an important way for us to test how services handle slowdowns and network problems as well.
It was really simple, it just dropped some configureable percentage of packets, which would force a resend, causing the other side to get packets delayed and out of order.
It ended up finding a lot of problems in our error handling code for network access.
I'm convinced 90% of webapp bloat would go away if the people building them didn't have a gold plated Cadillac computing experience.
https://firefox-source-docs.mozilla.org/devtools-user/networ...
Admittedly this only works for front-end, browser based testing.
It’s like the safety car in Formula 1: the cars will all slow down to crawl, but they won’t change order because they can’t overtake.
Better to test a scenario like in F1 where one car hits another, and then takes a few more out with them. That will really scatter the order and timings of things. That’s more how the real world of Wi-Fi and cellular connections works, so being able to handle that level of unpredictability is going to lead to way more robust apps and a better UX.
I have a dream of opening a co-working space for developers where you can connect to “Free Airport Wi-Fi” to really test your web application is going to survive in the real world.
More app developers could help others by testing with simulated intermittent connectivity.
From "Toxiproxy is a framework for simulating network conditions" (2021) https://news.ycombinator.com/item?id=29084277#29088775 :
> Many apps lack 'pending in outbox' functionality that we expect from e.g. email clients.
> - [ ] Who could develop a set of reference toxiproxy 'test case mutators' (?) for simulating typical #DisasterRelief connectivity issues?
1. Wait until you have enough data to send in at least one packet (maximize throughput).
2. Send data as soon as you have it, even if it is less than a packet's worth. (minimize latency).
If you want to send a file, for example, you should use the first choice. A common or naive implementation might be to read the file in chunks of 9kb (about the size of a jumbo frame), send it to the socket, and then flush it. There are multiple problems with that approach.
1. Not everyone has jumbo frames, so if we had a common MTU of 1500 bytes on our link, that means you'd actually send ~6.25 packets worth of data.
2. Even if you were using jumbo frames, TCP headers use up that space. So, if you flush the entire frame's worth of data, you'd send something like 1.1 packets.
3. The filesystem might or might not give you the maximum bytes you request. If there is i/o pressure, you might only get back some random amount instead of the 9k you requested. So, if you flush the buffer to the TCP stack, you'd only send that random amount instead of what you assumed would be 9kb.
These mistakes generally send lots of small packets very quickly, which is fine when the link is fat and short. As soon as one packet gets lost, you're still sending a bunch of small packets, but now we have to stop and wait for the lost packet again before we can start handling more of them. So, if you are sending to someone on, say, airport/cafe wifi, they will have atrocious download speeds even though they have a pretty fat link (the amount of retransmissions due to wifi interference is a large % of the bandwidth). This is very similar to "head of line blocking" but on the link level.
I only had to learn about this fairly recently because my home is surrounded by a literal ton of wifi access points (over 20!) which causes quite a bit of interference. On wifi, I can get around 800mbps on the link but about every 10th packet needs to be retransmitted due to having so many access points around me (not to mention my region uses most of the 5G band for airport radar, so there are only a few channels available in that band). So, when applications/servers don't fill the buffers, I get about 50kpbs. When they do fill the buffers, I can get around 300mbps maximum, even though I have an 800mbps link.
Hope that helps.
Would there be a benefit to autotuning with e.g. sqm-autorate in the suggested environments?
``` # Setup pipe sudo dnctl pipe 1 config bw 1Kbit/s delay 800
# Setup matching pf rule echo "dummynet out proto tcp from any to 127.0.0.1 port 11211 pipe 1" | sudo pfctl -f -
# Turn on firewall sudo pfctl -e
# Test time nc -vz 127.0.0.1 11211 Connection to 127.0.0.1 port 11211 [tcp/*] succeeded! nc -vz 127.0.0.1 11211 0.01s user 0.00s system 0% cpu 1.333 total ```
It turns out to be also a very good way to test a networking library by implementing it. Since your stack needs to be able to basically handle most adverse events properly.
The idea behind 'chaos engineering' is cool.
I am developing a progress bar for a web crawler, and testing on localhost is too fast to notice if there is any issue. With speedbump, I just do `podman run --net=host kffl/speedbump:latest --latency=1s --port=8001 localhost:8000` and test my crawler on http://localhost:8001
Neat tool.
Toxiproxy primarily serves the purpose of integration into tests. It proves particularly useful when testing features like the progress bar directly from code. Notably, for Go applications, integration is seamless, eliminating the need for an additional application. Instead, you can run the Toxiproxy server part directly from the codebase. An illustrative example can be found at https://github.com/Shopify/toxiproxy/blob/main/_examples/tes....
He was a joker that guy.
I'd love to jettison our hacky custom code and use something off-the-shelf instead.
Far too many realtime things, like games and videoconferencing (both of which frequently use TCP, despite it being badly suited to the application), perform really badly when the bandwidth increases and decreases, for example as I walk around a building with wifi and pass concrete pillars.
I want not a single dropped frame in those circumstances. Sure - you can have some lower res frames, but I don't expect to see 30 frames dropped in a row and a big glitch.
If there’s real interest I could dust it off… maybe rewrite in Rust.
https://www.pico.net/kb/how-can-i-simulate-delayed-and-dropp...
(just a short example)
PS. My knowledge of tc is very rudimentary. I only had to use it once.
[0] https://mininet.org [1] https://mininet.org/api/classmininet_1_1link_1_1TCIntf.html#...
I wrote a tool like that a few years back[0], but didn't publish most of what I discovered about e.g. how different CDNs tuned their TCP stacks (just a couple of anonymized examples[1]).
[0] https://github.com/teclo/flow-disruptor
[1] https://www.snellman.net/blog/archive/2015-10-01-flow-disrup...
This is quite difficult, I think. You only know how bad the connection is after dropping a frame or two, and the intervening network will detect it before the endpoints do.
From an application level, I suspect that on the socket layer level you can't do anything about TCP anyway, the underlying layer will retransmit after a delay.
[EDIT: Just realised you meant video frames, and not network frames. I think my comment still applies though - dropping a single transmission unit that is part of an I video frame still means discarding that entire frame, whilst dropping a B or P frame is going to result in glitches.]
For the traffic from the remote end, data packets can consist of "here is the low res data" and another packet for "and here is the extra data to turn the low res into higher res". The high res would be sent with QoS markers so that network hardware knows that if anything is to be dropped, drop that first. The different QoS streams could even be sent with different transmit powers and modulation schemes on the WiFi network.
Network speed is a function of link speed AND the success of packets to reach you. When you are on wifi, your device gets a very tiny slice of time to send packets. If you have more devices, your device gets less time to send packets. Further, a "loud router" nearby (or self-interference ... or radar in 5G space -- or a microwave, or Bluetooth in 2G space) can cause a frame to be interfered with and never be received. Thus it has to be retransmitted.
RSSI is only a very small part of your overall network conditions and doesn't really tell anyone anything.
What Zoom does is notice this and play the video frames in its buffer back more slowly in an attempt to smooth over the "no data coming from the Internet" period. I don't find this that helpful because if you comment on something someone just said, they've moved on to a new topic by the time Zoom unbuffers and displays all the frames. Clever hack, I prefer the complete dropout.