The bandwidth of a Boeing 747 and its impact on web browsing
blog.cloudflare.com
blog.cloudflare.com
Australia for a long time was a worst-case major economy networking situation with few, slow, and high-latency links (and still suffers from high costs due to limited links and telco monopolies).
For an understanding of how limited international data links were as recently as the early 1990s, one of the key public data links between the US and Europe at the time was a 9660 baud link according to a dated, but fascinating, compendium of networking at the time (lost somewhere in the Krell stacks, I'll dig a reference on request).
Jon Bentley's classic Programming Pearls discusses the problem of transferring graphical data images between Mountain View and Santa Cruz, California, in the 1970s. The high-bandwidth, cost-effective, sufficiently reliable solution? Carrier pigeons. Really.
I, for one, would be fascinated to see this.
But the best thing is that this was actually implemented at least once: http://www.blug.linux.no/rfc1149/
But the you'd wish the Krell would kill you.
http://en.wikipedia.org/wiki/Sneakernet
Never underestimate the bandwidth of a station wagon full of tapes hurtling down the highway. —Tanenbaum, Andrew S. (1996). Computer Networks. New Jersey: Prentice-Hall. p. 83. ISBN 0-13-349945-6.
How far we've come. Now we use 747s as the ironic competitor to digital transfer.
The main example I've heard is special effects companies like Weta Digital in New Zealand wanting to deliver video to the US across a 200ms latency link.
There are a few open protocols in various stages of development or you can pay for a product like FileCatalyst http://www.filecatalyst.com/ ( some interesting case studies on their website ).
It's actually quite elegant; a control connection over SSH plus a custom UDP protocol (FASP) for data. It does work as advertised and has some nice features that are handy in production, but licensing is very expensive ($20K for point-to-point at 1gb/s).
Can you, or somebody else please name some? It always irks me a little when someone just mentions that "these projects" exist, and then fail to mention any of them.
We offer a custom protocol file transfer as a service (uses UDP instead of TCP), but you can do a home grown solution yourself very easily.
If you have a big file of compressed content (video), do the following. If you have a lot of small files, tar and gzip them, then do the following.
md5 your file
split your file into chunks (eg parts _a - _z, or _aa - _zz)
use ncftp or similar capable of parallel transfers
optimize between 8 and 30 parallel depending on latency to saturate link
concat the file at the other end
check the md5
You can handily saturate a link from EU to Australia or India to US this way.That said, for Aussie customers sending us 1000s of movies, we FedEx or DHL pelican cases full of drives.
All engine data that the pilot can see is available realtime on the ground at the repair stations.
What this means is that an engineer can call up a GE-CF6 in the second position on the starboard wing and see how it is performing. While that might not sound all that interesting, there are literally tens of thousands of measurements taken a bunch of times a second and transmitted to the ground for each engine.
For instance, if the flight was 10 hours (and every package was 100 Kb) they'd send about 35GB of data per engine. This is just the engine measurements. This does not include avionics, communication, position.
Most of the information is propriety for how they store, transmit, and calculate these things so I do not know exactly how this works. I have seen one of these workstations and it is really impressive to know the exact RPM of the number two engine compressor is spinning at in real-time of a flight crossing the Atlantic.
Also, you can two-way text-message with the pilot from this workstation to pass perceived information about the engine performance.
The clever part was that most data would never be used - most of it was only needed if that engine developed a fault - so the database kept a central record of where each engine+date disk was in the world and only copied the data back to HQ if needed.
Now the plane has a sat data feed and the engine data goes live back to a monitoring station.
GE-CF6: A model of turbofan engine
I recall the US now has a relatively free hand in confiscating storage media for analysis/cloning during border crossings, and you might not get your copy back for a couple of months if they take a serious interest.
So, even though the aeroplane travel infrastructure itself is quite reliable, there are points of failure that would require another copy of the data to be created and shipped, with a much greater latency (It might take days to purchase more disks, transfer all the data onto them, find a willing courier, book their flight, etc).
This is also how astronomers work, radiotelescopes produce terabytes of data per day (the LCH produces ~15PB/year) so sending that over the wire makes absolutely no sense. Instead, they ship external disks for smaller datasets[0] and for bigger ones some go as far as shipping complete NFS/CIFS servers[1].
[0] https://science.nrao.edu/facilities/evla/data-archive/data-s...
[1] http://queue.acm.org/detail.cfm?id=864078 middle of the article
That is a great point though. A 4TB drive, hooked up to your internal SATA port, will do about 130 megabytes per second, which is around 9 hours for a full copy.
Drive speeds increase linearly with track density, but drive capacity increases quadratically, so this is only going to get worse. Hard drives are the new tape.
Hmm, not sure about those numbers.
Say we limit it to this trip, at 10 hours.
A Boeing 747 can carry ~30,000 kg (http://www.boeing.com/commercial/747family/pf/pf_facts.html) You can get get 3TB 3.5' drives now, and they weight about 1kg (http://www.amazon.co.uk/WD-Caviar-Green-WD30EZRX-SATA-600/dp...), so 30,0003TB is 90,000TB in 10 hours.
10 hours is 36000 seconds, so that means that the speed of the data on a 10 hour flight is actually 90000/36000 = 2.5TB/sec.
That said, if we were to account for the next generation jump (http://www.techspot.com/news/47860-seagate-60tb-hdds-possibl...) so 60TB for about a kilo, we'd get 6030000 / 36000 = 50TB a second.
Obviously, the longer the flight, the more inefficient this transfer becomes.
I was wondering if you derived the number via a similar calculation as above (a full 747), it's a direct reference to the size of the hard drive it's carrying, or just a simple typo?
When I transfer data from machine to machine, the data is immediately "usable" on the receiving machine.
When transferring data by physical means, the data needs to be brought to the final point of consumption, and possibly loaded into the host machine if it is not in some form that could be immediately connected/mounted.
So, the real comparison would be transferring data from the host machine to some form of removable media, taking that media to the airplane (early enough to meet all the pre-clearance stuff), waiting for the plane to takeoff, fly, land and then handoff the physical media at the other end.
In many cases, I'm sure the plane is still faster, but if we're talking about its theoretical bandwidth we need to look at the entire "trip" of the packets/data, not just one segment of it.
Assume a scenario wherein you change the destination to Mumbai. Your throughput will get reduced to 110 Mbps (assuming a 20 hour flight), while the bandwidth will stay the same as the networking medium (airline) has not changed, right?
Couple of other points -
* Two media with same bandwidth can have different throughput depending on the distance (i.e latency) or packet loss or other reasons. Consider two 100 Mbps broadband lines from SFO to London. One goes via the east coast, while other via Asia. Both will have different throughput; difference being roughly proportional to the difference in distances.
* 220 Mbps would imply that a maximum of 220 Mb data can pass through the medium, which is not the case, since we are loading TBs of data on the plane.
* The window size in case of an airplane would be the size of the data loaded on it i.e the size of the movie collection.
(and according to wikipedia a µSD card is 0.25g)
More like four times that: the 747-400 has a payload of 112 metric tons (248,600 pounds, for an MTOW of 412 metric tons).
However, it's a purely linear degradation with a hard upper limit.
Items detained indefinitely in the customs, items detained until bribe is paid, items lost, items stolen during load/unload, items broken (hits, pressure changes, magnetic gateways and handheld scanners etc.). Lots of problems in general, all are multiplied if your country customs and airport personnel are incompetent.
However, there is one thing that the delay (in the cited article) does not take into account is that the flights do not take-off all the time, so if you have data ready to be sent and no one to take it, then there is waiting delay before it can be sent. so if there is only one flight in a day then some data will have to wait between 24+10h to 0+10h before it can arrive. Since storage cost is very low, one could ship a lot of data but this data cannot be delay sensitive else.
on the other hand, using TCP one can prioritize data and send more important data before the rest. Of course there may be intrinsic problems with TCP receiver window and one of the proposal is to increase the initial window but the research community and the IETF is a bit split on the real benefits of it. The fear is that it may cause a congestion collapse. Moreover, the current buffer-bloat in the routers and the increase in the window size may only aid in causing a congestion collapse (http://tools.ietf.org/html/draft-gettys-iw10-considered-harm...).
And you need to sit in London traffic.
I guess fighting the speed of light sounds cooler, but that's not really the problem they are solving or speaking too, is it?
*edit just for fun: 5MB data storage transfer, 1956 http://mlkshk.com/p/AW8A (img)
Wouldn't doing it that way cut down on the latency signifigantly?
(or send the entire file, 1000s of packets) determine which don't make it and just ask for those?
This isn't my area of expertise, so I'm just curious.
It does send multiple packets (which helps throughput), but it doesn't cut down on latency. Round trip time is hard to improve.
1. TCP wasn't designed like that. It uses a rather simple scheme of sending back a byte count telling the other side up to what byte number it has received. So there's no provision for telling the other side that a bit in the middle is missing.
2. We don't send 1000s of packets at once because doing so would cause congestion on the Internet. So there's a whole subsystem in TCP that tries to determine the capacity of the link between two points and avoid creating congestion.
3. But the scheme you describe was actually added to TCP and it's called SACK: http://en.wikipedia.org/wiki/ACK_(TCP)#Selective_acknowledgm... Note that the original article isn't talking about the case where packets are being lost.
And the 1000's of packets thing isn't quite right either. With windows scaling (another option pervasively enabled) and a fat pipe (it does take some time for the algorithm to ramp up that far), you can have about 1GiB in flight on the wire at any time, that's hundreds of thousands of packets.
Because your LAN connection is probably 1000Mbps and your internet connection is likely <5Mpbs. Or the receiving sides internet connection is 2Mbps. Your computer only sees 1000Mbps and would dump a 1tb file out at line rate. Then it would ask 'what pieces are missing?' And have to send 99% of the file again. Even worse if multiple people are trying to send to the same recipient.
So instead we ask the receiving side how big its buffer is and only send that much before awaiting a response (in many cases more than 10 at a time). If there is network congestion that response will take longer automatically slowing the transmission rate. The problem is that TCP can't tell the difference between congestion latency and distance latency.
A better selling point for distributed delivery is the round trip when you deal with multiple transactions, for example requesting a web page and then the included objects, or submitting data and waiting for a response.
Modulo baggage reclaim, of course.
If you know the right folder you just copy it there, buy the game if you didn't already have it, and just wait a little for it to check the integrity of the install.
This article has me almost wishing to start a data courier service. Flying around the globe transporting data.
A rackspace portable storage center fits 12,000 hard drives (=36Pb) of data, into a 40ft shipping container.
A container ship like the Emma Maersk can carry 15,000 TEU (20foot containers) across the Atlantic in 5days
So 12,000drives * 3Tb * 7000 containers = 250 million Tb, and it takes 5days = 450,000 seconds
So a data rate of 250,000,000Tb/450,000s = 600 TBytes/s
Although you do have to watch out for pirates!
3TB drives weigh ~1 lb = 180,530kg/1lb * 3Tbytes / 3.5 hours in seconds ~= 235TBytes/s which is rather close. Especially when you consider how long it takes to load those container ships and how limited the destinations are.
PS: As to pirates your 84 million drives at 150$ a piece would be worth ~12.6 billion dollars which might actually be a tempting target even if you can only get 10% of their actual value in ransom.
Edit: Even this is small next to MicroSD cards. ~64GB/gram = 29Tbytes/lb = 2,270Tbytes/second across the Atlantic for ~12.64 billion.
[1] http://www.wolframalpha.com/input/?i=180%2C530kg%2F1lb+*+3Tb...
[2] http://www.wolframalpha.com/input/?i=180%2C530kg%2F1lb+*+3Tb...
PS: It's off by ~2.4 so probably did the kg -> lb conversion twice + some rounding errors.
Let's see if I can redeem myself Heathrow Airport averages a little over 1300 flights per day. If 1 flight can handle 1.2 million TB's http://www.wolframalpha.com/input/?i=180%2C530kg%2F1lb+*+3Tb... then Heathrow Airport in theory is a router than can handle 17,965 terabytes / sec http://www.wolframalpha.com/input/?i=180%2C530kg%2F1lb+*+3Tb... (Not that they are setup to handle freight, but at least flight time is not important only throughput.)
I was thinking that a gaziilion terabytes might be more tempting for the piratebay.org
= 1940.10 miles at 600MPH your at 3 hours 14 minutes not that there are any direct flights between those airports, but if your spending 12 billion on micro SD cards, I think you can charter an airplane and spend extra fuel going a little closer to top speed vs cruse.
It shows great circle routes for any pair of airports, with ETOPS, actual routes etc - great fun if you wanted to know why you are over Iceland
So next time you're considering physical transport include the time spent in the destination country's customs, with their highly incompetent and inefficient crews.
First off, do you really think it's more practical to ship data per airplane than a simple cable that's already connecting you? Not that bandwidth is completely for free, but most of the time you're paying a flat fee and transferring this 1TB won't cost you anything extra. I bet that Boeing is quite a bit more expensive.
Secondly, at some point the article assumes the 'transfer rate' of the Boeing to be 222Mbps. While this might be true for this practical example--let's ignore that you actually need to get the data on the plane in the first place--I bet that you could load it up from bottom to top and reach practically a speed of many gigabits ifnot terabits per second.
Thirdly, the Boeing flies with physical stuff. If the plane crashes, the data is lost. You would need backups, which take time. Over the internet it's not even possible to transfer the original medium; you'd need to first transfer and then delete it.
Fourthly, while it might be true that a packet gets corrupted 28 times during this 1TB transfer (I myself do not have statistics on this, but let's assume), on the internet you can retransmit a packet. Try retransmitting that one sector on the harddrive in under a second.
Fifthly, the article goes on about how TCP needs to wait for an acknowledgement (or 'ack' in short). I've made this mistake before when choosing for UDP instead of TCP because "it would not have to wait for acks and thus make my game much faster". It's wrong because there is a so-called congestion window. I do not know the exact figures, but it's like this: The server sends packet one. Packet one is in transit. The server sends packet two. Packet one and two are in transit. The server sends packet three. Packet one arrives. Packet two and three are in transit. The server sends packet four; the client sends ack 1. Packet 2, 3, 4 and ack 1 are in transit. Etc. The thing is that the server can send up to X packets ahead, without receiving an ack.
Sixthly, the article finally concludes that this is why you should choose for Cloudflare--so it was an advertisement after all. That explains why it's told in this story-like way and why it contains so many technical errors.
1. Yes, physical shipping is more practical than using the Internet for large transfers and this is done frequently. For example, Amazon Web Services offers the possibility to send them disks: http://aws.amazon.com/importexport/ Shipping of disks is also common in the movie industry. For example, a lot of digital movie distribution to cinemas relies on shipping of hard drives.
2. I used a statistic of 0.1% packet loss on the Internet and 28,000 flights per day to arrive at 28 flights lost per day.
3. There are two interacting things happening in TCP: the congestion window (which limits the number of outstanding _packets_) and the receive window which limits the number of outstanding bytes. This blog post does not address the congestion window at all and I assume that the receive window can always be filled. Thus is _overestimates_ what might be possible on a real link.
4. For very large transfers there are improvements that can be made using UDP and there are protocols that do just that. See, for example, TixelTec.
As well as sciences based on massive amounts of data (astrophysics and particle physics for instance, where experimental systems can produce terabytes per day), financial data analytics, or petroleum seismic surveys.
You should calculate how large a window you need to fill a 100mbps pipe between SF and London. Then calculate how likely it is to remain there before you drop a packet, and how long it takes to detect a dropped packet with typical timeouts, and so on.