A Desperate Plea for a Free Software Alternative to Aspera (2018)
ccdatalab.org
ccdatalab.org
There is a lot of good science behind fasp. An advantage it has over IETF protocols is that both ends trust one another. Another advantage, until recently, was out-of-order delivery.
The protocol totally ignores drops, for flow control. Instead, it measures change in transit time. The receiver knows, the sender needs to know, but the useful lifetime of the measurement is less than the transit time. This should make an engineer think "control theory!", and did. So, the receiver reports a stream of transit time samples back to the sender, which feeds them into a predictor, which controls transmission rate. Simple, in principle, but the wide Internet is full of surprises.
If you think this wouldn't be able to go a thousand times faster than TCP, you have never tried moving a file to China or India over TCP. :-) (Customers used to report 5% drop rates.) Drops and high RTT are devastating to traditional TCP throughput on high-packet-rate routes; read about "slow-start" sometime, and do the math. Problem is that for untrusted peers, drops are the only trustworthy signal of congestion. Recent improvements where routers tag packets to say "I was really, really tempted to drop this!" help some.
Torrents get the out-of-order delivery and the lower sensitivity to drops, but its blocks are too big.. Others commented that opening lots of connections gets around some TCP bottlenecks, but that helps much only when the drop rate isn't too high (i.e. not to India or China).
Do you think the protocol might cause problems if it were more widely used?
Interestingly (at least, I find it interesting) fasp can happily use up all the bandwidth the various TCP streams aren't using without affecting TCP rates at all. It also backs off and shares bandwidth with other instances of itself or other well-behaved protocols. With some coordination between them, multiple senders can share the available bandwidth in any chosen proportion.
So, you can set up where A always gets 70% when it is sending at all, and B, C, and D share whatever is left when they have anything to send, but all without ever slowing whatever TCP traffic is running. An administrator can (literally) drag up and down the rates of ongoing transfers according to organizational priorities.
That sort of featurism is part of why Aspera can command crazy prices.
It sounds like people are getting good results with fdt. It leads me to wonder if fdt has any of this good-network-citizen capability. A lot of people using fdt and making the net worse for everybody else would be an unfortunate outcome. Aspera was pretty careful about it.
(I keep trying to write ftd, but that is the florist service.)
But out of order delivery and (I'm guessing) out of order retransmits appear to be unique to fasp. What did you mean about these being unique "until recently"? [edit - nevermind, saw below you were referring to sack]
PS. I remember looking at fasp ~10 years ago and it looked like a fantastic tech. Someone's taking things that many people only talked about and putting them into a product that really worked well. Cutting edge stuff. A job like - a dream for many :)
(I am looking at https://blog.apnic.net/2017/05/09/bbr-new-kid-tcp-block/ .)
Fasp does not rely on RTT ACK timing, which is very noisy, but only outgoing delays.
The routers are applying their own algorithms to simulate sane queue timing, but they are doing all kinds of crazy shit under the hood to get better link performance. The queues, IOW, are a fiction provided to the endpoints to simplify their job and generate understandable, thus tunable full-route dynamics.
But the ACKs get completely different treatment, because traditionally nothing cared about ACK timing.
So, "onset of queuing", e.g., is pretty meaningless in practice, and fasp makes no attempt to detect it.
So it'll be interesting to see if BBR+QUIC can be made to do what FASP does.
One problem with UDP protocols over the open internet is that you cannot trust UDP packet contents, which could come from anywhere.
Some customers (e.g. DOD) sent in clear, unsigned. They trusted their networks.
Wow. I had no idea. This may be one of the dumbest things I've ever read. Encryption is free.
A brilliant woman from, where, Slovenia? figured a way to use SSE instructions to encrypt just as fast as the built-in instructions, I think 2 cycles per byte.
I guess DJB's ciphers run much the same way, nowadays.
There are pretty good references starting from https://crypto.stanford.edu/vpaes/ .
So they started using Aspera to send video from the UAVs over Iraq etc. to the Pentagon. Before that, they were literally flying boxes of videotapes (or CDs? Tapes seem hard to believe, except this was govt.) to a warehouse, where analysts could check them out, like from your neighborhood library.
Probably now it's all realtime. But they have to stream it, somehow, so it is likely still going via Aspera to Nevada or wherever they work the remote pilots.
On first read I thought you meant Aspera ran on the UAVs :) but I now realize the high unlikeliness of that...
Jihadists could tune in an ordinary scanner, and watch the video coming from them. Embarrassing.
It's kind of amazing really that the near future will probably be about removing massive amounts of code & design from all sorts of incredibly carefully engineered systems like the linux kernel in favor of a pile of linear algebra that can just figure things out. resistance is futile.
(Less flippantly - neural nets are mainly for the case where you can't effectively do manual feature engineering because your inputs are too varied and complex, you don't particularly care about the mathematical properties or guarantees of the solution, and where you don't even really know what the solution should look like [so are forced to randomly initialize a huge nonlinear system and optimize it until useful behaviors pop out]. Something like predicting path saturation would be far better suited to traditional approaches.)
Edit: By the way, I accepted your initial opinion about traditional approaches being better at face value but that was like ~45 minutes ago. After thinking about it a bit more I realized that there is alot of interesting network data that could be fed into such a NN, moving it (ding!) back into the nonlinear/complex system column. Suddenly my approach is looking competitive to also being the optimal one (excluding hybrids).
Care to elaborate?
Can you elaborate on this? Which recent development are you referring to? Thanks.
I've seen rather good speeds over some bad links from my bittorrent server using large window sizes (up to 9), BBR flow control, ECN fully enabled, and SACK.
Would you be willing to talk more about this, for someone not as intimately-familiar with the details?
But, imagine the route as a series of packet queues. Packet is sent, arrives at the first router, spends time in the queue, gets sent, put on the next queue, waits, goes, until it's delivered. If the total time taken, compared to the last packet's, is more, that means the queues are growing. If the total time is less, the queues are shrinking. You adjust your rate to stabilize them. Other flows are being added and removed all the time, so the optimal rate varies in steps over the course of the session.
This involves control theory, Kalman filters, PID, the whole business.
:)
How good is GNSS at simulating latency, buffer saturation, etc on a simulated network and how much complexity would I need to contemplate factoring in to build an accurate model of the average organically evolved, ad-hoc, multi-WAN-spanning sprawling corporate network?
I can't think of a more concrete, instructive starting point for someone interested in figuring out an alternative/reimplementation.
You can get 80% there quickly, 90% there in two years, but the rest is just hard. A versatile simulation environment feeding the algorithm with synthetic time events according to a configured schedule is probably the best way to ramp up quickly, but I don't know anything about GNSS.
One tricky phenomenon is bimodal and even N-modal delay samples from traffic split across different routes.
Around 2014, Aspera replaced most of the code wrapped around the algorithms to become much more nimble about applying them in more circumstances.
Hearing I can get to 80% quickly is encouraging from a MVP viability point of view, thanks. (And also from a fail-fast perspective; I've now gotten to wondering how much additional performance might be eked out of start-stop style traffic, and it's good to know it won't take long to find out whether the pursuit is worthless or not. Yay.)
I read the second half of that second paragraph as describing a network model simulation that "compiles" a particular routing graph/topology into a set-in-stone sequence of packet events that you then later analyze...? (I'm interpreting "the algorithm" is my target application code, and "synthetic time events" and "configured schedule" as hints at non-real-time pregeneration. This may be incorrect.)
Eek, the delay distributions you describe almost sound like they might be NP-complete to solve.
Routes tend to start monomodal and split and rejoin at specific times, so it is not usually as bad as you imagine. Also, the number of routes does not get large.
(Raptor codes (a type of FEC) can essentially transfer over UDP at the underlying line rate, even if packet drop is high. In other words, if it is a 1 gbit/s link, with 5% packet loss, you will be able to send over 900 mbits/s without needing any TCP features). This is also helps high latency scenarios, as you don't need test/back off how much data to send...you just send as much Raptor encoded data as you want down the pipe, and as long as you send enough recovery packets with it, you will be able to reconstruct the original flawlessly).
(Also see https://par.nsf.gov/servlets/purl/10066600
http://www1.icsi.berkeley.edu/~pooja/HowUseRaptorQ.pdf
https://a1f9fb7d-b120-4c73-98e3-5cdb4ec8a2ab.filesusr.com/ug...
Erasure coding helps, but it demands much more CPU than simply having a very wide window for retransmits. For bulk data transfer of files, retransmitting any part of a 1GB window is trivial.
Used fdt to transfer a 6TB archive out of AWS very smoothly at full speed.
I tried downloading the file using axel -n 32, and it took 1:28, which is only 9x slower. Curious though, when it started out with all threads running it was transferring data at 29594.2KB/s and hit 50% after only a few seconds, but by the end with only a few threads left running it was only doing 5625.8KB/s.
It looks like the performance varies a lot, possibly due to multiple interfaces or wan links being used.. or issues with their server.
Using axel -n 64 over http I was able to fetch it in 50s, which is only 5x slower, and most of that time was spent from 98 to 100% as the last few connections finally finished.
A smarter client that more aggressively re-fetched chunks being downloaded over slow connections would likely match the same performance.
I worked in VFX, so I've been using aspera since before it was owned by IBM.
Depending on what line speed you have, but for a gig link we made a simple protocol using parallel TCP streams.
Basically, it chunked up the large file into configurable sized chunks, and assigned a chunk to each stream.
Another stream passed the metadata.
This has all the advantage of TCP, with less of the drawbacks of a custom UDP protocol.
For transferring files from london to SF we were getting 800mbit/s over a 1 gig link.
If Aspera is as good as the parent suggests it's probably because they are operating a bittorrent-like network of geo-distributed peers. It would be very cheap to emulate that using cloud computing providers, especially if you are discarding the data after transfer.
The special sauce is measuring the packet loss quickly enough to make sure your not over saturating the link
Apart from that, its a fairly simple protocol.
It absolutely does not measure packet loss at all. The only response to a dropped packet is to request a re-send.
The protocol is simple but depends on very smart sending rate control. Much of the complexity of the current version is to perform well when the receiver has a slow disk, or the link goes through a satellite.
Zstd compresses better than Zip, and unpacks many times faster. It can re-use a custom dictionary saved off from a previous compression, so you can get quick startup time for later fragments.
From your point of view was the implied statement something akin to: "In the context of a high-throughput network transfer program choosing DEFLATE over an algorithm more well optimized for computational weight wouldn't be a good idea." ?
It is little different from building with ARMs instead of Z80s. Where Z80s were once a good choice, now they are almost always a bad choice. Obsolete. USB sticks vs floppies, OLED vs fluorescent-backed LCD, valacyclovir vs acyclovir, Google vs Altavista, Git vs Subversion vs RCS. Need I go on?
The parent would probably be best off using one of the specialty genomics compression tools, that's why I linked to them.
Writeup on zip file Appnotes: https://entropymine.wordpress.com/2019/08/22/survey-of-zip-a...
Actual PKWare AppNotes: https://support.pkware.com/display/PKZIP/Application+Note+Ar...
Section 4.4.5 is "compression method"
Aspera was faster on a single stream, but if you have a lot of files to move around (you usually do) you can just multiplex a bunch of tcp streams to get the same throughput.
So like in the example on their page, if you have 20 wget's hitting their ftp and there are no other bottlenecks, the throughput will be similar to using Aspera - and you have just a greater variety of free tooling for tcp based those protocols...
Looking at the problem holistically - if you're going to spend money on Aspera to fix situations with high packet loss, maybe you're better off spending money on a higher quality connection and using tcp.
1. Measure the existing performance on specific data sets.
2. Understand how many of the bottlenecks come from the network latency vs. I/O limits vs. CPU bottlenecks (if using compression).
3. See if any domain-specific compression is needed.
4. Document the typical use cases. Understand why the current solution sucks (e.g. requires redundant user actions). Write down user interaction scenarios. Design the UI to be as efficient as possible for those scenarios.
This is a non-trivial amount of work that would require a lot of back-and-forth interaction and on-the-go requirement changes and I don't think it's entirely honest to ask someone to do this work for free in the name of cancer research (after all, you are not donating most of your paycheck to charities, are you?). If the existing solution by IBM sucks, how about making a Request For Proposal [0] and seeing if smaller software vendors could offer something better given that you are actually willing to pay for the work?
P.S. A student/hobbyist can probably whip out some sort of a parallel TCP-like thing with a large window for free, but you would get the same performance by just cranking up the TCP window size via sysctl and using multiple HTTP threads (htcat was suggested earlier in the comments).
Is it possible some legacy systems that produce and consume these massive files on-site would more sensibly run in the cloud, directly & selectively accessing the data chunks they need over fast backbone connections?
Also, is there any room for an rsync type approach, sending compressed deltas rather than naively sending huge files that may be redundant?
Not to say disintermediating a BigCO expensive patented vendor-locked-in MLPOS (Market Leading Piece Of Shit) doesn't sound exciting -- it does. I'm just curious and a bit skeptical that it is always necessary to mass-copy all this data over and over.
Also, you would still need to get the data into the cloud in the first place.
Beyond that, compression of sequences by encoding only the differences from a specific reference sequence somewhat ties you down to using a specific reference, since that reference sequence is required for decompression. This is inconvenient because the reference sequence for a species is updated over time, and if you want to use the new reference, you'll need to decompress with the old reference and then re-compress with the new one. Do you want to do that for all your data, every time the reference is updated? Reference-based compression also means the compressed files are no longer self-contained, which may not be acceptable for certain use cases.
None of these issues are fundamentally impossible to address, but the point is that it's not nearly as simple as it sounds, and anything fancy you do to try and compress a sequence file will generally impose some additional requirements on what the receiver needs to do in order to read the file. And many receivers are not very technically inclined researchers whose plates are already full of other things they need to be doing besides figuring out how to decompress your new unfamiliar sequence compression format,
https://en.wikipedia.org/wiki/Compression_of_Genomic_Sequenc...
Not yet. Famous last words?
I honestly don’t care about about their proprietary UDP protocol, it’s nothing special, just another way to copy bits onto a wire. Dime a dozen.
The true value of Aspera is they provide an integrated browser plug-in that lets the technically challenged reliably upload large files. If the transfer is interrupted or either side changes addresses it deals with it gracefully. It’s also a bridge to AWS S3.
I spent some time looking for a replacement and while there are numerous download managers that facilitate people downloading large files, I couldn’t find any upload managers with the same level of integration and polish.
About the closet thing I could find is Cyberduck, but it’s not integrated with the web browser, not as easy for technically challenged people to use, and there is no support (community support but seems really hit and miss). However it does make good use of Amazon’s multiple upload api and will happily fill whatever wire it’s connected to.
Torrent software has largely the same pros and cons Cyberduck does.
License costs are pretty brutal and it does need decent amount CPU to make the most of it.
Some use a family of different approaches to cope with different performance domains. Some use intermediate relays to cut the bandwidth-deat product by reducing delay.
But the big key is this: standard tools aren’t even trying to be good at this. They’re aimed at making an Internet work, not at optimizing and particular flow.
Basically it chunks up data and pushes it out as quickly as possible with UDP then retransmits whatever is needed. It works great in enterprise networks where you have high packet latency due to shitty firewalls and inspection appliances.
I haven’t used it in awhile, but I think they also have a proxy like service that lets you do managed file transfer directly from your datacenter network instead. It’s possible that I was using a complimentary product to do that.
Its only faster over large high latency links. Instead of using a single TCP stream to push data (with that exponential roll off) it has a TCP session to do accounting, and a bunch of UDP ports to stream the data.
for links around 16ms latency, its not that much faster.
For a 10 gig link with a latency of 100ms, it runs pretty much at wire speed. (assuming you have a machine that can handle it.)
But the alternative is good, too. Worth the price? Depends on non-technical factors.
Also, this senior dev never wrote unit tests.
On the other hand, maybe bbcp will work, and the senior dev just didn't know what he was doing. There was another senior dev who, seeing what was coming down the pike, left his job, and on his way out he was like, "yeah you can do it with bbcp, just you gotta tweak your tcp congestion rules on all the hops (which we could do, but is probably not an option for OP)"
I mean... based on your description it sure sounds like he didn't know what he was doing.
For real transfer with big (> 500TB) Aspera will deliver what it says over WAN (over the Atlantic).
If you have better connections and not so much data syncthing will probably work if you can have someone manually take care of all exceptions.[2]
If you know your data and you can build quite a lot of stuff yourself you can get almost the same speed as Aspera with UDT.
There is also a go implementation of UDTs used by kcptun[3] but I haven't tried that.
1 https://github.com/syncthing/syncthing
2 I don't intend to disrespect syncthing here, but when your dealing with TB scale 24/7 things are never really up all the time, the network is unreliable, the sender host filesystem is corrupt and the receiver host filesystem is full or broken or ...
But it looks like GridFTP and Globus are basically dead https://opensciencegrid.org/technology/policy/gridftp-gsi-mi...
Tsunami sounds good but maybe needs some updating? I haven’t looked closely...
It's not free for "managed" (Enterprise) applications, though, which is what the OP seems to be looking for, and I'm not sure it's the best choice for high-speed.
Something that I think is free is the CERN FTS [2], which can use a GridFTP back-end, so possibly you can roll your own high-speed big-data infrastructure that way.
And personally, I'd like to see more people getting a Globus subscription. It's cost-effective when compared to tools like Aspera, and helps fund the development of Globus software and features.
I mean, with some reverse engineering, mentioned later in the post, it’s really not.
It's being run out of UChicago, and I know they have at least one large company using them. Feel free to reach out to the Globus team, or shoot me an email!
Or anything with a >0.1% packet loss.
But, I think Aspera uses some specialized compression for genomics data, and genomics compression is not easy.
My suggestion would be to talk to the telco that provides your wan. If you are using a direct internet connection my suggestion would be : don't.
Aspera never did in-flight compression when I was there. In-flight compression got even less interesting around 2010 when the long fiber links got faster than customers' (shared) storage systems. Users have good reasons to compress at rest, before starting a transfer.
You have to buy access, but they do sell it.
Latency isn't the problem here - it's throughput.
Latency turns out to have a very great deal to do with how hard it is to get the nominal throughput you pay for.
They did not put any data in the neutrino channel, to my knowledge, but in principle they could have.
You might think you would need quite a deep chord to beat microwave propagation over the surface, because the dielectric coefficient of rock reduces the speed of light there. But that wouldn't affect neutrinos, which definitely go much faster through any kind of rock than light ever goes down a fiber. True fact.
That is not to say there will never be any.
Thanks for the info.
But you have to turn the neutrino flux off and on, or at least vary the intensity, to have something to measure.
You might have been able to get tens of bits per second with their apparatus. Or per hour.
Look up "superluminal neutrinos" and work your way past the scandal to the actual experiment..
[1]: https://en.wikipedia.org/wiki/UDP-based_Data_Transfer_Protoc...
Previously https://news.ycombinator.com/item?id=8381480 (2014)
Limited to work with CIFS / SMB / NFS and S3 compatible / OpenStack Swift compatible / Ceph / Azure storage.
Commercial model is server based with no additional bandwidth charges .
Data stored in a non proprietary fashion so bi-modal access to data possible.
Not sure if it is working with QUIC or not as it is mentioned in blog posts.
Seems heavily focused on the media and entertainment industry.
It's unfortunate how many researchers don't do such research...
A good setup here would be to pipe data to `tar`, then `zstd`, then `ssh`/`scp`, and do the reverse on the receiver side.
From an end-user standpoint, all this could be abstracted away if one could get `wget` (named in the article) to support zstd compression natively, which is possible since zstd is also a web compression standard.
Such compression support is, by the way, available in `wget2`.
Maybe we can talk about compiling a test to see how it compares?
I'm curious. How fast are these compared to netcat/socat/mbuffer?
https://medium.com/nknetwork/nkn-file-transfer-high-throughp...
The postal system or commercial courier services seem like better options than buying plane tickets..