Tell HN: AeroFS - File Syncing Without Servers
aerofs.posterous.com
aerofs.posterous.com
1) Does it work well with huge files? 1GB+ etc? Will a 1-byte change mean complete download to all devices?
2) Does it work well with 100k small files in deeply nested folders?
3) Will you charge for software and/or support?
4) What happens when one of the devices doesn't have enough storage? 4GB SSD laptop vs. 100GB HDD.
5) Will any of my computers have to be up 24/7?
2. It should
3. We're not 100% sure how we're going to charge for it yet, but anyone who signs up for the beta will be grandfathered into whatever system we end up using (i.e. you won't have to pay)
4. I'm actually going to address this in a separate post, but we've designed it in such a way where you'll actually stream files from one device to the other (based on a least-recently-used policy), so in effect you should have access to ~104GB of data.
5. This relates to #4, but is up to you, largely. If you have computers that are up 24/7, great. If you don't, you can leverage our cloud servers for better availability.
I have, say, 500GB of music. It doesn't fit on my laptop, so it "lives" on an external HD. Could it sync the laptop with the HD so that I always still have access to a set chunk—say, 10GB—of my most-recently-used music files? This is a capability I've been waiting for something to support ever since I bought the drive.
Will your program be open source?
Now, I am not asking you to implement this peer type (although that would rock my world if you did), but would it be possible for someone to implement it themselves? In other words will you be providing a 'peer API'?
This means you should get more like 10-20GB per 100GB you commit, otherwise the cloud simply will not have enough space.
Then if you consider that even with many nodes containing your data, there is a decent chance all of them go offline at a certain time. You have to have many, many nodes for the odds to be small enough. Which means the best solution is to use the storage you committed as one of the nodes, so it is always available to you. Then, it really transforms into a cloud backup system, rather than a cloud file system.
You are talking about RAID5. However, RAID5 is useless if more than a few disks go offline at the same time.
RAID1/10 is most useful when there's a higher chance of multiple disks failing at a time, or when the odds of multiple disks failing in your RAID5, while low, are unacceptable.
Of course there are other things at work when you talk RAID0/5/10, but this is a large part of it.
* not completely decentralized or open source. If wuala goes out of business, your data may not be recoverable.
* web interface is lacking (poor folder navigation/listing)
* doesn't work without X on linux
* even with X, not all features are available through the command line or API, though some are
* the interface that is provided is clunky
* the status messages leave me wondering where in the process an update is. If a piece of software can't reliably tell me where it is in a process, I can't trust that the process is happening the way I expect.I suppose Linux definitely isn't their main market and while the Web Interface is lacking at the moment they are working on an overhaul.
[1] http://eu.techcrunch.com//2009/03/19/wuala-merges-with-lacie...
Concerning it not being completely decentralized - I consider it a plus since that fact ensures greater reliability.
Although I have no good ideas where it would be beneficial apart from being fun.
That suspiciously reads like the AeroFS people get a copy of your key. If that's the case then it's only marginally more secure than DropBox. Hope I'm reading that wrong...
The overlay network layer presents to the data management layer a transport-agnostic view of the Internet, and addresses peers using network-independent identifiers. In this way, data management can talk to any peer regardless of network topologies and firewall restrictions, as if the world is flat :)
The data management layer controls data versioning and update propagation in a fully decentralized way. As I described in another comment, we use version-vector-like data structures to track versions and mange conflicts. We use modified epidemic algorithms (http://portal.acm.org/citation.cfm?id=41841) for fast update propagation. AeroFS distinguish between peers and super peers. Super peers can help update propagation and peer communication in many ways.
There are some features we're going to implement down the road that can be done better with p2p solutions though (aggregated storage across devices, for example), so I hope you give us a chance! :)
Could be quite good for the use cases they specify.
I have computers in multiple geographically-diverse locations and need a large amount data (terabytes) to always appear in each location. Other requirements:
1. Direct sync between my devices, with no third-party cloud involved
2. Fast local sync when two devices are on the same subnet are detected
3. Since individual files can be 20 GB in size, interrupted synchronizations should automatically resume when the connection is re-established (without having to start over from the beginning)
4. Encrypted transmission of data, but not encrypted on disk
5. When renaming or moving a file from one folder to another, the system should be smart enough to detect that there's no need to re-transmit that file (i.e., it just needs to rename it or move it to the new location on all other devices)
6. Ability to throttle upstream/downstream bandwidth on a per-device basis
Neither Dropbox nor CrashPlan -- nor any other tool -- has been able to meet all of these requirements.
In short, this is a very exciting and welcome development. I sincerely hope that this problem will soon be solved!
If you want to chat more about your particular needs/use case, give me a shout at yuri@aerofs.com, I'd love talk!
1) How does this system handle two devices behind separate NATs? (aka a work device and a home device.)
2) What is the conflict resolution protocol if a file is modified in two or more locations? (Newest wins, automatic duplication for manual resolution, etc.)
However, having looked at the wikipedia page on Version Vectors it appears that is a protocol for detecting conflicts. I was interested in how you resolve them.
A simple example is a zip file that I add file A to on one computer and later file B to on another computer. When I sync up do I end up with a zip containing no new files, file A, file B, both files or a corrupt zip file. (Does the answer change if the zip file is encrypted?)
We will formally describe meta conflict resolution in a separate post. Because resolution for data conflicts is very application specific, we will publish an API to allow application developers to write their own conflict resolvers. Meanwhile, we will try to provide resolvers for popular file types by default.
From the end user's view, in most cases conflicts are automatically resolved without being noticed. User intervention is required if automatic resolution fails or the user wants to manually merge.
I'd guess anyone who spent time figuring out how to do Windows filesystems would want to be paid for the trauma.
Thanks for the pointer and good luck with the project, I've signed up for an invite. :)
EDIT: Thanks guys.
Edit: This is a reference to the picture on their signup page.
Later on we may enable them based on use cases and user feedback. Our API will include ACL management as well. Currently files are read/write accessible once shared.
BTW, loved the reference to HN karma in the screenshot!