The AeroFS Private Cloud
blog.aerofs.com
blog.aerofs.com
This filecloud is interesting, thanks. Any other product that you are aware of? I thought I had found them all, but you never know :)
Being able to work on a totally disconnected network is different. Maybe if you could support Eucalyptus's S3 or another backing store like that it would be equivalent (and that might be interesting).
I'm sure plenty of people are fine with an Internet-connected, AWS S3 backed system, though, especially with additional crypto.
I'm building a FreeNAS server with 6x3TB hard drives in a RAID-Z2 config. My goal is to allow my Mac to use it for Time Machine backups, but to also use AeroFS as my file sync mechanism when I'm both in the office and on the road.
Hopefully it works out smoothly. I'll have to figure out how to access the machine from behind my router, and I'll have to determine how to get it to automatically back up to S3 + Glacier. I think there's going to be a lot of details that I'll have to research here.
Note that this is completely unreliable. Time Machine isn't built to work over the network. The network support is badly hacked together and will destroy your backups on network faults. They basically treat the network volume as a block device that gets corrupted when the connection fails. Using it over wifi makes this painfully obvious.
Deleted comment
I ran into more issues than most because the particular AP I have at home doesn't work well with the macbook air I was trying to backup and resulted in a lot of disconnects. So it was a bit of a stress test but the issue is real and well documented. You can find instructions all over the web on how to fsck sparsebundles to try and get back your backups[1].
This is particularly annoying because Time Machine is supposed to be a user friendly version of versioned backups like you can do with rsync+hardlinks. I have a script to do that for me on my personal linux machine and would love to be able to deploy TimeMachine+AFP for the rest of the family. Unfortunately apple decided it really needed to have hard-linked directories for just this case and so decided that the way to do that over the network was to have a single binary blob with a HFS image on the server instead of the proper files in the filesystem (filesystems don't usually support hard linked directories, HFS was hacked to do it). That way you not only lose direct access to the individual files on the server you also lose the whole backups when there's disk corruption...
I should probably have a look again to see if iSCSI is a viable solution to this.
[1] http://www.garth.org/archives/2011,08,27,169,fix-time-machin...
For backup you can use one of many solutions to archive to S3. I wouldn't recommend Time Machine.
(sigh, it seems the tech community has become blind to anything that isn't heavily blogvertised.)
There's so much released and so much noise, it's hard to keep track of all the good things that come out - especially if they're not mentioned within their tribe of friends/followed ppl.
I don't know how to get round that - the web is huge and no one can keep track (or have a network that touches) all of the open source releases.
I won't know until I try it though, I could just as easily change my mind and use BTSync.
I have a 11x4T raidz2, and will be migrating to 3x4x4T raid10. My last array was 2x4x1.5T raidz1, 3 drive failures in its 5 year lifetime, but not concurrent so no data lost.
I use CrashPlan to get data from one place to another; I back up NAS to CrashPlan, PCs to CrashPlan, and PCs to NAS. If you don't use CrashPlan's app to back up to their cloud, it doesn't stop you from using it to back up to other machines you own. And it works on Windows, Mac, Linux, Solaris, and with a bit of hacking, FreeBSD.
Are you using Crashplan+ or any of the Pro/Enterprise plans? What sort of transfer rates do you get? And were you able to complete a full backup of that 11x4T array?
I do see CrashPlan updating at a leisurely rate, but that's OK with me, as my upload bandwidth is limited. I'd be interested to see if it completely chokes past a certain point, like you say it does for you. But it's not yet at the point where I need to worry about it.
If a drive develops a speed problem, I'll replace it.
Suppose that once one drive has failed, the chance of another drive failing (at least an URE, but potentially worse) while reading it entirely (as necessary during a scrub / resilver) is 10%. I don't think this is very implausible, as RAID HDD failures are often correlated, and not only because it's inconvenient to distribute HDD purchases.
The chance of successfully recovering in a mirrored scenario is 90%. You only need to read one drive.
In raidz1 with 6 drives, your probability is 0.9^5 - you need to get lucky 5 times, because you need to read every other drive fully to recover. That brings you down to less than 60%. Raidz1 with 5 drives is 0.9^4 - a bit over 65%.
Raidz2 for 6 drives combines these two stats. The probability of one failing disk while reading is 1-0.9^5. But you can recover from that. And the probability of that recovery is 0.9^4. So your consecutive probability of failures is (1-0.9^5)(1-0.9^4). Subtract that from 1 and you get your probability of getting lucky.
Work out the numbers, and you'll find it's only 86%. It's less than mirroring.
There are a couple of simplifications in the maths above (don't forget, a second drive failure means another drive to resilver, and a restart to the procedure), and a big assumption based on reliability %. But frankly, unless you think HDD failure is already fairly rare - and it's not - mirroring is usually ahead on safety, and the lead only increases with number of disks in the array.
I used to think like you do, that I'd feel more safe if a single disk failed, knowing I could still tolerate another failure. That's why my array is currently raidz2. But I worked out the numbers, and it changed my mind. And that's why my next array is going to be raid10.
Not to mention that it's hard to get much more than 150M/sec for any realistic I/O with raidz2. I have over 1G/sec of potential HDD bandwidth. Unless I'm doing a scrub, I can't touch the potential performance. And of course all disks act like they're on the same spindle, there is zero random I/O advantage for raidz.
The product has come a long way in the past year. In particular, our Private Cloud offering is deployed at a number of large organizations, all of which are quite happy.
Regarding comparison with BT Sync -- I think their product is great, but we try to go beyond simply syncing files between devices by allowing for a lot more administrative control to the IT organization while still giving the users a really simple Dropbox-like syncing experience. Things like remote wipe, version management, conflict resolution, and so on.
But YC did not directly help us land these customers in the forms of introductions. For the most part, they've come to us through word of mouth and internal employee references.
However, if you want tight admin control, you may have some issues. For example: If an employee leaves, we're not sure how to revoke privileges for them once they have access to the folder.
So, I recommend very small companies use BTSync since it's easy and free. But if you're a larger company, you have probably outgrown BTSync and should get something like AeroFS.
The challenge is that sales cycles are long and potentially high touch. One way to mitigate that is by getting sales distribution through 3rd parties. But avoid integrating with a bunch of 3rd party storage platforms unless you get commitments for leads from the vendors. In other words view integration efforts as an engineering to sales arbitrage.
I guess on the plus-side, you get phone support, if that is what you need.
Owncloud is open-source and really quite powerful, I am surprised it was not mentioned more in these comments.
I am curious, what other sales mechanisms and/or software packages have you tried before you settled on private cloud w/out touching your servers idea?
This may be counter-futuristic, but what if you sold them servers, along with your software? What if you gave your clients the best in classes storage, coupled with the best way to manage it? I suspect you've thought about it before and I want to know what the reaction was like.
Enterprise clients are a black box for me, so anything else you share would be interesting to know.
I've been advising my customers to look at it for quite some time now :)
The (optional) AeroFS Team Server is what would store file data in the company if you wanted to, but many of our customers actually just end up using the direct peer-to-peer syncing without a team server.