Dropbox Project Infinite
blogs.dropbox.com
blogs.dropbox.com
For example if you have a directory that is all stored in the cloud you can `cd` to it without any network delay, you can do `ls -lh` and see a list with real sizes without a delay (e.g., see that an ISO is 650 MB), and you can do `du -sh` and see that all the files are taking up zero space.
If you open a file in that directory, it will open, even from command line, then do `du -sh` and see that that file is now taking up space, while all the others in the directory are not.
You can right-click to pin files and directories to be stored locally, and right-click to send them back to the cloud so they don't take up space.
This is actually very different than traditional network file systems like SMB, NFS, WebDAV, and SSHFS. With a normal network file system over the WAN you would have major latency problems trying to `cd` and `ls` the remote file system. Most of them also don't have any ability to cache files locally when offline, or the ability to manually select which files are stored locally and which are remote. AFS does have some similar capabilities.
That seems like a lousy tradeoff.
In any event, if they were to index and provide search as a service as well, I wouldn't think it's something they do quietly. It would most likely include it's own huge marketing campaign.
Does not mean the files will get indexed, but there is no chance that Spotlight will trigger a unexpected terabyte download in the background.
I don't know a common search/find system that open()s or read()s files during the search by default. AFAIK Spotlight and Windows search are indexed searches. As for the indexing operations, I don't know how that is handled, they could disable indexing for remote files, or they could somehow integrate with indexing.
Based on my testing of a pre-released version of the feature (it isn't released yet), if you were to do something like `find ~/Dropbox -type f -exec md5 {} +`, it would download files.
As a user it did exactly what I expected. I was truly amazed. It was totally seamless and amazing.
Compared to the complexity of what has already been implemented, solving the problem of "I want to recursively open/read every file in my Dropbox, but I don't want it to download terabytes of data and fill by hard drive" seems fairly simple. For example there could be a setting for the maximum about of space Dropbox will use up, e.g., 40 GB, plus Dropbox could be smart enough to detect disk usage. If you `grep -R` it may download/open/read the files, once you reach 40 GB or near your disk capacity, Dropbox could start removing local copies of files that are not pinned to be local, i.e., remove the files that were downloaded because of the open()/read(), not the files you explicitly told it to keep local. I don't know how the team will choose to implement these features, but I'm confident that it will be well-thought-out and tested.
Remember, Dropbox is the company that especially monkey patched the Finder to get the sync icons (http://mjtsai.com/blog/2011/03/22/disabling-dropboxs-haxie/). They will go to great lengths for a seamless user experience, and do a ton of testing. I have no doubt that when Project Infinite is widely available it will be amazing, seamless, and have functionality many people thought wasn't possible or only dreamed existed.
... grep?
I confess, I did not watch the video, and only briefly skimmed the announcement.
Edit: I'm asking a genuine question of real technical interest here. How can this be implemented with no latency and real file sizes immediately available for inspection, while taking up no disk space? I went back and read the announcement again, and there are no hard details I can see that I missed in my initial skim. There has to be something stored locally, right? Hell, I'm running gigabit fiber here, and I still notice latency in the CLI for anything that requires a network connection. Perhaps I misunderstood the parent?
So I would see all my pictures list, if I decide to open the first one, it would take a few seconds to download, and then I start browsing the pictures, it would figure out I plan to look at all of them, and pre-fetch them from the cloud, so there would be on average no perceived latency.
They recently shut down their free plan.
That seems wrong to me. It would violate the assumptions of software that does stat() on directory entries and not only verifies presence but also non-zero size.
So it's risking buggy behavior to gain a latency edge over other networked filesystems. I think smart prefetching while preserving correctness would be better.
$ ls
$ truncate -s 1M foo
$ ls -lhp
total 0
-rw-r--r-- 1 catwell wheel 1.0M Apr 26 18:39 foo
$ du
0 .Is that still true for sshfs ?
People used to ask us if they should rsync to us directly or sshfs mount and then rsync to the mount, and we told them not to do that since the original rsync stat would basically download all files simply to look at them / size them.
But I don't think that's the case anymore. I think sshfs (or perhaps something about FUSE underneath) is smart about that now ... isn't it ?
Are you guys allowing full access to the machine now through LXC containers or some sort of VM?
No - it is the customer, on the client side, that creates an sshfs mount representing their rsync.net account.
It works very well and it is very nice to have a plain old mount point that represents your rsync.net account - especially since you can just browse right into your historical ZFS snapshots, etc.
But in the past, people did that and they got the bright idea to rsync to that local mount point, to do their backups, and that didn't work well.
But my understanding is that nowadays it would work better - you wouldn't download every single file that rsync simply stat'd or listed ...
We still don't recommend it, though. No reason to add that complexity.
Given a laptop with an SSD<1TB , I still want to be able to make use of the 1TB I'm paying for. Its currently possible by letting an entire folder upload to DB, and then unsyncing it. It stays on DB, but gets removed from the local drive.
I would love to be able to see those unsynced folders locally, with the same cloud icon announced here.
"A user who wants to back up and sync lots of media, and selectively offload it onto cloud storage to save local disk space" is probably not such a rare animal. And if I'm not feeling charitable, I would guess that this isn't explicitly supported because if it was, many more Dropbox users would start eating a lot more of their quota than they currently do.
Then I saw that it's not available on Linux. Which is really surprising, since supporting arbitrary filesystems is IMO dead-easy on Linux, and the easiest of the three desktop OSes. And since our company uses only Linux workstations (except for the designers), this is immediately unavailable to us. This is disheartening.
This is low level enough that it'd have to be implemented as a FUSE layer, which could add quite a bit of complexity. I wonder how they're going to deal with Linux. Will it still just sync everything?
And then they could have used the same codebase for Windows with Dokan (which has FUSE compatibility) and OSXFuse, although those two aren't installed by default. (FUSE is).
The mere use of FUSE doesn't make an fs complicated or things "severe". It's just a really simple API made available to userspace processes. In fact, you can build an fs in FUSE that does overlaying itself, using multiple, isolated processes. Or you can use a single, simple, small process to implement a simple fs. FUSE is literally just the glue that would allow something like dropboxd to send/recieve file info/data to/from the VFS.
Plus, when it comes to data/file sync tools like Dropboox, we need interoperability with existing tools: File Managers, filesystems etc. Not GUIs. Electron et al. are good for cross-platform GUIs, not CLIs or daemons.
So if Dropbox doesn't keep supporting Linux, I'll drop them like they are a hot potato. I do hope they won't make that mistake, because their continued Linux support has been the main sign that they care about me, their customer. And I don't want half-baked support either. I want it to be a first class citizen. Because I'm paying 14 EUR per month for my Dropbox account, amounting to 168 EURs per year, which isn't cheap at all. Yeah, yeah, it's the price of 5 coffees, but if I'd pay that price for everything I use, I'd be broke.
> Electron and nw.js
These would take care just for the UI, which for the purpose of doing file synchronization amounts next to nothing.
Until just recently, when I found odrive[1]. Not affiliated with them in any way, just a fan - if you don't want to wait until Dropbox decides to launch this publicly, check it out. It works with Dropbox/Google Drive/OneDrive/CreativeCloud/what-have-you and does exactly what Dropbox Infinite promises to do, except that it's seasoned, working, practically bug-free (that I can tell), and, well, available now.
For me, it's still more than just a feature addition though - I use odrive for all my cloud storage accounts. It's all under one roof. Your mileage may vary, of course.
Dropbox's response could be, "Hey we have training seminars for that".
There's even a little Explorer-native file overlay for such files (a black clock): https://blogs.msdn.microsoft.com/oldnewthing/20030827-00/?p=...
I wonder if Dropbox could just make use of this (e.g. by registering itself as an HSM backend provider to Windows somehow) rather than doing its own logic.
The problem with our current shared team folders is that you need to make a deliberate effort to share something with the team. That means when you want to pull information from the shared space, you're very likely to get an outdated copy. Thanks to 1.) no storage constraints and 2.) deep OS-integration, all files can actually be always up-to-date with Dropbox Infinite. That really sounds cool.
I don't think this solves the problem of taking forever to sync. Nothing has changed there. You are still bound by the same data retrieval and network latencies to get the file stored in a datacenter somewhere to you. That one is a harder problem to crack because it needs a lot of infra investment in expanding your content delivery footprint and replace SSD's with something much faster, like flash.
I already do this because I have to change my selective sync set quite often.
Yes and no. I have about 80 GB in my dropbox. This is no problem on my two workstations or my big laptop. Unfortunately my Macbook Air has a 128 GB drive. That means I keep having to exclude and include folders to be synced to my laptop or it will run out of space. That is a pain in the ass I would love to avoid.
Sure there's selective sync, but that's an all-or nothing approach.
I like the way Google approaches this with their Google Photos service. X most recent photos are stored locally, and everything else is pulled down on the fly as you search/browse for them.
Which app are you talking about? The desktop app is merely an uploader. It doesn't really sync photos.
it randomly indexed all my wallpapers and threw them all into my timeline... i'm still not finished with the resulting cleanup.
though it might have been Google Picasa desktop app, i used both back then
That's not what we're seeing. Her phone gets to be 100% full, and deleting a photo frees up space. When we delete a photo, we get dire warnings that "this will delete the photo on all of your devices, and in the cloud too." I download them to my computer before doing this, so nothing is really lost. But it's annoying.
Settings > Photos & Camera > Optimize iPhone Storage √This is why I was amused by the breathless articles last week about how to "free space on your iPhone" -- essentially by clearing the cache by golly!
(If the 16GB is filled up 100% and things fail, that's an actual bug though. You might really have filled it up with stuff that doesn't use iCloud, leaving too little for the iCloud apps to work with)
BTW all of the above is orthogonal to "optimize for space / use lower resolution on portable devices" that others have helpfully mentioned.
Theroretically this is nice, practically it will render som unpleasant surprises.
You take for granted that Dropbox hasn't considered these issues. Why?
Starting from https://tahoe-lafs.org/trac/tahoe-lafs/wiki/FAQ#Q23_FUSE with links to https://plus.sandbox.google.com/108313527900507320366/posts/... and http://lists.alioth.debian.org/pipermail/freedombox-discuss/... on the subject.
Linux is also a much larger target than Mac OS or Windows ... different filesystems, desktop environments etc. etc.
After all you can bypass the problem by running a Windows VM in VirtualBox with guest additions and a shared filesystem. Is another way to map the Windows Dropbox's network filesystem to the Linux one, whatever it is.
Still, very cumbersome and probably slow. I remember not stellar performances with sharing files between host and guest OSes in that way.
Hopefully Dropbox will release the API and somebody will write a user mode filesystem to interface this new service.
You're looking at the wrong end of it. Dropbox doesn't have to deal with any filesystem, Dropbox instead needs to build a filesystem.
It doesn't have to deal with any filesystem because Linux abstracts it. The same abstraction allows it to build its own filesystem really, really easily. How easy? I wrote, mounted, and used a filesystem in less than fifteen minutes. In node.js. It's that easy.
> desktop environments
Dropbox doesn't need to support various desktop environments, just a stable API for the desktop environments (actually, just file managers) to use. Isn't this how Nautilus, Thunar, Dolphin et al. already support Dropbox's existing features?
Not saying they're doing that here, I'm sure it'll take time to implement.
The problem is when that remaining 10% takes way longer than expected.
Looking at you Google Drive.
Especially since this is only part of their Business product for now.
That's it.
Just never give me the option of permanently erasing an entire folder of photos from every machine I've ever owned just because I mess up your wacky Selective Sync workflow.
I understand that there exist people who actually want all their machines to sync destructive changes, but I've never met one of them. More often, I see people lamenting that they deleted something important a year ago, thinking it was still backed up, but now it's gone.
So yeah, you're not a backup service, you're just sync. I get that. But please please be a little less syncy and a little more backupy and you'll fit the worldview of a lot more people.
It sounds like this is a step towards this (and four thousand steps past it), so at least we're getting close.
This seems vaguely astroturfy.
- on a Mac, if I press the spacebar to show the quick preview and start navigating down the list, will this trigger downloading of all the files?
- how are older files cached? Do we simply unselect devices to remove them locally? At least with Selective Sync, I can remove files eating up my hard drive (very important on devices with 128GB SSDs/HDs or less). It's much clearer in terms of what's on disk and how to manage it with Selective Sync. With Project Infinite, I'm wondering how one would remove files locally without fully deleting the file across the Dropbox folder for every other team member.
Keybase Filesystem [0] has been also doing this from the beginning.
>“Distributed” means you can access it from any device.
> “Filesystem” means that there is no sync model -- files stream in and out on demand. Among other things, that means that files on KBFS don’t take up space on your devices.
Plus the files are automatically signed, and end-to-end encrypted if in /keybase/private :D
I want to be able to tell when a command is going to block, or respond immediately with an error so that I can re-issue the command in a deferred manner. Having the command line lock-up because of filesystem latency is a bad user experience.
Yes, because everyone knows it would be amazing to have this kind of functionality just work for everyone, regardless of their personal OS and technology/device choices.
Bazil[1] has similar goals, but wasn't ready enough for me.
I do very much like sibling /u/pritambaral's suggestion of OwnCloud support. I tried OwnCloud briefly in the past, but ended up switching away. A feature like this might get me to try it again and stay.
[0]: https://github.com/burkemw3/syncthingfuse [1]: https://bazil.org/
A couple days ago I contributed https://github.com/DanielDent/git-annex-remote-rclone to git-annex, so in theory you could use Dropbox as a backend for git-annex if you wanted to (I haven't personally tried it with Dropbox).
[1] https://www.dropboxforum.com/hc/en-us/community/posts/201812...
- On the Mac I'm worried about using fuse in general. Stability, reliability, Mac OS X being a hostile environment as its not "blessed" by Apple.
- On Linux I'm more worried about the general nature of SSH as the basis for a filesystem. Which also applies to the running it on a Mac.
The reason I ask is because theoretically you could spin up a t2 micro with a 1tb EBS volume and forget the whole Dropbox thing, mount it wherever you please, whatever.
Obviously there are drawbacks, it's not local (fast), it doesn't have all the collaboration features of Dropbox, etc. but for relatively cold storage (colder than I need on a local ssd) that I want to access somewhat frequently, SSHfs seems like it could be fine.
I'm not sure how easy it would be to corrupt, etc and I'm not sure about the two things I mention above. Hoping someone here knows.
http://www.expandrive.com/apps/expandrive/
It works well, though there can be noticeable latency sometimes, even when saving very small files. It can also mount S3, FTP, Dropbox, etc.
https://www.eff.org/who-has-your-back-government-data-reques...
EFF has in no way verified the privacy policy is actually honored or that back-doors aren't installed in the companies' datacenters.
And that Dropbox was one of the companies in the whole PRISM scandal.
I still use it, mostly to transfer files between my devices, but there's no way I'm gonna be using it for anything sensitive. Everything sensitive goes to my organization's ownCloud installation which I completely trust since... well, I'm the one maintaining it.
As so many have stated before me, they better release this for everyone. It isn't a revolutionary feature, it's pretty basic, even if the behind the scenes technicals are complicated.
> all mounted file servers export the same file-system-like interface, regardless of the implementation behind them. Some might correspond to local file systems, some to remote file systems accessed over a network, some to instances of system servers running in user space (like the window system or an alternate network stack), and some to kernel interfaces. To users and client programs, all these cases look alike.
Next step: migrate 100% of the OS in the Cloud?
CODA [1] implemented parts of this as well.
Unfortunately, Dropbox has unencrypted access to all files and is too expensive, otherwise I'd be all over this. I'd rather allocate funds towards an open source implementation (I think something interesting would be done based on idea from CODA and NFSv4).
edit: working at the FS level according to https://news.ycombinator.com/item?id=11571640 - good job Dropbox! Anyone know if it's also supported on Linux?
(Now, if they can only implement a .dropboxignore feature...)
Because dropbox on phones (iOS atleast) has had exactly this forever, and it makes me wonder about the technical differences.
- I see all the files, and can open any of them at will.
- They don't take up space on my iphone unless cached
- I can select files that i want to be available offline
Dropbox has full control over which files it downloads and which files it needs to keep around.
On the desktop, the file manager and every single application (by virtue of their file open dialog) get access to any file at any time. You are supposed to be able to open any file in your Dropbox without going through a dedicated Dropbox app.
We've never had chance to push this though...
Actually, I think for older software it is fine for some reason. But in this case it looks strange.
The project is currently unmaintained, and we may open-source part of it.
How do you get a hold of a high level salesperson at Dropbox?
Would be great if, once a certain threshold is reached, the oldest downloaded files would revert seamlessly.
Im glad Dropbox is releasing this, I think it's very useful, but man is it trumped up. They have officially crossed chasm into the mainstream, when there is more marketing hype than technological innovation in a new release.
I consider Dropbox's feature a huge improvement over that.
This looks like an open source project that provides something similar for google drive: https://github.com/google/google-drive-shell-extension
On linux this could probably be even easier, just mounting a directory to a remote NFS server.
[1] https://msdn.microsoft.com/en-us/library/windows/desktop/aa3...
[2] https://en.wikipedia.org/wiki/Hierarchical_storage_managemen...
Also interesting thing to note is how everybody came down on keybase FS because the sync is on-demand. Most people citing dropbox and things like "if it was a viable option dropbox would do it and not have full sync". Well now... :D
Wish Dropbox had better Office integration. Currently, it just keeps making copies of a file if two people edit it.
Generally, cloud storage is done best by Dropbox. I'd love to be shown to be wrong.
Is there any particular reason you move it to Dropbox afterwards? (Or put another way, is there anything with Google Drive that you feel is missing, in your use case?)
But it was just an early one (1988!), many others followed. Coda, SFS, recently BackFS for FUSE, Sun's CacheFS and Red Hat's FS-Cache for NFS etc etc.
* Chunked uploads. Try to modify a 1GB file in OneDrive. The complete file will be resynchronized. Dropbox just uploads the chunks that have changed.
* LAN sync. If you are sharing files with family/colleagues that are on the same network, Dropbox shares file chunks peer to peer. This usually results in much faster synchronisation than downloading from the cloud.
* Dropbox has file requests. People can upload files to your Dropbox, but what others upload is not visible.
* In my experience, the Dropbox client is a lot more stable than the OneDrive client on OS X. OneDrive crashes or sometimes cannot sync certain files.
To me the speediness, file requests, and more reliability is well worth whatever the price delta is.
(Note: I have both, but I only use OneDrive as an endpoint for Arq backups.)
We'll soon have many TB SSDs for cheaper than traditional hard disk drives, so I'm not convinced this solves a real future consumer need. Maybe there are other uses beyond space saving that I'm missing.
cd dropBoxDir
grep -nr someStr .
https://torrentfreak.com/dropbox-scores-patent-for-peer-to-p...
My hybrid cloud storage team here at Microsoft in Silicon Valley (which was a startup called StorSimple that Microsoft acquired) already built this feature for our hybrid cloud storage virtual appliance that released a few months ago and its already in production and being used by customers.....
We called it Just in Time Item Level Restore but it's the exact same thing........and looks exactly the same in behavior as well.....and we have various optimization to make it efficient and deal with latency and stuff.....
(Btw we have nothing to do with one drive or one drive for business. That'a completely separate team. We only build enterprise storage devices)
Ugh this is why you want to be in the consumer space , not the enterprise space.
If a consumer product and enterprise product make the exact same thing/product, most people remember/see the consumer product.
The race conditions and clever data structures to represent the file blobs and file system structure in the cloud to do this efficiently is actually very cool......I'm not sure how they (Dropbox) implemented it but I wonder how similar it would be to how we did it.
Oh well this is my last week here in Microsoft and I'm moving to Apple in a different space so I guess I don't really care about storage anymore lol.
At least at Apple the cloud services and applications will be the backend for consumer used front end applications.
People will think Dropbox was the first one to do this because it's in consumer space so it's more well known......
Well actually I can link all the public announcements and stuff that was dated long before this video on the feature.
We should have made a public video like this. We have videos but they were all for internal demos....
Click the Preview link (skip to item level recovery section but the PM explanation doesn't do it justice so you won't see it as the same from a technical feature point of view but it's exactly the same and behaved the same zero storage on disk but the file system thinks it's there)
To accomplish all our feature we actually have low level stuff as well but I can't go into the details since I'm not allowed to so I'll leave it at that.
But Dropbox is a great company and I'm glad this feature is now available to Consumers not just enterprise.
https://azure.microsoft.com/en-us/blog/announcing-the-storsi...
https://azure.microsoft.com/en-us/blog/announcing-general-av...
(He reacted to my post saying he spent countless hours making videos so he was offended I say we didn't make videos. )
Ooops I didn't know. He should have linked them to hacker news.
(Start from the 3:00 where we already populated the share, and then watch till the end)
If it were as simple as you imply it is then surely we'd see this change made trivially to SMB, NFS, WebDAV, or SSHFS.
I agree with you that in hindsight this is an obvious feature, but they actually decided to implement it, and I bet it wasn't entirely trivial.
The biggest problem, as usual, is invalidating the local cache. For Dropbox, who own and implement the authority on the shared drive state, it is a _very_ (relatively speaking) simple problem.
Other protocols like SSHFS have to deal with filesystems of all kinds. Many of those do not support anything like inotify, and polling over huge directories would be horrible experience or performance wise (long delay or slowing down the whole host machine).
In hindsight? It's about time! I've been waiting for this for years...
I first had a good look at caching distributed filesystems (and tried to build one) in my undergraduate thesis in 1999.
Like SSHFS?
Edit:
More details here - https://news.ycombinator.com/item?id=11571640
But for me it's too late.
However, from what I recall of how NFS worked it didn't sync anything locally - if it was remote it stayed remote. This seems to be an interesting hybrid.
Plus all the other basic Dropbox features