Why Fred Wilson is wrong – files aren’t dead
blog.zamzar.com
blog.zamzar.com
I wished more people would see the endgame in situations like these: you're going to be paying through the nose for something that was already yours. Hosting your own data is trivial, the only case where I can see your data moving to the cloud is for backup purposes, off-site is better than on-site in case of disaster recovery.
So, have a good and long look at that firewall that protects you from the big bad cyber terrorists out there. That same firewall that stops the bad guys from coming in (and your ISP by blocking access to port 80 and a couple of others) are what keeps the peer-to-peer potential of the web from being realized.
So when VCs start trumpeting the 'end of files' make sure you realize what you're giving up.
You can buy MS office just the same as you can before, but consumers want to rent it. I balk at shelling out hundreds and hundreds for a fully featured copy of Office which will be out of date by the next release. Paying £50 a year or so for an always up to date copy (with a bunch of extras like Skype credit) is a much more attractive alternative.
> Rent this movie
Yeah, because when has that ever been a thing?
> Rent this song
And when will this ever be a thing (Renting a song is not the same as renting access to a library of thousands of songs).
Personally, I both make purchases and stream and I expect we'll continue to see a combination. And probably subscription plans that include some number of included downloads. (I think I've seen such plans in the past although I'm not aware of any major services offering this today.)
Subscriptions work better for software that people use day in and day out. The big issue for me is with things I use now and then and don't really care if I'm up to date or not. Or I need the software to access archived files in a particular format even if I'm not actively creating new ones any longer.
Same for movies and music. Yes, you've always been able to rent them. But in the future, you won't be able to own them forever.
What makes you say that?
I have to say that Microsoft is being very aggressive trying to get new customers by providing good services for a low price. I was accepted into their BizSpark program last year providing free Linux VPS and other web services for three years (while I am working on a new business idea). Also, there $100 per year per household service is really a good deal.
I'm curious about what tools/software you use to manage things.
Hosting your own data is trivial
No it's not. For a reason you've already recognized but glossed over. Backups. Even really tech savvy users have a hard time keeping backups going correctly. And not doing so means you have to worry about hardware failures and software errors accidentally wiping out your data. Not to mention the growing industry of somebody taking your data hostage.If you want to remove the incentive for people to move their data onto a Provider who manages the whole protection, backup, and availability side of their data for them you will have to provide a user friendly peer to peer solution for them somehow that doesn't require calling the tech-savvy {brother,sister,uncle,grandson,...} to set it up for them.
It's a testament as to how hard that is that no one has figured it out yet. Dropbox is probably the closest in that they keep the data on you box for you and allow you to pay directly for the service. But they still have a copy of the data that they can access if they should have a change in priorities.
Then they explicitely state that they store the data on their own servers. Is there another btsync?
EDIT:
Actually rereading it states they allocate space on their servers which sounds like they store data on their servers but I suppose it's also possible they might not.
A has files and gives B the key to the sync location. B enters it into btsync and chooses a folder. B downloads A's files over an encrypted channel directly from A's computer.
Add more people to get more to get multiple seeds.
At least, The internet was supposed to be peer-to-peer. It was for a while. Hosting data yourself - possible because your local computer was an equal peer on the network - worked wonderfully.
Sharing some photos with a friend was as easy as sending them a URL[1] to your local server. Yah, it slowed down your internet every time someone wanted to access the data, but that was fine because you only sent it to a couple friends. Better hosting was only needed if you wanted a larger audience or if you had some special requirements.
Unfortunately, a lot of people - either intentionally or from ignorance - ended up conflating "NAT" with "firewall" and generally popularized the idea that giving up your status as an equal peer on the internet was important "for security". This became the excuse to put off IPv6 until v4 addresses started to run low which took over as the reason to hide everybody behind a NAT.
A common response to this is that setting up a server is still technically complicated - the average user still won't understand how to setup Apache or how to configure DNS. Even with proper addresses, these technical issues will stop self-hosting from becoming popular! This used to be true, but... software gets easier to use over time. The only reason we haven't seen progress easier to use server software is.... who would run it? The people who understand NAT enough to bypass it don't need "user friendly" server software, and everybody that needs it cant access their server form the internet.
sigh - Walker's imprimature[2] was right. NAT was (or will be) what reestablishes barriers to publication and turns the internet back into cable tv.
[1] Traditionally, the "URL" was actually an IP address, and the file server was probably anonymous fTP.
This very scenario occurred in 2014 with Bitcasa, when they increased their prices from $99/yr for infinite storage to $999/yr for 10TB [1], and only gave 3 weeks to migrate files [2]. I don't think I've ever cancelled an account so quickly & with such anger.
Luckily I'd only stored ~100GB of backups so far on there, but lesson learned - the cloud is for sharing/collaboration & backups, not primary storage, always have local copies as well & always have contingencies so you can switch online providers immediately with little impact.
[1] http://techcrunch.com/2014/10/24/bitcasa-no-unlimited/
[2] http://blog.bitcasa.com/2014/10/23/important-we-are-upgradin...
I have a feeling we are getting backwards from the efficiency point of view and am thinking this won't be a sustainable way forward - even in distributed/parallel algorithms you try to keep locality to improve performance and save resources, not to transfer everything back and forth via some middle man just because some business guy came up with some "genial" idea how to milk money.
Internet companies have a few chances to earn money, subscription being one of them, but massive push into useless cloud for their particular business cases like in the case of Adobe CC or Microsoft Office just leaves bad aftertaste. Also, Dropbox/GDrive don't allow incremental update API calls for 3rd party apps, which wastes precious upload bandwidth, and in addition end to end encryption fully controlled by user is not provided which would justify full uploads. All of this just screams of artificial constraints which do not benefit anyone, in the long term not even to those companies.
How many people have this or a similar problem though? I would guess that most people deal with word documents, jpegs, and mp3's on a daily basis and rarely encounter much else.
Probably this is why YC funded companies such as http://www.wireover.com/
You need 4k+ camera for acceptable quality 1080p final result, 16k+ camera for good 4k, 32k+ for 8k, and this just brings to their knees the most powerful Intel processors or storage devices, not mentioning network... :-(
Fred Wilson is a smart guy, and knows that even iOS, Android, etc run on file systems.
I disagree files are dead too. But file systems is clearly not what he was talking about.
I think the core argument still holds though - which is that some people (Fred included ?) see files dying out as inevitable, whereas I'm not sure that's true (for the various reasons I cite).
So sure people want to download streaming videos. But chances are the result will be put into something that lets them view them by director, title, genre, newest, etc.
Generally it seems more natural and user friendly to have your photos, documents, music, and videos indexed. That way you don't have to remember the name, directory, or even what you computer you used to access it. You just look it up by when you used it last (i.e. I edited it yesteday) or by whatever metadata you remember.
So if a user's main method of access is by index of metadata not available in the filesytem, why even have a file system? A database seems more natural. After all why is /directory/fi lename more important than being able to look based on arbitrary metadata?
None of this implies that the underlying file systems aren't really useful in this situation. In fact, proper use of a file system would include separation of the media indexes from the media themselves, so that if you want to migrate to a better index, you can do so easily.
Nearly all real-world data has many-to-many relationships - for example, in music, we have several artists for the same song, and of course, artists perform multiple songs. The container formats we use for music are really shoddy and try to force these organic relationships into a limited set of tags, and then our filesystem doesn't even have knowledge of this information unless we add plugins for specific file formats - another messy area.
MusicBrainz and similar services organize music "the right way" in a RDBMS, and also includes support for things like multi-language tagging, which is sorely lacking in our filesystems. Instead of trying to come up with "one name to rule them all" for a piece of media, each object is just given an identity, a uuid, then all the metadata is related to the uuid.
If we wanted to keep metadata and media separate, the obvious solution here is to dump the media in blobs in the same database - give each piece of media a uuid, and create a new relationship table to link the uuid to its metadata. I'd personally like this solution for my music. To do the same for files, we'd have an explosion of symbolic links to map the relationships, and no tool to really navigate through them effectively because we're missing a query language. If we sat these indexes on the filesystem, we'd have a big loss of performance because of the extra layers of indirection and searching, for which no optimization is done.
A filesystem is really just a limited kind of database, but where we have a more powerful tool available, why not use it? Well, one reason is backward compatibility - our programs are written to look for files on filesystems, rather than streams from abstract sources. Perhaps what would be ideal in this kind of situation would be a FUSE layer which can expose the MusicBrainz database as a filesystem, but the underlying storage be the postgres db.
From my experience having moved from PC to Mac, I find iPhoto trying to hide my files and abstract things away a pain in the arse. On the old computer I had them in folders labelled 2012, 2013 etc. I now want to move some previous years to another disk to save memory and it's hard to figure that now. Previously I would just have dragged a folder. I guess my point is that fancy database structures may be hard for the human brain to conceptualise.
Fine, as long as that database is structured so I can arrange my own hierarchical structure of items, and manage them with something with the UI of a file manager.
I've yet to come across any of these "treat my collection of X as a database" tools that have been satisfactory enough for me to maintain that collection only through that tool. Not one.
That makes me doubt we're particularly close to doing away with more traditional databases.
In terms of my own data, I've moved more and more away from databases. E.g. my blog, my personal wiki etc. used to be in databases, but are now plain files because it's far less painful to work with.
Yes, I realise I'm not a regular user, but I've also seen enough "regular people" build deepl, complex hierarchies of files to realise that while some people may be satisfied with databases, many are not.
With those settings off the iTunes is just another media player, with the convenience of streaming to a couple AirPort Express devices.
But otherwise, yes, I agree, it makes a mess. I've forgotten to make sure these are off on a new machine once and a fresh install once, and rather than unmangling what iTunes does I just restored meda audio library from backup.
I always thought that the right answer would be a system based on tags, which, if you allow for the tagging of tags, is fairly hierarchical. As a metaphor, the tag is put on data that exists in an otherwise flat unorganized space, and you would place tags on the data after the fact (now you create folders, and "move" data into those folders).
Filing cabinets are awful analogies for file systems.
> Web directories failed for this reason, most people want to find stuff, not file stuff.
Completely different use-cases. Web directories "failed" because the web outgrew our ability to keep the data organized manually, and because once data exceeds a certain volume and you don't personally know how it is structured, search becomes far more efficient.
Most peoples data does not exceed the size where it is manageable, and the person doing the management is generally the person that does most of the lookups. The act of placing the data also serves to strengthen the memory of where to find it. This is also why I prefer spatial file managers - I remember locations of data by locality (and not just in files, but e.g. paragraphs in documents, or parts of source code too)
But the point is that file systems serve both needs, while databases, unless they "emulate" file systems, fail to serve the subset of users who does keep data organized.
> I always thought that the right answer would be a system based on tags, which, if you allow for the tagging of tags, is fairly hierarchical.
This, to me, makes no sense when you have a modern file system. A file system that supports symbolic links and extended attributes has far richer and more flexible semantics than every tagging system I've worked with.
From a UX perspective, most users don't care about the features of a modern file system. They never bother using symbolic links, the hierarchy only goes one level deep if that. It isn't about technology but about people, what works for them? Think like a UX designer and not a programmer to solve this problem.
I didn't think databases were the answer either. And I haven't seen a tagging system that is worthy of replacing a user facing file system yet (no one has allowed tagging of tags yet).
File cabinets are not infinitely hierarchical in any meaningful way (you can try to stuff files within files, but it breaks down so quickly in practice one doesn't do that, while that is exactly what makes file systems powerful), and does not have links. We learned in the early 80's how bad that analogy was after assorted attempts to take it literally (and make "file managers" that depicted actual visual filing cabinets).
> From a UX perspective, most users don't care about the features of a modern file system
Most users don't, and for those it doesn't matter that the applications operating on their data is layering its functionality on top of a file system.
Some users do, and for those of us who do, an application loses a lot of value if we lose the ability to control and structure our data hierarchically within a file system.
> Think like a UX designer and not a programmer to solve this problem.
The reason I am terrified of people taking data out of the filesystem is because I've seen what UX designers tries to do when they are given these tasks. It's not that it's "bad". It's that it satisfies only a subset of users, and that I'm pretty much never in that subset of users.
So no, we should not think like UX designers to solve this problem, because it is not a UX problem. It is an information organization and retrieval problem that the typical UI designer is in no way qualified to solve. A specific UI layered on top of this storage to provide a specific set of users a means of accessing that information is a UX problem.
But that only works if the data is stored in a way that allows those of us not served by those UIs to get at and operate on the data with other tools.
Subsidiarily, having control over one own data sets, means that you can process them with various tools, and therefore that those data sets are independent from the tools. (I'm pointing the finger at you, iOS and Android).
How this data is refered to, organized, and stored is irrelevant.
The classic notions of file, hierachical directory, file systems are very useful and practical, because they promote full data set ownership and processability by user controled programs.
Other kinds of systems may provide other notions (such as that of catalogs of objects, in capability based operating systems). But as long as the user has control over his own data, and can process it orthogonaly with his own programs, it doesn't matter much how it's stored.
The problem indeed is when corporations lure users into giving up control of their data and processing with convenience. Users should be better educated.
Oh, and developer should also try to provide convenience while preserving user ownership and control.
(Unfortunately, it's very hard to do on platforms like iOS and Android, notably thru their app stores).
I think that your concept of data ownership can be broken down into two related but independent concepts: access and control. Access means the ability to read; control means the ability to delete. Right now, if I upload content to almost any cloud provider I give access and control of my data to that provider: using G+ gives Google the ability to view every photo I upload, and to delete them all at a whim.
In a hypothetical cloud provider which enabled me to encrypt data locally, they would have no access to my data, but would still be able to delete it.
I would actually like the ability to grant control over a copy of my data without granting access to it; the two should be independent.
I've had a vague idea for some time of system which enabled me to store encrypted data and share the keys with my friends but not with the storage provider. There's been some interesting work in indexing encrypted data which could be pertinent to this, enabling a storage provider to offer search capability while still being unable to read the data itself.
I guess it turns out to have been cheaper to build their own almost-as-good Dropbox replacement than to partner with them.
I think that devices with a relatively small amount of SSD storage contribute to cloud and web service providers getting more control of people's data. Selective sync in services like OneCloud, Google Drive, and Dropbox allow users to just keep local copies of files they need in the near future. At least I do this.