Zero Data App
0data.app
0data.app
1. Autonomous Data
2. remoteStorage
3. Solid Project
4. Unhosted
5. Fission
[1] https://noeldemartin.github.io/autonomous-data/introduction....
In practice, Solid is only the technology that enables control of your data; it doesn't provide the incentives. It's still up to apps to actually provide you with that control, and reasons for it doing so can include customer demand or regulatory pressure.
But Solid can't force apps to give you control - after all, they might just choose not to use Solid in the first place.
The solution, help, I'm looking for is to be able to easily integrate this into development - to give users control of running the apps with their data but also for the ability to give access/send the data to a centralized platform - or request/command for it be removed; obviously that decision will be if they trust the platform, the governance and leadership of it.
Although it's not Solid, an older project is https://remotestorage.io, which is a library that needs to be explicitly used by apps supporting RemoteStorage, but does allow users to store the data on Google Drive or Dropbox too, IIRC.
It would make great sense to have a non-hosted backend option, though, for power users who prefer to self-host their data.
Decoupling data storage from the application layer has some advantages, yes, but keeping the data in the clear server-side brings along many of the problems of the original:
- Data is still centralized in large repositories maintained by a company, and these repositories are still valuable and still open to attack, from both inside and outside of the company.
- There is a pinky promise between the data storage company and the consumer to treat their data properly and not to look at it or sell aggregated versions of it, but it is, at best, a pinky promise.
I agree that the open nature of the protocol making "self-hosted" an option is absolutely fantastic. But until that is accessible and easy for every single person to use, then only smart tech people will truly "own their data". That's really all I'm proposing: self-hosting being considered the default.
Imagine the device has a (non-encrypted) database, and the app runs locally and interacts with that database, like normal. Think localStorage if you are web-oriented.
The sync would be a separate background process (i.e. managed by the "Solid server" part) that handles encryption and decryption. As for how to manage a "circle" of devices that share the key without revealing it to the server: you can add a device via a key-exchange with the untrusted device asking a trusted device for the key. You can perform a key roll to remove a device. This can all be done automatically, though, where all the user sees is a control to add or remove a device. The hard part is key escrow (you throw all of your devices in a lake), by password protecting a copy of the key. Apple uses HSMs and Signal uses SGX to prevent brute-forcing this backup key.
I am a huge fan and even started looking at Svelte to contribute to https://nomie.app. It is going back to syncing with your own CouchDB instance or just locally. Right now it’s local or to some blockchain company. I don’t remember the exact details any more but it seemed like no one had access to the data.
With Devonthink for Apple ecosystem, I synced between LAN/bonjour only for a while before switching to iCloud sync this year.
1. Local data gets lost, so you need to sync for backup. But when your device gets lost or when you need to access it from elsewhere, how would you login to the said backup service? The most intuitive option is still a username and a password.
2. It's hard to transfer data from one device to another. And to asynchronously share data with friends and family. Solutions like ipfs exist, but in practice they rely on many centralized services (eg: for pinning). And besides, securely sharing private data is not an option on any of them.
3. Lay people being able to safely handle cryptographic keys is an unrealistic assumption.
So IMHO, perfect 0-data is not going to happen. It used to work earlier (pre-internet) because we had only one device, didn't have a lot of media, and could carry floppy disks for the rare occasion.
What I think would make a difference is: browser apps being able to do TCP/UDP (with safe-guards) and sandboxed local disk-access. This allows web apps to do things like SSH, and login to the vast storage infrastructure that already exists - such as GitHub. Data remains fully under your control and becomes portable. Git IMHO is a better path to where Solid wants to go. Especially once SHA256 commit hashes land - to become your personal, consensus-less blockchain.
Most people dont know what NAS is.
Why do you automatically assume that someone uses BitTorrent, for example? Your comment reminds me of the dark time in our past when people who used Visual Studio just assumed that everyone else used it, too.
Sadly the code quality is prototypical at best and while I still use the app multiple times per week, I never released it and have quite a few leanings which I want to implement (like rewriting the sync logic in a separate library, changing the server side data model to improve performance, completely reorganize the user interface and add more features to the app), but aren't coming to due to other priorities.
However, what I also learnt is that it is possible to build something like it, even if your use-case includes advanced topics like offline collaboration. My hope is that one day we will have a library that make these things for developers as easy as firebase and every app will have it out of the box.
a) what's a thumbdrive?
and
b) what's local data?
Each answer resulted in even more questions. Why isn't it in the cloud? What do you mean I'd have to copy it again? I can run out of room?
I could understand needing to get across to some people that not all data ends up synced (eg, saved memes are not necessarily automatically backed up to Drive)... but reaching a point to where people could be completely unaware of USB storage media actually shook me to my core.
People now expect it all to just "be there" after they log in.
I very much do not want to have to go back to that crap. Having it all "be there" is exactly how I want things to be.
I believe one of the reasons for phones nowadays that come with 64GB or less storage, is you can push this limitation with cloud backup.
Backups are handled (by the user) - the user can purchase cloud storage through a provider (like https://inrupt.com/products/enterprise-solid-server/, though there are other free services - or you can host your own wherever you want). Users don't have to safely handle cryptographic keys or deal with device syncing.
Otherwise I’d say it’s clear you’re in the space and I agree with your assessment.
I think being in this bubble of working in and thinking about tech 24/7 can often make us miss the point - these apps are just meant to be tools to do a task.
I have some vague plans around building a job portal just for startups looking to hire ex-startup founders where the data is not owned by my platform and freely available for everyone to leverage as per user preferences.
It's true that most users don't want to configure remote storage for each app. That's solvable though. There should be a way to store your remote storage preferences and just give each app permission to store there in one click.
I’ve been thinking a lot about the difference between a product and a project. I’m launching an open source project and we need to make it a bit of a product, but I’m hoping to maybe grow a following on patreon or open collective so we don’t have to focus on selling things directly.
Anyway this seems like a cool open source project that might not be trying to be a “product” for the masses.
Just imagine. You use Duolingo to learn a language and it shuts down. What do you do? What's the export format. Will the new product have the same levels / trees/ etc? Very unlikely. You'll have to start all over again. It's not about the raw data, it's about the data that your work has generated.
If you're doing machine learning, then there's the raw data, but there's also all the other data resulting from your hard work: training jobs, experiment tracking, parameters and artifacts and metrics, models.
The question becomes... First, how important is your product in the user's life? If it's important, they'll rely on it a lot and use it a lot to do things that are important for them, or for their organization.
Then, the service needs to provide export options and send you a shutdown notification. These are not always true but, assuming they are, you'll have to scramble to exfiltrate your data.
Then when you exfiltrate your data, it's an export. You'll have to find another product or service, and then somehow hope you can import that data to the new product, and then set up another workflow.
And here I'm talking about consumer products. For enterprise, there are many parameters including who the vendor is, procurement, team size. Meaning if you are Google/Amazon/Microsoft/etc, then a lot of people will treat you like a utility that will continue to be there, even with the track record of some of these of shutting down products and services (how many products has Google shut down over the years).
Look at this thread[0] I posted about our machine learning platform and the conversation that followed. The person objected that as long as it wasn't open source, they wouldn't use it. But digging a bit deeper, it turns out what they wanted wasn't really open source, but open source was one solution/implementation to their underlying problem. Trust you'll be there or control.
- [0]: https://old.reddit.com/r/MachineLearning/comments/kolobf/p_c...
It's true that the limited period of time between end of life and servers going offline may be a problem, but that tends to be a long period of time (3-6 months in general from the services I've used that have shutdown).
Valuable applications more often than not simplify complex things or do hard things behind the scenes. Their value is in the workflow/experience or the processing that takes place on the data, including APIs and integrations to unlock users' creativity.
>It's true that the limited period of time between end of life and servers going offline may be a problem, but that tends to be a long period of time (3-6 months in general from the services I've used that have shutdown).
Again, have you used these as an individual relying on them lightly, or heavily, or as a business/organization where your work relied on these? Have you used them as a user, or as the person who was involved in making the purchasing decision for the whole team/organization?
For business applications, the conversation may be different, but for the consumer ("average non-techie" as the original comment said) I believe what I said holds true.
A simple "backup/restore to/from <x>" service/computer often will suffice.
People could do a lot with their photos, videos, music and movie files that were stored in standardized formats on their hard disks.
Witness the plethora of music and movie players, the tools formed around images and video browsing and manipulation, etc.
Even non standardized stuff like word and excel documents could be shared, backed up, organized using these files.
Yet you couldn't collaborate on or generate a link to these files - sharing them with your family or a coworker most likely meant starting an email chain mailing the file back and forth with changes.
You are correct of course about apps that were able to interoperate on standard file formats, but I think most of those workflows were fairly complicated for typical users.
My mother has burned family photos to optical disk 20 years ago, we have physical photos that lasted generations. Tell me which app or cloud product I can trust with data for my grandchildren?
Which app will not dissapear, be discontinued, ban you or delete my data if I stop paying?
I uploaded my medical records to microsoft health, now it's all gone.
Paying with credit is convenient, voting in demagogues and populaists is convenient, blind patriotism is convenient, signing contract without reading is convenient, but they all have consequences.
"In May 2017, Google announced that Google Photos has over 500 million users, who upload over 1.2 billion photos every day."
Also many phones come woth google photoes preinstalled and enabled by default, I am counted towards that statistic because it uploaded a few things before I disabled it. Even if they really use it, I am betting most of them Also use store same photos as normal files.
I would like to see fresh numbers now that you have to pay for google photos.
As an aside, you may want to back up your mother's optical discs; apparently those are starting to fail.
The only way to make privacy happen is if it’s dead simple, works out of box, and not only doesn’t regresses your experience, but meaningfully enhances it. I’ve failed to see projects that would do that so far, and same here - it adds extra complexity for end user, with no tangible benefits to the experience.
However, I completely agree with the second paragraph.
Dramatic and exaggerated hyperboles only help to turn conversation into flame wars.
The removal of that feature certainly wasn't accidental.
That in the end the platform have to provide a better experience to the end user and also give more power/flexibility to the developer, otherwise no one will use it.
That's the reason i think it requires a lot more work to make this feasible. Only "iterational innovation" wont do it.
I hope that what i'm almost about to launch here, might reach those goals.
do you know about the safe network? that seems relatively easy. the way its being designed is that you will have to log-in to access the network (which does add complexity compared to now) but the idea will be that you can sign up for websites and services using your log in ID. so you dont have to share your name or email address and there will be no need for a service to store your password
Go explain the concept to someone unfamiliar with it with a nonsensical name like this.
All terms aren't coined exclusively to speak to developers. ;)
Are they the target audience?
I understand that, and in fact most developers do. However if they invent a new term, it should be something that makes sense to everybody (not just developers or large tech entities), or at least that doesn't require a complicated explanation to define it.
It's nothing new either, the trend for self hosting and offline first apps have been going on for a while. In my opinion they failed at finding a good term that encompasses all this.
Liberdata / Liberitadata?
- third party manages encryption keys and data custody, users manage none (dropbox, G drive et al)
- third party manages data custody, users manage encryption key (e2e encryption, icedrive, pcloud, 1password etc)
- third party manages none, user manages data custody (and eventually encryption) (0data and more generally "storing files in your computer")
the 0data is just like going back to what we used to do a decade or two ago and we all know it has its drawbacks (data can be lost, stolen, corrupted, difficult to move)
the most popular data custody model where a third party has total control over our data but we don't also has its drawbacks (data misuse, data breach, data mining, data transfer etc)
the second approach which i am surprised not a lot of providers adopt is where we delegate data custody to a third party but we still have e2e encryption over the data contents also has its drawbacks (data can be lost if keys are lost) but it's what i think it's more compelling compared to this 0data philosophy.
If you have terabytes of data to store, it can work out significantly cheaper.
1. Each user would have a "data pod" configured in their browser, storing has as much structured information about the user as the user wants (can be empty, or it can have all the structured data fields you want to insert).
2. The user can update any fields at any point and how access to the data pod is done.
3. The user can setup a BID MINIMUM or MARKET value for access to its data pod, perhaps even having different bid values for each set of data. For example, an advertiser wants to know my name? $0.000001 per request. You want to know my address and what TV shows I like? $0.001 per request. Want my bank data? $1000 per request.
Further this data could be authenticated cryptographically by certain authoritative entities. My government could authenticate that I am indeed from country A, and my pod's data would be signed by them (netflix and spotify could authenticate my media consuming history, etc). From that point onwards advertisers know that this field has been validated and can be incentivised to pay more. This should get rid of the incentives where everyone will self-report as being a US citizen just so their requests have larger bids.
What have I missed?
Depending on "data pod" implementation you could also have the "netflix.com" managed fields only be editable by a call from "netflix.com" API, which I then decide to approve for bidding or not and at which price, without me being able to directly edit those fields. Basically write-only from the vendor side to prove authenticity.
It is write-only from vendor side seems like vendor will sign something for authenticity. Something like token signature.
So it has to have my "pod ID", otherwise I can replay this data, with another "pod ID".
Ofc netflix or your pod, can rotate this ID, but that also requires netflix etc to constantly sign new IDs.
This is gonna happen sooner or later, but I really don't like it. Once most users have a cryptographically signed national ID on their PC, a lot more websites will require you to provide it. Sites like Netflix that region lock media will force you to show your ID just to sign up
Won't be long before companies get away with the invasive crap they tried and failed to do in the past: "Users Revolt Over Blizzard's Requirement Of 'Real Names' In Forum Comments": https://www.techdirt.com/articles/20100708/03054610123.shtml
I bet that'll come with an uptick of piracy then.
I imagine that as everything moves to “streaming first”, and old licenses expire, we ‘ll leave región locking behind. Video games are not región locked, even though prices are adjusted for different markets.
Movies and series should move more and more towards a global market. When I was a kid movies took months to arrive to theaters in my country (for various reasons, like reusing celluloid or making better predictions of performance)... now big blockbusters premier simultaneously everywhere because a good portion of marketing is global.
So this system will mostly be used by the people with little disposable income, enthusiasts who enjoy gaming the system, and scammers with tricks to maximize earnings. Those people are not going to spend much money on buying expensive products, so the advertisers will find out that this bidding thing is just not bringing any results.
I tried one of the projects listed there, and it required me to log in. ¯\_(ツ)_/¯
For example, it takes about 3 lines of JavaScript to store and retrieve a base64-encoded, stringified JSON object in the URL fragment (the part after the # – https://en.wikipedia.org/wiki/URI_fragment).
They want the app to work forever, even if the underlying cloud storage service gets killed.
Hence, being able to transparently transfer your data to another underlying storage service without the app changing at all is what most normal people want.
They want a box saying "This app's storage expires on 20th of April. Please choose where you want to transfer your data without disruption of service (and where it will be hosted from now on): Dropbox, Google Cloud, OneDrive, pDrive, Mega, ..."
I've spoken with a lot of regular folks. They don't care about privacy or self-hosting or decentralization at all. They never will. To them technology is a tool with which they make their lives easier. We should take a page from them. :)
(Although to make it 100% clear, I do care about privacy and anonymity very much; but it shouldn't be idolized and put on a pedestal, otherwise people like me will be made irrelevant with time -- one could argue this has already happened).
---
TL;DR: Sites strictly advocating for some tech principles lost most of my vote a while ago. We the techies get too distracted by our own shinies and must make a come-back to pragmatism and serving the regular people. It's OK to code stuff as hobbies but promoting them as universal values gets a "nope" from me. And that comes from a guy who wants to retire at his own house with 3 internet connections, and work on Tor-on-steroids and automatic replication of encrypted data until he dies.
I'm glad to see so many applications follow such a principle, though whenever I use something like it I remain scared of not doing backups right.
Login is not required by the way, You can choose "Device" or "Decide Later" when presented the save options.
The best term I've come up with so far is "data ownership". I like this, but the principles are broader than just data storage. Having control over your computation can also important.
Any suggestions for a better term?
> as a user I wouldn’t want to use the app anymore
I still have a working copy of Corel Draw 10 from like 2005.
https://github.com/whyboris/Video-Hub-App <-- MIT open source
The saved "hub" (file/database) is just a JSON formatted text file renamed to `.vha2` and all the images are stored as JPG in a folder near the file. Yay!
I pay iCloud and Dropbox monthly so I don't need to manage my own data. I know its secure with Apple and Dropbox.
To me privacy concerns and hypotheticals are less important than losing data outright to attackers or my own mistakes.
So people don't want 'xzy app' with privacy, they really want gmail with privacy.
Or is the second diagram just about being allowed to access your data whenever/however you like. Eg if SocialNetwork just made an API that let you access 100% of any data they have relevant to you?
But how do you attract users without these things already in place?
There needs to be a killer service that offers something better than Google Drive, Dropbox, etc can provide. They're unlikely to be able to compete on storage price, so they'll need to differentiate in some other way.
I'm bullish on Solid in the long run, but it's a tricky problem.