Database-less torrent website
boredcaveman.xyz
boredcaveman.xyz
That aside, would this method scale? 135,000 torrents doesn't seem comprehensive, so I would expect real world use to have many more. Maybe a different SQLite db for different categories?
(My guess would be yes as I believe it breaks up the file for distribution, but have no idea if it’s exposed in the IPFS.js api)
EDIT:
After a quick scan of the docs, I think you can do this (but I certainly don't know enough). With at least the "The Mutable Files API" which "is a virtual file system on top of IPFS that exposes a Unix like API", you can provide an `offset` and `length` to `ipfs.files.read(path, [options])`. I don't know if that then translates to only downloading that part of the file from IPFS or not.
https://github.com/ipfs/js-ipfs/blob/master/docs/core-api/FI...
EDIT2:
In fact the `ipfs.cat` api that the OP is using takes `offset` and `length` parameters too.
https://github.com/ipfs/js-ipfs/blob/master/docs/core-api/FI...
It turns out that project was inspired by a few others [1], two of which [2][3] implement a vfs for SQLite on top of bittorrent doing exactly this suggestion, but with a bittorrent hosted file rather than IPFS. They only download the parts of the SQLite files needed when querying.
0: https://github.com/phiresky/sql.js-httpvfs
1: https://github.com/phiresky/sql.js-httpvfs#inspiration
Seems like a small price to pay.
I can see new stuff living in torrents that then get committed to a sqlite database that is then deployed every now and then.
Isn't this just a log (or a stream like Kafka or Kinesis)? In fact you might even be able say every database already has this ;) (binlog, oplog, etc)
> not the whole db
If rows are immutable, what part of the db is left not being immutable? If rows were immutable, doesn't that imply that any existing "ranges, searches, indexes, etc" are immutable too?
If you going to the effort of making a decentralized database, why not also decompose all of these parts from one another... no reason the tables (logs), indexes, search, etc have to live in the same database, they could be spun off as completely different parts. Basically a something that indexes logs and then something else that takes those indexes and makes the searchable.
In fact this is all the rage right now with centralized databases... all of the work being done with streaming systems just seems to be an effort of decomposing and inverting databases... every old is new again.
Anyways, I agree with the idea and don't really know enough about decentralized systems to really understand why such a distributed database can't or hasn't already been built.
The torrent is merkle-tree based, and for each change in the SQLite memory page it also updates the merkle tree giving you a different torrent infohash at the end.
So as databases are published together with the applications the initial torrent info is the same, so theres a swarm of initial peers. Once it gets updated and the torrent "turns into another", the solution is to use the RPC api (that my solution provides) over the initial swarm of peers and use some sort of distributed algorithm like raft or gossip to define how to deal with the subsequent changes over the peers.
The important part is that all other changed torrents from all the peers are available to all peers in the network, and with a RPC working as a abstract interface to organize them, it will work for a lot of different scenarios.
Also, how do peers find each other? And finally, do you intend to open source it?
No, thanks for pointing that out, will take a look into it.
> Also, how do peers find each other?
The idea is to use the torrents as a common shareable resource where given the peers have the same interest in that torrent, lets say the torrents works as a "meta-database" with just enough immutable metadata info, giving the developers of that application a list of peers that will have the same RPC service you designed, with a interface that suits your purpose according to your application goals (a distributed Youtube for instance), and giving you know and designed that API yourself whatever you want from each node, lets say a piece of data, be it a file or a database key-value range, you can ask your API for it, combining the torrent peers and whatever distribution combo you need.
> And finally, do you intend to open source it?
I've just did
https://github.com/mumba-org/mumba
It's badly documented given i'm on the final touches before a proper launch, but the "documentation" of the storage layer is planned for today, giving how important it is for the whole thing.
The first thing i've tried to get right was the storage layer, and the capacity to use mutable sqlite databases over torrent (together with files which are simpler given they are meant to be immutable).
Given theres no proper doc yet, i can point out to the source at
https://github.com/mumba-org/mumba/tree/main/lib/storage
where:
https://github.com/mumba-org/mumba/tree/main/lib/storage/bac...
is a modified chrome cache storage layer (which is the real underlying disk storage) that abstract the files and databases storages that from the bit-torrent layer perspective are on the disk.
https://github.com/mumba-org/mumba/blob/main/lib/storage/sto...
is the front end
and the:
https://github.com/mumba-org/mumba/blob/main/lib/storage/tor...
is the main abstraction that may be a "fileset"(collection of files as in torrent) or a dataset/database, which in this case you can get the sqlite db handle from the torrent object and deal with it as normal database.
I've take care to enable a key-value store over the sqlite btree, so both form of databases are possible, a key-value and a normal SQL database.
Key-values are important for the distributed case where you may want the nodes to have partial data and abstract a SQL layer (or whatever) on top of the distributed nodes, which is a better solution for distributed storage.
For the database distribution, a 64k SQLite memory page maps to a torrent piece of the same size, which can be synchronized over other peers that knows what's the root of the merkle tree of the given database is (you can use the RPC layer to coordinate this or use the bit-torrent DHT updating your slot with the new merkle root)
BTW This is what i'm using to distribute the applications, the DHT which points to a "database torrent" which in turn is a index to other files and database torrents. The application owner have write permissions over that particular DHT slot and can change the root merkle anytime the application itself changes (a new version, or some of the assets).
What im finishing right now is exactly the higher level layers that automate this whole thing and make it work "under the hood" without the users or applications distributors need to understand how it is implemented.
Of course the developers will have access to iterating over the peers of a giving torrent to group them over RPC interfaces, and the idea is also to give access to the DHT so solutions that need a access to mutable ever-changing merkle root is also possible (even though the torrent based solution cover most if not all of the ground), but the idea is to give the devs the tools to be creative about their solutions.
(The solution is a whole browser-based UI/web applications platform, that use torrents and the torrent DHT for the p2p storage layer)
Counter to Moxie's argument[0], I think a lot more people would be willing to host a server, if all they had to do to host a server is to seed a torrent. We already know many people are willing to do that.
I think this would be interesting if the application itself was just a Docker container that could output to a browser, similar to many other local-hosted approaches[1].
Thanks, its nice to hear that after all those years working on it and keep going by understanding how important is for all of us to get out of this rigged game where the cloud computer(+ web clients rulers) FAANG's are controlling where we are heading and a future where we(hackers) have less options technologically, not having enough freedom to program services that can ignore the mothership and create the same sort of services only relying on peers, is a goal being pursuit by them.
> I think this would be interesting if the application itself was just a Docker container that could output to a browser, similar to many other local-hosted approaches[1].
What my solution do is that the application have a "service" process that works like daemon to the application, being able to resolve route requests both locally (when over IPC) and remotely (when over RPC).
The routes are meant to serve html or any other content to the UI application (which is also another process akin to the Chrome renderer).
But the fact that every application might have a ever-running process, means that in some cases it can create other kind of services. For instance:
You can create a PostgreSQL wrapper application where you present the the database SQL api over RPC as a service, and manage the PostgreSQL instance with your service process.
The same works for Docker or QEMU for instance. Your application can be docker-based and can be deployed only on Linux.
Your applications (the service and the UI process) are natives by the way, and even the UI application controls the Web layer (rendering, layout, etc) nativelly talking directly to it (as a first SDK enabled over Swift for now).
This architecture will allow you to do those things, and present a way for the users to interact with your solution in a UX that is packed and develop together with the "backend".
With this hole thing you program the frontend and the backend as a whole, in the same programming language.
https://flyingzumwalt.github.io/ipfs-tutorials/curriculum/fi...
[0] https://en.wikipedia.org/wiki/Distributed_hash_table
[1] https://yggdrasil-network.github.io/2018/07/17/world-tree.ht...
So what I propose is this: There is a central source of truth hosted on IPFS via IPNS or mutable torrents. The data format can be whatever is more efficient to update and store, not easiest to read and query. (An event log is fine) Then, there is a reader application, like an RSS reader. The reader application can be a desktop application or it can be a centrally hosted web app that keeps a conventional SQL database behind the scenes. The idea is that there will be multiple reader providers like there are competing email providers.
Select users who have the decryption key (communicated out-of-band) might be willing to pin the encrypted data for the benefit of others who also have the key. There are also for-pay pinning services who probably wouldn't care whether the data was encrypted. And whether it's pinned it or not, as long as it's downloaded the data remains available to be shared (at least until it's GC'd).
> Uncaught ReferenceError: passQueryToResultpage is not defined
Yep.
"The underlying mistake is thinking about this like a game where you can make up rules for the government to follow." - https://news.ycombinator.com/item?id=29913036
Also the assumptions here kind of remind me of domain fronting: https://www.zdnet.com/article/amazons-aws-latest-to-give-up-.... Basically, it's assuming IPFS will protect it from the authorities, but what might actually happen is it makes IPFS a target of the authorities. Now that probably won't happen with this because the authorities don't actually care that much about torrenting/piracy, but the error is still there.
> The "rules for the government to follow" are the laws, and these aren't being made up, but rather pre-exist (safe harbour laws) with some pretty strong backers (goog et al)
That's true, but to elaborate on that quote a little bit: you can still make the mistake the it describes by making up your own interpretations of "the laws," rather than making up rules out of whole cloth. That's extremely common (especially with Constitutional law), and probably the most frequent way of making that mistake.
- I can't because I am not hosting it
- But you are displaying it on your website, take your website down or filter what your website displays to avoid showing copyrighted material
And that's the end of it. It doesn't matter where the data is stored. It does matter, a lot, where the data is displayed. If your website displays illegal content, you're facilitating access so you're breaching the law. And, honestly, I'm not sure there's an easy solution to that other than using onion or I2P sites which are only marginally more safe against copyright takedown requests because, due to their very nature, they are obscure and difficult to access.
I was curious so I compressed that torrent db with a few different methods:
11.1MB 11116544B dump.sqlite
10.2MB 10155419B dump.csv
6.6MB 6573399B dump.sqlite.gz
6.6MB 6565771B dump.zip
5.6MB 5616842B dump.rar
gzip is certainly suitable to be used in this situation, I stand corrected.The 10MB estimated size came from [100 bytes per row] * [100k rows].
50 of the bytes per row were "description", which should compress well (2-3x, I'd guess).
40 bytes per row were the IPFS ID/hash, IIUC. I assumed this is like a Git hash, 40 hex chars, which is really just 20 bytes of entropy.
He also estimated 14 bytes for the size (stored as a string representation of a decimal integer, up to 1e15 - 1, or 1PB?). That's about 50 bits or 6-7 bytes, as a binary integer. Sizes wouldn't be uniformly distributed though so it would compress to even fewer bytes.
So if SQLite was smart (or one gzips the whole db file, like you did), it makes sense that a factor of 2 or so is reclaimable.
I'm pretty sure it was posted on hw, but I can't find the link anymore.
On that page I do inspect, got console and do:
loadDBAndExec('SELECT * FROM my_table LIMIT 10;')
I get
Uncaught ReferenceError: loadDBAndExec is not defined
Why doesn't this work?
The live demo uses a webworker with different code: https://boredcaveman.xyz/demo/megacat/database-app.js