Willow Protocol
willowprotocol.org
willowprotocol.org
> This makes Willow a higher-order protocol: you supply a set of specific choices for its parameters, and in return you get a concrete protocol that you can then use. If different systems instantiate Willow with non-equal parameters, the results will not be interoperable, even though both systems use Willow.
Help me out here - isn't the point of a protocol that two independently developed systems don't have to agree on how to implement the protocol? What value does Willow have if two systems that both purport to be "Willow-compatible" aren't compatible with each another?
This is sort of how ActivityPub is a thing, but it underpins multiple, sorta-but-not-really interoperable systems like Lemmy and Mastodon.
"higher order" is some nonsense, and would make me shy away from using it...
So, perhaps it is strong language, but I think it is a reasonable reaction.
A protocol has the property that it is implementation independent, but that it has a defined interface (i.e. it is immediately usable).
This is neither (not a defined interface, implementation dependent). If it doesn't share either property with a protocol, then you can't claim that it is truly a protocol, "higher order" or otherwise.
This confused verbiage is what should be cause for concern - note that I can claim a "higher order" protocol with JSON or gRPC - its all the basic building blocks for a protocol, just both sides need to implement the same stuff!
Except, neither JSON nor gRPC are crazy enough to claim to be a "higher order" protocol, which to me puts this in the rubbish bin of over-complicated technologies looking for a problem, like SOAP, JavaBeans, OSGi - all of these could also be claimed to be "higher order" protocols as well.
The term is meaningless, and so I assume, is this project.
This part looks useful.
But is it useful to be a "willow" family of protocols? Probably not.
Their claims on the front page are extraordinary. Extraordinary claims require extraordinary evidence, and heading the page with nonsense is not a good start.
That you can claim to support Willow, or be Willow-compatible, without actually having to interoperate with your competitors. See e.g. the usage history of x509 in the nineties.
You can make two data structures with a const parameter.
If the parameter is not the same, they are not compatible (not the same type). The parameter can be tuned according to the specific needs of the application.
For example, SSH only works if both sides support and can agree on the same cryptographic algorithms, which is something that SSH is parametrized over.
In that sense it is even more high level and abstract than parameterized protocols like SSH.
GET / HTTP/1.1
Host: example.com
Could have the same semantics in any format: {
"method": "GET",
"version": "1.1",
"headers": [{"name": "Host", "value": "example.com"}]
}
Of course a server accepting the former won't be able to communicate with the latter, but that's just an implementation detail that Willow does not want to commit to at this stage, and does not make it any less complete. Just a bit impractical.True, but the industry has been moving away from "you have to configure both ends" towards autonegotiation.
e.g. RS-232 you had to configure data rates (e.g. 9600 baud) and encoding (e.g. 8N1 - 8 data bits, no parity bit, 1 stop bit), and if you didn't have the same config at both ends, communication wouldn't happen. USB, the primary successor, determines stuff like that by negotiation. Or similarly, with Ethernet autonegotiation, you no longer have to worry about manually configuring speed/duplex on each end–which I can remember being a big drama 20 years ago.
Probably not the best example, I just happened to have the PTP spec open at the time!
Say that I want to implement something Figma-like for designing drug-runner operations, Willow seems to be an excellent building block (Yes, the example is kinda out-there but it's meant to indicate that genericity intended here).
As far as I can tell, this is primarily a crypographic specification, like Noise (http://www.noiseprotocol.org) except for a stateful key-value store instead of stateless connections
More practically/less absurdly: is SSL useless because side A doesn't support the same ciphers as side B.
Perhaps a Willow negotiation protocol would be needed to reconcile, but it's a bad idea from a security perspective because it enables downgrade attacks.
A protocol is allowed to have presumptions, and then provide an interface.
Its not allowed to have no interface at all and no presumptions (A void).
As a construction kit, it has value for people who want to make protocols where they'll control both ends, but don't have to re-implement basic table stakes.
According to them, data on IPFS is immutable, stateless, and globally-namespaced, whereas data on Willow is mutable, stateful, and conditionally-namespaced. I interpret Willow as an authenticated, permissioned, content-based, globally-addressed, distributed database system, where an address has the hierarchy and expressiveness of a URL.
One particularly nice feature about the documentation: if you hover over an underlined word (https://willowprotocol.org/specs/data-model/index.html#data_...), a pop-up box provides a definition or explanation. Importantly, some terms in the pop-up are underlined themselves, so you can dig down into the terminology with great ease. More projects should implement this functionality.
_Surely_, they didn't write this from scratch, did they? _Surely_, there is a tool that they used that I can use, too. Right?
I personally found IPFS very disappointing in practice, so I'm very hopeful for a successor.
(The promise of IPFS is great, but it is excruciatingly slow, clunky, and buggy. IPFS has a lot of big ideas but suffers from a lack of polish that would make Augías look clean. And as soon as you scale to larger collections of files, it quickly crumbles under its own weight. You can throw more resources at it, but past some point it just falls over. It just doesn't work outside of small-scale tests.)
It is a set of open source libraries for peer to peer networking and content-addressed storage. It is written in rust, but we have bindings to many languages.
One part of iroh is a work in progress implementation of the willow spec. The lower layers include a networking library similar to libp2p and a library for content-addressed storage and replication based on blake3 verified streaming.
Most iroh developers have been active in the ipfs community for many years and have shared similar frustrations... See this talk from me in 2019 :-)
These were the most asked for bindings (python for ml, golang for networking and swift for ios apps).
We are using uniffi https://mozilla.github.io/uniffi-rs/
Would you need C or C++ bindings?
It seems uniffi does create C compatible bindings in order to make bindings for all these other languages. But these are internal bindings that are ugly and not intended to be used externally.
They seem to do their own streams, while we are adapting QUIC to a more p2p approach. Also the holepunching approach seems to be different. But I would love to get more details.
this was the presentation at DC'31. i will also check out iroh! thanks for working in building something in this space, it is much much needed!
Regarding holepunching, our approach is a bit less pure p2p, but has quite good success rates. We copy the DERP protocol from tailscale.
I am confident that we have a better story regarding handling of large blobs. We don't just use blake3, but blake3 verified streaming to allow for range requests.
Also I wrote my own rust library for blake3 verified streaming that reduces the overhead of the verification data. https://crates.io/crates/bao-tree
I tried to get on their discord at https://veilid.com/discord, but I get an invalid invite. You know a better way to get in touch?
thanks for the links, i will get in touch personally when i try ir0h :)
We would like to take our rust willow impl and separate it a bit more from our code base, so that iroh documents are just users of the willow crate.
Although I am concerned that while dumbpipe does mention cryptography, sendme's webpage makes no mention of it (?).
Which works against it. E2EE is a requirement today.
These tiny tools are basically one week projects to show off the tech, but they try to be useful on it's own as well.
200,000 files could take a while to advertise, but from memory it should work, should hang for less than 15 minutes. But depending on your hardware, file size, quality of connection to your peers, alignment of planets, etc.
If you add one order of magnitude above that, it starts to become tricky. Manageable if you shard over several nodes and look for workarounds for perf issues. But if you keep growing a bit past that point, it can't keep up with publishing every small chunk of every file one by one fast enough.
But it's also very possible perf has improved since the last time I tried it, so definitely take this with a grain of salt, you might want to try installing and running the publish command and see what happens.
Even if there's a lot of sharding and propagating and whatever to do, it should happen in the background, and never interfere with user experience.
From your description, it seems their implementation has serious issues.
I'm not sold on IPFS and will look at Willow and IROH.
(Mathematically, a name-addressable system is actually a superset of content-addressable systems as you can always use the hash of the content as the name itself.)
It's a superset in that sense but not a superset in another sense.
In a content-addressable system, if I post a link to another piece of content by hash, then no one can ever substitute a different piece of content. Like, if I reference the hash of a news article, no one can edit that article after the fact without being detected. This is a super-useful feature of CAS that is not a feature of NAS. Other implications:
* I can review a piece of software, deem that it's not malware in my opinion, and link to the software by hash. No one can substitute a piece of malware without detection.
* Suppose you get a link from a trusted source. Now you can download a copy of the underlying content from any untrusted source, without a care about authentication or trusted identities. This describes BitTorrent.
^ which you, by definition, already have
The problem with IPNS is that the performance is... not great... to put it politely, so it is not really an useful primitive to build mutability.
You end up building your own thing using gossip, at which point you are not really getting a giant benefit anymore.
IPNS uses the IPFS kademlia DHT, which has some performance problems that you can argue are fundamental.
For solving a similar problem with iroh, we would use the bittorrent mainline DHT, which is the largest DHT in existence and has stood the test of time - it still exists despite lots of powerful entities wanting it to go away.
It also generally has very good performance and a very minimalist design.
There is a rust crate to interact with the mainline DHT, https://crates.io/crates/mainline , and a more high level idea to use DHTs as a kind of p2p DNS, https://github.com/nuhvi/pkarr
It would be possible to add a layer on top of IPFS to include some context with every hash lookup so the search can be more focused, but the original design just assumed it was ok to do a log2 search for every chunk.
Using it to publish every tiny chunk of a large file is a horrible idea. It leads to overwhelming traffic.
If you publish a few TB of data, due to the randomness of the DHT xor metric you have to basically talk to every node on the network. Add to that the fact that establishing a tcp libp2p connection is much more heavyweight than sending a single UDP packet like in the bittorrent mainline DHT, and you are basically screwed.
In iroh we don't publish at all by default. But if you have to use a DHT, the fact that we have a single hash for arbitrary large files due to blake3 verified streaming helps a lot.
You still get verified range access.
Syncthing is designed specifically for file system sync (and does a very good job). Willow could be used for file system tasks, but also for storing app data that is unrelated to file systems, like a KV store database.
You should be able to write a good syncthing like app using the willow protocol, especially if you choose blake3 as the hash function.
This is a protocol for generic shared information spaces, where each person still owns & can manage permissions for their pieces of data in the space. It's a general idea that's present & implicit in most existing online spaces.
What does it mean that Earthstar will become a Willow protocol? Isn't it an implementation of Willow?
- One in typescript - One in rust
willow was not developed in a vacuum.
the willow folks have worked with us while we have implemented many ideas from willow, starting with range based set reconciliation ( https://arxiv.org/abs/2212.13567 )
they have been open to removing parts that have turned out to add too much complexity to implementations.
We've been working with the willow team as we go & giving feedback on the spec.
disclosure: I work on iroh.
Quick googling did not give me a proper grasp of the use cases for iroh/IPFS vs iRODs.
Would you be willing to list the benefits of iroh vs IRODS?
Seems like one would want IRODs if they have massive amounts of highly sensitive data that needs fine grained access control. You would want iroh if you're building an app that uses direct connections between end-user devices to scale data sync
Iroh is named by a certain fictional character that likes tea. Any similarity is a coincidence.
But it seems like iRODS is much more high level than iroh. E.g. iroh certainly does not contain anything for workflow automation. You could probably implement something like iRODS using iroh-net and iroh-bytes.
"The file was in my sleeve the whole time!"
What's the purpose of having separators in the keys?
Please try to follow RFC3339 when writing dates.
E.g. 11/01/2024 is ambiguous, as it could be January 11th or November 1st, whereas 2024-01-11 is RFC3339-compliant and does not exhibit this problem.
Still useful otherwise if not… assuming there’s an actual client/server for it on mac/linux…
E.g. in iroh-sync (which is an experimental impl of the willow protocol) you are not concerned with global scaling. You care only about nodes that are in the same document.
So while if you request hash QmciUVE1BqKPXMSvTTGwHZo1ywYdZRm9FfBvEJkB6J4USb via ipfs, you are trying to globally find anybody that has this hash, which is a very difficult task.
If you ask for some content-addressed data in an iroh document, you know to only ask nodes that participate in this particular document, which makes the task much easier.
Edit: regarding clients, iroh is released for osx, windows and linux. Iroh as a library also works on ios. Download instructions are here: https://iroh.computer/docs/install
That does not go anywhere.
This is disappointing. What's been read can never be un-read; to say otherwise is deceptive.
In this case, yes, it's impossible to guarantee that some malicious peer doesn't ignore my "plea to delete". But combined with the fact that my data will only be replicated to/by peers I already have a trust relationship with (as opposed to e.g. on a blockchain) it provides another layer of protection that a system without deletion simply doesn't have. Not perfect, but not useless either.
The project's goals are hard and noble. It would be better to under-promise and over-deliver than to make everyone question their claims. Maybe I'm just a grumpy old man at this point, but there are already too many caveat emptors in computing. They could have said "better" erasure of data.
You're right that it's nigh-meaningless for a public cluster.
Attackers are a whole other matter, and their existence doesn't make the feature pointless, for the above reasons.
This is a good point, and "GDPR compliant erasure of data" would be a great way to explain it. As a user I can guess what that means, and as an engineer it doesn't sound like magic.
I appreciate people trying to do something hard and noble.
Maybe "total erasure of data" is a too strong promise, but the fact that you can not force nodes that you don't control to unsee things is common knowledge, so in my opinion this does not need a qualifier.
Willow's claim has to do with erasure of the _networked_ data. It doesn't claim that copies people make are destroyed. Almost everyone understands and expects that if you can view data, you usually can somehow make some kind of copy of it. The question usually comes down to: how good of a copy?
Perhaps the best way to prevent perfect copying of data is to prevent someone from viewing it on a device they control.
This is true for the target audience of the article, but certainly not for people in general. It might be true for people in their 20s, but I strongly doubt it's true for any other age range.
The CALM theorem would like a word with you.
You simply can't have consistent non-monotonic systems.
Forgetting is ok, deleting is not.
> This community has not been reviewed and might contain content inappropriate for certain viewers. View in the Reddit app to continue.
Wow, what absolute horseshit. The march continues to acquire marketing signals at any cost.