Distributing NixOS with IPFS
sourcediver.org
sourcediver.org
I hadn't even thought about using the FUSE integration of IPFS, but it makes a lot of sense. Nix is a lazy language, and the nixpkgs repository basically defines one big value: a set of name/value pairs for every package it contains (as well various libraries for e.g. working with Python packages, Haskell packages, etc.). The only difference between installed/uninstalled packages is whether anything's forced the contents to be evaluated yet.
Likewise, an IPFS FUSE mount conceptually contains the whole of IPFS. The only difference between downloaded/undownloaded files is whether anything's forced the contents to be evaluated yet.
Perhaps a nail in the coffin of one of Nix's biggest absurdities.
One benefit of schemes like this that people don't talk about much is that, by no longer downloading from an expected place, you're removing the possibility for a compromised developer or server operator to selectively serve up malware to a targeted user. Instead you're getting the file over bittorrent and checking its hash, and you could gossip with other bittorrent clients to confirm that everyone's trying to get the same hash.
Compare with the state of the art in most software updates, which is that you connect to some download server and it could serve signed malware to people on its target list and probably no-one would notice.
(Schemes that use some of these techniques to take out the single point of malware-insertion have been called "Binary Transparency" schemes, as an analog to Certificate Transparency.)
https://www.gnu.org/software/guix/manual/html_node/Invoking-...
Basically, since builds are reproducible, you can automatically build from source and see if the hash of the binary you built matches the one you are downloading.
Obviously, source can be still compromised. But that's probably something IPFS won't fix unless wherever you get sources from is also on IPFS.
Where your distrust of keys would have more merit is the storage and strength of said keys. Eg If they dont have a strong passphrase and stored on an NFS / CIFS share then one could argue that they're no more secure than a bespoke build script.
I think we can do the same thing for software. Why not try?
Yes there are tools one can run to protect themselves against the aforementioned, eg application firewalls, sandboxing, network firewalls, etc. But at some point you have to trust that Microsoft Office / Firefox / your favourite Linux distro / whatever was built honestly from the outset as software - even with the source code available in the case of OSS - is far too complex to reliably vet before running in production.
The issues with storage and transport (IPFS, HTTPS, etc) is a different matter because they're to protect against external attacks rather than corruption within the company itself. This is where the issue of software signing might fall short. Not because of disreputable people within the business but more because of negligence (eg certificates not being stored securely so attackers can inject malware into the software and then sign it themselves)
So I'm not against criticisms regarding software signing; I just don't agree with your points regarding the motives of the signers. Simply put, if you cannot trust key people within a business to write and release software honestly (negligence aside), then you should not be installing nor running their software to be begin with.
For example, everybody hosting a file in IPFS could automatically become a BT seeder for that file. And searching on Pirate Bay for a certail file path / hash could make the file accessible to people who have a BT client but do not (yet) run IPFS on their machine.
BitTorrent would for that purpose be harder to modify, and claim backwards-compatibility, so I think they went all-in, and bitswap - with pluggable "trading strategies", ie how to decide which peer to share with and how much - how to avoid free-leach and so on. There was a mention of combining ipfs with filecoin - bitswap would make that integration easier.
Individual IPFS nodes are certainly blindly trusting the developer's signature as a stamp of approval. Adding more nodes doesn't make that problem better. It makes it worse by providing a greater false sense of security.
In the case of a compromised server operator, as long as hosting company X is smaller than Amazon, it's always better to use Amazon's cloud service to mitigate the possibility of server operator tampering.
Trust isn't really something you can algorithmically fabricate. At a certain point it always reduces to a tautology: "I trust this thing because I trust it." Distributed compiled code, because of its opacity and complexity, is an excellent example of exactly how hard it is to kick that bootstrapping tautology further down the road.
Distributing binaries via IPFS is functionally identical to distributing signed binaries from a central server, provided clients always check the signature. Now, that last bit isn't necessarily always true, but if your problem is "why aren't my clients checking their signatures", solving it with IPFS just doesn't make sense. It's like saying "This person isn't PGP signing their emails, so I'm going to download all of my emails using Bittorrent."
When will the IPFS people finish up https://github.com/ipld/cid so we can link whatever content addressable data we want?
I'd use git tree objects, despite SHA-1, because it's widely supported. Or do a format identical tree objects but with the IPFS's multihash and SHA-1 banned.
Point is, underlying protocol should be agnostic to hashing scheme, we should have a trait/type class like
/// Node in try
trait Payload {
type Hash: HashingTrait;
fn unpack(Payload) -> (Vec<u8>, Set<Hash>);
fn pack(Vec<u8>, Set<Hash>) -> Payload;
// Implement either and get the other for free!
fn hash_packed(p: Payload) -> Hash { hash_unpacked(packed(p))
fn hash_unpacked(p: (Vec<u8>, Set<Hash>)) -> Hash { hash_packed(packed(p)) }
}
any `(Hash, Payload)` than can define a `(binary blob, Set<Hash>) -> Hash` and Payload function should work.The cid stuff has been implemented and initial support for it is being landed in our 0.4.5 release (which will be soon, hopefully release candidate within a week).
With that and IPLD, you can craft arbitrary objects in JSON or CBOR (theres a 1 to 1 mapping, objects are stored as cbor) and work with them in ipfs. For example, i could make an object that looks like:
{
"Contents": {"/":"QmHashOfPkgContents"},
"Compression": "bzip2",
"NarSize": 12345,
"References": {
"foo": {"/": "QmHashOfFoo"},
"bar": {"/": "QmHashOfBar"}
},
"Signature": "signature info, or a link to the signature",
}
(please excuse my attempt at recreating a nar file in rough json).This object could then be put into ipfs with:
cat thing.json | ipfs dag put
And you would get an ipld object that you can move around in ipfs, and do fun things like: ipfs get <thatobjhash>/Contents
to download the package contents, or: ipfs get <thatobjhash>/References/foo
to get the referenced package (or open that hash/path in an ipfs gateway to browse the package graph for free in your browser :) )IPLD, last I checked, supports relative paths (which can make certain cycles), and not every node child gets its own hash. This is too much flexibility for my purposes (Nix or otherwise).
Also, when interfacing with legacy systems like git repos, one needs to dereference a legacy hash without knowing what it points to, which is easiest done with custom schemas.
Now, granted, customs schemas aren't a super fine-grained solution as every node in the network that cares about the data needs to implement the schema, but they are useful tool for these reasons (and that downside doesn't apply to private networks).
The CID (address format) in IPLD doesn't represent types of systems, but it represents types of data structures. E.g. in the case of git, it's not "git,$thehash", but instead "git-tree,$thehash" or "git-commit,$thehash".
That way you know which code you'll need to run once you have the object's payload, or you could have datastores that simply pull blocks out of a git repo.
Is this getting closer to what you mean?
While I'm not opposed to treating git that way, do note that git hashes are specifically constructed by prefixing the serialization of blobs, trees, and commits separately so that collisions are not likely.
https://github.com/multiformats/multicodec/blob/2725f3c5cd7b...
The big takeaway here is a really like the idea of IPFS, and want to be a full fan, but everywhere I look I see dubious interfaces. I see what already looks like legacy cruft, and they haven't even hit 1.0!
* CID is finished and live in go-ipfs@master and js-ipfs@master. We haven't announced it widely because go-ipfs@0.4.5 is still to land. (ooof)
* IPLD spec needs work, but work continues.
I wanted to add that:
* please contribute to CID to get it where you need it to be.
* i am personally very interested in defining IPLD data structures and their operations in a good language. This will probably be transpiled to the IPFS impl language, or compiled down to WebAssembly and run in a small WA VM (the web of datastructures)
Structs/Traits already exist in some form that is not well defined yet, what you are describing is a generalization of what is happening with Ethereum right now, we have eth-block being a IPLD "Format", which is basically a struct with some particular characteristic where the parser instead of being written in a IPFS VM language, it is written and executed as part of the daemon.
The idea that you have described above or a subset of it, it's part of the plan!
Please do participate in the issue that I linked you to!
I'll know real progress has been made when my browser can resolve something like:
cas://sha256/2d66257e9d2cda5c08e850a947caacbc0177411f379f986dd9a8abc653ce5a8e
Of course, there's also IPFS, Zeronet and Freenet which all address this exact issue in slightly different ways, all more web-targetted.
* https://github.com/ipfs/js-ipfs * https://addons.mozilla.org/en-US/firefox/addon/ipfs-gateway-...
We want to make it easier though, and haven't really got there yet. But if you have questions, #ipfs on freenode is a good place where core devs and the community hang around usually.
I had the feeling NixOS has a bit of a hard time get users and prove that it's a superior solution to ansible/docker/chef/etc. probably because of it's mediocre UX, haha.
But this would add another killer feature to it.
BTW, there is a small typo:
IPFS is aims to create the distributed net.
It should be: IPFS aims to create the distributed net.edit: Didn't mean to hit reply. Sorry.
comment1
[deleted]
comment3
However, you can delete comments only until a certain amount of time. If you wait too long, you can't delete anymore (and this is what happened here)."Ask HN: Where can I follow the changes of HN itself?"
Eg, IPFS isn't permanent hosting - it's purely hosting as long as there are seeds, like bittorrent. Hypothetically if a package is very old there may be no seeds for it anywhere. Someone (NIXOS/etc) will still have to pay for hosting.
Perhaps an equivalent thing for "someone will still have to pay for hosting" is that although that's true, anyone can put money into the pot to keep it going or bring it back, it's not reliant on the original creator to keep paying for it.
Sure, point your domain to your IPFS hash and use dnslink, it's quite reliable already actually. That's how ipfs.io is hosted for example, and we haven't hit any issues so far.
> Using IPFS' DNS (IPNS) means you have to keep the IPFS daemon running constantly, or else the files will be purged within an hour
So, the files won't be purged, but the record you push out with IPNS won't be valid after 24 hours. You can solve this easily by using /ipfs/:hash instead of /ipns/:id and it wont disappear.
Yeah, it certainly is possible to host static sites on IPFS -- I have been testing it on my site[0] just for kicks. However, since I really enjoy using my domain name, rather than "ipfs.io/ip{f|n}s/$hash", I'm reluctant to try adapting my site to IPFS. I am aware that this is an alpha product, though, and I can't stress enough how cool it is for files containing the same data to be given the same ID (hash). That way you don't have to run shasum ever again :D
(on a random note: what's up with /blog returning the IPFS blog, instead of my own content? did i misconfigure something?)
edit: And the part which turns /ipns/ipfs.io into an /ipfs path is called dnslink: https://github.com/ipfs/go-dnslink -- it resolves to what's in TXT _dnslink.ipfs.io.
> (on a random note: what's up with /blog returning the IPFS blog, instead of my own content? did i misconfigure something?)
I'm so sorry, that's a bug I put into the nginx configuration -- will fix it!
Also, it would require a preimage attack against one of the hashed items to be useful which SHA-1 will likely be resistant to a long time (though decreasing with the number of items hashed) and SHA-1 is unlikely to be vulnerable to a preimage attack in the near future based on what we know so far.
The signature and certificates that are used to validate the top-level index can be based on a far better hashing algorithm independently of the content-based hashing.
I'll take the construction "other people are inventing <the thing i invented some time ago>" to mean that you think you came up with this first, or at least prior to Nix or IPFS. And thus ":-)" to be a bit sarcastic and unhappy, instead of genuinely happy.
Similar times:
* http://appfs.rkeene.org/web/timeline?c=78c60b0c9e7da1c9&unhi...
* https://gist.github.com/jbenet/8f000606f2009495c56177f6ca2c1...
* https://github.com/ipfs/ipfs/commit/8004db75262fcd29399d6c7f...
* https://github.com/jbenet/random-ideas/issues/19
* i think the Nix people probably came up with this way before either of us
* and i am willing to bet everything that at least 100 other people (maybe thousands) have came up with this exact same idea over a decade before any of us...
In fact, most of the best "original ideas" in the IPFS body of work were probably first discovered decades prior.
I repeatedly see people succumbing to sadness over multiple discovery. It shouldn't be sad, it should be a happy event, as it confirms our own thoughts and presents an opportunity for collaborations. :)
Further thoughts here: https://gist.github.com/jbenet/8f000606f2009495c56177f6ca2c1...
2. The items with similar timelines are not comparable. The link to AppFS is a link to working code being written and the other links are to documents regarding a desire to have working code written someday, maybe.
3. It's a good idea, that I did not claim to come up with. It's a refinement of how things were done on UNIX campuses long ago where the was a common NFS-mounted directory with applications installed on it. This is definitely not original out claimed to be do. Then the same model that Sun/Oracle uses for Solaris packages in "ipkg" in Solaris 11 is mixed in with that. The result is something very similar to something old as well called 0install, back when it used LazyFS. Again not original. The thing that AppFS adds beyond just the merging of these ideas is the additional idea that the filesystem should be writable. This is not original exactly either since the same thing is done for ClusterNFS (except per-system changes are preserved instead of per-uid).
Given that I've obviously used all these technologies to achieve similar goals it should be obvious that I don't think AppFS is original -- quite the opposite, I've been using it for decades.
The intent of my post, which you seem to have missed, is to point to something else that does the same thing as this new thing. The reason that I would do this is because the ideas a are old and we can improve upon them by looking at similar implementations. I certainly spent a lot of time with 0install's LazyFS (probably one of the heaviest packages).
As an aside, I found your post condescending in tone. Given that you acknowledged that all your statements were based solely on assumptions for things I did not write, I do not think this was effective in communicating your thoughts over this medium. Overall, I'm not sure what your message would contribute if these assumptions you made were wrong. My guess (assumption) is that you did not consider what your message was if your assumptions were wrong. If my assumption/guess is wrong here, please disregard this section and do note that this message still has content that communicates meaning.
Also: I think some people are unaware that Nix hashes are not content addressable. The best solution (which OP is proposing) is probably to use the .nar hashes in IPFS which is content addressable.