If you preserved an immutable tag of the source code and all its dependencies, a copy of the compiler version used and all build flags, then you’ve still got some big holes in your ability to reproduce a binary:
1. OS version & patches installed
2. OS configuration
3. Hardware used (processors can have weird subtle bugs, microcode can affect execution behaviour, etc etc)
4. Transient issues - the golden copy to be reproduced for some post event investigation might have contained a bit flip leading to impossible to reproduce verification signatures
Etc etcOr is reproducability just a spectrum and you try to get further along it with some careful attention to detail, rather than an absolute to be achieved?
If the latter are you not cheaper just archiving binaries and tagging them with the source + deps + compiler + arch used to build them? Thats a 5 minute job to setup in your CI process and costs comparatively little to maintain vs wasting expensive human brains chasing down a futile goal.
Look for example at guix, that bootstraps explicitly with what they call "the maxwell equations of software".
Imposing a web-based build system, into the kernel of all places, is a kick to the face to people who care about reproducibility.
For something like Go it's trivial though. If you're using pure Go code then reproducible builds are pretty much as simple as "compile using Go 1.xx".
I don't understand why you think 3. and 4. are issues. Bit flips are very unlikely at compilation scale, and the hardware you use to compile something shouldn't affect the output. In either case you can just compile it twice and on different hardware and compare the result to confirm.
It is totally possible to use cargo without ever touching crates.io, but if you need any dependencies you'll of course have to provide them through some other means (local file system, git repositories, or a custom package repository).
The point is: all options are available for all systems, so suggesting any workflow doesn't work with any of these is just incorrect.
And I never said any workflow doesn't work with Cargo. I said I wished it didn't download code (by default) and that it didn't care how code ended up on my drive, that it just worked like a normal build system. That you can use Cargo in that way doesn't help.
How does the code get on your hard drive? I'd imagine from downloading it from some random website (arguably _more_ random since it's less likely to be a centralized place like crates.io). Or you could vendor the dependencies so that they're included when you get the source code for the thing you're working on, but as mentioned throughout this discussion, cargo lets you do that too.
cargo sounds pretty much like a new bespoke build process that doesn't even work across different languages on the same OS
I still think cross-language efforts (like bazel and I'm sure there are others, maybe nix?) seem generally better, but I suspect there is some fairly good reason they are less widely used.
I think the problem is that bazel and nix are just not that easy to get started with. If you're using them you're likely to have a team (or at least one person) working full-time on them, since bending everything to a singular worldview involves a lot of work.
I'm honestly astonished that programmers of a language that is deemed to be "safe by default" thought that this behavior was acceptable in any form, not to say the default. If downloading things at build time is somehow necessary, it should be an obscure option behind a flag with a scary name, like --extremely-unsafe-i-know-what-i-am-doing, that prompted the user with a small turing test every time that it is run. Cargo is just bonkers, it doesn't matter at all if it is "convenient" or not. Convenience before basic safety and reproducibility is contrary to the spirit of the language itself.
It's as if bounds checking in the language was deferred to a third party that you need to "trust" in order to believe that you won't have segmentation faults.
Edit: for things like the kernel, vendoring dependencies is still probably not a bad idea, of course
What happens when a given dependency adds new kernel-inappropriate features? Are kernel devs going to act like distro maintainers and decide between forking, maintaining patch sets, etc.?
A dependency veering off in a direction you don't like is one of the risks of using someone else's code instead of writing it yourself. Cargo makes it easy to use forked dependencies, and forking a dependency is almost always less work than if you'd never used it and written the code yourself from the beginning. (And to be clear this is only a problem for future evolution; a crate author cannot remove or modify an already-published version of their crate.)
I can grab the kernel sources from 1997 and build them today. Will I be able to build rust code from 2022 in 2047, because the 1997 kernel will still build at that date.
Where would you be grabbing it from? ...From a website? "Websites shut down, large websites with big storage demands are especially vulnerable to attrition. Who wants to pay the mounting bill for keeping decades of revisions of historical Linux kernels online?"
- Archiving the complete history of all crates in crates.io is perfectly feasible today for an individual. Over time that might change.
- Setting up a mirror is straightforward, should you want to do so: https://github.com/rust-lang/crates.io/blob/master/docs/MIRR...
- crates.io is financed by the Rust Foundation and is at no risk of disappearing, it is a very well funded effort.
- Using cargo with an alternative repo is not difficult, requires some one-time configuration.
- Vendoring your dependencies is supported.
- cargo hits the network to look for semver compatible updated versions of your dependencies on specific moments if you don't have a Cargo.lock file.
- Not updating your dependencies stops you from getting the rug pulled from under you if an unwanted change happens, but it also stops you from getting any desired changes including security vulnerability fixes.
- Even if you vendor all of your dependencies, you still have to audit them the first time and every time you update them. Are you? Most aren't. Code you haven't written yourself can't be assured not to be malicious, and code you've written yourself can still have exploitable mistakes.
It's easy enough to keep your own website up as long as you want to, the liability is other projects and services, especially when the scope of those services is "archive everything for everyone forever".
I'd like to see that data if so -- I have pretty big doubts that your statement has merit without some sort of evidence.
Kernel.org's repository is also of major versions, not every minor release and patch. That really wouldn't do for cargo. If it has ever been released, it needs to be kept in storage for as long as the rust ecosystem exists. That's decades, maybe even centuries of passing on the torch and hoping the next guy accepts the responsibility. Hoping you can find a next guy.
Now the lifetime of the dependency is that of your project. There's even tooling 'cargo-vendor' to help manage this setup.
Alternative of course is implementing it all yourself, which cargo doesn't prevent.
Can you? Do they still compiler with current compiler? You'll probably need to find a compiler of that time... And also all the interpreter for all the build scripts. Was that using bash or some old Perl? Maybe something more esoteric like m4 or tcl?
The point is that it always had many external dependencies to bootstrap. And adding one is not such a big deal, it just add another thing to archive among the many other things. The crates.io archive is probably not even that big.
But even if it has broken, I can just download an old linux distro. They effectively form a cohesive snapshot of the state of the toolchain whenever they were assembled. Slackware 3.1 from 1996 might be appropriate.
You will also need era-appropriate hardware to get that software to install.
At any rate, we are indebted to the future to preserve the present, as our past has been preserved for our benefit.
What happens when a crate version has to be removed due to a critical CVE or court order (IP Law violation, perhaps)? There may come a day where crates.io becomes torn between not breaking Linux source and not hosting actively bad source code.
Note that some of those concerns do apply to vendoring source as well, but the additional download step also removes options that the kernel maintainers have as long as they ship all the source for the kernel in one tarball. Like more control over the timing of inevitable decisions.
CVE = The Yank flag. Cargo will refuse to add new yanked packages to a lock file, but if a yanked package is already in the lock file, it will still build. The package is not actually deleted. https://doc.rust-lang.org/cargo/commands/cargo-yank.html
Legal = Hard delete. Nobody will go to jail just to avoid breaking your build. Of course, since crates.io and kernel.org are in the same legal jurisdiction, is there any actual difference here?
That's not just a rhetorical flourish, I'm actually curious what the answer is. As far as I know, (1) it almost never happens and (2) when it does, the change is made in upstream repos and as a practical matter, everyone downloads those changes and their up-to-date local copies lose that code.
The previous tarballs still work and contain the relevant code. Your build wouldn't rely on hosts complying with court orders in countries you might not live in.
If the code isn't vendored, just referenced with URLs, the old tarballs stop working.
But I think the kernel would vendor crate dependencies, partly so that people can build without accessing the network, simply because that's policy in many places.
To the second set of questions, how is this any different than any other dependency the kernel has? If the answer is "the kernel has no dependencies" then yeah, I'm very sympathetic to the argument that bringing in rust libraries is not a good reason to start having dependencies when none previously existed at all, but is that the case?
You must specify --locked to get that behaviour
A cargo build ends up there calling into the resolver’s resolve_ws_with_opts() which would refresh the lockfile.
Not resolve_with_previous() which would use the lock file as-is.
The only reason this sticks in my mind is i ran into an issue building bat after i made some changes, i obviously assumed it was my changes so went through the process of debugging and backing out my changes until finally i was back to a virgin branch and still failing - passing —frozen —locked fixed it.
That's exactly what it does. The developer is not really expected to thoroughly review the codebase of every dependency.
Just like javascript, all sort of supply chain attacks are made possible.
A single malicious library can sneak into large ecosystems easily.
Wait till you find out about java ecosystems
I know investment bank dev teams pulling whatever they need from maven central with no oversight or introspection.
This is borderline inevitable for most modern development stacks, though .lock files can definitely help, even adding hashes to check against if you care about your dependencies being the same as when you first download/add them to the project and/or inspect the code.
As for worries about the things in those URLs disappearing, in most cases you should be using a proxy repository of some sort, which i've seen leveraged often in enterprise environments - something like JFrog Artifactory or Sonatype Nexus, with repositories either globally, or on a per-project basis.
The problem here is that all of these repositories kind of suck and that the ecosystem around them also does:
- for example, Nexus routinely fails to remove all of the proxied container images and their blobs that are older than a certain date, bloating disk space usage
- when proxying npm, Nexus needs additional reverse proxy configuration, since URL encoded slashes aren't typically allowed
- many popular formats, like Composer (or plenty more niche ones) are only community supported https://help.sonatype.com/repomanager3/nexus-repository-administration/formats (nobody will ever cover *all* of the formats you need, unless you limit yourself to very popular stacks)
- many of the tech stacks that have .lock files may also include URLs to the registry/repository from which they're acquired, so some patching might be necessary
- in technologies like Ruby, actually setting up the proxy isn't as easy as running something like "bundle install --registry=..." as it is in npm
- in other technologies, like Java, you get into the whole SNAPSHOT vs RELEASE issue and even setting up publishing your own packages to something like Nexus can be a bit of work; the lack of proper code libraries for reuse and abundance of code being copy-pasted that i've been being a proof of this in my mind
Of course, i'm mentioning various tech stacks here and i don't doubt that in the long term Rust and other technologies might also address their own individual shortcomings, but my point is that dependency management is just a hard problem in general.So, for most people the approach that they'll take is to just install stuff from the Internet that other people trust and just hope that the toolchain works as expected, a black box of sorts. I've seen plenty of people just adding packages without auditing 100% of the source code which seems like the inevitable reality when you're just trying to build some software with time/resource constraints.
There's even a tool called cargo-vendor that does this for you!
The difference is that no C or C++ package management features are proposed for incorporation in the Linux kernel SDLC.
Sounds like they did a decent job anticipating that use case!
Some of it was a bit awkward to actually use this way in the early days, but those harsh edges have been since sanded off.
Addendum: Until quite recently, this was quite cumbersome. It also meant that all cargo invocations (by the same user) would use that override, always. It meant that compiling someone else's project became quite the hassle. Or compiling a project that mostly uses system dependencies, but some crates.io deps. But the situation is improving.
Cargo binary projects have Cargo.lock by default with checksums of all dependencies used. crates.io doesn't allow changing past releases, and has a policy of not deleting any crates unless legally required (user-accesible "yank" hides, but doesn't delete). Crates.io index is a git repository with full history of all changes to the registry, so you can recreate its state at any point in time (in case you lost your Cargo.lock, you can reliably remake one from the past).
And on top of that there's `cargo vendor` command that makes a local offline copy of everything you use, so you can fully archive a Cargo project and rebuild it any time later.
The answer there is often "use FFI". But if we're all going to use C APIs anyway, then shouldn't our package managers support C?
But C has probably the most awkward build system culture of any language, complicating the job of packaging quite a bit. A lot of language specific package management systems get lift from offloading the ugly C support problems to other layers. See python and "manylinux" for instance.