The new PostgreSQL 17 make dist
peter.eisentraut.org
peter.eisentraut.org
Generated output, vendored source trees, etc. aren't, or can't be, meaningfully audited as part of a code review process, so they're basically merged without real audit or verification.
My personal preference is never to include generated output in a repository or tarball, including e.g. autoconf/automake scripts. This is directly contrary to the advice of the autotools documentation, which wants people to ship these unauditably gargantuan and obtuse generated scripts as part of tarballs... an approach which created an ideal space for things like the XZ backdoor.
This speeds up CI (the generation path can be done in parallel) and most local development.
The one catch is that it relies on mostly trusting whoever has a commit bit. But if you don’t have that and any part of the build involves scripts that are part of the repo itself, then you’ve already lost.
Would the comparison not show that the person you're trusting goofed or is being malicious?
If the dev goofed, then good thing it got caught.
If the dev is not trustworthy, then you have evidence of such untrustworthiness.
Bingo. This is what I am working towards convincing people to adopt at my current job. It's a long road.
I would be very interested in how seeing how other people are doing it.
Thanks!
The real work is being able to transform the generation task into a reproducible step that be run consistently anywhere. Containerizing those steps can help but it’s not strictly required nor is it enough if the “inputs” are a non-seeded random or the current time.
It relies on your generated artifacts being deterministic, which is a design goal of that particular project so works fine there.
The inputs and the generation will obviously be defined.
If the generated files are what you say? Well, just embed the generation step into the build system. A simple approach like that is easily made reproducible, and we avoid introducing noise into the repository.
And yes, those have to be deterministic with regards to inputs, it does not make sense otherwise.
If you make a large but simple refactoring, like renaming a frequently-used function across a large repo, nobody is going to audit that diff and check for extra changes.
Things don't have to be this way, Google's source control systems apparently has tools that can do such refactorings for you in a centralized fashion, and one could make something like that for git.
That's not entirely correct. Indeed there was a part of the xz backdoor that lived in the configure script. However, that part was also included in the sources of the configure script as found in the tarball (and not in the git archive).
Thus regenerating the configure script didn't help, but regenerating the tarball did.
It adds unneeded complexity.
On the other hand, keeping tarballs close to the git tree makes it easy to reuse git archive and related GitHub features, provided the repo properly includes some kind of versioning information in tree.
I, as a developer, organize sources in a way that make it easy to work for another developer. My software will never be compiled by any user. All my users use build artifacts.
I might consider adding autogenerated code, but only when I'm like 99% sure that this code won't ever change. For example that's the case for integration with many organizations where WSDLs are agreed upon once and then never touched. Having Java sources regenerated every build just adds few seconds to every build time without noticeable advantages.
The fact that some Linux users prefer to build software from the sources and at the same time do not want to install necessary build tools is a bit strange situation.
May be containers should be better utilized for this workflow. Like developer supplies Dockerfile which builds a software and then copies it to some directory. You're running `docker build .` and they copying binary files from the container to the host.
As u/nrabulinski says, you can have the CI system generate and commit (with signed commits) autoconf artifacts.
The same can be said about autotools itself :/
Historical and current use indeed vary, and many times even using autotools itself isn't as appropriate.
There is a learning curve for either Nix or Guix that puts many off. However its not that steep, certainly it is many orders of magnitude easier than maintaining PostgreSQL, and once you are over that you no longer need to do things like keeping a dedicated clean machine just to pack a tarball. Write the derivation and anyone, anywhere, on any machine can generate the exact same tarball with a one liner
The barrier caused by the initial steps of learning Nix/Guix is a shame because once you are over it, it is difficult to see why software is built any other way (the same may apply to bazel, but i have no experience with that).
They just aren't by default (because they include a timestamp) and you need to jump through multiple hoops to get them there, consistently. (And things like "apk add" or "apt install" can't be used unless you're installing pinned versions)
A reproducible build is grand, but somewhat tangential to that goal, and hard to obtain in practice. Besides the timestamp problem already mentioned, you can't always pin the versions of system libraries and other distribution-provided software. The large long-term cost of hosting and geographically distributing content leads to many distributions, and especially their externally provided package mirrors, discarding stale versions from repositories. Often, the only available versions are the one included in the release plus the latest N, with N sometimes as small as 1.
If you're building a no-frills image for production deployment of a single piece of software, this problem can be bypassed thanks to distroless and other stripped-down base images, but "batteries included" images can't go this route.
Of course there is a balance here, there is a reason to pin versions. I'm stating why you shouldn't do that, but I cannot figure out all the pros and cons and how they should work out for your needs.
A variation of the above is reproducible builds are not that useful - sure you can prove the build is the same, but in the end you want the latest security fixes applies and so by the time you create the replacement build and verify it the build is obsolete.
Don't get me wrong, reproducible builds are important and do good things - but there are severe limits to what you can/should do with them and so while it is important to demand them, they are not important to use yourself.
# build initial images
# add semi-static inputs (mostly static config data, crypto data, signed inputs)
# add final watermarks
So each step can be verifiedAre you pinning your base image? Where did that come from? Are you pinning your packages? What about their dependencies? Are you locking down the hashes or just hoping that your distro won't replace a package in-place?
And that's before you get into crap like OpenShift certification that blanket requires a `dnf update` statement.
I can't imagine building such a monstrosity with anything else. And since the plugins are dependant on postgresql but not eachother I can add and remove them at a whim. Nix will create layers for me automatically.
And when I upgrade postgres I know I all packages will be built against the new postgres because Nix.
I think Nix could use list/dict comprehensions and some more devcandy sure, but it's really really great.
And at the end of the day, if you just go look at the source it's all there available to you, you don't have to wonder how Debian or RedHat built their golden postgres, there's no golden anything in Nix because if their hashes don't match mine I won't be pulling from their cache.
I think Nix biggest issue is that it doesn't attract promo skiddies the same way an imperative dirtbag like Salt or Ansible would, and most people can't even comprehend the things that open up when you can trust your shit.
Wanna write the hackiest perl script ever that'll never keep working? That's what activates most people's new NixOS generation still (there is a rewrite undergoing).
But back to point, Nix on Ubuntu patches /etc/{bash,fish,zsh}rc, creates the /nix top folder and that's it. It doesn't eat your system.
Yes, it has warts and they're big. But it's the only way forwards
You may be speaking from the perspective of using Linux because this here is some "you gotta be kidding me": https://nix.dev/manual/nix/2.18/installation/installing-bina...
Would you recommend using Nix even in that context?
I've also seen the bison and flex output included in VCS, and the same guidelines apply.
1) Do they commit the generated flex/bison to git now? So tarballs match git
2) Or do they now leave it up to the end user to generate flex/bison? And run their custom Perl scripts, etc.
Switching to `git archive` is fine, and you can add files to that, but https://github.com/postgres/postgres/blob/master/GNUmakefile... doesn't. So, I guess users now _have to_ run `autoreconf -fi`? No, because those are now committed in the source tree (https://github.com/postgres/postgres/blob/master/configure).
What is different about gzip and bzip2 that causes this?
Also Gzip includes metadata about the compression like a timestamp for when the file was compressed, and bits for which OS was in use when the file was compressed, etc. so the out put is never 100% the same, though it ought to be easy to work around that part.
This could be fixed regardless of git version by calling external gzip after generating a plain tarball (or even decompressing the bzipped tarball).
Hmm, I wonder which approach would actually be fastest. Does cache contention break the obvious use of `tee(1)`?
What packages are they referring?