This speeds up CI (the generation path can be done in parallel) and most local development.
The one catch is that it relies on mostly trusting whoever has a commit bit. But if you don’t have that and any part of the build involves scripts that are part of the repo itself, then you’ve already lost.
Would the comparison not show that the person you're trusting goofed or is being malicious?
If the dev goofed, then good thing it got caught.
If the dev is not trustworthy, then you have evidence of such untrustworthiness.
Bingo. This is what I am working towards convincing people to adopt at my current job. It's a long road.
I would be very interested in how seeing how other people are doing it.
Thanks!
The real work is being able to transform the generation task into a reproducible step that be run consistently anywhere. Containerizing those steps can help but it’s not strictly required nor is it enough if the “inputs” are a non-seeded random or the current time.
It relies on your generated artifacts being deterministic, which is a design goal of that particular project so works fine there.
And yes, those have to be deterministic with regards to inputs, it does not make sense otherwise.
The inputs and the generation will obviously be defined.
If the generated files are what you say? Well, just embed the generation step into the build system. A simple approach like that is easily made reproducible, and we avoid introducing noise into the repository.