On building 30k Debian packages
moyix.blogspot.com
moyix.blogspot.com
If you can get your hands on a buttload of memory, try using a tmpfs mount instead. For i/o heavy builds it's like turning on ludicrous speed.
> trying to automate everything in a language where failures are silent
set -u -e -o pipefail
> and can do exciting things like call "rm -rf /" when you meant "rm -rf ${foo}/${bar}"I mean, if you're calling any recursive-file-unlinking thing from Python you face the same problem. Same solution applies: normalize the path (readlink, realpath) and test if it's "/" before adding an argument to the command. (Also don't use -f unless you need to, it ignores errors)
set -e -u -o pipefail
_rm_r () {
declare -a rmargs
for arg in "$@" ; do
fullpath="$(readlink -f "$arg")"
if [ ! "$fullpath" = "/" ] ; then
rmargs+=("$fullpath")
fi
done
rm -r "${rmargs[@]}"
}
_rm_r "${foo}/${bar}"
But I agree with OP that you should use any language you're most familiar with to produce the best results possible in the least amount of time.The pipefail bit is a nice trick; I should remember to do this by habit when writing bash scripts (and to send the script through shellcheck)!
I did experiment with just preinstalling the top 300 or so most common build dependencies, but this ended up causing more problems due to conflicts (and not just version conflicts – some packages behave differently at build time depending on what's already installed; an example is any R package which will use my old nemesis xvfb-run if it's present on the system [1]) than it saved in time or space.
I do feel a bit guilty about how much I'm hammering the debian repos though, so I will probably set up a local apt mirror and point my sources.list to that instead next time.
[1] https://salsa.debian.org/r-pkg-team/dh-r/-/blob/master/dh/R....
Also, are you using ccache / distcc?
Not using ccache/distcc since I'm just running this on my own workstation. I'm not sure ccache would help here because each build runs in an isolated container environment, to prevent cross-contamination with other builds. I'm also instrumenting each build using bear to generate a compilation database, and I think distcc would mess with that? (I can't imagine distcc knows how to transfer over the bear wrapper binaries / shared libs?)
To be honest it's been... a very long time... since I've used these. But the theory is that it all "just works", in that if it can compile something with cache or remotely, it does, but if not, it just works as normal. So I think it would work for your use case? Maybe a Gentoo user can correct me :)
For source files, I'm thinking SSD, and leave some of the RAM for filesystem caches.
In addition, the Zsh equivalent of the Bash "strict mode" is a bit stricter. Here's mine:
emulate zsh
setopt \
err_exit \
pipe_fail \
warn_create_global \
warn_nested_var \
no_unset
https://git.sr.ht/~wintershadows/dotfiles/tree/master/item/l...>Avoid shell hackery. This is probably controversial
You bet. It's less than 5 minutes to learn how to make shell _not_ fail silently and an additional 10 minutes (I'm overstating the time) to figure out how to check variables for whether they will evaluate to empty strings -- to avoid deletion catastropes.
Not saying anyone should be "forced" to learn this. It's simply more accurate and helpful to say, "I don't know and I don't want to learn" instead of blaming a language.
NB: When I say "sufficiently bad at bash" I think of myself first. Just because I know about 'set -e' and friends doesn't mean I trust myself to get it right nor do I believe anybody else should trust me to either.
This includes both inline invocations like python -c, but more commonly simply a .py file in the same directory as the shell script.
So I think of shell as the main(), and then it often calls Python and R. (One big reason I started using shell was for preparing data in Python and then analyzing it in R.)
That said, I think shell needs better support for JSON and TSV so you can interoperate with languages like Python, JS, and R better. Things like the 'read' builtin aren't great for this.
I'm a staunch proponent of "anything over 100 lines should not be written in bash". You can do many more things easier and with less google-fu to find your niche SO answer.
Unless you have absolutely no clue about other scripting languages that already exist on Debian, just don't do it, there's so many better options that actually last and can even be read by other people.
Here's one tool I wrote in Bash. It's an alternative to Terragrunt. 616 lines. https://github.com/pwillis-els/terraformsh
Here's another one. A static-app package manager/versioner/pinner thing. 409 lines of POSIX shell (the code is intentionally ugly, don't take this as an example of good code). https://github.com/peterwwillis/clinst/blob/main/clinst
For both of these tools, 90% of what they do is just managing environment variables and calling other programs, and that's what shell scripts are best at.
If you don't have enough memory, using an nbdkit tmpdisk (https://libguestfs.org/nbdkit-tmpdisk-plugin.1.html) is a middle ground. It's backed by disk, but disables flushes which gives you a good 30% speed boost. (Obviously only use this for stuff you don't care about / can reproduce, because disabling flushes is dangerous and will lose data if the machine crashes).
* https://en.wikipedia.org/wiki/FreeBSD_Ports
* https://www.freebsd.org/ports/
There's already a 'pre-canned' bulk build infrastructure:
* https://docs.freebsd.org/en/books/handbook/ports/#ports-poud...
* https://wiki.freebsd.org/VladimirKrstulja/Guides/Poudriere
I would love to extend this even further to include projects on GitHub, but once you step outside the distro build systems it gets impossible to actually compile them in any standardized way. I might try to do some kind of best-effort (try ./configure && make, cmake, etc.) though.
FreeBSD Ports allows you to specify GH information to pull down code/tarballs in a standardized fashion:
* https://wiki.freebsd.org/Ports/SimpleGithub
* https://docs.freebsd.org/en/books/porters-handbook/makefiles...
Then you'd list build- and run-time dependencies.
The closest thing for a source hub and automated build system that is probably CPAN and the like (PIP, CRAN).
Another build infrastructure, focused more in the HPC niche, is Spack:
* https://spack-tutorial.readthedocs.io/en/latest/index.html
It would provide an ongoing and always up-to-date data set rather than relying on external entities interest, good-will and resources to do it (such as yourself) and could provide significant insights into the entire Debian project.
Is there some added value?
It helps building dependency graphs, so that you can run static analysis tools. Coverity IIRC does something like that to get the list of units, and build (preprocessor directives, etc) options.
- Is this code even used? (Maybe it's only included for some specific configurations, or #ifdef'd out)
- What does MAGIC_MACRO(x,y) do? Maybe this is only in a header file that's generated at build time, and you can't understand/analyze code that uses it without knowing its definition.
- What compiler version was used? Code may be interpreted differently depending on the compiler (particularly if gcc or clang-specific extensions are used, but also just due to things like compiler bugs or idiosyncrasies)
Coverity wrote a nice article about some of these issues back in 2010: https://cacm.acm.org/magazines/2010/2/69354-a-few-billion-li...
> Law: You can't check code you don't see. It seems too trite to note that checking code requires first finding it... until you try to do so consistently on many large code bases. Probably the most reliable way to check a system is to grab its code during the build process; the build system knows exactly which files are included in the system and how to compile them. This seems like a simple task. Unfortunately, it's often difficult to understand what an ad hoc, homegrown build system is doing well enough to extract this information
[...]
> The right approach, which we have used for the past seven years, kicks off the build process and intercepts every system call it invokes. As a result, we can see everything needed for checking, including the exact executables invoked, their command lines, the directory they run in, and the version of the compiler (needed for compiler-bug workarounds). This control makes it easy to grab and precisely check all source code, to the extent of automatically changing the language dialect on a per-file basis.
[...]
> Law: You can't check code you can't parse. Checking code deeply requires understanding the code's semantics. The most basic requirement is that you parse it. Parsing is considered a solved problem. Unfortunately, this view is naïve, rooted in the widely believed myth that programming languages exist.
> The C language does not exist; neither does Java, C++, and C#. While a language may exist as an abstract idea, and even have a pile of paper (a standard) purporting to define it, a standard is not a compiler. What language do people write code in? The character strings accepted by their compiler. Further, they equate compilation with certification. A file their compiler does not reject has been certified as "C code" no matter how blatantly illegal its contents may be to a language scholar. Fed this illegal not-C code, a tool's C front-end will reject it. This problem is the tool's problem.
compile_commands.json is a way to solve that problem in a "build system independent" way. It's just a list of exactly what commands where used for what files. As long as you generate that file correctly, you can use whatever build system you want (cmake, hand-written make file, 1000s of lines of Perl 5, a shell script from 1983, whatever) and your editor can correctly show errors, warnings, provide auto-completion, all that good stuff.
An alternative (at least for clangd) is to use compile_flags.txt, where you can just specify whatever flags you want the compiler to use
EDIT: if you're asking why this person in particular wants compile_commands.json files, it seems like he just wants a unified format for the compile commands of all the files in all the packages so he can use it for his analysis or whatever.
-rfakeroot has been the default since dpkg 1.14.7, released in 2007:
Most software which is maintained upstream and within nixpkgs is fine. But there's a bunch of "cruft" that you accumulate over time.
[0]: https://hydra.nixos.org/jobset/nixpkgs/trunk
P.S. The link above builds for x86_64-linux, aarch64-linux, aarch64-darwin, x86_64-darwin, which also contributes to the failure rate.
In practice a few percent of packages will fail to build (packaging is hard!) but it’d be good enough for getting large quantities of compiled code.