GNU make insanity: finding the value of the -j parameter
blog.jgc.org
blog.jgc.org
Quite frankly, I find this whole "look i did a horrible hack and let's see who can make the best worst horrible hack" thing quite stupid and silly.
GNU Make is free software released under a free license, my opinion is that instead of doing that crazy thing, the author could have just written a patch for GNU Make in order to make it export a "JOBS" environment variable to all its child processes.
Oh but yes, "I felt this feature was missing so I added it" is way way way less cool than "geez the gnu make folks are insane lollerplex they have no way to know how many jobs they're running".
ALSO: http://www.catb.org/esr/faqs/hacker-howto.html#believe2
Here it is a patched version of GNU Make, providing a #J variable that holds the number of jobs as passed via -j/--jobs:
http://santoro.tk/~manu/gnumake.png
As can be seen in the screenshot, it can be passed as environment variable to programs called by makeSource: https://github.com/esantoro/make
default: qualcosa.sh
JOBS=${#J} ./qualcosa.sh
Source for sample shell-script: #!/bin/bash
echo "According to \$JOBS environment variable, the argument of -j/--jobs is $JOBS"Of course you can patch the software and in the long term that will of course be the best solution. But it can take a long time until that patch actually makes it into the main repo, and even longer until it will be in the default `make` on all major distributions. So if you want to use this feature right now then writing a patch is useless for you, unless you only ever use it on your own machine.
The big advantage of the workaround is that it actually works with current versions of make.
Anyway, it was my small contribution to the discussion :-)
> I don't know how many hours the author has spent/wasted
> on this topic
The author, John Graham-Cumming, has been extremely active on the GNU Make mailing list, is the author of the "GNU Make Standard Library" of supplemental functions for Make, has written at least 2 books on GNU Make, and has developed commercial products that integrate with GNU Make (namely, "Electric Make").That is to say: He is one of the world experts on GNU Make.
I'm confident that he did this more "for fun" than because he actually wanted to get the number of jobs.
It took me about 15 minutes to come up with the solution. I wrote it up for fun because it illustrates some of the functionality of GNU make that people might not be aware of (order-only prerequisites, $(eval)) and using $(call) in a recursive fashion.
I actually spent more time on the blog post.
Asking for -j value inside the Makefile was a tell-tale sign of this sort.
Also: never, ever, build a make system that relies on recursive make invocations.
Automake is a particularly egregious violator is this rule (and other sanity rules) and has probably hurt make more than anything else.
-- Stuart Feldman
AM_INIT_AUTOMAKE([subdir-objects])
and reference your source with their relative path in the (only) Makefile.am in $(top_srcdir) bin_PROGRAMS = foo
foo_SOURCES = src/main.c \
src/foo.c src/foo.h \
src/quux/bar.c src/quux/bar.h \
src/quux/baz.c src/quux/baz.h
You can use relative paths like in most places automake accepts file names. The only time I've run into trouble using non-recursive automake (1.13.4) is adding flex+bison source. This requires BUILT_SOURCES, which is confusing, and automake is very picky and non-obvious about the names of the generated source files. I think it may not have been updated to support relative paths yet? It required some heavy workarounds: AM_YFLAGS = -d
# this was VERY important
# (must be abs_top_srcdir, not top_srcdir)
foo_LFLAGS = --header-file=$(abs_top_srcdir)/src/xyzzy_lexer.h
foo_SOURCES += src/xyzzy_lexer.l src/xyzzy_parser.y
EXTRA_DIST += src/xyzzy_lexer.h
# note the inconsistency here
BUILT_SOURCES += src/foo-xyzzy_lexer.c \
src/xyzzy_parser.h
# ...and this doesn't match BUILT_SOURCES
MAINTAINERCLEANFILES += src/foo-xyzzy_lexer.c \
src/xyzzy_lexer.h \
src/xyzzy_parser.h
Even with this annoyance, I would still highly recommend using this style instead of the traditional (recursive) automake. Autotools Mythbuster has more information on this other modern-autotools features:----------------------
PID := $(shell cat /proc/$$$$/status | grep PPid | awk '{print $$2}')
JOBS := $(shell ps -p ${PID} -f | tail -n1 | grep -oP '\-j *\d+' | sed 's/-j//')
ifeq "${JOBS}" ""
JOBS := 1
endif
all:
echo ${JOBS}
----------------------
quanticles@glados $ make -j8
8
quanticles@glados $ make
1
----------
edit: for silly mistake
grep -oP '\-j\x00\d+' /proc/$PID/cmdline
Edit: and no need for sed when a simple cut -b4- does the same job.If you don't have /proc but do have ps (1), you could get the parent pid like this:
PID := $(ps -p $$$$ -o ppid=)If so, $$ is the easiest way to get the PID of the current shell. That works in several other shells, too, but I am not sure sh is guaranteed to have it (FreeBSD's sh has it. See http://www.freebsd.org/cgi/man.cgi?query=sh)
$ cat Makefile
test:
+env | grep MAKEFLAGS
$ make -j4 test
env | grep MAKEFLAGS
MAKEFLAGS= -j --jobserver-fds=3,4
If you want to act like a sub-make, then parse MAKEFLAGS from your environment, and if you see -j and --jobserver-fds=R,W , then parse R and W as file descriptor numbers, block waiting for a byte from R before starting an additional parallel job, and write a byte to W when done with a parallel job... CMake? http://www.cmake.org/
I had high hopes for scons, but that turned into a quagmire. Any build system that's going to unseat autotools will need to obviate the need for almost all custom scripts. AKA, specify the 'what', and let the user apply a 'recipe' for the 'how'. Or something...if I knew what it was supposed to look like I'd build it.
SET(CMAKE_FIND_LIBRARY_SUFFIXES ".a")
SET(BUILD_SHARED_LIBRARIES OFF)
SET(CMAKE_EXE_LINKER_FLAGS "-static")
The bigger problems with building a static binary on Linux are that:
1. Most Linux distros don't install static versions of libraries. You can't link in a static library dependency that doesn't exist.
2. glibc can't "really" be linked statically because of the way it's designed. See here for the gory details: https://gcc.gnu.org/ml/gcc/1998-12/msg00083.htmlNaturally, you have problems #1 and #2 with autotools or any other build system as well, since they aren't CMake problems.
Truth is most autoconf setups are just arbitrary pastiches of other autoconf setups that the devs didn't really understand. Half the time a bunch of the test results aren't even used, or are testing for something that isn't an issue on any platform it's ever going to be compiled on.
And when that makefile breaks? At least you have a hope in hell of understanding why it broke. Autoconf breaks? Good luck.
This is probably also why libtool's configure probes no fewer than 26 different names for the Fortran compiler my system does not have, and then spends another 26 tests to find out if each of these nonexistent Fortran compilers supports the -g option.
I did this to see for myself if autotools bashing is justified and I can tell you this: If a configure script is that bad, then the project's developers did a poor job using autotools.
checking whether mkfifoat is declared without a macro... yes
Why? Nowhere in grep is there a single call to mkfifoat. Why does it care?I also liked this:
checking for MAP_ANONYMOUS... yes
checking for MAP_ANONYMOUS... yes
checking for MAP_ANONYMOUS... yes
checking for MAP_ANONYMOUS... yes
checking for MAP_ANONYMOUS... yes
checking for MAP_ANONYMOUS... yes
checking for MAP_ANONYMOUS... yesAnd the work has a bigger risk because you can't actually check on more than a couple of distributions before you ship.
And then I tried to run it. It took 6.5 seconds to run without any input. The "can't handle filenames with spaces" bug has been open for almost two years.
And then I went crawling back to make. Again.
Sometimes the modification time is not useful for deciding when a target and prerequisite are out of date. The P attribute replaces the default mechanism with the result of a command. The command immediately follows the attribute and is repeatedly executed with each target and each prerequisite as its arguments; if its exit status is non-zero, they are considered out of date and the recipe is executed."
When I want this in GNU Make I end up hacking together something resembling:
define test-exec
$(filter OK,$(shell if $1 >/dev/null; then echo OK; fi))
endef
define test-docker-image
$(if $(call test-exec,$(DOCKER) inspect $1,,docker_$1))
endef
target: $(call test-docker-image,foo)
$(DOCKER) run foo ...
docker_%:
$(DOCKER) build --force-rm -t project/$* $*
It seems this can be replaced with: docker_%:VPdocker inspect: %
$DOCKER build --fore-rm -t project/$stem $steminteresting academic exercise though. unless i'm mistaken, this will only detect up to 32 build jobs? i'd hate to see the Makefile that works for arbitrary values..
> you don't want the total number of threads spawned to be more than the machine has
that may not be necessarily true, and regardless, it might be a sane decision for yourself, but it might not be for others.
Or integrate with make's jobserver as someone else has described.
redo is the simplest, which is unsurprising given that it was originally proposed by djb. Alternatively, you could look into mk, which is the make replacement used by Plan 9, Inferno and derivatives.
[2] http://homepage.ntlworld.com/jonathan.deboynepollard/FGA/int...
I switched to ninja mostly for more complex projects and it's very nice, you can do anything in the generation stage and get all the power you need and then run it through ninja and get the full speed benefit of native code for the most common operation (build it!).
I used in the past both scons and waf and they are in a way like ninja but since the actual build process is still in python the time to build is longer. scons was a big pain in the ass for a large project and I regretted introducing it. waf was ok though.
I still do use make and makefiles for simple projects. It is usually simplest to get running and requires only a build stage without any configuration like ninja does. I also have the basic constructs hard-wired into my brain and can write that simple makefile faster than I can find the example code for ninja.
"Shake build systems are Haskell programs, but can be treated as a powerful version of make with slightly funny syntax. The build system requires no significant Haskell knowledge, and is designed so that most features are accessible by learning the "Shake syntax", without any appreciation of what the underlying Haskell means."
SET(i 0)
while(i LESS count)
LIST(GET extras 0 file)
LIST(REMOVE_AT extras 0)
LIST(APPEND extras "${root}/${F}/${file}")
MATH(EXPR i "${i}+1")
endwhile()
"Functional" with a FU. It's also buggy as hell.It does not do this. Been using it since 2.6. Get one of their super-tweaky ever-changing commands wrong and it doesn't generate dependencies AT ALL on some platforms. Sure, your project builds. And it's also a poisoned dart trap for anyone who attempts to modify your code.
A quick skim of the bug list is enough to make any truly-fastidious engineer's sphincter clinch. When you have over 1000 bugs that aren't even assigned to a dev ... in a dev tool ... hoo boy.
Er, what? Chromium has over 50000 unassigned bugs. Any software that sees a lot of use is going to have lots of bugs, and depending on the workflow sometimes none of them will be assigned.
I asked you "how is it buggy?" and your answer essentially is "It's buggy, trust me."... weasel answers, much?
I'm by no means a fan of cmake. But if I use and trust it, yet can generate much better criticism than you do, how are you going to convince anyone they shouldn't be using it?
"CMake is actually very reliable" --scrollaway
That's not criticism. That's just an opinion. No more or less valid than mine. But certainly less supported by evidence.
It still has its own arcane syntax rather than being pythonic, which I think is its greatest shortcoming - it emulates make syntax too much, where I think it would be nicer to just be a good proper python library that specializes in build scripts.
[0] https://twitter.com/Steve_Yegge/status/523367285658894337
For C++, and if you don't need anything really crazy. It's very readable (JSON-like), and custom rules can be defined in Javascript instead of "learn another language". It's also very fast - kind of highlights how slow Make is! No build system should be slow.
It's usable now but still kind of a work in progress.
1. Traditional Make files. Keep it simple, don't feel compelled to use all the dark corners of the language, and accept that it will be a little more verbose than some other solutions. And don't expect portability to Windows, since you won't get it (cygwin doesn't count as Windows). There are some portability problems between UNIX variants, but actually I think it's pretty easy to deal with these in plain old Make.
2. CMake. It makes the simple things simple and the hard things possible. It has its own simple scripting language which build files are written in, so you don't need to tear out your hair worrying about whether the the user's version of bash / python / perl / whatever matches yours. The language has the abstractions and functions that you need built-in, so most CMakeLists.txt tend to look the same (less wheel reinvention than you would get just using a general purpose language like Python or Perl for Makefiles.)
CMake is built with backwards compatibility in mind, so you can easily use old projects with newer versions of CMake. This is something you just don't get with autotools, where you have to have the correct version (not newer, not older) of autotools installed to build the project.
CMake does the detection of which header files are needed to rebuild which .c or .cc files which plain old Make doesn't do (without adding clunky extensions). CMake is also portable to Windows, which autotools is not. (again, limping along under cygwin doesn't count.) CMake can generate visual studio projects which can be used to directly build software.
3. Visual studio. If you only care about Windows, this is a sane choice.
On the other hand, reading the command line isn't sufficient. The -j argument does not require a job count, defaulting to "as many as possible".
Also, you would have to check your OS, too since the -j argument is ignored on MS-DOS (https://www.gnu.org/software/make/manual/make.html#Options-S...)
Best you can do is to run make -n target to probe, but it's possible to write Makefiles that run code even though -n is used, which will defeat such probes at distribution scale. It quickly becomes an exercise in heuristics and output parsing.
You need to write a full parser to discover all the targets.
$ make -pqrf <(curl https://gist.githubusercontent.com/jgrahamc/2cc6df7fd5cd61c4d93f/raw/parallel) | grep '^[^ ]*:' | cut -d : -f 1
par-%
par-22
par-30
par-5
par-27
par-31
all
FORCE
/proc/self/fd/13
par-13
par-9
par-23
par-19
.SUFFIXES
par-15
par-7
parallel
par-29
par-28
.parallel
par-4
par-25
par-6
par-14
.DEFAULT
par-24
par-0
par
par-1
par-10
par-16
par-2
par-11
par-17
par-20
par-8
par-21
par-12
par-26
par-18
par-3I don't think grep '^[^ ]*:' really counts as writing a parser.
>
> Best you can do is to run make -n target to probe, but it's possible to write Makefiles that run code even though -n is used, which will defeat such probes at distribution scale. It quickly becomes an exercise in heuristics and output parsing.
The comment I replied to, chaosfactors, specifically specified grepping the makefile. You're arguing that you can find all the targets and I never said you can't. I simply said you can't find them all by only grepping the makefile.
"At distribution scale"... judging by your username, I take it your use case is something like finding out if a source tarball provides the standard targets like "install" and "distcheck", for all packages in debian? I can see how that would be rather painful.
I wonder if there's room for adding a "--dry-run-yes-really-dont-run-any-recipes-at-all" flag to make. As far as I remember the main use case for the current behaviour is (a) to run recipes that create makefiles that are then included into the main makefile, used in build systems that automatically figure out header dependencies and things like that; and (b) process makefiles in subdirectories for "recursive make" build systems. I wonder if you don't do those steps, would you actually miss out on any significant information re. top-level targets? I imagine that use case (a) is used to add prerequisites to existing targets, rather than creating new targets.
Or maybe a more general solution, instead of "--dry-run-yes-really", would be providing access to make's parser as a library, instead of having to parse the output of "make -p". I'm sure IDEs could make good use of that. Who's got time to implement that though.
319 on Mac OS X 10.8
337 on Centos 6
CC=script/send_to_cluster gcc
Then specify the number of jobs you'd like to run simultaneously on your cluster with the "-j" flag.