HNHacker News
TopNewBestAskShowJobs

jkool702

110 karma · joined July 14, 2025

submissionscomments
jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
> I'm curious about the performance is of forkrun "echo ." in a billion jobs vs. say pure C

Short answer: in its fastest mode, forkrun gets very close to the practical dispatch limit for this kind of workload. A tight C loop would still be faster, but at that point you're no longer comparing “parallel job dispatch”—you're comparing raw in-process execution.

Let me try and at least show what kind of performance forkrun gives here. Lets set up 1 billion newlines in a file on a tmpfs

    cd /tmp
    yes $'\n' | head -n 1000000000 > f1
now lets try frun echo

    time { frun echo <f1 >/dev/null; }

    real    0m43.779s
    user    20m3.801s
    sys     0m11.017s
forkrun in its "standard mode" hits about 25 million lines per second running newlines through a no-op (:), and ever so slightly less (23 million lines a second) running them through echo. The vast majority of this time is bash overhead. forkrun breaks up the lines into batches of (up to) 4096 (but for 1 billion lines the average batch size is probably 4095). Then for each batch, a worker-specific data-reading fd is advanced to the correct byte offset where the data starts, and the worker runs

    mapfile -t -n $N -u $fd A    # N is typically 4096 here
    echo "${A[@]}"
The second command (specifically the array expansion into a long list of quoted empty args) is what is taking up the vast majority of the time. frun has a flag (-U) then causes it to replace `"${A[@]}"` with `${A[*]}`, which (in the case of all empty inputs) collapses the long string of quoted empty args into a long list of spaces -> 0 args. This considerably speeds things up when inputs are all empty.

    time { frun -U echo <f1 >/dev/null; }

    real    0m13.295s
    user    6m0.567s
    sys     0m7.267s
And now we are at 75 million lines per second. But we are still largely limited by passing data through bash....which is why forkrun also has a mode (`-s`) where it bypasses bash mapfile + array expansion all together and instead splices (via one of the forkrun loadable builtins) data directly to the stdin of whatever you are parallelizing. If you are parallelizing a bash builtin (where there is no execve cost) forkrun gets REALLY fast.

    time { frun -s : < f1; }

    real    0m0.985s
    user    0m13.894s
    sys     0m12.398s
which means it is delimiter scanning, dynamically batching and distributing (in batches of up to 4096 lines) at a rate of OVER 1 BILLION LIONES A SECOND or at a rate of ~250,000 batches per second.

At that point the bottleneck is basically just delimiter scanning and kernel-level data movement. There’s very little “scheduler overhead” left to remove—whether you write it in bash+C hybrids (like forkrun) or pure C.

jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
So I thought about this for a bit, and this actually doesnt surprise me all that much. This makes sense when you consider the following 2 things:

First, 14k items in batches of 100 are only 140 batches. 140 batches in 160 ms is not even 1000 batches per second. For reference, parallel tops out at around 500 per second (but is dreadfully slow) and forkrun, in its normal "passing quoted arguments via the cmdline" mode, can do about 10000 batches per second. I have no doubt rush is far more capable of distributing batches quicker than parallel, so theres a good chance that "how fast the parallelization engine can distribute work" isnt the main bottleneck for either frun nor rush for this particular workload.

Second, the way frun distributes batches is very efficient but requires setting up a substantial amount of supporting machinery. This puts (on my system) the "no-load run time" of forkrun at about 80 ms.

    time { echo | frun :; }

    real    0m0.078s
    user    0m0.027s
    sys     0m0.064s
And this 80 ms difference is pretty close to the time difference you are seeing. Id bet the "minimum no-load time" for rush is considerably lower - perhaps a couple of ms.

forkrun is optimized for plowing through MASSIVE amount of very fast running inputs...it is capable of plowing through a billion (empty) inputs a second in its fastest mode. 14k inputs just isn't enough to amortize the startup of all the lock-free machinery.

I would venture to guess that if you repeat the same test but with 100x more inputs, the relative difference between frun and rush would be considerably less.

jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
I appreciate the high praise re: forkrun.

forkrun's NUMA approach is really largely based on the idea that, as you said, "real workloads mostly wait for memory". The waiting for memory gets worse in NUMA because accessing memory from a different chiplet or a different socket requires accessing data that is physically farther from the CPU and thus has higher latency. forkrun takes a somewhat unique approach in dealing with this: instead of taking data in, putting it somewhere, and reshuffling it around based on demand, forkrun immediately puts it on the correct numa node's memory when it comes in. This creates a NUMA-striped global data memfd. on NUMA forkrun duplicates most of its machinery (indexer+scanner+worker pool) per node, and each node's machinery is only offered chunks from the global data memfd that are already on node-local memory.

This directly aims to solve (or at least reduce the effect from) "CPUs waiting for memory" on NUMA systems, where the wait (if memory has to cross sockets) can be substantial.

jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
Im happy to hear forkrun is working well for you!

I'll have to look into what would be required to package forkrun for the various distros. I'll try to make it happen in the near-ish future.

jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
I hate to say it but forkrun probably wont work in cygwin. I haven't tried it, but forkrun makes heavy use of linux-only syscalls trhat I suspect arent available in cygwin.

forkrun might work under WSL2, as its my understanding WSL2 runs a full linux kernel in a hypervisor.

jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
forkrun complements things like SLURM (and even MPI). forkrun is intra-node, and is all about utilizing all the resources any given node as efficiently as possible, including when the node has a deep NUMA topology (e.g., it's EPYC-based). This allows SLURM and MPI to focus on inter-node work distribution and coordinating who gets to run things on which node and things like that.

tl;dr: forkrun takes over the "last mile" of actually running things on a given single node, so SLURM can focus on what it does best: efficiently allocating and distributing work to different nodes across the cluster.

jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
whoopsie...

    code_to_parallelize "$nn"
should be

    code_to_parallelize "$nn" &
jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
The difference you are seeing in this specific usage is because frun is dynamically adjusting batch size and worker count (by default it always begins at a batch size of 1 and using 1 worker). It is pretty darn good at dynamically pinning these down pretty quickly, but with only 14k total inputs split you are probably ending up with 2-3 times as many jq calls as you do setting the batch size to 100 inputs from the start, and you may not be fully spawning 32 workers.

If you want an apples-to-apples comparison, try running the following. This tells frun to use 100 lines per batch (-l 100), to use 32 workers (-j 32). Please let me know how this one compares to the rush invocation in terms of runtime.

    printf '%s\n' 0* | frun -l 100 -j 32 -- jq -rf my_program.jq
side note: you should be able to use a space as a delimiter (-d ' ') and run

    echo 0* | frun -l 100 -j 32 -d ' ' -- jq -rf my_program.jq 
NOTE: when I posted this reply using a space as a delimiter was broken. I just pushed a PR to the forkrun main branch that fixes this. If you re-download frun.bash and source it in a new bash instance, then the above space-delimited command should work as well, and is the most direct apples-to-apples comparison to your rush command.
jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
parallel works fine so long as the time per job is on the order of seconds or longer.

Let me give you an example of a "worst-case" scenario for parallel. Start by making a file on a tmpfs with 10 million newlines

    yes $'\n' | head -n 10000000 > /tmp/f1
So, now lets see how long it takes parallel to push all these lines through a no-op. This measures the pure "overhead of distributing 10 million lines in batches". Ill set it to use all my cpu cores (`-j $(nproc)`) and to use multiple lines per batch (`-m`).

    time { parallel -j $(nproc) -m : <f1; }

    real    2m51.062s
    user    2m52.191s
    sys     0m6.800s
Average CPU utalization here (on my 14c/28t i9-7940x) is CPU time / real time

    (172.191 + 6.8) / 171.062 = 1.0463516152 CPUs utalized
Note that there is 1 process that is pegged at 100% usage the entire time that isnt doing any "work" in terms of processing lines - its just distributing lines to workers. If we assume that thread averaged about 0.98 cores utalized, it means that throughout the run it managed to keep around 0.066 out of 28 CPUs saturated with actual work.

Now let's try with frun

    . ./frun.bash
    time { frun : <f1; }

    real    0m0.559s
    user    0m10.409s
    sys     0m0.201s
CPU utilization is

    ( 10.409 + .201 ) / .559 = 18.9803220036 CPUs utalized
Lets compare the wall clock times

    171.062 / 0.559 = 306x speedup
Interestingly, if we look at the ratio of CPU utilization (spent on real work):

    18.9803220036 / 0.066 = 287x more CPU usage doing actual work
which gives a pretty straightforward story - forkrun is 300x faster here because it is utilizing 300x more CPU for actually doing work.

This regime of "high frequency low latency tasks" - millions or billions of tasks that make milliseconds or microseconds each - is the regime where forkrun excels and tools like parallel fall apart.

Side note: if I bump it to 100 million newlines:

    time { frun : <f1; }

    real    0m4.212s
    user    1m52.397s
    sys     0m1.019s
CPU utilization:

    ( 112.397 + 1.019 ) / 4.212 = 26.9268 CPUs utalized
which on a 14c/28t CPU doing no-ops...isnt bad.
jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
> And if you’re objective, what could be done to other tools to make them competitive?

I wanted to reply separately to this bit, because I needed a bit of time to think about and respond to it.

To be frank, parallel optimizes for "breadth of features" and has, for example, the ability to coordinate distributed computing over ssh. But it fundamentally assumes that the workload itself will take dramatically longer than the coordination.

To really be competitive in "high-frequency low-latency workloads", where you have millions of inputs and each only takes microseconds, you would need a complete rewrite with an entirely different way of thinking.

Let me drop a few numbers just to drive this point home. Parallel is capable of batching and distributing around 500 batches of work a second. forkrun, in its "pass arguments via quoted cmdline args" mode is capable of batching and distributing around 10,000 batches a second. This is mostly limited by how fast bash can assemble long strings of quoted arguments to pass via the command line. In forkrun's `-s` mode, which bypasses bash entirely and splices data directly to the stdin of what you are parallelizing, forkrun is capable of batching and distributing over 200,000 batches a second.

The biggest architectural hurdle most existing tools have that makes it impossible to achieve forkrun's batch distribution rate is that almost all use a central distributor thread that forks each individual call (which is very expensive) and that is ALWAYS the bottleneck in high-frequency low-latency workloads. Pushing past this requires moving to a persistent worker model without a central coordinator. This alone necessitates a complete rewrite for basically all the existing tools.

That said, forkrun takes it so much further:

* It uses a SIMD-accelerated delimiter scanner + lock-free async-io to allow for workers to not only execute in parallel but to read inputs to run in parallel.

* It doesn't just use a standard "lock-free" design with CAS retry loops everywhere - it treats the problem like a physical pipeline of data flow and structurally eliminates contention between workers. The literal only "contention" is a single atomic on a single cache line - namely when a worker claims a batch by running `atomic_fetch_add` on a global monotonically increasing index (`read_idx`).

* It doesn't use heuristics - it uses a proper closed-loop control system. There is a 3-stage ramp-up (saturate workers -> geometric ramp -> backpressure-guided PID) to dynamically determine the batch size and the number of workers that works extremely well for all input types with 0 manual tuning.

* It keeps complexity in the slow path. Claiming a batch of lines literally just involves reading a couple shared mmap'ed vars and an `atomic_fetch_add` op in the fast path, which is why it can break 1 billion lines a second. The complexity is all so the slow path degrades gracefully, which is where it smartly trades latency for throughput (but only when throughput is limited by stdin to begin with).

* It treats NUMA as 1st class and chooses the "obvious in hindsight" path to just put data on the correct NUMA node from the very start instead of re-shuffling it between nodes reactively later.

I could go on, but the TL;DR is: to be competitive, other tools would really need to try and solve the "high-frequency low-latency stream parallelization" problem from first principles like forkrun did.

jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
10 years represents going from

    maxJobs=$(nprocs)
    while read -r nn; do
      code_to_parallelize "$nn"
      (( $(jobs -p | wc -l) > maxJobs )) && wait -n
    done < inputs
to a NUMA-Aware Contention-Free Dynamically-Auto-Tuning Bash-Native Streaming Parallelization Engine. I dare say 10 years is about the norm for going from "beginner" to "PhD-level" work.
jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
> I ask because I see multiple em dashes in your description here, and a lot of no X, no Y... notation that Codex seems to be fond of.

I asked a few LLM's for tips on writing the HN post. The post is my own words, but their style may have rubbed off on me a little bit. I'm admittedly better at the technical aspects than I am at "writing good catchy posts that dont turn into 20 pages of technical writing", so...

jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
curl isnt required - you just need to source the `frun.bash` file. Downloading frun.bash and sourcing it works just fine. directly sourcing a curl stream that grabs frun.bash from the github repo is just an alternate approach. It is not "required" by any means.
jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
How did it work for you?
jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
Theres no "install" - you just need to source the `frun.bash` file. Downloading frun.bash and sourcing it works just fine. directly sourcing a curl stream that grabs frun.bash from the git repo is just an alternate approach. It is not "required" by any means.
jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
So, in forkruns development there have been a few "AHA!" moments. Most of them were accompanied by a full re-write (current forkrun is v3).

The 1st AHA, and the basis for the original forkrun, was that you could eliminate a HUGE amount of the overhead of parallelizing things in shell in you use persistent workers and have them run things for you in a loop and distribute data to them. This is why the project is called "forkrun" - its short for "first you FORK, then you RUN".

The 2nd AHA, which spawned forkrun v2, was that you could distribute work without a central coordinator thread (which inevitably becomes the bottleneck). forkrun v2 did this by having 1 process dump data into a tmpfile on a ramdisk, then all the workers read from this file using a shared file descriptor and a lightweight pipe-based lock: write a newline into a shared anonymous pipe, read from pipe to acquire lock, write newline back to pipe to release it. FIFO naturally queues up waiters. This version actually worked really well, but it was a "serial read, parallel execute" design. Furthermore, the time it took to acquire and release a lock meant the design topped out at ~7 million lines per second. Nothing would make it faster, since that was the locking overhead.

The 3rd AHA was that I could make a very fast (SIMD-accellerated) delimiter scanner, post the byte offsets where lines (or batches of lines) started in the global data file, and then workers could claim batches and read data in parallel, making the design fully "parallel read + parallel execute"

The 4th AHA was regarding NUMA. it was "instead of reactively re-shuffling data between nodes, just put it on the right node to begin with". Furthermore, determine the "right node" using real-time backpressure from the nodes with a 3 chunk buffer to ensure the nodes are always fed with data. This one didn't need a rewrite, but is why forkrun scales SO WELL with NUMA.

jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
So...yes, the execve overhead is real. BUT there's still a lot you can accomplish with pure bash builtins (which don't have the execve overhead). And, if you're open to rewriting things (which would probably be required to some extent if you were to make something intended for shell to run in Go) you can port whatever you need to run into a bash builtin and bypass the execve overhead that way. In fact, doing that is EXACTLY what forkrun does, and is a big part of why it is so fast.
jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
So, there are a few reasons why forkrun might work better than this, depending on the situation:

1. if what you want to run is built to be called from a shell (including multi-step shell functions) and not Go. This is the main appeal of forkrun in my opinion - extreme performance without needing to rewrite anything. 2. if you are running on NUMA hardware. Forkrun deals with NUMA hardware remarkably well - it distributes work between nodes almost perfectly with almost 0 cross-node traffic.

jkool702··on Show HN: Forkrun – NUMA-aware shell parallelizer (50×–400× faster than parallel)
Hi HN,

Have you ever run GNU Parallel on a powerful machine just to find one core pegged at 100% while the rest sit mostly idle?

I hit that wall...so I built forkrun.

forkrun is a self-tuning, drop-in replacement for GNU Parallel (and xargs -P) designed for high-frequency, low-latency shell workloads on modern and NUMA hardware (e.g., log processing, text transforms, HPC data prep pipelines).

On my 14-core/28-thread i9-7940x it achieves:

- 200,000+ batch dispatches/sec (vs ~500 for GNU Parallel)

- ~95–99% CPU utilization across all 28 logical cores (vs ~6% for GNU Parallel)

- Typically 50×–400× faster on real high-frequency low-latency workloads (vs GNU Parallel)

These benchmarks are intentionally worst-case (near-zero work per task), where dispatch overhead dominates. This is exactly the regime where GNU Parallel and similar tools struggle — and where forkrun is designed to perform.

A few of the techniques that make this possible:

- Born-local NUMA: stdin is splice()'d into a shared memfd, then pages are placed on the target NUMA node via set_mempolicy(MPOL_BIND) before any worker touches them, making the memfd NUMA-spliced.

- SIMD scanning: per-node indexers use AVX2/NEON to find line boundaries at memory bandwidth and publish byte-offsets and line-counts into per-node lock-free rings.

- Lock-free claiming: workers claim batches with a single atomic_fetch_add — no locks, no CAS retry loops; contention is reduced to a single atomic on one cache line.

- Memory management: a background thread uses fallocate(PUNCH_HOLE) to reclaim space without breaking the logical offset system.

…and that’s just the surface. The implementation uses many additional systems-level techniques (phase-aware tail handling, adaptive batching, early-flush detection, etc.) to eliminate overhead at every stage.

In its fastest (-b) mode (fixed-size batches, minimal processing), it can exceed 1B lines/sec. In typical streaming workloads it's often 50×–400× faster than GNU Parallel.

forkrun ships as a single bash file with an embedded, self-extracting C extension — no Perl, no Python, no install, full native support for parallelizing arbitrary shell functions. The binary is built in public GitHub Actions so you can trace it back to CI (see the GitHub "Blame" on the line containing the base64 embeddings).

- Benchmarking scripts and raw results: https://github.com/jkool702/forkrun/blob/main/BENCHMARKS

- Architecture deep-dive: https://github.com/jkool702/forkrun/blob/main/DOCS

- Repo: https://github.com/jkool702/forkrun

Trying it is literally two commands:

    . frun.bash    # OR  `. <(curl https://raw.githubusercontent.com/jkool702/forkrun/main/frun.bash)`
    frun shell_func_or_cmd < inputs
Happy to answer questions.
jkool702··on Show HN: Timep – A next-gen profiler and flamegraph-generator for bash code
no problem. thanks for helping me discover and fix a bug in timep that all my test cases missed.

Sorry the overhead is too high (relative to cube.bash's insanely low avg command runtime of something like 1 microsecond) to be really useful as a profiler...hopefully itll still prove useful strictly for mapping the code execution/structure and seeing how many times a given function got called and things of that nature.

jkool702··on Show HN: Timep – A next-gen profiler and flamegraph-generator for bash code
> Thanks. I suppose this will depend on each script, as there is another commenter here claiming that the overhead is much higher.

The better way to think about overhead with timep is "average overhead per command run" (or more specifically per debug trap fire). this value wont change all that much between profiling different bash scripts

The code that commenter was profiling was a rubix cube solver that was impressively well optimized: 100% builtins, no forking, all the expensive operations were pre-computed and saved in huge lookup tables (some of which had over 400,000 elements), vars passed by reference to avoid making copies, etc. The overhead from timep was about 230 microseconds per command, but that code was averaging a microsecond or two per command.

To put it in perspective, bash's overhead any time it calls an external binary is 1-2 ms. so in a script that did nothing but call `/bin/true` repeatedly timep's overhead would probably be a little under 20%.

> Re: the binary, that's fine. Your approach is surely easier to use than asking users to compile it themselves, but I would still prefer to have that option.

I mean technically you can, but i'll give you that its not really documented unless you read through all the comments in the code. A makefile is probably doable and would make the process more straightforward.

that said, its on my to do list to figure out how i can setup a github actions workflow to have github automatically build the .so files for all the different architectures whenever timep.c changes. Perhaps that would alleviate your concern.

> After all, how do I know that that binary came from that source code?

I mean, you can say that about virtually any compiled binary. sure , some of them (like from your distro's official repos) have been "signed off" on by someone you trust, but that is a leap of faith you have to make with anything you install from a 3rd party. And, in general, i feel like "compiling it yourself" doesnt really make it safer unless you personally (or someone you trust personally) look through the source to check that it doesnt do anything malicious.

> I wasn't even aware that Bash was so easily extensible.

bash supporting loadable builtins isnt a well known feature. Its really quite handy when you want to do something (e.g., access to a syscall) that bash doesnt support and you cant / dont want to used a external tool for it.

IMO, the biggest hurdle behind using them is that they make scripts much less portable - you either need to setup a distribution system for it (and require the target system has internet access, at least briefly) and/or require the target system has a full build environment and can compile it. which are both sorta crappy options.

unless, of course, you were to base64 encode the binary and directly include it inside of the script. ;)

jkool702··on Show HN: Timep – A next-gen profiler and flamegraph-generator for bash code
It took some time, but I figured out what was causing the errors when profiling cube.bash - the code assigns huge (some >400,000 elements) associative arrays in a single command, and timep was taking the full (several MB) $BASH_COMMAND from those and trying to do stuff with it and bash couldnt handle it.

Try profiling cube.bash with the timep version in the "timep_testing" branch (https://github.com/jkool702/timep/blob/timep_testing/timep.b...). This is my "in development" branch that contains a handful of improvements, one of which is that timep will truncate the BASH_COMMAND at 16kb. On my machine at least that timep version successfully profiles cube.bash.

now - regarding efficiency. timep's overhead is more-or-less constant per command (or, more specifically, per debug trap fire). what "percent overhead" this equates to is entirely dependent on how long the average command being profiled takes to run. And for things like base.cube, that are virtually all builtin commands that dont fork, that time is low. For cube.bash it is really pretty remarkably low.

Looking at the "full" profile and stack trace that timep generated, running cube.bash involved running around 7150 commands. I also profiled a modified version that stops at the ` echo scramble: "$@"` line - that one was about 3350 commands. Meaning the timed part of cube.bash (where it actually solves the cube) represents about 3800 bash commands. This part of the code (when run by timep) took 870 ms or so on my system. which puts the per-command overhead at under 1/4 of 1 ms (about 230 microseconds to be precise).

230 microseconds per command is best-in-class for a trap-based profiler - many take an order of magnitude (or more) longer and dont collect cpu times or the code structure metadata needed to reconstruct the full call stack. To put it in perspective, bash (on my system at least) has 1-2 ms overhead every time it runs an external binary. Your cube.bash is just impressively stupidly fast, so much so that the 230 microseconds still introduces considerable overhead.

jkool702··on Show HN: Timep – A next-gen profiler and flamegraph-generator for bash code
also re: BATS

Im aware of it, but have never ended up actually using it. Ive heard before the sentiment you imply - that its great for fairly simple script...but not so much for long and complicated scripts. And, well, the handful of bash projects that Ive developed enough to have the desire to add unit testing for are all range from "long and complicated" to "nightmare inducing" (lol).

Its on my "to try" list one day, but I have a sneaking suspicion that timep is not a good project to try out BATS for the first time on.

jkool702··on Show HN: Timep – A next-gen profiler and flamegraph-generator for bash code
the LINENO is (mostly) reliable, so long as the code that is running comes from a file somewhere. timep runs the code that you want profiled by generating a script file that:

1. declares all the variables timep uses to track state and things like that 2. initializes all the extra stuff needed to make the instrumentation work (initial values for some variables, defining instrumented traps, setting `set -T`, etc.) 3. runs the code to be profiled. for scripts the script content is added here. For raw commands the commands are added. For functions the function is defined and then called as a function.

this is saved as a script in the timep tmpdir (by default a unique directory under /dev/shm/.timep), made executable, and then it is run. The setup is done like this for a few reasons, but the biggest of those is "to get correct LINENOs.

This makes it so that LINENO gives a correct result, only it is shifted by however many lines it takes to do steps 1 and 2.so the code records the lineno just before the "code to be profiled" starts running, so it can shift it back the right amount when it records it.

That said, timep has a handful of very minor bugs (4 to be precise) - one of them is that fir functions that call a subshell the lineno is wrong (it is shifted forward by a few hundred lines). So it isnt quite 100% perfect, but is pretty close.

(side note: the other 3 bugs have to do with command grouping in the output being ever so slightly off for deeply nested subshell + background fork sequences).

timep also includes a function that I wrote that tries to get the original function code (instead of the version that `declare -f` outputs), so that the lineno's on functions match up better with the original function definition source code.

jkool702··on Show HN: Timep – A next-gen profiler and flamegraph-generator for bash code
Yay, a comment!

>I find it hard to believe that there is minimal overhead from the instrumentation, as the README claims. Surely the additional traps alone introduce an overhead, and tracking the elapsed time of each operation even more so. It would be interesting to know what the actual overhead is.

note that timep does a 2 pass approach - it runs the code with instrumentation just logging the data it needs, and then after the code finishes it goes back and uses that data to generate the profile and flamegraphs. For "overhead" im just talking abut in the initial profiling run...total time to getting results is a bit longer (but still pretty damn fast).

if you run timep with the `-t` flag it will run profiling run of the code (with the trap-based instrumentation enabled and recording) inside of a `time { ... }` block. You can then compare that timed to the running the code without using timep, giving you overhead. I used that method while running a highly parallelized test that, between 28 persistent workers, ran around 67000 individual bash commands (and computed 17.6 million checksums of a bunch of small files in a ramdisk) in 34.5 or so seconds (without timep). using `timep -t` to run the code this increased to 38 seconds. So +10%. And thats really a worst case scenario...it indicates the per command overhead is around 1 ms. what percent if the total run time that translates to depends on the average time-per-command-run for the code you are profiling.

side note: in this case, the total time to generate a profile was ~2.5 minutes and the total time to generate a profile and flamegraphs was ~5 minutes

timep manages to keep it so low because the debug trap instrumentation is 100% bash builtins and doesnt spawn and subshells or forks - so you just have a string of builtin commands with no context switching nor any "copying the environment to fork something" to slow you down.

> The amount of patience required to instrument Bash scripts must've been monumental.

It took literally months to get everything working correctly and all the edge cases worked out properly. Bash does some borderline bizarre things behind the scenes.

> EDIT: I took a look at `timep.bash`, and I may have nightmares tonight... Sweet, fancy Moses.

Its a little bit...involved.

> Also, I really don't like that it downloads and loads some random .so files

So by default it doesnt do this. It has the ability to, but you have to to explicitly tell it to....it wont do it automatically.

By default, it will get the .so file using the base64 blob (in ascii string representation and compressed) that is built into timep.bash. This has sha256 and md5 checksums incorporated into it that get checked when the .so file is re-created using that base64 blob (to ensure no corruption).

the .so file is needed to add a bash loadable builtin that outputs microsecond granularity CPU usage time. Without this it will try and use /proc/stat, which works but the measurement is 10000x more coarse (typically it displays cpu time in number of 10 ms intervals). to get microsecond accuracy you need the .so file.

> Or better yet: make their source visible as well, and have users compile them on their own.

the source is available in the repo at https://github.com/jkool702/timep/blob/main/LIB/LOADABLES/SR...

compile instructions are at the top in commented out lines.

the code is pretty straightforward - it uses getrusage and/or (if available) clock_gettime to get cpu usage for itself and its children.

the `_timep_SETUP` function has the logic included (but not used by default - you have to manually call it with a flag) to let allow you to use a .so file that is in your current directory instead of the one generated from the builtin base64 blob. So, you are able, should you wish, to compile and have timep use your own .so for the loadable.

However, realistically, most people who might be interested in using timep wont do that. So it defaults to fully automating this and having it all self contained in a single script file, while allowing for advanced users to manually override it. I thought that was the best overall way to do it.

jkool702··on Show HN: Timep – a next-gen profiler and flamegraph-generator for bash code
Im currently working on adding the ability to record user/sys cpu time (in addition to wall-clock time) to timep. This will be coupled with a new modification to the timep_flamegraph.pl script that will control the flamegraph coloring saturation based on cpu time (sys+usr) - lower cpu time will be more faded and higher cpu time will be more vivid/saturated.

those interested can see the current progress in the "timep_testing" branch of the github repo. The trap timing instrumentation and the new flamegraph script are both done, but post-processing the new cpu times is a work-in-progress. The new flamegraph script is backwards-compatible with the current timep, so you can tryout that part now if you want.