Poop: Performance Optimizer Observation Platform
github.com
github.com
Memory layout of your code has such a big impact on performance on modern computers that measuring performance without removing that variable leads to wild goose chases, where you think you improved something, but in reality you incidentally got the compiler to move the code around a bit.
Emery Berger has an excelent talk on this [1], and a causal profiler that they developed called Coz[2].
Branch miss-predicts might be somewhat invariant to that, but still, one of the main points of the talk is that people do too much eyeball statistics, mistaking the variance of the underlying stochastic process with actual signal.
Coz is pretty trivial to set up with Zig too.
I've been working on algorithm benchmarking software, and I observed that small changes to algorithm code, even in parts that would not execute, would have a large effect on execution speed.
Even better they seem to have a solution to the problem. Going to explore further to see how we can use this technique.
If you're curious about the poop-making process, I created it and released it all on one 6 hour twitch stream [1].
Watched a few mins and bookmarked it for another day. There were many other name suggestions that are good but I'm glad you kept the initial name.
struct perf_event_attr attr = {
.type = PERF_TYPE_HARDWARE,
.size = sizeof(struct perf_event_attr),
.config = PERF_COUNT_HW_INSTRUCTIONS,
.disabled = 1, /* only need to read once, not updating values, but fails with 0 also */
.exclude_kernel = 1,
.exclude_hv = 1
};
int fd = perf_event_open(&attr, 0, -1, -1, 0); /* fails, returns -1 */
Microsoft's kernel appears to have the right CONFIG_* as well, so I feel like it must be supported in some capacity, even if limited.Imagine you are struggling with performance of some tool and a more knowledgeable senior comes in and says "you need to use poop" or "just poop it!".
> Woob (Web Outside of Browsers) is a library which provides a Python standardized API and data models to access websites.
Interesting idea -- basically an API on top of a passel of site-specific web scraper modules.
> "The MIT License (Expat)
Copyright (c) Poop Contributors
Permission is hereby granted..."To be clear there are definitely a few ideoms I'd need to learn before considering this readable (it ain't psudocode), but I generally see all the building blocks and recognize what they're achiving. Maybe I'd have put them together slightly differently, but that could be equally true in Python. I'd rather learn this next than C++.
The density of how it's been written reminds me of some lisp code I've read, some implementations of state machines can be hard to follow too (in the classic "but where does it do the thing?" sense). I attribute a lot of the stuff that I don't follow to the lower-level C influence; I'm working for the first time with memory managment, structs, static types, pointers, bitwise opperations, etc.
Think I made my point in there somewhere >z>
Would love to know where you're coming at it from, I'd imagine that the C or 'kinda functional control flow' aspects might explain what looks most forign to you too (unless you have broader experience with low level languages)?
Style-wise, looks fine. Code split up into functions where it makes sense, inlined where it doesn't. I've seen worse - for example, code that takes the equivalent of that, and splits it into 100 functions across 10+ files.
This so many times. Will never understand why people are so afraid of multi thousand lines of properly structured code but will happily giggle when the same structure is split across 10-20 files. And no, it’s not for reuse sake.
Encapsulation and modularity are good engineering practices. Personally, I will not contribute to any project that opposes such things
Where did I say anything about that?