Hi, it's my project. Feel free to ask me anything.
Hi, it's my project. Feel free to ask me anything.
(I love the name but am always curious what people's inspirations are for naming their project. I love when a project name is unique, creative, descriptive, and playful, and Bytehound nails all four IMHO)
Also, thank you for doing this and sharing it!
Imagine a set of workers that ingest data in parallel, would that work?
Currently it's pretty simple and i am spawning a process within the worker so it reads some stuff such as memory usage, cpu usage etc... But I would like to improve it.
It could probably be done, but the analyzer would have to be mostly rewritten. (Which I currently have no plans to do.)
If it's a native program - possibly, depending on whether it's possible to LD_PRELOAD on Android, but you'd most likely have to connect to it through SSH and launch your program that way.
(Sorry, I have very little experience with Android so I can't really be of too much help here)
Although I could probably recommend you a hands-on project/exercise to do:
- Write a simple memory allocator in C/C++/Rust/Zig/any similar systems language using raw `mmap` and `munmap` syscalls (run "man 2 mmap" in your terminal for details). This is how fundamentally almost every program allocates memory on the lowest level (with some exceptions, but I'm not going to get into that).
- Allocate a bunch of memory with your allocator without actually reading/writing from than memory and check the program's RSS, then write to it and check the RSS again. Try allocating more memory than you have RAM and see if it works. Run the program under `perf` and check the page faults counter - see how the page fault number changes if you a) never write to the memory you allocated, b) only write to a single byte per page, c) write to every byte you allocated, d) write twice to every byte you allocated.
- Play around with the `madvise` syscall ("man 2 madvise"), in particular with `MADV_DONTNEED`.
- Try `mmap`ing on a file on your disk. `mmap` it with multiple processes at the same time.
Is it correct to say memory behaviour of a program entirely depends on what/how the program allocates and its memory access pattern? Is there any hello world memory allocator out there?
File system moves blocks to memory, and kernel access them as pages, and kernel also moves them back to disk(using file system) depending on memory pressure. Given this is ture, I am assuming, when studying memory behaviour, virtual memory setup is good enough to ignore file system level details initially - but at what point these details become important? For example, just as cache miss, I am assuming, it is also very expensive for file system(paired with underlying physical medium) to go out and gather non-contiguous blocks and put them together to respond to a file access request.
Mostly yes, but also on the rest of the system.
> Is there any hello world memory allocator out there?
Yes. That's a naive mmap-based allocator. (: It's terribly inefficient (slow and wastes a ton of memory) but you could in theory hook it to any program and it will work.
> virtual memory setup is good enough to ignore file system level details initially - but at what point these details become important?
For the basics you can ignore swapping. In general you can probably ignore it altogether (RAM is cheap), unless you're studying memory mapped I/O.
> I am assuming, it is also very expensive for file system(paired with underlying physical medium) to go out and gather non-contiguous blocks and put them together to respond to a file access request.
Data in memory is accessed in pages (so reading one byte will bring it the whole page). Data on disk is also stored in pages (although they might be a different size than the page size of your CPU). So what happens is roughly that the kernel will read a page you've hit from the disk and map it into your memory, and then probably read more data in the background while giving you the control back.
Frame-pointer-based, I imagine?
The main two tricks are: it preprocesses all of the DWARF info at startup for faster lookups, and it dynamically patches the return addresses of functions on the stack injecting an address to its own trampoline, which allows it to skip going through the whole stack trace every time it needs to dump a backtrace. For example, if you're running a function nested 100 stack frames deep and that function calls malloc 100 times then Bytehound will only go through ~300 stack frames in total (~100 times for the first call then only ~2 frames for each successive call, if my math is right), while other similar tools will go through 10000 stack frames (going through all ~100 frames to the very bottom for every call).
TP 5.0 from 1988 was the first version that had it.
The idea was to make sure the code the CPU returned to would actually be in memory.
I'm pretty sure Windows 1.0 did something very similar.
I'm curious if you've done any benchmarking for your implementation as well?
Nice!
> I'm curious if you've done any benchmarking for your implementation as well?
Not in any detail; I just checked that it's significantly faster than doing it naively and left it at that since it was fast enough for my use case.
[1] https://github.com/koute/not-perf/blob/master/nwind/src/loca...
Also nice use of Gimli - did something similar to make creating stack traces on crash cheaper to symbolicate.
For performance profiling I find that `perf`-like sampling profiling works well enough to find the hot spots, and then Valgrind's Callgrind is great for micro-optimizing the hot spots code on the assembly level.
Of course, it would be cool to have a unified memory + performance analysis tool like this, but I don't think I can justify the time investment to write one in my spare time.
Yeah, I'm really happy that Gimli exists, considering the absolute insanity/complexity pit of DWARF.
However you then still need debug symbols of some kind to convert those to names.