Run Python Applications Efficiently with malloc_trim
reliability.substack.com
reliability.substack.com
And even if you return memory to the OS, that doesn't actually solve too-high memory usage.
Some things you can do:
The article mentions `__slots__` for reducing object memory use, and other approaches include just having fewer objects: for example, a dict of lists uses far less memory than a list of dicts with repeating fields. And you can also in many cases switch to a dataframe with Pandas, saving even more memory (https://pythonspeed.com/articles/python-object-memory/ covers all of those).
For numeric data, a NumPy array gets rid of the per-integer overhead for Python, so a Python list of numbers use way more memory than an equivalent NumPy array (https://pythonspeed.com/articles/python-integers-memory/).
https://thehftguy.com/2020/09/25/how-to-deduplicate-string-o...
They were suffering from GC pauses etc. on their ingestion hosts. They spent ages experimenting with various ways of tweaking garbage collection in Java, using different GCs, tweaking settings, etc. They even spent time experimenting with manual memory allocation, but found that to be extremely painful and somewhat fragile.
In the end they found that all they really needed to do was just produce less garbage, which was their ultimate "well duh!" revelation. They spent time looking at what was actually producing the garbage, and how they could avoid it. Got rid of a lot of standard coding patterns from their code in favour of new patterns that reduced object allocation, and away went all their problems.
(Full disclosure I’m the author of the project)
But I found the README quite confusing. I guess it's a new project so that's understandable.
Why use DataClassFrames and not pandas? Because it's statically typed? The comparison table puts pandas and dataclassframes on equal footing, and the rest of the README doesn't make much sense to me.
[1] and I thought about putting them in my shell project: http://www.oilshell.org/blog/2018/11/30.html
Although QTSV is probably in the more immediate future: https://news.ycombinator.com/item?id=25022836
We run a linter that confirms with the user before stripping them, which works out well in practice.
https://github.com/hakavlad/prelockd
https://github.com/hakavlad/memavaild
It can greatly improve responsiveness. Demo:
https://youtu.be/QquulJ06dAo - playing supertux + 12 `tail /dev/zero` in background
https://youtu.be/DsXEWvq60Rw - `tail /dev/zero`, swap on HDD, memavaild, no freezes
What about zram and zswap advantages?
BTW you're helping making the world to be a better place but only nerds download those tools. It would help even more if you could lobby such tools as default in distros such as manjaro, arch, Ubuntu, fedora, etc
I'd be pretty surprised if `malloc_trim` had a significant effect on cpython memory usage as most python memory gets allocated in 256KiB "arenas", which, what with fragmentation, are unlikely to _ever_ be reclaimed.
On the other hand, the article dismisses threads with some vagueries around the GIL, suggesting people need to reach straight for processes if they're serious. Really, unless your code has almost no I/O or C-accelerated, GIL-less sections, if you're not using both threads and processes, you're just burning memory unnecessarily.
(edit: oh and then there's async but I'm a bit old-fashioned for that)
See for example https://www.joyfulbikeshedding.com/blog/2019-03-14-what-caus... . In the end I think the consensus was that jemalloc was just better than invocing malloc_trim, but invoking malloc_trim now and then can certainly be a lot better than using neither malloc_trim or jemalloc.
The idea is to tweak when `mmap` or `malloc` are used by the Python interpreter. One allows memory to be released to the OS right away, whereas the other is not.
It is a useful trick if your application is generating lots of small objects.
Do you have some pointers about arenas always using mmap? I'd like to know how that trick can work if that was the case.
> For example, GC pauses are notorious in other managed memory languages like Java.
(includes link to 2017 article:
https://dzone.com/articles/how-to-reduce-long-gc-pause
)
The author might like to know that since 2017, Java has two of the most sophisticated Garbage Collectors on the planet.
1. Red Hat's Shenandoah GC
2. Oracle's ZGC
(I could also mention Azul Systems work on its CCCC)
Both Shenandoah and ZGC claim to run on multi terabyte heaps with 1 ms max pause times.
Adaptive process and memory management for Python web servers https://instagram-engineering.com/adaptive-process-and-memor...
Previously we used two per-worker thresholds to control respawn: reload-on-rss and evil-reload-on-rss.
...
However, since worker respawn is expensive (i.e. there are warm-up costs, LRU cache repopulation, etc..)
For a smaller scale, I always wondered if FastCGI being the "default" would have saved a lot of headaches. Your workers just get recycled automatically all the time.
If you can make your startup fast enough (and I think most apps can), then you can just let the OS do its job. Although it's true that Python can be really slow to start if you import many modules...
glibc SONAME is different on some architectures.
You can use ctypes.CDLL(None) instead, which should work everywhere.