Python consumes a lot of memory – how to reduce the size of objects?
habr.com
habr.com
Those are in package `array` and they are a very barebones version of numpy's ndarray - with just one (explicit) dimension and no overloaded operators. But if you just want to keep a bunch of numbers in a contiguous array, they can save you tons of memory.
(I know the purpose of the article is to describe a more complex data structure, but arrays can still get you very far.)
It just goes to prove that you can indeed write FORTRAN in any language!
Here are my code and results:
https://gist.github.com/danudey/94b5442b617a734a2366e824be15...
Is there an approach that you use that produces better results than this? Because my naive approach is unmanageably worse for both performance (not recorded) and impact on Python's memory usage.
The people who lament Java's memory usage never complain about Python or Ruby because... if you're worried about the JVM you would never consider Python!
In my experience many dev think the scripting language heritage means these languages have smaller/lighter runtimes
Sure, memory consumption and CPU perf are both pretty bad in Python, but latency and memory footprint of the runtime itself are pretty good, so it’s ideal for tooling, crons, lambdas etc. JVM is comparatively efficient once you have the JVM ready, but that sure takes tons of resources.
I’m hoping Graal changes this. I don’t like JVM-based languages, but I do like technological progress.
- https://mail.python.org/pipermail/python-dev/2014-May/134528...
- https://mail.python.org/pipermail/python-dev/2018-May/153296...
If your script is expected to be run interactively on a frequent basis, then Go, Rust, C++, or even Bash (for simple stuff) will give you much lower user-perceived latency than Python.
JVM starts really quick (~100ms on my machine) and doesn't use much resources as long as your app is small. that's... Uncommon in Java land though. Even simply apps pull in Guava/Apache Commons and a few client libraries. This can easily be thousands of classes. Nobody thinks about it because runtime cost for loading shitloads of code is so low. But you can improve this a ton by using ProGaurd and stripping out stuff you don't need
If I can do it in numpy, I can probably get a lighter, faster implementation with Python than in Java. More generally, if I can do it with a Python package that's actually a fairly thin wrapper around a C, C++ or Fortran library, then Python also has a decent chance of being the easy winner.
If none of those situations apply, then yeah, typically Java ends up being more efficient.
(I've never successfully used Node so I have no idea what that's like in memory use.)
Depends on the algorithms that use that code. Two separate arrays is not a bad way to represent a vector-of-tuples. And so if you think of everything as a vectorised operation (and if that makes sense for your use case) then the code can be clean enough.
Efficiency-wise it depends on how the locality of reference falls out. There are cases where it's more cache-efficient to store the pairs next to each other, but again for big vectorised operations you might lose nothing by using two cache-ways instead of one.
I would say that Python is the right tool for some things, not for others. Why not use Go, or Nim or C to store/access/return these data in a data structure that transforms the values from/to python?
Python is a great high-level wrapper around low-level programs (c.f. SciPy)
The PyPy site links to this blog post (from 10 years ago) with some info: https://morepypy.blogspot.com/2009/10/gc-improvements.html
And a quick search found this relatively recent post that does some measurement of it: https://dev.nextthought.com/blog/2018/08/cpython-vs-pypy-mem...
So don't use a lot of objects. In C# you use a struct and arrays of them to avoid creating an object on the heap per array entry. In java you have to resort to SoA, in Python it's the same. Just because you have objects doesn't mean everything should be an object. Even languages that follow an "everything is an object"-design usually have an escape hatch such as plain value arrays.
The fact that it's unboxed is an implementation detail.
https://docs.microsoft.com/en-us/dotnet/api/system.int32?vie...
For efficient access, the array must be a consecutive array of primitives. This is the case in both C# and java for integers. In C# it’s also the case for an array of Vector2 with 2 primitives each, which isn’t the case in java.
My point is this: avoid heap allocating many things in collections. They must be raw (primitive, consecutive) data without per instance overhead and of course without heap alloc/GC cost.
A small service that landed in my lap needed to read(only) a data source that was roughly 10k lines of yaml. No way that was going to be in any way efficient so I asked for suggestions. All the work-a-day devs(I am not) instantly said the same thing without a single real thought about it: Make it a database duh!
Long story slightly less long, Loading up the libraries to interface with a database ate between 2 and 3 times the memory(depending on the DB and lib) that simply loading the entire 10k line yml ate and offered slower performance and required more code.
SQLite was pretty darn close but in the interest of saving developers from themselves vis a vis parameterized queries, or the need to queries all together for that matter, increased the required code for zero benefit.
The service still hums along with a 10k line yaml in memory. "Worse is better" indeed.
It can be hard to beat an in memory data structure when your data set is small enough, true.
However, "better" is dependent on the situation. If the data is unlikely to change, stay the course. If not, then while a database might be more maintainable over the long term, even if not as efficient.
The Cython and Numpy cases directly store the actual data and this has the larger effect to reduce memory.
To the best of my knowledge, Python isn't one such language with a standard, is it?
https://en.m.wikipedia.org/wiki/X32_ABI
As effectively many Python workloads are usually object oriented business logic and objects are mostly pointers, setting up x32 user space "halved" the memory usage. It also made execution performance faster because of better CPU cache utilisation.
Sadly, x32 was "very custom" and very hard to support. Last I heard x32 is being phased out from Linux kernel.
> The best results during testing were with the 181.mcf SPEC CPU 2000 benchmark, in which the x32 ABI version was 40% faster than the x86-64 version.
Does 64-bit instruction set provide some segments or functionality form this? How about "native" pointers coming from glib and such?
If there has to be base + offset translation on every pointer access it is way too slow.
I would also assume JavaScript VMs in browsers would be already utilising this, as web page workloads are not gigabytes (hopefully).
It does do this, but it's not too slow - the overhead of the translation is lower than the benefit of reduced memory transfer, increased cache space, etc. Obviously - otherwise people wouldn't be doing it.
The memory access instructions in the x86 ISAs can do base + offset in a single instruction.
It's time to update your intuition: memory access is usually slow, much slower than simple arithmetic.
I read some longer article about the evolution and options for memory addressing in JVM, if only I could find it... but here's another one: https://wiki.openjdk.java.net/display/HotSpot/CompressedOops
There are variations on this method that can be applied to other memory management pools. For example, let's say are allocating many objects for processing, need them only within a certain span of time and have to free all of them when it's over. Well, rather than individually tracking allocations you can allow a single perhaps growing memory area, store the start offset somewhere and reference them with a shorter reference relative to that area.
> Sadly, x32 was "very custom" and very hard to support. Last I heard x32 is being phased out from Linux kernel.
This is really sad to hear. Do you happen to know of any kernel discussion threads on this besides this one[1]? (I'd love to chime in, in support of the X32 ABI.) I'm a little surprised honesty, as Linus has mentioned multiple times how important it is to him that the user-facing ABI remain stable.
Second Edition of High Performance Python in the works
I'm very pleased to say that Micha and I have started work on a Second Edition of High Performance Python with O'Reilly, planned for early 2020. This book will use Python 3.7 and will add tools that barely existed 4 years ago including Dask and probably some Tensorflow, amongst lots of other goodies. I'll let you know how the book progresses via this list. So far I've updated the Profiling chapter, added some advice on 'being a highly performant developer' and have rebuilt the Cython/Numba/PyPy code for the Compiling chapter.
If you have to read everything into memory, because you're doing some sort of transforms like a list to a dict or such, explicitly deleting variables can help, as well, rather than waiting for them to go out of scope, but this should be pretty rare.
Python is good enough 90%. You get faster 2x 10x, etc code by picking better algorythms or solutions. In Python, you should not be caring about 10% or 30% speed improvement. It's not worth it, It's not Python's strength.
When you need faster you go to C based libs (the python libs that need to go fast are already in C), https://en.wikipedia.org/wiki/Cython https://en.wikipedia.org/wiki/Numba etc.
[1] https://www.pypy.org [2] https://www.graalvm.org/docs/reference-manual/languages/pyth...
There are other approaches which can work too, like compiling C to LLVM IR using Clang so a function can be inlined by an LLVM based JIT at runtime.
What is the downside to the __slots__: trick? In what cases do you need the dict and weakref?
When I last did it, I was wrapping a C++ API that I needed to compare against another dataset for a merge. The C++ API (which, admittedly, I also wrote), didn't have equality or hash operators defined on the Python objects. So, I monkey-patched them in for the keys I needed. It was actually the most elegant solution I could come up with as I could then naturally use the objects in sets and easily get differences to update the appropriate dataset.
As an aside, when I wrote the python wrapper around my C++ API, I purposely didn't define equals and hash operators for the objects, despite having full information to do so on the natural keys, because I wanted the flexibility to do the monkey-patching and override how the objects were compared depending upon circumstance.
A lot of it is due to the standard patterns for small Flask apps (e.g. instantiating your app directly at the module level rather than in a factory function, putting all of your business logic into ORM classes) causing increased complexity and unintended side effects down the line. Flask apps also struggle quite a bit with performance, and the only way to do things asynchronously is with something like celery.
For Python, aiohttp is also a really interesting web framework, and the asynchronicity should be easy to grok for JS devs.
Personally, as someone who has worked with Python/Flask professionally for years, I’d pick TypeScript and any JS web framework any day!
I present to you, Quart[0], a flask-like web framework using Python 3.7's asyncio framework.
It works pretty well in the single production application I have tested it in.
The biggest one I run into is that __slots__ does not play well with multi-inheritance:
https://stackoverflow.com/questions/472000/usage-of-slots
Also, slots somewhat contributes to obfuscation since slotted objects don't work in the exact same ways as regular objects do so developers may run into weird stack traces they might not immediately recognize.
Truth.
To me anyway, debugging code is massively harder than writing it in the first place. Nothing worse than trying to fish out some odd issue that plagues you after spending two hours writing some code, and then 4 hours sorting it out.
1> (defstruct blank ())
#<struct-type blank>
2> (pprof (new blank))
malloc bytes: 16
gc heap bytes: 32
total: 48
milliseconds: 0
#S(blank)
32 bit: 2> (pprof (new blank))
malloc bytes: 8
gc heap bytes: 16
total: 24
milliseconds: 0
#S(blank)
The structure instance has a pointer to its type, followed by a numeric ID (which is also in the type, but is forwarded to the instance for faster access). The ID is combined with a slot symbol to perform a cache lookup to get the offset of a slot. The numeric ID is a fixnum, which leaves a few spare bits for a couple of flags: struct struct_inst {
struct struct_type *type;
cnum id : sizeof (cnum) * CHAR_BIT - TAG_SHIFT;
unsigned lazy : 1;
unsigned dirty : 1;
val slot[1];
};
If someone wanted to shrink this, they could patch the code so that the dirty flag support, and lazy instantiation of structs is made optional (as in compiled out), and so is the forwarding of the inst->type->id to inst->id. This struct type is not known outside of struct.c, which is only some 1700 lines long. If you take out the id member from struct_inst, the C compiler will find all the places that have to be fixed up; literally a 15 minute job.I can also think of a more substantial refactoring that would eliminate the type pointer also. There is no room in the heap object handle to store it directly: heap handles have four words in them; the COBJ ones used for structs have the type field, a class symbol, a pointer to an operations structure, and a pointer to some associated object (in this case struct_inst). All structures share the same operations structure. However, if we dynamically allocate that operations structure for each struct type, we could stick the type pointer in there, taking it out of the instance. Thus an instance size could literally just be sizeof(pointer) * no-of-instance-slots.
Running Mailman it will not run in 1GB of ram (the US$5 VPS) so I had to give it 2GB. 2GB? For a mail list server?
Had the same problem running motioneye. It will not run on a Raspberry PI Zero. It is advertised to run on a Zero (Apparently I could squeeze it on by doing some Python magic....)
What total crap is Python! What waste, what hubris, what technical failure!
I suspect the problem is actually Django - what hubris! NIH! lighttpd is running sweetly, with some rust templating in a acceptable fraction of the memory.
No, what he actually bought for his $60/year was license to complain about mailman & python :)
What is a waste is writing a poor quality greedy web server when Apache/Lighttpd/Nginx all exist and are much much better.
What really cost me was the two weeks I spent trying to make motioneye (python) and motioneyeos (not) work as advertised. As advertised! It is not only that the python is too bloated that stops it working, it does but there are other reasons. Very clearly no body has really tested it. It is a symptom of python culture. Shoddy and arrogant.
I feel terribly burnt by it. Twice. So my stack of desire in this domain is quite small and simple: Never to hear again from python! A very personal POV, be happy in your python hacking. I saw the opportunity to have a public rant and I took it!
> ...which received a rating of [stackoverflow] (https://stackoverflow.com/questions/29290359/existence-of-mu... -python / 29419745). In additioobjects liken, it can be used to reduce the size of objects in RAM...
Advice in many of the comments about other languages sounds so tone deaf. The question is not about using less memory. It’s about getting _exactly_ the feature set of Python while using less memory.