Copy-on-write friendly Python garbage collection
engineering.instagram.com
engineering.instagram.com
As a related metric, sometimes it’s not your average response times and scale that matters but how much headroom you’ve got for a sudden influx of traffic. If memory/pagefile thrashing and hard faults bring your system screeching to a halt, you won’t be happy no matter your scale when you can’t benefit from the eyeballs and pageclicks coming your way.
Cache friendliness is a double-edged sword. There are plenty of use cases where a GC can be more cache-friendly than manual memory management, in particular if you have a generational GC with a bump allocator.
[1] What makes life (considerably) harder is when you have multiple threads sharing a heap, but that's not at issue here.
OCaml has had one since the 1990s [1]; it's a fairly standard generational, compacting collector with incremental collection for the old generation. Lua has had an incremental GC since version 5.1 (in 2006); as it's frequently used as a scripting language for video games, it's safe to assume that pause times aren't much of an issue.
The problem is that virtually every language since then has pretty much decided to have all threads use a global shared heap. Once you do that, you run into all kinds of challenges, such as root discovery from thread stacks without stopping the world. That said, there are plenty of languages that have this option, anyway; it's simply more challenging, not impossible.
Languages that are single-threaded maintain thread-local heaps don't have the problem. Python and Ruby (unlike Lua) have issues for historical reasons (they started out with basic reference counting and mark-and-sweep collection, respectively, and then had to maintain backwards compatibility [2]).
Intermediate designs (having both thread-local heaps and shared heaps at the same time) are also possible, but that design space hasn't been explored much.
[1] http://prl.ccs.neu.edu/blog/2016/05/24/measuring-gc-latencie...
[2] I think that in principle one could make the cycle detector in Python incremental (it's basically a form of trial deletion); Ruby eventually got an incremental GC for its major generations in 2.2, but I believe there are still some inherent limitations due to the lack of write barriers in C code.
It's not. The cases where it happened to me during the last 10 years and dozens of web projects spanning on multiple frameworks were generally me doing something stupid:
- using a mutable object as default value or in a class attribute
- doing something in __del__
- keeping DEBUG=True for Django (which is very well known for causing memory leaks)
> Is this just life now?
Nope. I have big Django apps running (like 500k users/day streaming video sites) and it doesn't happen.
> Are there any good tools out there for figuring out what part of the application keeps requesting objects that don’t get destroyed?
Yes for course: http://tech.labs.oliverwyman.com/blog/2008/11/14/tracing-pyt...
This does cause memory use to increase with the number of queries issued, but a debug log is not really what I'd call a leak (and genuinely leaking memory in Python takes some effort).
I’m not a pythonista, so please forgive my ignorance here, but: is this just the default configuration for Django? Does it ship out-of-the-box with adapters to log that to MySQL or similar instead? Why does it not log the last n connections for some sane (and configurable) default value of n instead?
It does that only when DEBUG=True, which is not the recommended production configuration.
That is what the "debug=True" variable does. Sets things up for running in a minimal development environment. Another feature? It restarts the server when any source files change.
Suffice to say, it's really not meant to be run in anything like production.
https://docs.djangoproject.com/en/dev/ref/settings/#debug
>Never deploy a site into production with DEBUG turned on.
>Did you catch that? NEVER deploy a site into production with DEBUG turned on.
>Still, note that there are always going to be sections of your debug output that are inappropriate for public consumption. File paths, configuration options and the like all give attackers extra information about your server.
>It is also important to remember that when running with DEBUG turned on, Django will remember every SQL query it executes. This is useful when you’re debugging, but it’ll rapidly consume memory on a production server.
You mentioned SQLite - a second “querylog.db” with a table of queries and a table of instances (query time, parameters, whatever) would be an obvious option (at the expense of slightly exaggerating your db response times, but I’m sure debug mode is already slower as is).
The reason it logs those queries is just for debugging. There are a lot of different ways they could do it, but they don't matter because you can just turn that feature off, or set it up to log in a different way, or whatever else.
Logging to a DB would be worse, because then it stores a bunch of useless queries that all look very similar. That logging is used for debugging, and wouldn't be something you'd want to use in production.
Why hasn't anyone already done this? Because it's trivia, not an actual problem that needs fixing.
Also objgraph: https://www.darkcoding.net/software/finding-memory-leaks-in-...
If you are lucky to use Python 3.6+, you got tracemalloc: https://docs.python.org/3/library/tracemalloc.html
Mind sharing your stack? I'm currently torn between Laravel and Django. It won't be the first time I've used an MVC framework (CodeIgniter, some Rails) but I like working with Python. However, Laravel also seems to be heavily used. I don't have anything against PHP but wanting to compound my familiarity with Python is one big reason for leaning towards Django...
Sometime a bit of crossbar.io.
Laravel is a nice framework. I prefer it to symfony, espacially since they love Vue.js which is now my fav front end lib.
But the definitive advantage of using Python is that you get a much more versatile language than PHP. Python is not just good for the Web. It's heavily used for scripting, automation, sysadmin, data analysis, GUI, pentesting, etc.
People don't just use Python with HTTP in mind. You find it at Apple, NASA, at Sony, Fedora, in the state of Geneva, in french schools, in 3D (Maya, Blender) or geography (most GIS)...
Basically, you get a tool that you can use for a lot more stuff, because the ecosystem has been developed for a lot more types of tasks.
> gunicorn
Any particular reason you use it over uwsgi? I last did real research in 2013 and have not revisited since then.
Python noob here. Can you expand on that a little? Or perhaps link a post? I can't think of an example or how it'd impact memory
The following is some pretty good reading on the matter, and the example given is also one that would lead to a memory leak:
https://stackoverflow.com/questions/1132941/least-astonishme...
Dropbox implemented a custom allocator with different types and sizes getting their own arenas, leading to long term memory stability. Even controlling all their own code, processing millions of file paths, metadata, and buffers inevitably lead to memory leaks.
I can't talk about general situation but in this specific case it's because they've turned off GC previously, so growing infinitely was expected. Details are in the first blog post they've linked from the OP.
I watched the RubyConf AU version: https://www.youtube.com/watch?v=nAEt36XNtAE&t=2482s
But looks like there was a version at Rubyconf as well that may or may not be better/different: https://www.youtube.com/watch?v=8Q7M513vewk
On the other hand, I am disappointed that this change goes into upstream Python instead of properly solving the problem by making reference count implementation really CoW friendly so that all applications would benefit from it without the need for a careful use of a special function.
This change allows you to run more workers with less memory.
Instagram has had a few good posts about how they've approached this problem, here's another: https://engineering.instagram.com/dismissing-python-garbage-...
Wut ?
When I have performances problems, my RAM is full with varnish cache, redis stored objects and postgres buffer.
Python is like, low, low on the list.
Again, most people are NOT instagram or Google. They have a very atypical load.
And sure, most people aren't instagram. This is a python-at-big-scale problem, when the cost benefits of fitting more requests onto fewer servers actually matter.
One of the biggest reasons this actually makes a difference is that, in SoA, internal API scaling against your monolith can become expensive.
And, again, can we stop with these pointless comparisons? Past 1 front end server the only cost that matters is how many requests your one node can handle and how much that one node costs. If you’re not a VPS and you have your frontend http cache correctly configured, then it doesn’t matter how much smaller than Google you are, comparisons are valid (although, of course, the fewer servers you have the more you can afford to spend on them; though you probably aren’t making as much money as IG/FB/Google either...).
I call it write on assign, WOA. A normal GC leaves it alone and does not force fresh memory maps for these.
When forking a process and performing lots of allocations in the parent process (e.g. lots of imports & objects creation before forking) and doing so on a CoW-forking OS (but I'm guessing that's most modern POSIX systems).
It's doesn't have to be. You can use asyncio and deal with many requests at a time by one process.
> I love python, but it is not suitable for backend work at scale.
What's at scale ?
Most projects I work on IRL, including in banks and administrations, never reach the level of scale this even remotely a problem. Too many people think they are Google.
Bank sites are not youtube.
Besides, operations on your bank account don't even hit the Python backend, but a dedicated system. Usually some COBOL dinosaure they froze, wrapped into a Java service and exposed through a RESTish API so that the rest of their system can use it without ever having to touch it again.