You Should Compile Your Python and Here’s Why
glyph.twistedmatrix.com
glyph.twistedmatrix.com
At least this is honest.
No matter what the script does, no matter how fast the libraries, the Python intepreter has a slow startup time.
On multiple occasions I have seen people commenting on HN argue that Python is not slow.
I think for these commenters Python is "fast enough".
For others, like me, it may not be "fast enough".
IOW, the question is not whether Python is "fast" or "slow", but whether it is "fast enough" or "too slow" for a given user.
For some users, it might not be fast enough. And that's fine.
(Not a card carrying Python fan or anything, haven't used it in a decade, I just like this architecture.)
> IOW, the question is not whether Python is "fast" or "slow", but whether it is "fast enough" or "too slow" for a given user.
100% this. There isn't REALLY any such thing as good/bad/fast/slow/cheap/expensive, there's just "meets spec better" / "doesn't meet spec as well". Everything is relative to your requirements.
Quite a lot of programs are almost entirely bopping along from one OS or numpy call to the next, with a tiny bit of business logic thrown in. You've got a better chance of making your users happy by thinking about that business logic at a high level, than trying to optimize code that's only doing 1% of the actual work.
My use for compiled code is when there's no OS, and every nanosecond matters, e.g., working with microcontrollers.
This really depends on your perspective. Sure, it's a lot slower than running a native binary. But it's still fast enough for interactive tools that you run often. If you compare that to java tools, for example, which take seconds if not dozens of seconds to start (gradle, I'm looking at you!), python is far better.
gradle has pretty much the worst ergonomics of any developer tool I've ever used[1], so that's a pretty low bar.
[1] e.g. "Something went wrong. Re-run with -debug or -info. Now here's 5000 lines of useless stack traces that won't help at all." Nevermind an ecosystem of plugins where it's not clear where the DSL ends and API starts, and makes it virtually impossible to locate the actual source of an error message, if you get one at all.
Gradle's entire architecture is the best argument I have ever seen for why you should never use a big blob of shared mutable state when all you really needed to do was pass discrete values up and down the call stack.
Compare with Bazel. I appreciate why people bounce off of many of its strictures. But when you just look at the architecture and what it takes to develop for it, it's amazing how much easier it is to grok, while achieving a similar level of flexibility.
Oh really?
https://www.mercurial-scm.org/wiki/OxidationPlan
> chg's very existence is because we need hg to be a native binary in order to avoid Python startup overhead. If hg weren't a Python script, we wouldn't need chg to be a separate program.
The problem with this workflow is that it means you end up rewriting large chunks of your code over several years in systems that have fairly different idioms. If you suspect you might in the future be performance bottle-necked, writing your code in a faster (but still productive) language to start can be a lot better because you can then improve performance more incrementally without multiple full rewrites.
Also I really dislike python optimization tricks like `output = stdout.buffer.write`.
Grab a web page and parse it for some data. Suddenly, C looks like trash. The network latency negates any speed advantage from C and the fact that C has to manually manage memory everywhere makes your C code look like a dumpster fire relative to the Python code. You can always pick a task that favors one language or the other.
What people forget is that fast/slow ALSO includes "time to develop the code". If I can write the Python code in 1 hour and the C code in 10 hours, the code has to be a lot slower or be run a lot more times before the time wasted writing the C code pays off.
here's a list of 41 issues where either the user or the mypy devs are saying "use type: ignore as a workaround for now":
https://github.com/python/mypy/issues?q=is%3Aopen+label%3Afa...
1) During AoC I mostly used cpython and numpy. One day my code was running too slow so I fired up pypy. It ran even slower. Turns out numpy and pypy do not play well together in terms of performance.
2) I use matlab in work, and had some code that was slower than I liked, despite efforts to profile and optimise. I used the coder tool to compile it to C and got a 32x speedup. I was surprised, as matlab has a jit, so didn't expect a 32x performance delta, but there it is.
It's a Janus of a programming language. There are lots of complexities that alternately tie software engineers up in knots and let them go into raptures of blissful hackery. But they generally aren't all that visible in the face it presents to data scientists and devops.
And it's hard to make sweeping statements about "big or performance critical". A couple times in the past few years years I've seen or been involved in projects that successfully replaced big Java applications with Python implementations that were 1% the SLOC and had better performance. They were special situations, to be sure, though more so, I think, in terms of explaining the performance improvement than the cost savings.
I worked on a codebase many years ago when I was an intern. That code would read frames off a CAN bus in a car. For some reason, it could only handle about 100 frames per second before running out of CPU. It turned out someone had done the same; they turned every 8 byte CAN payload into a string of 64 characters, sliced the relevant fields out, and then converted back to integers. There were about 2000 frames per second coming in, and it was just way too inefficient.
In my usage, it will maybe never happen, outside of a unit test I wrote for it. I am really tempted to rewrite it tomorrow to use but shifts...
The novice, in his frustration, struck the side of his computer. The master walked over and asked what he was doing. The novice exclaimed, "my computer is not working and I do not know why!" The master admonished him, saying "you cannot solve the problem by striking the computer without knowing what is wrong." Then the master struck the side of the computer, and it worked flawlessly.
(I love this one because one time I was doing a gig and my computer refused to boot. Suspecting that the vibration from being lugged into the car and taken for a long drive had unseated something slightly, I gave it a sharp smack and power cycled it, and it started up fine.)
> A novice was trying to fix a broken Lisp machine by turning the power off and on.
> Knight, seeing what the student was doing, spoke sternly: “You cannot fix a machine by just power-cycling it with no understanding of what is going wrong.”
> Knight turned the machine off and on.
> The machine worked.
How about a list of integers:
$ txr
This is the TXR Lisp interactive listener of TXR 274.
Quit with :quit or Ctrl-D on an empty line. Ctrl-X ? for cheatsheet.
This could be the year of the TXR desktop; I can feel it!
1> (digits 37)
(3 7)
2> (digits 37 2)
(1 0 0 1 0 1)
3> (reverse *2)
(1 0 1 0 0 1)
4> (mapcar (op - 1) *3)
(0 1 0 1 1 0)
5> (poly 2 *4)
22
:)Similar, in awk:
https://gist.github.com/jaysoffian/e41ca479d70e60efe59fded93...
The tone and empathy with the target audience ('s thought process) is on-point.
In terms of building a team at a company? In my experience it's been arbitrary. Manager / Lead has experience in language X, decides to hire people who also know it. Or all the other teams are already using language X. Or Manager / Lead has heard "X language is good at Y" and decides to go with that. Or there's simply 10x more engineers (and cheaper) available for language X than Y.
The times I've seen a language picked for a particular purpose:
- Perl/Python used for web apps. It's interpreted so you can just upload your source code and refresh your browser. Faster and simpler than having to compile/package it, lots of useful frameworks and modules, lots of developers.
- Erlang/OTP for telecom. Kind of on the nose, but there you are.
- C and C++ for embedded applications.
Go compiles so fast and reads cleaner imo, that I never saw the benefit of python. I always thought debugging python was clunky, could be unfamiliarity with interpreted languages. Also a few times I needed to use an ODBC driver with python left a bad taste in my mouth. But overall I get what you're saying, my new team writes all of their lambdas in python so I'm going to have to learn.
1. Most of what I am doing is calling other programs, often over ssh. The speed difference between the two is going to be tiny (where it is not, I use Golang)
2. Python is a more ergonomic language to work in a lot of the time. Wrapping a series of steps in a try block and then handling the errors in one go (very common in the things I am doing) is just easier to write, easier to maintain, and easier to read than what I have to do in Golang. And there are a ton of modules in Python that are just easier to work with than in Golang. And the resistance to the 'while... else' and 'for... else' patterns just saddens me. And the 'for' implementation in Golang often makes me want to tear my hair out... why make it hard to do this by reference?
3. Even in places where Golang should be much better, sometimes it is a challenge. For example when I am trying to parallelize something, but want to only have N number of workers going at a time. In Python I just use the worker pattern and I am done, I can even feed from one set of workers to another. In Golang I have to use a limited channel, and be careful that I block on that channel before I do anything that is going to eat a lot of memory (otherwise I have protected the CPU, but not memory resources). It feels like I am fighting the language there.
The common use of Python is a huge issue.
Although, looking now I'm seeing a short support window for go releaI. That can hurt uptake.
Like go 1.18 was released 2022-03-15 and will EOL Q1 2023.
Lots of projects want long term support so that churn isn't desired. I'm curious how teams handle that.
I've got a few sysadmin type scripts I wrote in perl that I'm currently rewriting in python because other people will use and need to update those things and they will be more likely to know python than perl. I can't imagine they'll be more likely to know Go either.
I'm not sure I understand that part of the mypyc documentation (https://mypyc.readthedocs.io/en/latest/introduction.html#why...). Does that mean that you can't use something like numpy at all?
A book, a website, a cheat sheet or even a MOOC or part of a MOOC?
This is even more relevant if you are still learning the language. Focus on learning it, and leave arcane always changing implementation details for after you know it well.
Also for all languages where you want to do performance work: Learn to use a profiler, so you can find the things that matter. And then selectively look what you can improve there. Even in Python, some hacks or slightly unergonomic patterns in a really hot loop can be worth a lot.
[1] https://docs.python.org/3.11/whatsnew/3.11.html#faster-cpyth...
- Get the overall system structure in Python. Get the architectural design right, and the big-O stuff right.
- If there are bottlenecks, re-code those directly in C, cython, or similar. Or better yet, find libraries.
Python is great for expressing high-level operations and system design. It is also very easy to integrate with native code. I've never had much happiness in optimizing Python itself beyond that. Broadly speaking, code falls into three categories:
- Most code: Instant. Performance doesn't matter. Python
- Some code: Big-O(lifespan of the universe). Not worth building.
- Narrow slice of stuff in between: Don't do in Python, but use Python as the glue.
Anyway, here it is, this is the second edition of the book https://www.amazon.com/dp/1492056359
I read the first one at work a few years ago and was fascinated with all the goodies it shared with us.
In my humble opinion, it's a must read by all Python developers, experienced or not; everyone will learn something new just by reading it.
A couple of advises:
- the right algo will go a long, long way.
- know your data structures. E.G: assigning to a slice is ridiculously fast (even while unpacking), memory views may save a lot on byte heavy workloads, heapq and deque are underrated, etc. Also check out https://wiki.python.org/moin/TimeComplexity for big O notations on python builtin types common operations to understand what you pay for.
- know the stdlib. collections, itertools and functools all contain incredible gems.
- delegate. Python is a fantastic glue language, use it for what it's good at. Your database, your numpy arrays, your cache are all amazing are what they do, no matter the language. Let them do the heavy work.
- don't kill good perfs by ignorance. I regularly see people casting a generator, iterating on a dataframe, calling readlines() on a file or doing something else that is destroying the otherwise excellent perfs of their program.
- know the ecosystem. There are some very good fast libs out there: diskcache, sortedcontainer, scipy, uvloop...
- use threads to avoid blocking a GUI, multiprocess to share work between CPU and asyncio to speed up network operations. Each tool has a sweet spot. But threads are underrated, they work well for a couple of hundred parallel network operations, and most C libs will actually release the GIL, so they can use several CPU more often than you'd thin. Also, use pools if you can, shared_memory in 3.8 or mmap.
- sometime the dirty solution is just faster, like subprocessing to ffmpeg.
- The more recent, the slower. Sure, I love statistics, pathlib and dataclasses. In a regular code they are great. On a bottleneck however, they are very slow.
- the array module is not supposed to speed up the code, only save memory. But sometimes it does.
- printing to the terminal is limiting. Sometime your program is doing fine, the display is preventing it to go faster. At least check the flush.
- comprehensions are faster than alternatives.
- pre-allocating lists and dicts can help.
- measure. The austin profiler is your friend.
- rewriting the hot path in a faster language is likely more interesting than writing the whole program in it. A bit of nim or rust is easy to call from python.
- some python distribution are faster than others. I don't mean just pypy, but also regular cpyhon that have been built specially for some architectures. E.G: https://www.intel.com/content/www/us/en/developer/tools/onea...
- nuitka code compilation can speed up the code by a factor of 4, it's nice, especially for startup.
Do you have any numbers to back this up? Why would Pathlib be slow?
TL;DR: Path() initialization is very heavy compared to simple strings, and subsequent path operations produce new Path() objects every time.
The video then illustrates the point with a patch to the black formatter cache yielding a 40X speedup.
"Our speedup on ARM (30% on a Graviton EC2 instance) is comparable to our speedup for x86 (34% on an Intel i7-6700)." https://blog.pyston.org/2022/04/01/pyston-v2-3-3-arm-support...
2021.aug.30 : "Pyston Team Joins Anaconda to Expand Open-Source Project Development" https://www.anaconda.com/blog/pyston-team-joins-anaconda
Roadmap: https://blog.pyston.org/2021/10/26/pyston-roadmap/
Github: https://github.com/pyston/pyston
Other than that, I think the python code is quite readable and understandable with the context provided in the post.
Very well written post!
However I was unable to compile from my mac, could only compile using linux. but that could be as I'm using older version of dev tools.
If you just have function and class definitions, this isn't too bad, but when you start doing things like setting up caches, reading files, and testing whether or not there's a GPU at the top level, you can add several seconds to the startup and balloon your memory usage. One wrong import somewhere in your code base, and your application will craw to a standstill on launch.
I find Python to be a great language to describe business logic, and honestly it should end in 10 years to be within at most 2x slower than Javascript.
I see a bunch of Python code that is just a massive dump of custom classes on top of classes, inheriting other classes.
The amount of pointer chasing to get to the actual primitives can be mind-bogling. Add that to a big nested loop and you've got a performance tragedy.
I think a mypyc version would be an order of magnitude harder to debug.
At least for now, I wouldn't use it in production.