The performance of pure python code is orders of magnitude worse than non-interpreted languages, there's no point trying to shave off 0.5% off a 5000% difference.
The performance of pure python code is orders of magnitude worse than non-interpreted languages, there's no point trying to shave off 0.5% off a 5000% difference.
Also Python internals do have some guarantees but in this case it's a semantic guarantee. Because builtins can be shadowed you'll always have to pay the performance cost of looking them up which isn't true for {}.
A 2x speed improvement is very significant if it happens to be in the critical loop of your code. You want a 5000% difference? Just find 6 tweaks like these and you might just get there.
Of course in most cases whatever you're doing to build the dict is more likely taking up most of the time, not building the dict itself, but understanding why one option is 2x as fast is still important. Though one can only hope that it will soon be irrelevant when JITed python becomes a thing.
Optimizations don't compound like this.
I think outside of that niche, though, there are places where people are writing heavily CPU bound code in Python, because it is so easy to become CPU bound. Case in point: I recently sped up an ML ingestion pipeline by multiple thousands of percent by switching from a pure-python PDF library to one that wraps a .so written in C.
So my point, restated: if you're CPU bound in pure Python code and you have time to try and optimize it, just rewrite the critical section in C and use the FFI. This is how 90% of "Python" libraries get implemented anyway. Compared to this, trying to make Python code more CPU efficient is a waste of time.
By the way, I have done work on CPU performance optimizations in Go, Rust and C, and the things you'd typically do are not possible in Python anyway. You're basically left with randomly tweaking the code until the benchmark gives a thumbs up, because it hit on some cpython idiosyncrasy that will completely change around a few versions later.
1) Tweak the code without understanding* of how it'll affect performance, because every single line of code hides behind it such complexity that any improvement you find is almost certainly overfitted to your version of Python. Eventually a benchmark will spit out a nice number. If successful, make a modest improvement, maybe 30%.
2) Rewrite the critical section in C. If you need to optimize further, you are now able to draw on 50 years of know-how in a well-studied field. The improvement will be on the order of thousands of percent.
They both take about the same amount of time. Why should you do (1)?
* There is a difference between random tips like "replace {} with dict()" and fundamentals like how the CPU cache works, or the branch predictor. The former is almost certainly a quirk of the current Python version, the latter has been the same for the past 20-30 years. If you do work on software performance, you rely on your knowledge of the fundamentals and a tool like `perf` to make educated guesses about where you can save some cycles. These fundamentals are basically irrelevant to Python code and so you have to make effectively random attempts.
Changing version or even platform can change the absolute efficiency.
In this case the speed difference will likely always be there, since `dict()` can be overriden (monkey-patched) - hence the interpreter needs to resolve it at runtime.
You could probably override it using Python’s extension system.
Forgive me, I know Python very well, but not C/C++. I’ve learned to make no assumptions.
(My current day job is doing GPU optimization on iterative medical image reconstruction computation. But of course, I don't do that in Python.)