This is where CPython lives. CPython has been getting sped up for the last 15 years. The end result is not a blazing fast language knocking the socks of the competition; the end result is that as they've added features, they're been successfully able to more-or-less tread water or make a bit of progress, and they have remained merely one of the slowest languages in common use, instead of unusably slow.
Picking off a few improvements in instruction dispatching in CPython wouldn't have changed Python reposurgeon's capabilities much at all.
The gulf between Python and Go isn't a matter of a few dozen percent, it's a matter of a few dozen times in terms of speed and generally substantial integer multiples on memory use too.
As I said yesterday, many people interpret this as an attack on Python. It isn't. It is just something you need to know to make good engineering decisions. I've used Python many times in the past several years, and I've never been bitten by its performance issues... because I think about it before I use it. It is definitely something you need to think about before you reach for it. It actually was OK for reposurgeon for the most part... it just couldn't handle the largest repos, but those are the exceptions, after all.
Otherwise, your argument isn't very compelling, especially after having admitted that you haven't even used Python in ways that would warrant any assessment of performance for those use cases.
So, if that version was 40x slower than C (and that is very optimistic), the latest version is still 30x slower. That is the point that the GP made.
I would not call that substantial at all. I'd prefer to keep the CPython interpreter simple, people who want faster Python can use pypy or other things.
The real work in scientific Python will always be done in C or CUDA extensions. No 50% speedup will change anything.
X’ X’/Y X’/Y
0.5 = – = –––– = ––––
X X/Y 40
X’
=> – = 20
YAs a baseline, my code can do 100 operations / second.
With 50% speedup (+50%), my code can do 150 operations / second.
If C was 40x faster than the baseline, it could do 4000 operations / second. After the speedup C can still do 27x as many operations per second. (4000/150 = 26.66..)
The evidence that Python is still slow is abundant and overwhelming. It's not like it's a subtle thing. It's a huge gulf. If you don't understand this, the problem isn't that jerf on Hacker News didn't spell it out to your satisfaction; the problem is that you are not evaluating things calmly and sensibly from an engineering perspective, but are using your political brain to evaluate programming languages. Until you stop doing that, you will never understand what is going on; once you do stop doing that, the truth will be so obvious you won't need me to spell it out for you.
I think here is your problem. You're expecting dynamically typed, interpreted language to be comparable to statically typed compiled language.
Python is best used as a glue language, or when performance is not critical.
That doesn't mean you can't have fast code with python. I for example was able to encode a live video into using x264 then steam it over the Internet in python, then on the other side receive the data, decode and display it.
Though the encoding/decoding was done by a C code that was interacted with through python package.
I suppose I could have the whole thing done in C, but then my code would be more complex and longer. It will also be harder to iterate on it.
But C-extensions don't require the small speedups and additional interpreter complexity that this project brings.
Maybe all those debug logging statements in the last Python version [2] that, unlike the built-in logging framework, use the string formatting operator to construct the log message regardless of whether debug logging is enabled.
I can't imagine that encoding everything to latin-1 as a workaround for Unicode support is very good for performance either, unless he was only using the Python 2 version.
And maybe, for example, the empty_comment function doesn't need to call the str.strip method twice on the same string? Maybe that could be more efficient?
[1] http://esr.ibiblio.org/?p=8161&cpage=1#comment-2065946
[2] https://gitlab.com/esr/reposurgeon/-/blob/911d5c1168f7839855...