It states that python is faster then c, that is not possible since python is build with c. There could be other reasons such libs or implementation.
Also note that the issue he had was not resolved.
The comment was about that python is seen as slow. But that is not always the case.
Once a dev is able to understand the difference between the python and c parts. Python can be quite performant, and efficient with memory.
But if one would actually create a application that does more then just read a file it will be slow again compared to c and rust.
The root cause is not about page alignment. In fact, all allocators are aligned.
The root cause is AMD CPU didn't implement FSRM correctly while copying data from 0x1000 * n ~ 0x1000 * n + 0x10.
> Other questions never answered, why did pyo3 add so much overhead? it was over half the difference between the two.
OpenDAL Python Binding v0.42 does have many place to improve, like we can alloc the buffer in advance or using `read_buf` into uninit vec. I skipped this part since they are not the root cause.
Other way around: with glibc it was page-aligned; with the others, it wasn't.
This weird Zen performance quirk aside, I'd prefer page alignment so that an allocation like this which is a nice multiple of the page size doesn't waste anything (RAM or TLB), with the memory allocator's own bookkeeping in a separate block. Pretty surprising to me that the other allocators do something else.
From the article.
> In conclusion, the issue isn't software-related. Python outperforms C/Rust due to an AMD CPU bug.
In a really strict sense it's impossible to talk about the speed of languages, since any turing complete language could be implemented in any other. In practice when people say X is faster than Y, they mean in practice as actually used; it's completely possible, for instance, that if you ask a large pool of C programmers to... I dunno, sum ten billion integers, and the same to a large pool of Python programmers, most of the C devs will reach for a `for` loop and most of the Python devs will reach for numpy and get vectorization for free, and if that's the case then it's reasonable to say that Python is faster. Or in the actual case at hand, writing the same(ish) program in Rust and Python on the same hardware does result in the Python version being faster, even though it's a bug from that exact hardware not getting along with something under the hood in the Rust version.
Instead, it maintains its own memory space. Consequently, transferring data from the Python environment into NumPy or vice versa is relatively slow.
The process of opening a file and travesing its data within Python relies heavily on the C code behind the scenes, resulting in near-C performance.
However, if one were to write an algorithm along the lines of LeetCode - one that has a time complexity of n*2 - Python's performance will be slower compared to other languages. This difference could range from a factor of one to potentially even a hundred.