Typically from a user perspective, the initial starting time is either manageable or imperceptible in the cases of long running services, although there are other costs.
If you look at examples that make the above claim, they are almost always tiny toy programs where the cost of producing byte/machine code isn't easily amortized.
This quote from the post is an oversimplification too:
> But the program will then run into Amdahl's law, which says that the improvement for optimizing one part of the code is limited by the time spent in the now-optimized code
I am a huge fan of Amdahl's law, but also realize it is pessimistic and most realistic with parallelization.
It runs into serious issues when you are multiprocessing vs parallel processing due to preemption, etc .
Yes you still have the costs of abstractions etc...but in today's world, zero pages on AMD, 16k pages and a large number of mapped registers on arm, barrel shifters etc... make that much more complicated especially with C being forced into trampolines etc...
If you actually trace the CPU operations, the actual operations for 'math' are very similar.
That said modern compilers are a true wonder.
Interpreted language are often all that is necessary and sufficient. Especially when you have Internet, database and other aspects of the system that also restrict the benefits of the speedups due to...Amdahl's law.
In summary, it depends. I am talking about compute performance, not I/O or general purpose task benchmarking. Yes, if you have a mix of compute and I/O (which admittedly is a typical use case), it isn't going to be 20-100x slower, but more likely "only" 3-20x slower. If it is nearly 100% I/O bound, it might not be any slower at all (or even faster if properly buffered). If you are doing number crunching (w/o a C lib like NumPy), your program will likely be 40-100x slower than doing it in C, and many of these aren't toy programs.
Python isn't evaluated line-by-line, even in micropython, which is about the only common implementation that doesn't work in the same way.
Cython VM will produce an AST of opcodes, and binary operations just end up popping off a stack, or you can hit like pypy.
How efficiently you can keep the pipeline fed is more critical than computation costs.
int a = 5;
int b = 10;
int sum = a + b;
Is compiled to: MOV EAX, 5
MOV EBX, 10
ADD EAX, EBX
MOV [sum_variable]
In the PVM binary operations remove the top of the stack (TOS) and the second top-most stack item (TOS1) from the stack. They perform the operation, and put the result back on the stack.That pop, pop isn't much more expensive on modern CPUs and some C compilers will use a stack depending on many factors. And even in C you have to use structs of arrays etc... depending on the use case. Stalled pipelines and fetching due to the costs is the huge difference.
It is the setup costs, GC, GIL etc... that makes python slower in many cases.
While I am not suggesting it is as slow as python, Java is also byte code, and often it's assumptions and design decisions are even better or at least nearly equal to C in the general case unless you highly optimize.
But the actual equivalent computations are almost identical, optimizations that the compilers make differ.
> A compiler for C/C++/Rust could turn that kind of expression into three operations: load the value of x, multiply it by two, and then store the result. In Python, however, there is a long list of operations that have to be performed, starting with finding the type of p, calling its __getattribute__() method, through unboxing p.x and 2, to finally boxing the result, which requires memory allocation. None of that is dependent on whether Python is interpreted or not, those steps are required based on the language semantics.
i.e.
if(a->type != int_type || b->type != int_type) abort_to_interpreter();
result = ((intval*)a)->val + ((intval*)b)->val;
The CPU does have to execute both lines, but it does them in parallel so it's not as bad as you'd expect. Unless you abort to the interpreter, of course.