One major difference here is that in Python you can overwrite min(), so it has to access this function through more abstraction, but not in PHP. I bet PyPy is close to PHP here however.
One major difference here is that in Python you can overwrite min(), so it has to access this function through more abstraction, but not in PHP. I bet PyPy is close to PHP here however.
When I avoid the min() call like this ...
min2.py:
i = 10_000_000
r = 0
while i>0:
i = i-1
if i<500: r += i
else : r += 500
print(r)
I get a speedup of about 50%: time python3 min2.py
4999874750
real 0m1.338s
Still 4x slower than PHP though.With PyPy I get an amazing speedup of 96%:
time pypy3 min.py
4999874750
real 0m0.088s
4 times faster than the PHP version!I'm not sure what the state of things is regarding PyPy and running a web application. Doing a bit of googling, it seems that calling C functions from PyPy is slow, so one would have to check if the application uses those (for example to connect to the database).
def spam():
t1 = time.time()
i = 10_000_000
r = 0
while i>0:
i = i-1
if i<500: r += i
else : r += 500
t2 = time.time()
print(r)
print("Time:", t2-t1)
reports 0.87 seconds outside a function and 0.43 seconds when inside a function.(Using the shell's time, of course, also includes the Python startup time. When I use it I get 0.94 and 0.49 seconds, indicating a 0.06s startup penalty.)
Also with pypy3, putting it in a function about doubles the speed:
time pypy3 min_in_function.py
4999874750
real 0m0.045s def f():
from math import *
return cos(3.0)
https://peps.python.org/pep-0227/#import-used-in-function-sc... notes "The language reference specifies that import * may only occur in a module scope. (Sec. 6.11) The implementation of C Python has supported import * at the function scope."Python 2.1 added static nested scopes (see the PEP 227 link), which enabled a much faster locals lookup. With that one case excluded, all of the locals can be determined at parse/byte-compile time, stored in a simple vector, and referenced by offset:
>>> z = 0
>>> def f(x):
... y = x + z
... return y * 2
...
>>> import dis
>>> dis.dis(f)
1 0 RESUME 0
2 2 LOAD_FAST 0 (x)
4 LOAD_GLOBAL 0 (z)
14 BINARY_OP 0 (+)
18 STORE_FAST 1 (y)
3 20 LOAD_FAST 1 (y)
22 LOAD_CONST 1 (2)
24 BINARY_OP 5 (*)
28 RETURN_VALUE
The "LOAD_FAST" uses offset 0 to store "x" and offset 1 to store "y", while "z" required a LOAD_GLOBAL to find z in the globals() dictionary.https://x.com/marekgibney/status/1817495719520940236
I was surprised that you cannot access dynamically created local variables even though they exist.
Looks like the local dictionary still exists. Just that whether accessing a local variable looks it up in the local or global dictionary is determined at compile time.
In Python 3.12, quoting https://docs.python.org/dev/whatsnew/3.13.html
> PEP 667: The locals() builtin now has defined semantics when mutating the returned mapping. Python debuggers and similar tools may now more reliably update local variables in optimized scopes even during concurrent code execution.
with more details at https://docs.python.org/dev/whatsnew/3.13.html#whatsnew313-l... , which adds:
> To ensure debuggers and similar tools can reliably update local variables in scopes affected by this change, FrameType.f_locals now returns a write-through proxy to the frame’s local and locally referenced nonlocal variables in these scopes
def f():
exec('y = 7') # will still create a local variable
print(locals()['y']) # will still print this local variable
print(y) # will still try to access a global variable 'y' Python 3.12.2 ...
Type "help", "copyright", "credits" or "license" for more information.
>>> def f():
... exec('y = 7') # will still create a local variable
... print(locals()['y']) # will still print this local variable
... print(y) # will still try to access a global variable 'y'
...
>>> f()
7
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "<stdin>", line 4, in f
NameError: name 'y' is not defined
It does print 7 in Python 2.7.See also:
https://x.com/marekgibney/status/1817495719520940236
and
Messing with locals has been a long-time danger zone. I don't know how it is supposed to interact with exec - another danger zone - and I don't know enough about the changes in 3.13 to know if it's resolved.
EDIT: This uses a Nim program https://github.com/c-blake/bu/blob/main/doc/ru.md run as `ru -t`, but for the very fast variants you can get a more precise wall time from https://github.com/c-blake/bu/blob/main/doc/tim.md
pypy3_10-7.3.16_p1 - (first form,2nd form) x same in a func
TM 0.045421363 wall 0.034517 usr 0.010855 sys 99.9 % 56476 mxRS minCall
TM 0.043447211 wall 0.029492 usr 0.013763 sys 99.6 % 56124 mxRS ifStmt
TM 0.033408064 wall 0.020605 usr 0.012758 sys 99.9 % 55960 mxRS minCallInFunc
TM 0.033115636 wall 0.026059 usr 0.007018 sys 99.9 % 55956 mxRS ifStmtInFunc
python-2.7.18_p16
TM 1.755861105 wall 1.751740 usr 0.001991 sys 99.9 % 5944 mxRS
TM 0.940162912 wall 0.935362 usr 0.003988 sys 99.9 % 5916 mxRS
TM 1.155755936 wall 1.151529 usr 0.002993 sys 99.9 % 5928 mxRS
TM 0.432215672 wall 0.430851 usr 0.000998 sys 99.9 % 5800 mxRS
python-3.11.9_p1
TM 2.678891504 wall 2.675141 usr 0.000994 sys 99.9 % 7992 mxRS
TM 1.348009364 wall 1.344612 usr 0.001995 sys 99.9 % 7864 mxRS
TM 2.018986091 wall 2.014702 usr 0.001994 sys 99.9 % 7972 mxRS
TM 0.768633122 wall 0.766915 usr 0.000997 sys 99.9 % 7980 mxRS
I recall when Python3 was first going through its growing pains in the mid noughties that promises of eventually clawing back performance. This clawback now seems to have been a fantasy (or perhaps they have just piled on so many new features that they clawed back and then regressed?).Anyway, Nim is even faster than PyPy and uses less memory than either version of CPython:
doIt.nim:
proc doIt() =
var i = 10000000
var r = 0
while i > 0:
i = i - 1
r += min(i, 500)
echo r
doIt()
#TM 0.024735 wall 0.024743 usr 0.000000 sys 100.0% 1.504 mxRM
In terms of "how Python-like is Nim", all I did was change `def` to `proc` and add the 2 `var`s and change `print` -> `echo`. { EDIT: though if you love Py print(), there is always https://github.com/c-blake/cligen/blob/master/cligen/print.n... or some other roll-your-own idea. Then instead of py2/py3 print x,y vs print(x,y) you can actually do either one in Nim since its call syntax is so flexible. }It is perhaps noteworthy that, if the values are realized in some kind of array rather than generated by a loop, that modern CPUs can do this kind of calculation well with SIMD and compilers like gcc can even recognize the constructs and auto-vectorize. Of course, needing to load data costs memory bandwidth which may not be great compared to SIMD instruction throughput at scales past 80 MiB as in this problem.
I believe PyPy is clever and checks for new or removed elements in module and builtin scope. If they haven't changed then lookup can re-use previously resolved information.
> clawing back performance
On my laptop, appropriately instrumented, I see Python 3.12 is faster than 2.7. (Note that they use different versions of clang):
% python2.7 t.py
version: 2.7.17
module scope: 1.0321419239
function scope: 0.575540065765
% python3.12 t.py
version: 3.12.2
module scope: 0.9053220748901367
function scope: 0.44064807891845703
Python is far from speedy, yes. For my critical performance code I use a C extension and compiler intrinsics.On a different CPU i7-1370P AlderLake P-core (via Linux taskset) I got a 1.5X time ratio (py3.12.5 being 0.371 and py2.7 0.245). Anyway, I should have qualified that it's likely the kind of thing that varies across CPUs. Maybe you are on AMD. And no, I sure don't have access to the CPUs I had back in 2007 anymore. :-) So, there is no danger of this being a very scientific perf shade toss. ;-)
Same pypy3 also showed very minor speed-up for being inside a function. So, I think you are probably right about PyPy's optimization there and maybe it does just vary across PyPy versions. Not sure how @mg several levels up got his big 2X boost, but down at the 10s of usec to 10s of ms scale, just fluctuations in OS scheduler or P-core/E-core kinds of things can create confusion which is part of the motivation for aforementioned https://github.com/c-blake/bu/blob/main/doc/tim.md.
I should perhaps also have mentioned https://github.com/yglukhov/nimpy as a way to write extensions in Nim. Kind of like Cython/SWIG like stuff, but for Nim.
Let's test templating in Python and PHP!
template.py:
def main():
t = 'Number {NR} is cool'
i = 10_000_000
r = 0
while i>0:
x = t.replace('{NR}', str(i))
r += len(x)
i = i - 1
print(r)
main()
template.php: <?php
$t = 'Number {NR} is cool';
$i = 10_000_000;
$r = 0;
while ($i>0) {
$x = str_replace('{NR}', $i, $t);
$r += strlen($x);
$i = $i - 1;
}
print($r);
Results: time python3 template.py
218888897
real 0m2.854s
time php template.php
218888897
real 0m0.994s
time pypy3 template.py
218888897
real 0m0.795s
So for this use case, PHP is 3x faster than CPython and PyPy is 20% faster than PHP.It’s a fun sport though. I love it.
So yes, the comparison of OP is CPython vs PHP.
Those systems are image based, at any time, you can eval code, break into the debugger, change code and redo the interrupted code, affecting everything that the JIT already improved on the whole image.