Pycraft: Minecraft engine in Python
github.com
github.com
It seems like a harmless enough thing to do but that code is effectively creating a new array with values [1, 2, 3, 4, 5] and testing if the result of len() is in that array by iterating over it. This check is happening multiple times for each draw call every frame.
Sadly I see this kind of stuff in Python all the time and it just adds weight to the argument that Python is not a performant language. 1 <= len(vals) <= 5 would have been more pythonic and certainly more efficient (and obvious) but I have to wonder if under the hood it's just doing the same inefficient operation.
You can actually get a very good idea for what Python is doing under the hood using the dis module:
dis.dis(lambda: 1 <= len(vals) <= 5)
1 0 LOAD_CONST 1 (1)
3 LOAD_GLOBAL 0 (len)
6 LOAD_GLOBAL 1 (vals)
9 CALL_FUNCTION 1
12 DUP_TOP
13 ROT_THREE
14 COMPARE_OP 1 (<=)
17 JUMP_IF_FALSE_OR_POP 27
20 LOAD_CONST 2 (5)
23 COMPARE_OP 1 (<=)
26 RETURN_VALUE
>> 27 ROT_TWO
28 POP_TOP
29 RETURN_VALUE
Compare to: dis.dis(lambda: len(vals) in range(1, 5))
1 0 LOAD_GLOBAL 0 (len)
3 LOAD_GLOBAL 1 (vals)
6 CALL_FUNCTION 1
9 LOAD_GLOBAL 2 (range)
12 LOAD_CONST 1 (1)
15 LOAD_CONST 2 (5)
18 CALL_FUNCTION 2
21 COMPARE_OP 6 (in)
24 RETURN_VALUEThis output doesn't really provide any clarification.
You have two function calls in the range version and only one in the non-range version. You are just hiding some of that code in a dynamic function call. Also, your compare_op consists of two <= in the non-range version and an "in" compare_op for the range version (which probably hides the element by element comparison or hash/bloom if it is O(1).
Fewer lines of code does not mean more efficient. Regardless of whether it is source code or bytecode.
It appears to me they went ahead and posted the bytecode for both because they had spun up an interpreter to get the bytecode for one.
$> python -m timeit "1 <= 4 <= 5"
10000000 loops, best of 3: 0.0877 usec per loop
$> python -m timeit "4 in range(1,5)"
1000000 loops, best of 3: 0.413 usec per loop
$> python -m timeit "4 in (1,2,3,4,5)"
10000000 loops, best of 3: 0.0776 usec per loop
Tested in python 3.5.0, Even with the improvements to the range function, the comparison is almost 5x faster.Agreed that 1 <= len(vals) < 5 would be more Pythonic.
"x in range(10)" will operate in constant time and memory in Python 3. Whether it is actually more efficient than "0 <= x < 10" is another matter, of course. I would expect the call to range() to dominate here, and indeed, it is much slower in this microbenchmark:
$ python3 -m timeit -s 'x = 8' 'x in range(10)'
1000000 loops, best of 3: 0.351 usec per loop
$ python3 -m timeit -s 'x = 8' '0 <= x < 10'
10000000 loops, best of 3: 0.0732 usec per loop
Even aside from this, I find the "0 <= x < 10" syntax to be clearer.From a formal CS perspective, that's why this is wrong; It's because 'in' is an o(n) operation, while the two comparisons are constant time. It may be abstracted away by the range object overloading, but that's a fairly narrow optimization it's able to pull off.
On modern CPUS, branching operations like compare usually are your problem; An unconditional function call is easy to optimize, but a branch in a looping construct is guaranteed to lead to one or more mispredictions. Trying to minimize the surface area of comparisons is an important part of performant code.
Also, as noted in other comments, range object has special cased O(1) implementation of in for integers (range_contains_long() in Objects/rangeobject.c)
CPython bytecode interpreter tries to minimize amount of branching it causes by using some non-obvious tricks, but still by definition it is going to cause at least one essentially unpredictable branch per interpreted instruction, so optimizing for number of branches in user python code is mostly pointless endeavor.
python3 -m timeit -s 'x = 8' 'x in range(1000)'
1000000 loops, best of 3: 0.405 usec per loop
python3 -m timeit -s 'x = 8' '0 <= x < 1000'
10000000 loops, best of 3: 0.0742 usec per loop
python3 -m timeit -s 'x = 8' 'x in range(100000)'
1000000 loops, best of 3: 0.405 usec per loop
python3 -m timeit -s 'x = 1123' 'x in range(100000)'
1000000 loops, best of 3: 0.415 usec per loop
python3 -m timeit -s 'x = 1123' '0 <= x < 100000'
10000000 loops, best of 3: 0.0745 usec per loop
While 0<=x<10 seems to be 5 times faster, both seem to be independent of the actual chosen values.If it doesn't affect performance then who's thinking is "wrong" here?
Example in this thread: half of the commenters consider that it means 1 <= len(vals) <= 5 and the other half consider that it means 1 <= len(vals) < 5...
There are already some people talking about Cython for some of the fundamental graphics stuff.
Basically, this is just a base so that people can see what they're working towards.
Not sure if you'd be interested, but I have some content I'd be happy to add for background type information (i.e. degrees of freedom, calculating movement, frames of reference, etc). Might make a good starting point for contributors but not sure if that's the direction you see your documentation taking.
That said, they also promised a modding API too...
[0] http://web.archive.org/web/20100301103851/http://www.minecra...
Still, as long as the idea's out there... I heard the codebase isn't actually that great...
[1] https://pyglet.readthedocs.org/en/pyglet-1.2-maintenance/pro...
Do do that, you first have to write code for pulling levers, which will first require some digging around with how pycraft handles textures to make it more generic.
Check back in a few months.
A lot of people have been doing pull requests for unit-tests, documentation, and continuous integration. So I'm not too worried.
We're setting up a policy for requests, but that policy is living in code.