Python Bytecode Explained
github.com
github.com
There is a second part that may be of interest: here a python tracer is implemented, one that shows the side effects of each line, as it is executed (it shows the effect of all the various load and store instructions). The objective was to get something that is similar to the set -x built-in of the bash shell.
https://github.com/MoserMichael/pyasmtool/blob/master/tracer...
And it's all part of this advanced python course: https://github.com/MoserMichael/python-obj-system/blob/maste... (well, I am still working on it)
And me is also looking for a job again ;-( I need a new job in April. So here is my linked-in profile. I also do C++ and Java/Scala. Available on-site in the Tel-Aviv area, considering remote only jobs anywhere else. https://www.linkedin.com/in/michael-moser-32211b1/ E-mail address is in my HN profile.
https://twitter.com/phil_eaton/status/1482801489907273739
https://www.linkedin.com/posts/phil-e-97a490178_really-fanta...
We're also in the Tel Aviv area (but remote is an option if you prefer) and we've been looking for someone like you who explains technical topics in simple terms.
We have some low level stuff (e.g. I wrote a python debugging tool for Kubernetes which injects debugby into target processes using gdb [1])... but also a lot of higher level stuff in our python framework for k8s automation.
Hope I can interest you
That's also how dropbox used to obfuscate their client when it was python. They would ship only pyc files, which is just bytecode. But they would change around the opcodes, map multiple numbers to the same opcode, etc. Then also stream encrypt the pyc file and hide the key inside of it.
The "Looking inside the Dropbox" paper where some researchers reversed engineered it is interesting: https://www.usenix.org/system/files/conference/woot13/woot13...
Sadly `-mdis` requires feeding a file by path or data through stdin, so for mucking around it’s not the best.
This is incorrect. Python bytecode files are versioned alongside the interpreter, so when CPython finds a __pycache__/*.pyc file which is the wrong version, it will just ignore it and won't cause any problems.
Can someone elaborate on this? Having separate stacks makes sense for coroutines, but does this mean that a normal Python function call allocates a private stack for that function?
[1]: https://github.com/MoserMichael/pyasmtool/blob/master/byteco...
That also means the bytecode works solely within its own stack segment.
All it means is that python bytecode is stack based where most instructions pop arguments and push results on operand stack. In contrast with register based VMs. When implementing a VM it makes sense to store call stack and operand stack separately so that you don't have to mix types. You probably don't want to allow function to uncontrollably modify operands in lower frames as in most cases that would be either a bug or vulnerability. Having separate operand stack for each frame also makes any kind of analysis much easier. Call instruction can be viewed as a fat instruction which pops some amount of arguments and pushes single result back. Once you restrict cross frame operand stack access, whether it's stored in single or multiple arrays becomes an implementation detail. Many other VMs do more or less the same JVM, AVM2(flash), CIL(C#). It doesn't necesarily mean that the stacks are separate after JIT but from the perspective of bytecode instructions operand stacks are separate.
Isn't compiling and then immediately running the code exactly what a just in time compiler is? Or do I have a misunderstanding of the term?
"Classical" (again, every term is fuzzy) JIT compilers either do this machine code compilation after seeing a good candidate _entire function_ or a good candidate _section of code within a function_. Good candidates are often areas of code that are executed a large number of times and with consistent internals (e.g. iterating from 0 to 10000 with variables inside that have provably fixed types).
But there are infinite variations of JIT compilation.
In any case, CPython doesn't do that switching from bytecode to generated machine code. Pypy does do that. As does V8 and the JVM and so on.
A JIT would compile the bytecode to machine code then run it directly (at least for frequently executed code paths). There is no "switch" anymore. Each bytecode instruction has already been replaced by the corresponding machine code.
Is there any reason why official python doesn't have any JIT option? Would that be too fastidious to develop?
Desires to keep the implementation simple and approachable (relatively), as well as avoid issues of performance cliffs and such.
Also the C API has historically been extremely broad and provided large access to what amount to implementation details, making this keep working properly with a jit is difficult (at least for anything but a simplistic macro-ish JIT).
as for why cpython doesn't use a jit? most likely to prevent any breakage with c modules
Reasoning given for course correction was (AFAIR) that Python really could be faster for things like data science or ML.
Because compilers are complicated and have trade-offs.
> and had to break backward compatibility anyway
A compiler shouldn't break backward compatibility.
I don't understand what you mean by this in context (since they introduced a new language in python3).
That question doesn't make sense, because a compiler shouldn't have any impact on your compatibility.
They can introduce a compiler without breaking compatibility, so they don't need to do it with a new language version.