Inside cpyext: Why emulating CPython C API is so Hard
morepypy.blogspot.com
morepypy.blogspot.com
For the polar opposite of this, consider the Lua API. You don't get pointers to any VM data structures except a single pointer to the "Lua state". You do not perform any manual memory management.
Lua's approach has yielded amazing results. LuaJIT is not only source-compatible with pre-existing Lua 5.1 extensions, it is binary compatible with them. You can take a .so that you built before LuaJIT ever existed and it will work with LuaJIT without recompiling. This is astounding to me.
Moral of the story: keep your interfaces as narrow and encapsulated as possible!
[1] Lua 5.3 dropped the generational GC because real-world results didn't justify the complexity, but apparently for 5.4 they came up with a better design.
So CPython is actually several things layered on top of each other.
- PyObject-based object model; this includes PyObject, PyTypeObject, PyUnicodeObject... and I think that's it? This is the equivalent of COM IUnknown. It doesn't actually know anything about Python proper, but it defines the operations in terms of which language itself will later be defined (like the idea that objects have a refcount-centric lifetime, and have attributes, and operations like "call" and "add" etc).
- A bunch of standard data structures built on top of that, like PyLongObject and PyListObject. This is just a pure extension of the above - again, defining more terms in which the language is defined, like what happens when you add two ints.
- A bunch of specialized data classes which store Python bytecode and provide the framework for its execution, like PyCodeObject, PyFrameObject and PyFunctionObject. Note that these don't know anything about how the bytecode is produced, nor about how to actually execute it. But they do know about things like local variables (so PyCodeObject will store the list of locals, and PyFrameObject will allocate space for them), so Python-the-language starts creeping in here.
- The parser which produces AST (which is itself a bunch of Python objects), and the bytecode compiler that produces PyCodeObjects out of that AST.
- And finally, the actual interpreter, that ties it all together by providing semantics for the bytecode contained in PyCodeObjects.
This layering is even visible in the Python source itself (https://github.com/python/cpython/tree/v3.7.0) - the first three things live under ./Objects in the source tree, and the parser and the interpreter is under ./Python (with some bits under ./Grammar and ./Parser). So, roughly speaking, ./Objects is the object model, and ./Python is the language proper. The headers are interdependent, unfortunately, but it's not that hard to break them apart if anyone cared to.
In a year. In a project which could put Python next to JS, for the last pain-point that prevents it.
Python - one of the top 10 most popular languages - community and all its industrial user, including some of the most successful companies on the planet, can afford to put 3 person-months of work for that feature.
There has to be something else at play here that I'm missing. Well, other than missing that "donate" button for a tad too long...
I think what they are saying is just: We got a lot of work done on this during those two sprints. It's not a statement about the work that is remaining.
I didn't quite get as far as I wanted, since the module system still relies on conservative stack scanning to find C-extension GC tools (because everyone else wasn't sold on JNI-like explicit local references), but it's still much more tightly specified than the Python API.
The Python extension API has another problem: it relies on FILE* and other assorted bits of the C runtime. That's mostly okay on POSIX-y systems where it's common for a whole process to share one C runtime, but on Windows, where different modules can come with different C runtime versions, this kind of leaking of objects across an interface boundary really hurts.
Does this seem reasonable? Is this even possible? I don't know much about the internals of PyPy...
What if we wrote that in pure python instead? What if we moved the computation to a GPU c extension? What if we use a diffent GC strategy?
Maybe I'm a sucker for the underdog, but in my mind, PyPy could save the Python ecosystem from irrelevance. People are looking at Rust and Go with the excuse of performance, and they are now the new Hip language. Even Ruby is catching up.
Without an answer, Python could be side-tracked for a increasing number of scenarios.
https://www.boost.org/doc/libs/1_61_0/libs/python/doc/html/t...
Are there any changes which would make Boost-based extensions better integrated/supported by PyPy?
The linked-to document only talks about Cython and cffi.
Thanks for the pointers to what's going on in the C++/python integration layer. I'll experiment with it.
For this, I think the Python will follow the same fate as Perl did.