https://www.reddit.com/r/ProgrammingLanguages/comments/81wkg...
I didn't know about it when I did the initial work for OPy in 2017 though.
I find the "expression-based" style quite interesting and I suspect it will help me understand the "compiler2" code better. I plan to look at it in more detail as I'm optimizing Oil.
This post links to a few posts that mention "byterun". These two pieces of work are probably what pushed me over the edge to apply to Recurse Center! I'm going from May-August this year. I'd love to connect with anyone interested in this kind of thing (e-mail in my profile)
-----
And if you have any advice on how to compile Python to more optimized code, I'm interested. I assume this will involve creating some new VM instructions (i.e. it can't just be done with the bytecode compiler)
This is a very concrete task -- the OSH parser is around 5,000 lines of code that would be very annoying to port to another language. Essentially, it's 3 interleaved recursive descent parsers and a Pratt parser.
I already have benchmarks that show it's 40-50x too slow:
http://www.oilshell.org/release/0.5.alpha2/benchmarks.wwz/os...
Leaving aside the rest of the shell (which is not big either), I think it's an interesting question if you can recover that factor of 40-50 without rewriting the code. It's written in a pretty "static" style without much dynamism. You don't need any special language features in a recursive descent parser.
I already did something like this. I wrote a whole bunch of Python regular expressions for the lexer, then compiled it to C code via re2c:
When are Lexer Modes Useful? http://www.oilshell.org/blog/2017/12/17.html
re2c code: http://www.oilshell.org/blog/2017/12/files/osh-lex.re2c.h.ht...
(And to anticipate a question from passers by: it does not make sense to use a parser generator here -- I wrote about this extensively on the blog, e.g. http://www.oilshell.org/blog/tags.html?tag=parsing#parsing)
For optimization I guess it comes down to making productive restrictions on the Python dialect to rule out some of the extreme dynamism. This sounds like a really cool project, one that's too big for me to have much idea what'd help without investing more time. What you're doing in stripping down CPython reminds me a little of how Luke Gorrie's started adapting LuaJIT to his own purposes: https://github.com/raptorjit/raptorjit
I wonder whether it might be feasible to import a Python module (running all the initialization code) and then walk the reachable object graph to serialize it into code in a more static subset.
Oil is filled with this pattern: do a bunch of metaprogramming at startup to make some data. Then use that immutable data for the rest of the program. There are very much two stages.
I think Lua-Terra might be closest to the thing I want, although I haven't had a chance to play with it:
I mention Bob Nystrom's language Magpie, which has an interesting model. Do the type checking at main(), not after parsing! But everything that happens before main(), at import time, is metaprogramming!
A Problem with Type Checking http://www.oilshell.org/blog/2016/11/30.html
I think I want to do compilation/optimization right before main(), not just type checking. It's all a bit vague right now, but I think OPy can go in this direction. Having the compiler written in its own language facilitates this. You can run code first, and then compile.
This post is also related:
Type Checking vs. Metaprogramming; ML vs. Lisp http://www.oilshell.org/blog/2016/12/05.html
Also note that C++ is a two-stage language too. In fact Herb Sutter just proposed that they unify the two languages. Like you can use STL with constexpr at compile time. Link on this page:
https://github.com/oilshell/oil/wiki/Metaprogramming
I think your suggestion is exactly what I've been thinking, so if you want to talk more about it / work on it, feel free to mail me :) The code needs a bit of work but I think it's a promising direction.
Yeah I agree that you want to remove dynamism. For something like BINARY_ADD, the operands will only be strings or numbers in a parser. You can cheat and just assume there is no operator overloading.
Although I don't know how much that will actually speed things up! I should make a profile at the C level. I have profiled at the Python level and sped things up 6-7x already.
The other thing I think will help is using something like spans/slices instead of strings. Parsing creates all these tiny string objects. And also as I mentioned changing the representations of the nodes, which is a fairly naive Python representation now. Python objects are huge!
-----
I didn't know about raptorjit, but I had heard of Snabb Switch awhile ago on Hacker News. It looks interesting!
It reminds me of the Dart language being inspired by v8. The problem with v8 is that you can change one line in JS and your performance will just fall off a cliff. It's hard to detect unless you write benchmarks, which most people don't. So Lars Bak started Dart to remedy that problem, designing the language around stable JIT performance!
There's a good chance I'll still be around in late May. See you then!
I mapped out the work the other day and it looks fun. It's just at the edge of my knowledge.