Prototype PHP interpreter using the PyPy toolchain - Hippy VM
morepypy.blogspot.com
morepypy.blogspot.com
It would be nice if there were comparable benchmark results available, or a discussion of what is different between both approaches/implementations.
EDIT: bb link didn't work, replaced it with ACM portal link.
They were using an old version of PyPy and did not use some of the advanced features of the JIT generator.
It's interesting that you mention the advanced features. I looked at Hippy and the most interesting JIT feature Hippy uses is _virtualizable2_, which virtualizes all function locals and unboxes them. We tried using it ourselves, but it forces each function to have a static list of variables and no dynamic variable accesses (like $$x). It looks like Hippy falls back to the regular implementation for dynamic variable accesses, where the entire list is stored in a dictionary. Now I'm wondering how much this happens in real-world code, we assumed it does happen enough times.
Also, I'm working on posting a publicly-available version of the paper. I'll post a link when I do that.
In any case, keep up the work!
I've been toying with the idea of building a little language, just for fun. The mess of building an interpreter has always kept me from it, but as I've spent some time looking at PyPy, I think I might just have a go.
http://morepypy.blogspot.com/2011/04/tutorial-writing-interp...
http://morepypy.blogspot.com/2011/04/tutorial-part-2-adding-...
http://tratt.net/laurie/tech_articles/articles/fast_enough_v...
One of the big goals of the project is to create a general tool for people to implement interpreters and VMs. The intention isn't to just provide a JIT vm for python, but to allow all sorts of new or existing languages to be (re)implemented easily. Don't be surprised if you start hearing about other languages being implemented in RPython.
I used ply to write the parser (aiming at 1.9.3, which has a lex/yacc parser), and that was the first time I leveraged a LALR parser, so it's very experimental, and as such code quality is lacking to say the least. Also, since ply uses (at least) kwargs, it won't pass through pypy build step. Still, early performance between cpython and pypy gives a clear edge to pypy.
To give you an idea of how much experimental it is, it's not even in my ~/Workspace (which is usually the step before github), but still in ~/Sandbox/pypy/pyby. I actually had a hard time finding it again.
In addition, an implementation in pypy should be able to support eval, which is impossible (or perhaps extremely difficult) with hiphop.
I was referring to the benchmarks in which they measured Python code in pypy running against pure C code (see some examples below). While I'm not familiar with the benchmarks you're referring to, I doubt anybody implemented all (or even parts) of Django, genshi or html5lib in C.
http://morepypy.blogspot.com/2011/02/pypy-faster-than-c-on-c... http://morepypy.blogspot.com/2011/08/pypy-is-faster-than-c-a...
If I understand it correctly, most of their backend services are implemented in some other language, and their PHP code is mostly used as a more flexible templating language. If that's the case, it shouldn't be all too hard to migrate away from PHP, if the chose to do so.
It seems like they spend a lot of engineering effort in optimizing PHP, and probably also a lot of CPU cycles in executing it. I assume there is a tipping point somewhere, when the investment in PHP stops making sense, even given effort to port legacy code.
There are a lot of arguably nicer alternatives to PHP, and I bet they are not benefiting from the one really superior feature of PHP (easy deployment on shared hosting).
The standard execution model for PHP apps is, request comes in, code is executed, code goes away, resources returned. PHP is optimised to suit this model. For nearly anything else, you have one or more long-running processes which take requests, thus skipping the code loading/interpreting/compiling stage, which tends to be expensive.
It's quite easy to do python CGI on shared hosting without concerns, but it's not going to be fast. Once you start looking at things like FastCGI or WSGI, you have a process which, at least some of the time, is persistent.
http://corp.galois.com/blog/2010/11/30/galois-releases-the-h...
"the benefits of the JVM only apply to languages that map well onto Java concepts"
With InvokeDynamic (InDy) this is no longer true at all. InDy lets you manage the entire lifecycle of method dispatch at a particular call site in a way that HotSpot can optimize across. This lets you define your own dispatch semantics that can potentially do things like complicated argument conversions (e.g. rolling all of the arguments up into an array) and it will still dispatch as fast as Java. This is true because InDy lets you define your own polymorphic inline method caches at call sites, so you can have a slow lookup for a method handle which is cached and invokes as fast as InvokeVirtual the next time it's called.
HotSpot can also inline code using InDy just like it can inline virtual method dispatch in Java. This lets the JIT compiler optimize across several methods, potentially implemented in many languages, as if they were a single body of code.
InDy will play a particularly important role in JDK8, where it's used to implement lambdas.
As for JVM, which is definitely "done", you don't get access to low level concepts and you don't get support for most dynamic language concepts. For example in Jython escape analysis does not work at all, because everything escapes via frames (which in python is accessable from application level using sys._getframe). Invokedynamic only helps marginally here - you still need a JIT that can optimistically remove unlikely paths of the execution or a very smart compiler-to-the-JVM, which is unnecessary in PyPy.
The other element is low level stuff. JVM is opaque, you can't code in terms of C structures. RPython as well, but it allows you if you insist and sometimes there are very good reasons to insist. Look in rpython/ directory in the hippy checkout for an implementation of an ordered dict with all it's oddities. In the JVM if such a primitive does not come (and it's unlikely enough), you're out of luck.
Eg. class loading in the JVM is (afaik) still static, classes once loaded cannot be modified, which is a problem for languages such as Ruby and JS, the basic data types of the JVM might not match your language, etc.
Not sure if that makes a big difference, in essence you'll have to implement those semantics in either case, be it in RPython for PyPy or Java for the JVM.
http://docs.oracle.com/javase/1.4.2/docs/guide/jpda/enhancem...
To be eligible to start a Kickstarter project, you need to satisfy the requirements of Amazon Payments:
—You are 18 years of age or older.
—You are a permanent US resident with a Social Security Number (or EIN).
—You have a US address, US bank account, and US state-issued ID (driver’s license).
—You have a major US credit or debit card.
So it's a big no-no for a lot of people who happen not to be residents of the US, like me.Not sure if any of these are worth looking into. You definitely lose out on the brand weight that comes with kickstarter.
It's absolutely not my call to say whether making hippy complete is easier than making hiphop fast. I can provide you estimates on the former, if requested, I cannot potentially provide estimates on the latter.
In short - LLVM is great if you want to write a language like C, PyPy is great if you want to create a dynamic language VM like Python or PHP.
http://morepypy.blogspot.nl/2011/04/tutorial-writing-interpr...
However, I hope to have some time in the fall in which to pursue this. If anyone else takes it up in the meantime, I'd love to contribute.
Interestingly, a native CoffeeScript interpreter already exists: Poetics, which runs on the Rubinius VM. Very cool project. Another very interesting CoffeeScript project is the MoonScript project, which compiles a variation of CoffeeScript into Lua.