PHP Next Generation
php.net
php.net
But I wish it was done a little nicely from the community point of view (the >4gb strings stuff).
I'm considerably less talented than Dmitry and I spent several months of my life trying to unsuccessfully write a JIT for PHP - most of what stopped me was the rest of the Zend engine itself.
While I was maintaining PHP-APC, I spent many weeks trying to write a basic block JIT for php, when Zend is using the CGOTO core (FYI, if you are still using APC, switch to Zend OpCache).
This would compile code which didn't have any jumps into a native chunk and swap out the opcode's handler location into my native chunk.
The little I did actually do ended up being fairly involved assembly rewrites of the inner loop.
http://notmysock.org/blog/php/optimising-ze2
No matter what I did, the issues of the bytecode organization (the ->result reference) and the lack of type verifiability in the code generation resulted in me slowly throwing away every prototype somewhere between the for loop and running the default benchmark.php.
I haven't read through all the changes yet, but IMHO the Zend engine will be an absolute pain to deal with until we get to type inference/verifiability into the bytecode format so that integers get integer register ops in the JIT instead of always being zval_* based.
But a cleanup was due. And a faster VM (either HHVM or PHPng++) is good news for the regular PHP users.
The size_t/int64 refactor has been going through several iterations and RFCs since 2013.
They could have been integrating the changes in their code as they went, and then launched with 'Hey, we have this cool new project, including the optimization work that people have been doing'. Instead, they kept their work secret, knowing that their work was incompatible with another large refactor that people were contributing to, and then said 'hey, abandon all the work you were doing because who cares'.
In particular, their primary advertised concern was on the 4% increase in memory usage - but the whole purpose of phpng seems to be to make PHP more compatible with JIT compilation later on - and that seems to almost universally requires large amounts of memory to work really well.
The thing is, that is all being achieved despite how badly the internals is run. Seriously, I tried to get involved and was turned off entirely by how it's run as an "old boys" clique where new developers and changes are considered enemies. It's really sad, and I think PHP could become a much nicer language overall if that changed. Some people who are far smarted than myself have the same opinion too, so it's not just me...
Then, the arguments made in the discussions were basically besides the point. Overall memory usage is not really important on modern hardware, since memory is cheap. The real concerns are time performance - if the performance is degraded, it would be a valid reason to reject some of the changes. The vast majority of criticisms were just along the lines of 'but why do we even need 4GB+ strings' - I can see exactly why Pierre would get frustrated at having to parrot the same line. Performance in current php was clearly fine, and performance in phpng was only ever mentioned in completely vague terms - the most empirical it got was 'i guess 20-30% worse' which isn't really a trustworthy figure by any means.
The burden there was up to the phpng team to recognise that there were changes in the standard php pipeline that would potentially invalidate some optimisations. Coming in at a late stage and saying that it would ruin everything, having made no effort to stop someone wasting much time, was at least inconsiderate and probably rude also.
Then attempting to effectively filibuster the situation by repeatedly firing irrelevant arguments at Pierre, then recruiting randoms to try and vote against it as some kind of 'we are being undermined' campaign, then trying to change the voting rules, it all just smacked of a very amateurish and/or rude community. Yes, they were mostly by the same people, but at some point someone could have very easily come out and pointed out (for example) the inherent fallacy in the 4GB+ argument, or the fact that the phpng developers should have informed Pierre earlier that their new optimisations might conflict. It seemed like a number were happy letting some fight their preferred argument with the wrong reasons.
The arguments of 'wasted data' aren't convincing - the arguments of 'wasted data causes poor performance' would be convincing were there any data whatsoever to back it up, and if those who argued actually worked from that standpoint.
My problem was (as someone who read the mailing list) was simply that the very real concerns were brushed under the carpet, and that the situation was allowed to develop in the first place.
Anyone have any pointers to the conversation they are participating in with this post?
(mavci's link is an O'Reilly summary of the state of the PHP ecosystem and I didn't see any mention of PHPng.)
UPDATE: Perhaps this link is what started the confusion the php.net post appears to be trying to address: http://grokbase.com/p/php/php-internals/1455aesx7r/phpng-%04...
Otherwise, yeah, Hack is an improved PHP. To be clear, it adds additional functionality to PHP, but is still entirely backwards compatible. PHP >= 5.4 still has solid bones.
80% of the reason I use hacklang is just so I can extend whatever base PHP classes my CMS provides, but implement my methods without having to add the piles of manual type-checking that PHP requires.
Actually Java works in the same way; the confusion stems from imprecise use of jargon. Java tends to use the word "type" for things which would more precisely be called "tags". Tags are the run-time information which distinguish different values of the same type, and can't be erased in general.
Java adds a bunch of rules about handling tags, for example object values have a "class" tag; class tags must be statically specified (AKA "type signatures"); class tags can be pattern-matched automatically (AKA dynamic dispatch, method overloading and inheritance), allowing functions to be defined in separate chunks (AKA methods); functions can only be applied to arguments which will match a pattern (AKA "type checking"), etc.
These rules are checked at compile time as well as the types. Unfortunately all these different concepts tend to be grouped under the umbrella term "type checking", which makes fine-grained discussion and comparisons to other languages more difficult.
This separation is different than many other languages, and lots of folks have found it confusing. It largely exists for technical reasons, and we should be able to have a better UX for end users; this is something I've been thinking a lot about and want to try to improve in the language going forward. (I work full-time on the Hack team.) Please do give it a shot and let us know how it goes!
I'm really interested to see what you end up coming up with UX wise. What sort of edge cases do you think the `hh_*` apps will miss? Running `hh_client` on save isn't a bad thing, IMO, and having it output JSON for easy integration is amazing :) But I'm curious what the type checker will miss currently?
Is there a public mailing list where these sorts of discussions happen or is it more internal to Facebook? I'm really interested in Hack (as is where I work, we push PHP to it's limits so we're keen to use something that can catch even more bugs from day 0) so thanks very much for pushing PHP even further forward!
Are the technical reasons for missing type errors that wouldn't be triggered at run-time insurmountable? Or is it more a "finding the right UX to expose it"?
There isn't a mailing list right now. A few discussions are just in-person since we sit next to each other, but we're trying to make things as transparent as we can. Lots of stuff (and hopefully more moving forward) happen on #hhvm on Freenode or on GitHub issues -- basically the same channels as the HHVM project itself. There is also nontrivial discussion during code review, which is very unfortunately all internal right now; we hope to have that all moved external as soon as we can. (There are a lot of tricky integrations with internal tools that need to happen for that to work.)
Most frameworks work with hack now, so your right it would be easier if there where just 1 version of "php".
https://github.com/facebook/hhvm/issues/1787
I guess the class has been renamed in trunk, but according to this, the syntax differences still remain:
https://github.com/facebook/hhvm/issues/1871
And while we're at it, might as well point out the missing Closure::bind and Closure::bindTo from PHP 5.4:
- it runs really well on expensive, FB hardware. there's no consideration given to anything else (performance-wise) in development. that's not to say it's slow, but it works best on 64GB servers with fat SSDs.
- it's a huge moving target when it comes to php5 compatibility -- there are a few intentional inconsistencies and a lot of unintentional ones. fixes for zend incompatibilities have a whack-a-mole effect: the typical patch fixes one inconsistency and introduces a few more.
- bad documentation and bad code quality
- it is open-source, but only in the most superficial sense. it's really an FB internal project, so good luck making any changes that help you but don't help FB
it's still better than zend PHP, and, to be fair, most of the incompatibilities are in cases where zend behaves stupidly. source: i was an HHVM contributor. (edited to space out my ascii list)
the fact that no one who doesn't work at FB is allowed to merge code (correct me if I'm wrong) made me feel like I was helping FB more than the OSS community. I get that HHVM is business critical to FB, but if you want it to truly get community support it needs to be spun off from FB.
edit: and I will add that my language was way too harsh. it's not hard to get in a PR provided it doesn't break anything in an FB internal test suite which I can't see (which never was a problem for me, but I could see it being one) and didn't degrade FB performance.
The code quality, however, isn't 'bad' by any stretch of the imagination. The FB team managing it is very knowledgeable about PHP internals, and are appropriately strict about determining what zend-compatibility fixes to commit. Their code's overall quite well-written, which makes me care a whole lot less about how poorly it's documented.
sparsely commented
I already like the sound of it :)So Facebook or any other company running code on their own server can shift to HHVM with (relative) ease but open source projects (like WordPress) and libraries can only switch once everyone else switches, Which leaves us in a catch-22, so Hack remains forever in a symbiotic/parasitic relationship with PHP.
From comments it seems like some people are afraid that FB will keep driving HHVM project in it's own way. But this fact is dependent on contribution. FB may loose grip on HHVM if If people outside FB start contributing more on it.