PHP: The Right Way
phptherightway.com
phptherightway.com
Missing test and QA tools though. Probably an oversight, since the author does suggest following Derick Rethans and Sebastian Bergmann.
And another one: read up on the SPL library before you start re-inventing that wheel.
I've learnt JQuery and Python lately, but my Swiss army knife is still PHP. So this has already fixed a lot of incorrect techniques I've been suing in a side project.
It is fun how much a very old moralist sentence by Confucius applies well to PHP. He said 其本亂而末治者否矣 which can be translated, in software language development context, as "Build a nice, reliable language on shitty definition? Bullshit!".
(The word-by-word translation is "your - root - messy - and/but - leaves/result - governed/orderly - have ? not - hey!")
This sentence also resonates with the odd Turkish encoding PHPbug recently reported here: if you have so many wrong architecture decisions in the core of your product, no amount of patching and "under the rub" filling will save you.
Edit: added word-by-word translation + typo fixes.
A good portion of why people like to mock PHP is due to the way people end up using the language rather than the language's ugly parts. This guide is a great step toward pointing people in the right direction with the language. PHP has come a very long way and continues to improve. A big chunk of the battle is having decent guides to replace all the crap that's out there showing new PHP devs the wrong way. So why not talk about the merits of the guide rather than taking an easy opportunity to shit on PHP. It's really gotten old.
Hey, remember when JavaScript was like the worst language ever and everyone felt the need to remind everyone about that constantly? I do. And now all of them are writing blogs about how awesome node.js is.
Jeff Atwood says PHP sucks
|- Says we should make other languages fill in on what PHP does best
|- Post goes on HN
|- PHP apologists say "No! PHP is fine, I've made a career out of it!"
|- 'Atwoodians' continue to reject PHP, ruby/python need to be more accessible
|- PHP apologists make reasonable arguments for why PHP isn't as bad as everyone says, while admitting its probably worse than python/ruby
There remains a conspicuous lack of response from the 'Atwoodian' camp regarding what could be done to add PHP's quick-start strengths to other languages so that PHP developers can develop similarly to how they currently do, just with a less wonky and problematic base language.By no means, as an outsider, do I perceive the PHP apologists as having redeemed PHP to the point that the discussion should be about 'how to fix PHP' rather than 'how to discourage its propagation' as was originally put fourth in a fairly reasonable way by Atwood.
I realize why saying 'PHP apologists' is problematic, but I'm mostly seeing responses that are in themselves critical of PHP, while espousing its widespread use, quick-start capabilities, and developer intertia--and not espousing why the fundamentals of the language are better of the equals of ruby/python.
Having said that, it seems to me that there's certainly a fair amount of work being done to make deploying apps in other languages pretty simple. With mod_passenger, for instance, setting up a Rails app under Apache isn't notably more difficult than setting up any other Apache virtual host. Deploying Python web applications isn't quite as simple yet (at least in my experience), but it's nothing that should be beyond the ken of someone who's figured out how to write a Python web application in the first place. And that, in turn, isn't really beyond the ken of anyone who's learned how to write reasonably good PHP MVC code.
However, I appreciate very well the need for such a guide when you have no other choice. I did not read all of it, but it is certainly necessary to have strong coding conventions and a defined set of "must do/musn't do". I remember we had even built a code checker that was forbiding the use of many php functions and syntax, and had a set of wrappers around basic functionalities like string functions or date functions.
He is sharing some useful guidelines to approach it more properly and help people that are still using it.
I, for one, am a fan.
Yeah, that doesn't mean anything. In fact, to be frank, when I read that, I fill in the underlying context:
"I used PHP for fives years professionally, and the code was horrible. Granted, it was code I wrote because PHP let me write code that way..."
or
"I used PHP for fives years professionally, almost a decade ago..."
> I remember we had even built a code checker that was forbiding the use of many php functions and syntax, and had a set of wrappers around basic functionalities like string functions or date functions.
Yeah, that didn't do much to change my mind.
If you take it personally, thats your issue, not mine,
Also, magically, I can't use empty(functionResult()); but I can say $result = functionResult(); empty($result);
There are poorly designed parser rules but the namespace operator is not due to that. It's due to the fact that a single PHP file will be compiled to byte code without knowing what the symbols represent until runtime. The separate operator is needed because names are resolved before execution begins and before the symbols are known.
> Also, magically, I can't use empty(functionResult()); but I can say $result = functionResult(); empty($result);
Empty is identical to the not (!) operator except that empty() can operate on undefined variables, undefinted array keys, or undefined properties. empty(functionResult()) doesn't make any sense because functionResult() can never be an undefined variable. You just use !functionResult().
Scheme has:
call-with-current-continuation http://people.csail.mit.edu/jaffer/r4rs_8.html#IDX509
Tail Call Optimisation
A numeric tower http://people.csail.mit.edu/jaffer/r4rs_8.html#SEC50
A Code/Data equivalence and therefore macros
Ports
Javascript has: Prototypal inheritance
Every 'object' is mutable bag of string indexed properties
There's a pile of differences, even just at the semantic level before you get into the syntax or the equivalence semantics.I understand that if you take a specific subset of javascript, you can write code as-though it has some scheme semantics (minus call/cc and TCO), but that's true to a similar extent of any language with lexical scope, closures and anonymous functions.
JavaScript could stand to be a bit less verbose (look into CoffeeScript for some relief if you value your sanity and time) in some areas and type coercion should just go away; but if you use strict mode and strict equality, a huge amount of pain is alleviated.
If one take the time to learn about constructor functions, the prototype chain, and own properties; they are 70% there. Most people can get by on just that.
Are you really going to trot out that old lie?
Why not read this, instead? http://me.veekun.com/blog/2012/04/09/php-a-fractal-of-bad-de...
First you demand some "foundations". But if you look on virtually all existing popular languages, none of them were designed exactly in the form they are now. Java had tons of changes and additions, Python had object model change and now has new version that changed so much that it's not even backwards compatible, etc. etc. Does it mean they lacked "sane foundations"? No, it means requirements changed, so did they. PHP changes too. You want formal definition of PHP? But what would be the benefit of it? Who would benefit from it existence? Without answering these questions it is hard to expect anybody would create it.
Then you proceed to focus on some obscure bug that 99.9999% people couldn't care less about, with implication that since this bug is not fixed whole language is crap. I don't even know how to address this - do you seriously expect this be taken as an argument?
1. Ad hoc language design process (if you can call it that at all) resulted in a serious lack of forethought in the language's case sensitivity rules. Somehow they ended up with about half of the language being case sensitive and half being case insensitive. Somehow they ended up not considering the implications of "case insensitive" with regards to internationalization, and thus ended up with utterly bizarre rules where code works or doesn't work depending on the locale it's run in.
2. Extreme emphasis on backwards compatibility means an unwillingness to fix the bug by e.g. making the language fully case sensitive, or regularizing the rules for case-insensitive comparisons among language identifiers to be locale-independent. This could be phased in over a long period of time with progressive deprecation to ensure that anyone relying on the old behavior has years of notice to fix it.
3. Somehow, the broken case insensitivity of these identifiers also breaks case identical lookups for certain names. Despite the bug existing and being discussed for ten years, this still doesn't work. No matter how bizarre your case-insensitive comparison rules are, there is absolutely no reason whatsoever that looking up a class using the exact same byte-for-byte name should fail. That this hasn't been fixed means that either the PHP team doesn't care, has no idea how to fix this, or the code base is too broken and crufty to make a fix practical. None of these alternatives says anything good.
The issue itself is not that important (unless you're Turkish, or running on Turkish servers, or somehow end up running in a Turkish locale), but the fact that it exists and has not been fixed in a decade means something is seriously wrong.
>> Somehow they ended up with about half of the language being case sensitive and half being case insensitive.
Why you think it's because of the "lack of forethought"? Case-insensitive methods/classes were decided to be so for a reason (since it made easier to write bigger projects in PHP). You may not like this decision, but it doesn't mean everything you disagree with is because nobody thought about it.
>> Somehow they ended up not considering the implications of "case insensitive" with regards to internationalization, and thus ended up with utterly bizarre rules where code works or doesn't work depending on the locale it's run in.
Again, the only reason why that bizzare situation with Turkish locale exists is because it wasn't noticed. There's no deep and mysterious reason rooting in flawed PHP design why it can't be fixed - as far as I can see, the fix is not hard either, it just that nobody got to fixing it.
2. Changing language implementation with potential of breaking of 99% of existing code is not a "bugfix". If you don't understand this, maybe it's too early for you to lecture people on how to design languages. It can be done, but there's no reason to - it would lead to massive breakage with no upside in any functionality.
3. "The issue itself is not that important" - BINGO! You said it. If the issue is not that important, why you're wasting time dwelling on it and make it sound as if this is the most important thing ever? You wrote a long post supposedly about deep flaws in PHP design and process - but the only thing you actually wrote about is an obscure bug that you yourself admit is not important! Maybe it's time to discuss something that is important then?
is creepy. Never ever run other people's code without at least giving it a glance.
Or the second time that an IP address hits it.
Or something else that you haven't thought of.
Update: Such as doing something nasty to http://getcomposer.org/composer.phar instead - a 520KB php file that the first script downloads and runs without even md5ing first. Did you audit that too? Did you understand everything that it was doing?
A far cry from the safer/verified "download this and check it's MD5 checksum" method that I'd prefer.
Seriously, fixing package management so we can continuously integrate arbitrary code would be great. Getting arbitrary OS package creation to be almost as easy as pushing code to GitHub seems like a very worthy goal.
Also, pushing random code via apt-get is not going to win you any sysadmin friends. They like their servers to be stable, and their packages to be well tested.
Much better not to be broken in the first place.
Viable verification is based on an unbroken hash (preferably SHA-(256|384|512), but SHA1 still not broken even it it's definitely too short nowadays. And the hashes need to be distributed properly, either PGP signed or hosted on a trusted TLS protected site...
Yet many projects still distribute MD5 hashes, thinking it got any value.
wget mysqltuner.pl
perl mysqltuner.pl
(Yes, they actually have a .pl domain for it)
> Never ever (ever) trust foreign input introduced to your PHP code.
Where'd I put that sense of irony...
Unless you've read every line of the Linux kernel, we all succumb to the later point sooner or later.
'apt-get install whatever' is just as magically scary and dangerous.
No it isn't, there's GPG signing and things going on there.
That's really just Cargo Cult security, isn't it? Signed packages can just as easily be malicious. In fact a repository server could be a much worthier target for the injection of bad code than a single, relatively obscure web project.Protecting against that is the whole fucking point of using things like apt-get/PEAR and GPG/code signing.
In context, what you say makes no sense.
Regardless, https would be better, but if you are that nervous as you suggest you are, then https doesn't solve your problem either. Neither does some hash thrown up on some site for you to compare against.
aw3c2: 'curl | php' is creepy
timaelliott: apt-get is just as scary
kudos: no, it has signing
udo: that's cargo cult
me: no, it protects like https (meaning it authenticates and stops MitM)
I was responding perfectly in context to point out how the protections apt-get have are useful. I also tied it back into the original comment but you can ignore that part if you like. I don't know why you think I was 'trying to be witty' and ignoring context.Still, https doesn't validate anything. You could add a certificate, but then that only means you are talking to the server assigned the cert, not that the actual package is good.
I still stand by what I said: mountain and mole hills.
Mr. Cantor, I have added to the list of people to never hire I keep in my notebook.
I do not trust a random website that could easily be MITMd.
> curl -s http://getcomposer.org/composer.phar -o $HOME/local/bin/composer ; chmod +x $HOME/local/bin/composer
Assuming that $HOME/local/bin is in your $PATH, the current user will have a "composer" command available.All told, I love site, and I hope it keeps iterating. PHP may be ugly, but it's powerful, and most of its bad reputation comes from good coders having to pick up the pieces from bad coders.
"When various authors collaborate across multiple projects, it helps to have one set of guidelines to be used among all those projects"
Spaces, Tabs, OTBS, whatever. It doesn't matter which one was chosen, just that something was chosen. Ask the Pythonista's what PEP8 did for their developer's ecosystem.
Also it's important to note that that PSR-2 was based off already existing coding standards like Zend's and Symfony's.
I think it's silly to pursue per-language consistency; it's a pipe-dream that just leads to more religious wars over minutia. Per-project (or per-organization) more than suffices.
Reformat their code.
As for "per language consistency", Go does it just fine...
4.2. Properties
This guide intentionally avoids any recommendation regarding the use of
$StudlyCaps, $camelCase, or $under_score property names.
Whatever naming convention is used SHOULD be applied consistently within a
reasonable scope. That scope may be vendor-level, package-level, class-level,
or method-level.
4.3. Methods
Method names MUST be declared in camelCase().I've seen this happening in C, followed by C++, and now Java and C#.
They will do it regardless of the language being used.
Java itself can be a beautifully terse language.
Compared to almost every other language I've had exposure to, it's more verbose.
That said, it is difficult to bring a legacy code base in line with modern style, though you can improve it over time.
Also, this could benefit from some other gotchas, extremely surprising behavior and best practices for avoiding common pitfalls.
But it goes without saying that protection from SQL injection is a must, it is just PDO's statement binding is not the only way. MySQLi does support data binding, and for others it is easy to DIY using sprintf and whatever is the escaping method target native driver implements.
Edit: Everyone talking about databases: paramaterized queries, check them out.
SQL injections are very rare to non-existent in code where the programmers generally rely on parametrized queries.
SQL injection occurs when you're not escaping data while producing output, namely, an SQL query sent to the DB.
XSS attacks occur when you're not escaping data while producing HTML, but you don't need angle brackets to do it. <a title="$string"> allows for XSS injection with just a quote character.
Header injection attacks occur when you're not escaping data while producing HTTP/MIME headers, and all you need is a line-break character.
The escaping always depends on the output context, and the data in the database cannot be made safe for all these contexts. You will still need to allow people named "O'Brian", let people post "<_<" smileys, use Unicode in their names, etc.
No. The PHP world recommends, time and time again, using PDO and binding variables to queries. I've yet to meet an individual who does it the other way, other than people who are relying on extremely outdated tutorials (7+ years ago). Hell, even this document does. This document, unfortunately, uses the word filter in the wrong way, but the intention is still fine.
> I do an exact equality check against a white list.
Pretty much the way you are supposed to do it (except, I'd actually use the value in the white list rather than the matched input value, as the two should not be the same).
These functions should be used to "validate" input rather than "filter" it.
The filtering of input has nothing to do with security (that should be handled by output escaping and paramterized queries) and should only be used for improving the quality of the saved data.
I <3 O'Reilly books
you could pre-encode that as safe to paste into raw SQL, or as well-formed (X)HTML, but you can't do both simultaneously. Either encoding would end up distorting the content in the other context. You have to encode during output (and writing to a database counts) using the rules of the system consuming that output. Lots of crappy web forums visibly mangle punctuation in a futile effort to avoid this.Validating on the other hand is checking if the values satisfy the expected requirements.
You only want to apply certain filters if you know they apply the current output. (htmlentities if output as html, mysql_real_escape_string if embedded in a query).
(1) normalization (things like trim, etc. so you don't do stupid things like invalidate an email address just because it was input with trailing spaces). The result is passed to a validator.
(2) validation (if a number is expected, it's only valid if given a sequence of digits within the correct range). Valid values are passed on to the next phase (likely persisted).
Sanitization is generally not needed and in most cases likely to cause data corruption. I'm sure there are use cases for it but it should be the exception, not the rule.
Libraries with different code styles can be used together without problems. It seems like they are using the PSR to declare they're own style as superior.
My only complaint with it personally is that I Can. Not. Stand. putting opening brackets on their own line.
I'd MUCH prefer people spend time writing tests for their code vs debating/arguing/refactoring code style. If your tests and good and coverage is high, there's far less chance I should ever even have to muck around inside your libraries, much less modify them.
And again, one doesn't preclude the other, but I see so many people making a bit stink over style - tabs vs spaces et al years before the PSR stuff - yet rarely do I see even 20% of that effort spent on getting people to test.
I say this not as someone who tests everything 100%, but as someone who doesn't do much of it, and understands the importance of it. When I'm deciding to use someone else's libraries, I never make a decision based on the code style they've chosen - I base it on maturity, documentation, and tests/examples provided (when there's a choice, sometimes you don't have a choice, or you roll your own). The really good libraries? I never even have to look a their code - it might be all run together on one line for all I care - it just works.
I agree that, of the two, good test coverage is more important than pretty formatting, but a good code style is going to make the code easier to debug when those tests fail and easier to extend when the library needs to do more. It's a simple matter adopt a coding style that works and stick with it, so there's really no good reason not to do it.
By comparison, it does seem like a trivial thing to obsess over. There are a number of coding styles which are all fine in their own right. I think people should worry less about the specifics and more about the overall intent-to write consistent, clean code. How many blank lines you have after block of code or where you put the opening bracket doesn't really mater in the end, as long as you follow the same structure across the whole code base.
Standardizing formatting for a team/project is not a bad thing, certainly, but if there's competing styles on a team, but they're making good progress - hitting deadlines, high test coverage, good/deep engagement with stakeholders, etc., code formatting is just not something I'd bother enforcing. It may be a later step to go back during a project post-mortem and get the code ready for 'deep freeze', assuming the project is 'done', but how often does that happen?
My perspective is probably a bit different than some here - I freelance, and work in multiple languages for different clients, often concurrently. In a 6 month period I had 2 PHP projects, 2 Grails projects and a Rails project, each running for several months and overlapping. I worked with a different group of people on each project, each with varying skill levels and backgrounds. I adopted my style/technique as best I could for each project, and didn't harp on formatting with anyone, because... it just really doesn't matter. What does bug the heck out of me is someone rechecking out my code, spending time reformatting it, then checking it back in, and counting that as 'work', when there's many many many other issues needing to be worked on.
However, you miss the point. The point of this isn't really to define all the options. Rather, it's to spell out best practices (and PSR-0 is really a best practice, and all about interoperability) that the community generally agrees with. Better to have a standard coding style than to have none at all, and if you have to choose one, PSR-1 and 2 are an excellent choice (considering how it was devised).
The whole concept of PSR-0 is ridiculous anyway, because PHP supports registering multiple class loaders. If a project wants to use a naming convention that won't work with standard spl_autoload(), they can register their own autoloader.
The built-in spl_autoload offers better support for modern code than the PSR-0 recommended autoloader
Python has had PEP-8 for quite a while now, and it definitely helps in improving both code quality and understanding random libraries that you find on the internet. It doesn't catch everything, but then that's what Pylint is for.
That being said... Every language has its strengths and weaknesses. There is nothing fundamentally wrong with PHP (except maybe that if you look at it it's not a very exciting language, even a bit boring). Security for instance has nothing to do with the language itself. And as for PHP syntax, it's a blessing compared to some other languages; at least we don't have a semicolon debate in PHP land ;-)
There are a LOT of long term php developers who have never used them.
Edit: also might be worth adding a bit about steering clear of phpclasses.org as well as w3schools, they bother contain far more bad, than good code.
I last used PHP back with PHP3 (and then went C++ => Java => Python => Python/JavaScript => Ruby => Python/R), but a bunch of code I want to read at work uses PHP (with Zend). I no longer remember most of what I learned about PHP3, though obviously the PHP syntax seems to be at least somewhat readable as a sort of amalgam of Perl and C++ syntax and idioms. What does, e.g., Facebook use to get engineers who don't know PHP (but might know C++ or Python) up and running?
There is a lot of good code to look at and learn from, and the Bootcamp program teaches a bunch of our own library and some guidance on what good and bad standard library functions are for certain cases. We also have a "newbie" group where people can ask questions about the codebase. And if you do something without knowledge of a built-in (or Facebook) function/class and make things more complex than they need to be, your code reviewer will let you know how to do it easier/more idiomatically. Every once in a while there are language tech talks, and sometimes even whole-day training sessions (we did one on exciting new C++11 features).
Besides that, I definitely spent a lot of my time with php.net/function_name open and exploring when I couldn't remember all the functions available or the parameters and order.
A good example is: python-guide.org -- http://docs.python-guide.org/en/latest/index.html
user: kaolinite
20 year old Python/PHP/C developer
That font will be just perfect for you in about 15-20 years.Edit: Just checked, definitely just the stretching, though a bit smaller would be better for me.
Oh! The guide does not mention the case: if a class contains only one static method, please, use a function. It does not look as educated, but it's obviously a function.
Not PHP specific, I admit.
Refactoring to a single function took it down to under 20 lines.
I wish I knew what motivated people to do things like this.
Hell, I even went there myself at one point, at first doing it to DRY up a bunch of procedural scripts, but later introducing more abstractions to bend the existing code to my will.
Now, I prefer simplicity and will often scrap code that jumps through all sorts of hoops just to do something simple.
int main(int argc, char *argv[]) {
for(int i=0;i<10;++i)
printf("%d\n", i);
printf("%d\n", i);
}
The last line fails to compile—'i' is no longer defined after exiting the loop.Java:
public static void main(String[] args) {
for(int i=0;i<10;++i)
System.out.println(i);
System.out.println(i);
}
That also won't compile.Coming from a language that supports variables with block scope, PHP's behavior is very surprising indeed. If the first mention of $object is in the loop, I can understand someone expecting $object in the second loop to be a distinct variable that just happens to share the same name.
This behavior is even surprising coming from Perl, though for a slightly different reason:
@array = ('c', 'c++', 'java', 'perl');
foreach $item (@array) {
if($item eq 'perl') {
$item = 'php';
}
}
foreach $item (@array) {
print $item . "\n";
}
For that matter, PHP's reference semantics are a little surprising in general coming from any other language I've ever used. #include <stdio.h>
int main()
{
int arr[10] = {1,2,3,4,5,6,7,8,9,0}, *ptr;
for(int i = 10; i>=0; i--)
{
ptr = &arr[i];
printf("%d\n", *ptr);
}
printf("%d\n", *ptr);
}Also note that, in C, you have to explicitly dereference pointers:
ptr = 42;
You should get a warning from your compiler if you try that. PHP effectively goes ahead and changes it to: *ptr = 42;
Perl requires you to explicitly dereference references as well: @array = ('c', 'c++', 'java', 'perl');
$item = \$array[3];
$item = 'php'; # Does NOT replace the item in the array
foreach $item (@array) {
print $item . "\n";
}
PHP is weird in this regard: if you ever assign a reference to a variable, later assignments go through that reference automatically. To convert the variable back to a normal variable, you have to unset it. I don't know of another language that acts like this. Can you think of one?Combine that with the lack of block scope, and you get surprises like the example BadCRC linked.
(Well, ignoring operator overloading anyway… you could conceivably import PHP's concept of variables and references into C++.)
A few months ago, I got the opportunity to do a little PHP programming. I was wrong. I didn't have a bad impression of PHP simply out of ignorance. In fact, my opinion of the language is now much lower than before. I honestly can't conceive how anyone could legitimately defend it for anything other than its ubiquity.
Perl does have a similar concept, but it calls it aliasing. Perl references behave more like what you'd expect coming from other languages.
I wish PHP had called its references something else. They're not really references at all.
Just because I think PHP references are weird doesn't mean I don't know how they work. It does make me wonder, though, why PHP references work the way they do. What's the benefit of doing it this way instead of the way just about every other language ever does it?
Most of "other languages" similar to PHP don't even have concept close to references, except for C++. Similarity with C++ may be briefly confusing, but it's not unheard of that in different languages concepts meant to achieve the same thing work differently. Expecting PHP would match C++ in every detail or would have to invent completely new terminology altogether makes little sense.
>> What's the benefit of doing it this way instead of the way just about every other language ever does it?
"Every other language ever" doesn't do it in any particular way. Perl has pointer-like references and they are nightmarish to work with. Languages like Ruby or Java pass everything by value or by object reference depending on how you look on it, since everything is an object. Some don't have the concept of references all. Low-level languages like C/C++ have pointers. In some languages variables are not mutable, so the whole question is moot. Saying that "every other language ever" does some specific thing in this regard and only PHP does it differently is meaningless - different languages do it completely different, and PHP has its own way. It doesn't match your favorite one - fine, everybody is entitled to have one's favorite ways, but that does not make PHP "weird" or wrong in any way, just as it doesn't make C, Java, Python or Perl wrong.
Then it got pointers, and references suddenly weren't needed that much anymore. But it kept the old semantics, and still needs references for passing arrays around, since arrays are not objects, and PHP assignment is still by copy.
(Yes, the "what if you add another statement line later?!" trope has been registered.)