I think. I haven't seriously used php since 2003 or so.
If you want to avoid that possibility, you can use a non-PHP format (eg: YAML) and parse it.
A solution would be to have a .ini-like (or some other simple-to-parse format) config file and PHP code to read its contents. PHP code could be leaked, but config file contents wouldn't.
If that i?php blunder happened to me, my users would see 10 lines of code. One include for the framework (which lives outside the web root), plus three calls to get the framework to handle the current request.
Those 10 lines would be totally unproblematic
Of course, you are completely right, it should still be inaccessible from the web.
How about that gotcha though? A single character is mistakenly changed from "<" to "i" and that exposes the source code to the browser. Think about that.
No, joking aside, the normal way to handle this is to have the perimeter server (nginx or varnish) catch any 5* responses and turn them into a user-friendly error-page. That way you never expose sensitive stack traces to your users.
So, this is standard stuff and easy to fix. However who of us hasn't screwed up on a similarly trivial issue before? I wouldn't judge them too hard on this one, happens to the best of us.
If you have a separate (ini-style) configuration file, every time you get a new request, the file will have to be read in from disk and parsed. On heavily-loaded web servers this can be a significant performance issue.
Configuration stored in a PHP file will be cached by your opcode cache and so doesn't incur any per-request parsing/reading overhead.
The problem here is not that Tumblr stored their configuration in PHP. The problem is their lack of testing their changes.
(As I mentioned elsewhere in the thread, this particular nasty issue can be solved by simply enforcing that every .php file begins in '<?'.)
it will just pull it from the cache, not needing to be parsed or read from disk.
Configuration stored in a PHP file is bad. Changing working code simply because you want to add a new slave is a horrible idea. It would be like having to recompile Apache from source every time you want to add a new vhost because all the entries are written in C.
Interpreted languages like PHP, Python, Perl, etc all make it easy to put config options directly into the language. When your app is small, it's fine. But as it gets larger, move it out.
But yes, that they didn't even catch this is in testing is the real problem.
I don't believe PHP's ini_read caches anything. Sure, your OS cache is going to have that file in, but PHP still has to do the fopen(), fread(), parse, fclose() dance on every page hit. This does not happen with an opcode-cached source file.
I have actually benchmarked this. Amortized across billions of page hits per month it produces a notable saving.
Kudos points out that you could use APC - this is absolutely true but IMO it basically equates to doing the same thing in a more complex way.
> Configuration stored in a PHP file is bad. Changing working code simply because you want to add a new slave is a horrible idea. It would be like having to recompile Apache from source every time you want to add a new vhost because all the entries are written in C.
This is a false comparison between interpreted and compiled languages. Not to mention that plenty of C programs use #defines for configuration. Varnish even translates its configuration language into C and dynamically loads it as a module, which provides significant flexibility.
In a dynamic language, and where you trust the person doing the configuration, there's no reason why your configuration shouldn't be in a source file. Both Django and Rails do it this way.
so check if the ini file key is in memcache, if so get it from there, if not, ini_read it and populate the cache. Not hard.
Not true. In Rails, sensitive information like database passwords or AWS keys are not stored in the source, you store them in configuration files, YAML files by default. It's also widely recommended that these not be checked into version control.
I wasn't referring to straight PHP. There are multiple ways to cache the config file's contents so you don't have to hit the disk to read.
> This is a false comparison between interpreted and compiled languages. Not to mention that plenty of C programs use #defines for configuration.
Your misunderstanding me here. I'm not saying #defines are bad. However, they are a different level of configuration from vhosts. You aren't #defining your virtual hosts. Same with php. You have the C code with specific options being set, and then php.ini for the user facing stuff.
> In a dynamic language, and where you trust the person doing the configuration, there's no reason why your configuration shouldn't be in a source file.
Except we saw at least one reason today.
Also, it's trendy to drop unseasoned developers onto it and expect them to build complex applications that are publicly accessible yet 100% foolproof. Go figure.
If you asked me, I couldn't even tell. I've encoundered some in the past, sure, but found (in docs) both rationale and the correct way to use stuff.
There was that reference-vs-value matter when passing objects around, in PHP v4.x, but that's fixed by v5.
But I am talking specifically about something that would just dump the code of the php file to the browser.
That would have easily caught this problem, no drama.
I would put my money on something like this being the cause.
Edit: I'm wrong. It does not catch this.
Root of the issue is that by default php outputs anything not between php tags to the browser.
Perhaps some sort of buffer control at the top most include can disable output to browser until a later time. That's pretty much how templating systems work except they don't disable output afaik thus leaving possibility of this open.
ob_start(function() {});Forcing every file to start with <?php is just PITA for developers working on templates.
Whaddya know, PHP is a templating language, after all.
Also, it takes us less effort to ship product that matches performance requirements with `raw' PHP than with another layer of abstraction.
Yep, our templates were .tpl files, which is probably a good convention to have even if you use raw PHP as your templating language.
But you get the drift. This is something you have to deal with when you use PHP.
The hitch was of course that the file in question had a "i?php" (someone was using vi and hit i too many times?) instead of "<?php" which lead the php interpreter to output that section of code as if it were HTML. That feature is something PHP could provide an option to disable and only allow explicit echo/print/printf which should be what templating engines use.
This explains why it was indexed by Google.