But I am talking specifically about something that would just dump the code of the php file to the browser.
That would have easily caught this problem, no drama.
I would put my money on something like this being the cause.
Edit: I'm wrong. It does not catch this.
Root of the issue is that by default php outputs anything not between php tags to the browser.
Perhaps some sort of buffer control at the top most include can disable output to browser until a later time. That's pretty much how templating systems work except they don't disable output afaik thus leaving possibility of this open.
ob_start(function() {});Forcing every file to start with <?php is just PITA for developers working on templates.
Whaddya know, PHP is a templating language, after all.
Also, it takes us less effort to ship product that matches performance requirements with `raw' PHP than with another layer of abstraction.
Yep, our templates were .tpl files, which is probably a good convention to have even if you use raw PHP as your templating language.
But you get the drift. This is something you have to deal with when you use PHP.
The hitch was of course that the file in question had a "i?php" (someone was using vi and hit i too many times?) instead of "<?php" which lead the php interpreter to output that section of code as if it were HTML. That feature is something PHP could provide an option to disable and only allow explicit echo/print/printf which should be what templating engines use.
This explains why it was indexed by Google.
I think. I haven't seriously used php since 2003 or so.
If you want to avoid that possibility, you can use a non-PHP format (eg: YAML) and parse it.
A solution would be to have a .ini-like (or some other simple-to-parse format) config file and PHP code to read its contents. PHP code could be leaked, but config file contents wouldn't.
If that i?php blunder happened to me, my users would see 10 lines of code. One include for the framework (which lives outside the web root), plus three calls to get the framework to handle the current request.
Those 10 lines would be totally unproblematic
Of course, you are completely right, it should still be inaccessible from the web.
How about that gotcha though? A single character is mistakenly changed from "<" to "i" and that exposes the source code to the browser. Think about that.
No, joking aside, the normal way to handle this is to have the perimeter server (nginx or varnish) catch any 5* responses and turn them into a user-friendly error-page. That way you never expose sensitive stack traces to your users.
So, this is standard stuff and easy to fix. However who of us hasn't screwed up on a similarly trivial issue before? I wouldn't judge them too hard on this one, happens to the best of us.
If you have a separate (ini-style) configuration file, every time you get a new request, the file will have to be read in from disk and parsed. On heavily-loaded web servers this can be a significant performance issue.
Configuration stored in a PHP file will be cached by your opcode cache and so doesn't incur any per-request parsing/reading overhead.
The problem here is not that Tumblr stored their configuration in PHP. The problem is their lack of testing their changes.
(As I mentioned elsewhere in the thread, this particular nasty issue can be solved by simply enforcing that every .php file begins in '<?'.)
it will just pull it from the cache, not needing to be parsed or read from disk.
Configuration stored in a PHP file is bad. Changing working code simply because you want to add a new slave is a horrible idea. It would be like having to recompile Apache from source every time you want to add a new vhost because all the entries are written in C.
Interpreted languages like PHP, Python, Perl, etc all make it easy to put config options directly into the language. When your app is small, it's fine. But as it gets larger, move it out.
But yes, that they didn't even catch this is in testing is the real problem.
I don't believe PHP's ini_read caches anything. Sure, your OS cache is going to have that file in, but PHP still has to do the fopen(), fread(), parse, fclose() dance on every page hit. This does not happen with an opcode-cached source file.
I have actually benchmarked this. Amortized across billions of page hits per month it produces a notable saving.
Kudos points out that you could use APC - this is absolutely true but IMO it basically equates to doing the same thing in a more complex way.
> Configuration stored in a PHP file is bad. Changing working code simply because you want to add a new slave is a horrible idea. It would be like having to recompile Apache from source every time you want to add a new vhost because all the entries are written in C.
This is a false comparison between interpreted and compiled languages. Not to mention that plenty of C programs use #defines for configuration. Varnish even translates its configuration language into C and dynamically loads it as a module, which provides significant flexibility.
In a dynamic language, and where you trust the person doing the configuration, there's no reason why your configuration shouldn't be in a source file. Both Django and Rails do it this way.
so check if the ini file key is in memcache, if so get it from there, if not, ini_read it and populate the cache. Not hard.
Not true. In Rails, sensitive information like database passwords or AWS keys are not stored in the source, you store them in configuration files, YAML files by default. It's also widely recommended that these not be checked into version control.
I wasn't referring to straight PHP. There are multiple ways to cache the config file's contents so you don't have to hit the disk to read.
> This is a false comparison between interpreted and compiled languages. Not to mention that plenty of C programs use #defines for configuration.
Your misunderstanding me here. I'm not saying #defines are bad. However, they are a different level of configuration from vhosts. You aren't #defining your virtual hosts. Same with php. You have the C code with specific options being set, and then php.ini for the user facing stuff.
> In a dynamic language, and where you trust the person doing the configuration, there's no reason why your configuration shouldn't be in a source file.
Except we saw at least one reason today.
Also, it's trendy to drop unseasoned developers onto it and expect them to build complex applications that are publicly accessible yet 100% foolproof. Go figure.
If you asked me, I couldn't even tell. I've encoundered some in the past, sure, but found (in docs) both rationale and the correct way to use stuff.
There was that reference-vs-value matter when passing objects around, in PHP v4.x, but that's fixed by v5.
; This directive controls whether or not and where PHP will output errors,
; notices and warnings too. Error output is very useful during development, but
; it could be very dangerous in production environments. Depending on the code
; which is triggering the error, sensitive information could potentially leak
; out of your application such as database usernames and passwords or worse.
; It's recommended that errors be logged on production servers rather than
; having the errors sent to STDOUT.
; Possible Values:
; Off = Do not display any errors
; stderr = Display errors to STDERR (affects only CGI/CLI binaries!)
; On or stdout = Display errors to STDOUT
; Default Value: On
; Development Value: On
; Production Value: Off
; http://php.net/display-errors
display_errors = OffThe point is that no server configuration can save you from an error like this.
Huh? Yes you can protect against errors like this; http://news.ycombinator.com/item?id=2343675
A deployment strategy that requires testing that pages show what you expect them to show would also likely catch it.
Except for the obvious – never edit files on the live server – other ways to protect against this would be to have multiple opening tags (first line of file would just be <?php ?>, then another opening tag on the second line), have your VCS check that certain files begin with <?php, store the config in a non-executable way (in a YAML file, or in the server environment), or using a combination of file() and eval() to always prepend the '<?php'.
And people should really install a PHP error handler first thing - before they load anything else - that delivers errors with a HTTP code 5xx they can catch in their caching layer.
With other web frameworks it is generally encouraged to put source code and static content in different directories. With the code completely outside the web root. This is much safer: never put code (or passwords) in your web templates.
It is possible to do this with PHP, but in practice almost no one does that. PHP, in my experience, has a lot of these insecure-by-default issues.
Perhaps a more secure idea would be to require an opening stanza for the reverse - for content that should simply be printed to standard out? i.e. <?out to make content that should be outputed, not <?php for content that should be interpreted.
The easiest solution would be to define a new file extension/MIME type which is "PHP-by-default".
Unless you include it from somewhere in the web root, but that's the other insecure-by-default behaviour I was hinting at. With a secure-by-default web framework, it's not possible to get the code to show at all because it's not intermingled with the content.
Every block of PHP code must begin with '<?php', regardless of where it's located, or whether it's included from another file.
I do agree with you that this is a silly behaviour. But it's nothing to do with the web root.
- PHP and generally all the frameworks based on PHP are strongly encouraging putting code outside of the web root. Just take a look at directory structures in Zend Framework (http://framework.zend.com/manual/en/learning.quickstart.crea...) or Symfony2 (http://symfony.com/doc/2.0/book/page_creation.html#the-direc...)
- ...and almost everyone that uses PHP professionaly does that.
Yeah, the fact that there is no way to separate code and content. In my past days as a PHP dev, I've done this and similar things many, many times (also putting a space at the beginning or end, and having random other stuff fail because something can't set a header anymore).
It has always appeared obvious to me that PHP should define a "pure code" file, one in which "<?php" would be illegal syntax, and which would never have its content written to the client. That's how nearly all PHP is written, anyway, and it would eliminate a good deal of stupid errors.
Try including the sensitive connection protocols from a non-www directory?