I think tumblr has a huge security hole
pastebin.com
pastebin.com
But the whole self-righteous and incredibly passive-aggressive "man, i think these guys have a huge problem..." followed by the "man, these guys need to shape up" and "n00b mistake!" are unproductive to the extreme.
I mean, find a bug, report it, move on. Maybe i've got some unrealistic notion of karma or general human benevolence or something, but it seems hard to believe that this is such a difficult path to take especially when nearly everyone commenting has to deal with bs of this sort day-in/out.
If someones pants are sagged around their knees, I expect them to have noticed this themselves, and by walking around in public they've accepted the possibility of ridicule.
I consider having a non-beta site to be 'walking around in public', and revealing code and internal data (however small) when there is a malformed request to be 'walking around in public with your pants sagged at your knees'.
-----
That said, rappers don't have to buy belts and maybe the culture of the 'new economy' is such that some people do not have to value security.
Tumblr isn't guilty of criminal negligence, but they are guilty of a very serious failing of basic security precautions. Luckily there are other layers of security at play preventing this from being a catastrophic disaster for tumblr. However, if a group of thieves break into your bank and drill into your vault you do not go home and rest easy because they only managed to drill through two feet of your vault's hardened steel and there was an entire 3 or 4 inches more. Less so if you'd done something dumb like leave the keys to the vault in a coffeeshop.
This isn't to say security isn't important, but they're rushing to make Tumblr as fun to use as possible so they can survive.
Yes, they received money. They also have monstrous growth. Now they can afford to expand the engineering processes beyond, "get it working" to "make it work really well and securely".
Good things are still to come from Tumblr so let's go easy on them when they use duct tape instead an arc welder.
If they failed to take security into account in the early stages, never mind implement it at the beginning of development, then odds are they won't be implementing it effectively any time soon, especially with the rate at which they'll be expected to keep growing and adding functionality.
This kind of issue that they're showing now could (and probably should) have been detected and handled early on, even with a simple third-party code review.
And the fact that they are as big as they are, and growing as quickly as they are, means that they should have an increased sense of responsibility when it comes to security and protecting their users.
Good science would also suggest Tumblr should get some experts to help them discover anything else that might be lingering, which they're planning to do. Much like a peer review process.
Your attitude is important for those in the security industry as it pushes things forward, but remember that not everyone has the time to spend on it that you might. It can either be an asset or the bane of your existence. As an asset, you get paid for the things you understand because others don't. As the bane of your existence, you fight society for not knowing what you know.
Tumblr is hiring. Maybe you should apply and help them fix it?
I was, more than anything, responding to the parent's post regarding "they can just add security later" idea.
It's true that I tend to work on projects where security is a huge deal (online banking, government, global video game services including in-game payments, etc). As the architect of these systems, a key part of the design is security, and while other projects don't have to be quite as diligent, that doesn't mean they should just ignore it altogether.
I'd also like to hope that my attitude is not just for those of us in the security industry, but for everyone making web-based applications.
Personally, I think any online service does their current or potential clients a disservice if they don't take security into account early on.
As soon as you take money from someone, I consider that to be a responsibility that has been accepted to not only provide the functionality you offer, but to do it in an appropriately secure manner.
It's the classic techie vs. sales guy argument; we don't want it perfect, we want it on Wednesday.
The problem is that if even simple and effective security is overlooked or not dealt with early on, you'll almost always be forced to accept a compromise rather than take the required time to implement it properly.
As to the job, I'm already quite busy, thanks. Between implementing Oracle clusters and my new startup that's just closing on our financing, I've got my plate full.
Besides, I hate PHP. ;)
It's easy to get going quickly but becomes an issue fast. Smart people tend to move away from PHP, but startups find this difficult.
Hiring is then an issue and the bad practices persist.
You don't need someone full-time to help develop best practices, help design or architect code/systems, or to do code reviews.
Any startup that gets funding should, in my opinion, get a short-term consultant to come in, take a look around and offer suggestions and advice. Even if it's just for a day. These are resources that have been there, done that, and wouldn't be interested in a full-time gig with the company to begin with. Even if the company could afford them.
Hopefully this isn't too far off topic, but it's something that I see missing from a lot of clients that I get called into. (Not saying these guys haven't done this, either.)
I think there's a lot of value for a startup to validate their work with outside help, especially if they're relatively new to the game. Even if it's just pointing them to some articles or reading for them to follow up on, or just mentioning ideas of things they should look into; it can prove to be huge.
For instance, I'm currently mentoring a few developers/teams on a part-time, couple hours a week basis. Some of it is just being available on MSN to answer a quick question every now and then, other times it's doing a couple code reviews. Other times it's grabbing lunch/beer with them to discuss the concepts of things like how to implement continuous integrated testing, or managing other processes, or discussing new tech that they've heard of. For the most part, they're not paid gigs either. I enjoy helping people do stuff well, and the little bit of good faith help usually leads to some good future work.
Sometimes all it takes is asking the right question to get them to think about things in a different manner.
The last gig I got was for a major video game company. The job involved a .NET stack, which I knew nothing about and had never worked with, not even a little bit. I got the job despite that fact because in the interview I asked them general (admiteddly leading) questions, such as "how are you handling _____, or what is your plan for _____". Even though all of my experience had been in non-MS tech stacks, the technical director that was interviewing me was furiously taking notes. It was clear that they hadn't planned for some of the things I was mentioning, never mind thought of them. But once they heard the question, it was painfully obvious they should have. As a result, a 3-day consulting engagement turned into full-time for two years.
In my case right now, I've sought a mentor in some new tech I'm working with in my current startup. It's allowed me to hit the ground running, and take advantage of their experience. A half day one-on-one made all the difference for me, and will drastically improve the quality of what I'm doing.
Yet I have worked in a lot of different environments with PHP over the time and this never happened to me (but I was close to). It's a big, big mistake, not just a tiny error.
People, I know it's en vogue to bash PHP (just wait, in a few years it'll be Ruby and Python - remember, PHP was once hyped, too, and now it's going in the other direction) - but if you criticize PHP, could you at least try to sound like you've actually developed in PHP for more than a week?
Because most of the negative comments here about PHP have absolutely nothing to do with PHP as such - the Tumblr error in question has to do with incompetent programmers. If you read The Daily WTF you'll know that incompetent programmers can screw up no matter what language they're using.
This isn't a problem with PHP. It's a problem with ANY app storing configuration options inside the actual language. It would be like storing vhost information in C and forcing you to recompile all of apache to get a new site up.
Yes, config options are easy to add to interpreted languages. I see it happen all the time, and not just in PHP. Perl, Python, Ruby.
The problem is that to change the config options, you are essentially changing what amounts to the application, and if you break your config, you break your application.
The standard way of storing config options in PHP is with .ini files. At least, in my circle. Sure for quick scripts, we throw it in the actual code, but whenever we have to manage more than a few config options, it get's moved out.
I realize I'm expanding on what you said a bit. =) I just want to make it clear that this isn't a PHP problem. It's a basic programming problem.
Many other frameworks use the source code language for configuration without risking this particular error, because they keep templates separate. (In Django, for example, templates use a different language and are read out of template-specific directories.)
It's not even a problem unique to configuration files. A disclosure of source code would have occurred by replacing the first character of a PHP source file. The fact a configuration file was involved makes the effect both more severe and less likely to be caught ahead of time. More severe, because it discloses passwords, rather than just proprietary code. Less likely to be caught, because standard testing methods tend to rely on a staging or test config instead of the production one.
There are other posts disagreeing over various single villains, but a wider view is that Tumblr was bitten by a combination of a single-character typo, an otherwise useful PHP language design choice, the common and otherwise harmless practice of configuring via a programming language, and the blind spot typical test suites have for production-specific issues.
Judging by the mistake, I wouldn't be surprised if they edited the file live on the server.
For what may be a thoroughly tested and properly deployed release, what happens when a sysadmin needs to update a password for a database? In many cases, without custom tools or a formal process, they'll just pop the config file open with vi and rsync the change out to their web cluster. I'd bet this is what happened in Tumblr's case. A sysadmin did it. Probably to avoid doing a full production redeployment for a simple config update. They will put resources into formalizing this now that they've been bitten.
> For what may be a thoroughly tested and properly deployed release, what happens when a sysadmin needs to update a password for a database?
They coordinate the efforts with someone on the development team to deploy this. A sysadmin touching source code is as bad as a developer making changes to the networking side, especially if neither are talking back and forth.
This is how we work. Any changes made by networking are first vetted on by me, for example, for the systems I'm responsible for. I work with them to ensure that deployment is done at the proper time, and we handle any possible problems on our end. The networking team doesn't touch anything we work on, and vice-versa. Communication becomes key.
Most companies pin down their processes in response to issues that pop up.
I haven't read the update, so I really don't know.
I'm not suggesting they are morons. Rather, they made a series of mistakes. Every programmer makes mistakes. The key is to have a strategy to catch mistakes. Something as simple as a staging server where something like this could have been caught, and having a deployment strategy where you cannot get past staging without going through deployment is a good idea.
Just in case.
edit: oh man. Check out what Google has on this[1]
edit: in readable form on github gist[2]
[1] - http://www.google.com/webhp?hl=en#sclient=psy&hl=en&...
http://www.google.com/search?q=site%3Atumblr.com+m3MpH1C0Koh...
If you have a separate (ini-style) configuration file, every time you get a new request, the file will have to be read in from disk and parsed. On heavily-loaded web servers this can be a significant performance issue.
Configuration stored in a PHP file will be cached by your opcode cache and so doesn't incur any per-request parsing/reading overhead.
The problem here is not that Tumblr stored their configuration in PHP. The problem is their lack of testing their changes.
(As I mentioned elsewhere in the thread, this particular nasty issue can be solved by simply enforcing that every .php file begins in '<?'.)
it will just pull it from the cache, not needing to be parsed or read from disk.
Configuration stored in a PHP file is bad. Changing working code simply because you want to add a new slave is a horrible idea. It would be like having to recompile Apache from source every time you want to add a new vhost because all the entries are written in C.
Interpreted languages like PHP, Python, Perl, etc all make it easy to put config options directly into the language. When your app is small, it's fine. But as it gets larger, move it out.
But yes, that they didn't even catch this is in testing is the real problem.
I don't believe PHP's ini_read caches anything. Sure, your OS cache is going to have that file in, but PHP still has to do the fopen(), fread(), parse, fclose() dance on every page hit. This does not happen with an opcode-cached source file.
I have actually benchmarked this. Amortized across billions of page hits per month it produces a notable saving.
Kudos points out that you could use APC - this is absolutely true but IMO it basically equates to doing the same thing in a more complex way.
> Configuration stored in a PHP file is bad. Changing working code simply because you want to add a new slave is a horrible idea. It would be like having to recompile Apache from source every time you want to add a new vhost because all the entries are written in C.
This is a false comparison between interpreted and compiled languages. Not to mention that plenty of C programs use #defines for configuration. Varnish even translates its configuration language into C and dynamically loads it as a module, which provides significant flexibility.
In a dynamic language, and where you trust the person doing the configuration, there's no reason why your configuration shouldn't be in a source file. Both Django and Rails do it this way.
Not true. In Rails, sensitive information like database passwords or AWS keys are not stored in the source, you store them in configuration files, YAML files by default. It's also widely recommended that these not be checked into version control.
I wasn't referring to straight PHP. There are multiple ways to cache the config file's contents so you don't have to hit the disk to read.
> This is a false comparison between interpreted and compiled languages. Not to mention that plenty of C programs use #defines for configuration.
Your misunderstanding me here. I'm not saying #defines are bad. However, they are a different level of configuration from vhosts. You aren't #defining your virtual hosts. Same with php. You have the C code with specific options being set, and then php.ini for the user facing stuff.
> In a dynamic language, and where you trust the person doing the configuration, there's no reason why your configuration shouldn't be in a source file.
Except we saw at least one reason today.
so check if the ini file key is in memcache, if so get it from there, if not, ini_read it and populate the cache. Not hard.
How about that gotcha though? A single character is mistakenly changed from "<" to "i" and that exposes the source code to the browser. Think about that.
No, joking aside, the normal way to handle this is to have the perimeter server (nginx or varnish) catch any 5* responses and turn them into a user-friendly error-page. That way you never expose sensitive stack traces to your users.
So, this is standard stuff and easy to fix. However who of us hasn't screwed up on a similarly trivial issue before? I wouldn't judge them too hard on this one, happens to the best of us.
A solution would be to have a .ini-like (or some other simple-to-parse format) config file and PHP code to read its contents. PHP code could be leaked, but config file contents wouldn't.
Deleted comment
Edit: I've tested this:
$ php -v
PHP 5.2.6-3ubuntu4.6 with Suhosin-Patch 0.9.6.2 (cli) (built: Sep 16 2010 19:51:25)
Copyright (c) 1997-2008 The PHP Group
Zend Engine v2.2.0, Copyright (c) 1998-2008 Zend Technologies
$ cat test.php
<?php
require "/tmp/test2.php";
?>
$ cat /tmp/test2.php
i?php
define("TEST", "test");
?>
$ GET http://localhost/test.php
i?php
define("TEST", "test");
?>I think. I haven't seriously used php since 2003 or so.
If you want to avoid that possibility, you can use a non-PHP format (eg: YAML) and parse it.
If that i?php blunder happened to me, my users would see 10 lines of code. One include for the framework (which lives outside the web root), plus three calls to get the framework to handle the current request.
Those 10 lines would be totally unproblematic
Of course, you are completely right, it should still be inaccessible from the web.
Also, it's trendy to drop unseasoned developers onto it and expect them to build complex applications that are publicly accessible yet 100% foolproof. Go figure.
If you asked me, I couldn't even tell. I've encoundered some in the past, sure, but found (in docs) both rationale and the correct way to use stuff.
There was that reference-vs-value matter when passing objects around, in PHP v4.x, but that's fixed by v5.
But I am talking specifically about something that would just dump the code of the php file to the browser.
The hitch was of course that the file in question had a "i?php" (someone was using vi and hit i too many times?) instead of "<?php" which lead the php interpreter to output that section of code as if it were HTML. That feature is something PHP could provide an option to disable and only allow explicit echo/print/printf which should be what templating engines use.
This explains why it was indexed by Google.
That would have easily caught this problem, no drama.
I would put my money on something like this being the cause.
Edit: I'm wrong. It does not catch this.
Root of the issue is that by default php outputs anything not between php tags to the browser.
Perhaps some sort of buffer control at the top most include can disable output to browser until a later time. That's pretty much how templating systems work except they don't disable output afaik thus leaving possibility of this open.
ob_start(function() {});Forcing every file to start with <?php is just PITA for developers working on templates.
Yep, our templates were .tpl files, which is probably a good convention to have even if you use raw PHP as your templating language.
But you get the drift. This is something you have to deal with when you use PHP.
Whaddya know, PHP is a templating language, after all.
Also, it takes us less effort to ship product that matches performance requirements with `raw' PHP than with another layer of abstraction.
; This directive controls whether or not and where PHP will output errors,
; notices and warnings too. Error output is very useful during development, but
; it could be very dangerous in production environments. Depending on the code
; which is triggering the error, sensitive information could potentially leak
; out of your application such as database usernames and passwords or worse.
; It's recommended that errors be logged on production servers rather than
; having the errors sent to STDOUT.
; Possible Values:
; Off = Do not display any errors
; stderr = Display errors to STDERR (affects only CGI/CLI binaries!)
; On or stdout = Display errors to STDOUT
; Default Value: On
; Development Value: On
; Production Value: Off
; http://php.net/display-errors
display_errors = OffThe point is that no server configuration can save you from an error like this.
Huh? Yes you can protect against errors like this; http://news.ycombinator.com/item?id=2343675
A deployment strategy that requires testing that pages show what you expect them to show would also likely catch it.
Except for the obvious – never edit files on the live server – other ways to protect against this would be to have multiple opening tags (first line of file would just be <?php ?>, then another opening tag on the second line), have your VCS check that certain files begin with <?php, store the config in a non-executable way (in a YAML file, or in the server environment), or using a combination of file() and eval() to always prepend the '<?php'.
And people should really install a PHP error handler first thing - before they load anything else - that delivers errors with a HTTP code 5xx they can catch in their caching layer.
With other web frameworks it is generally encouraged to put source code and static content in different directories. With the code completely outside the web root. This is much safer: never put code (or passwords) in your web templates.
It is possible to do this with PHP, but in practice almost no one does that. PHP, in my experience, has a lot of these insecure-by-default issues.
- PHP and generally all the frameworks based on PHP are strongly encouraging putting code outside of the web root. Just take a look at directory structures in Zend Framework (http://framework.zend.com/manual/en/learning.quickstart.crea...) or Symfony2 (http://symfony.com/doc/2.0/book/page_creation.html#the-direc...)
- ...and almost everyone that uses PHP professionaly does that.
Perhaps a more secure idea would be to require an opening stanza for the reverse - for content that should simply be printed to standard out? i.e. <?out to make content that should be outputed, not <?php for content that should be interpreted.
The easiest solution would be to define a new file extension/MIME type which is "PHP-by-default".
Unless you include it from somewhere in the web root, but that's the other insecure-by-default behaviour I was hinting at. With a secure-by-default web framework, it's not possible to get the code to show at all because it's not intermingled with the content.
Every block of PHP code must begin with '<?php', regardless of where it's located, or whether it's included from another file.
I do agree with you that this is a silly behaviour. But it's nothing to do with the web root.
Yeah, the fact that there is no way to separate code and content. In my past days as a PHP dev, I've done this and similar things many, many times (also putting a space at the beginning or end, and having random other stuff fail because something can't set a header anymore).
It has always appeared obvious to me that PHP should define a "pure code" file, one in which "<?php" would be illegal syntax, and which would never have its content written to the client. That's how nearly all PHP is written, anyway, and it would eliminate a good deal of stupid errors.
Try including the sensitive connection protocols from a non-www directory?
I hope they've changed them.
right?
I hate to see someone else work in the clear like this. It’s like popping a zit before your first date. It’s painful and will show up for day/s afterwards. Now I know what will be today’s headline I can bypass techmeme.
Yes its a big boo boo. It’s a massive security risk and to some it may feel like the end of the world but by then it will be tomorrow. Passwords will be reset, keys will be replaced and the valley will be talking about something else. Hopefully it won’t be someone else’s mistake.
P.S Don’t forget to test your code before deploying – now you know why.
It does hammer home the point of staging before deploying. Also the point of making sure you vary your passwords between sites.
Although, on the plus side, having a site that mashes up tumbler as a content provider certainly has given us plenty of opportunity to fine tune and improve our caching strategy.
The dashboard controller alone has approximately 120 actions.
PHP doesn't do well on either count.
Writing secure PHP code isn't impossible, but it's tedious even for seasoned developers.
Their site has a front-end controller (www/dispatch.php) to which all requests get routed (by a mod_rewrite rule, perhaps), rather than having .php files in the web root for each page. This file sets up the environment -- among other things, it registers a custom error handler and includes a config file (config/config.php) that defines a bunch of constants. Then it dispatches the request to the appropriate controller, based on the URL.
Someone (probably a sysadmin) edited the config file and accidentally changed the opening tag (<?php). This caused PHP to output its contents, rather than parsing and executing it. Since there was no output buffer active, those contents were sent directly to the user, which caused HTTP response headers also to be sent automatically. That's the big first line you see in the output.
Since no error had actually been triggered by this point, execution continued. It tried to set an HTTP header ("P3P: CP="P3P_CP"", whatever that means). However since HTTP headers had already been sent, this did trigger an error, which was passed to the custom error handler, which sent some debug output (the rest of the output you see) and stopped execution.
Granted. This should move outside of the web root, but just looking at this one file that just sets up configuration values provides no ground to say anything about general code quality
No it doesn't. Having to modify live code just to change configuration options is bad. If I want to add a new slave, or change some configuration option, it should have no affect on running code. I shouldn't be changing any software logic, and that's exactly what editing this file does.
Their are standard ways to avoid this. But no, if changing your config file means changing what essentially amounts to changing the executable of your code, you're asking for trouble.
Just like you'd want to have a well described (preferably automated) way to restore from backups, you should also have one for resetting all passwords. Such a process is also useful for protecting against disgruntled employees.
Does anyone know how long was it actually in this state? (There's a heck of a lot of entries in the Google search quoted in another comment, but then how often does Google index tumblr?)
Did no-one at least press F5 or CMD-R after making the edit, let alone run tests? Quality control is the real issue. I can easily imagine myself making this mistake, typo's are the source of the majority of my bugs, but I find it hard to imagine taking more than 10 seconds to notice it.
I doubt anyone using PHP at Tumblr's level is mixing raw HTML into their PHP files and thus has no need for the php tags.
include '../app/front_controller.php';This is what happened with Tumblr.
I just tested a few variations on one of my old projects and the reason why I didn't see it is because I had ob_start() and was killing the buffer in my templating class (which is a good idea)
not to mention that the db passwords were in an external yaml file anyway (most of the frameworks do this) which would get loaded and cached
we should start a support group
if(getenv('tumblr_env') == 'production') {
//See: http://php.net/manual/en/function.register-shutdown-function.php
register_shutdown_function('five_hundred_err_func')
}We’re triple checking everything and bringing in outside auditors to confirm, but we have no reason to believe that anything was compromised. We’re certain that none of your personal information (passwords, etc.) was exposed, and your blog is backed up and safe as always. This was an embarrassing error, but something we were prepared for.
The fact that this occurred at all is still unacceptable, and we’ll be seriously evaluating and adjusting our processes to ensure an error like this can never happen again.
It's a mess! Very long, not dry, no wonder the dev fat fingered it...
Is the rest of the framework like this?
No headless output testing? Manual output checking (at the least)?
For the record, as much as I hate to admit it now I was a big long time PHP developer. My code is used round the web today...
Here is how Flask can do configuration (a Python micro framework):
<code>
class Config(object):
DEBUG = False
TESTING = False
DATABASE_URI = 'sqlite://:memory:'
class ProductionConfig(Config): DATABASE_URI = 'mysql://user@localhost/foo'
class DevelopmentConfig(Config): DEBUG = True
class TestinConfig(Config): TESTING = True
</code>Still, I think config files are often best when they're verbose and as straightforward as possible. There are fewer surprises this way. If I'm working on a project by myself or with a very small team of people I trust, I'd go the inheritance route. If the team is larger, I'd go the verbose route.
Edit: it's whoever dumped this as opposed to an attacker
I don't think tumblr have acted on this yet. The other exposed pages have S3 API secret keys, facebook api secret keys, the username and passwords for vimeo, clickatell (whatever that is), twitter oauth secret key, etc. they are going to have to revoke and re-setup each of these.