This is similar to requiring a whitelist for mass assignment (rather than making a blacklist optional): it calls out in the code "These are our weak points! Check them carefully!"
This is similar to requiring a whitelist for mass assignment (rather than making a blacklist optional): it calls out in the code "These are our weak points! Check them carefully!"
When a particular function or class is found to be a problem (or when it is known to be a problem in advance), it is (re)named something like "foo_POTENTIAL_XSS_HOLE" or "FooMapNonStlCompliant" throughout the code base (using pfff or similar tool), and then the potentially long process of removing the problem can go forward with some confidence nobody will add to the problem while you're working on it.
It is, but I think it ignores the reality of API design. Who would ever start with dangerous_load? If you knew it was dangerous, you'd probably fix it. No, you start with load(), then find out it's dangerous afterwards. The fix will be a breaking change, so you make safe_load().
"You should replace load()!", I hear you cry. But that removes backwards compatibility. And people don't update legacy software. So then old software becomes vulnerable to who knows how many other security problems because of it's obsolescence.
When I wrote the Texcaller library, which simplifies compiling (La)TeX code, I disabled dangerous features such as "write18" from the very beginning.
I you don't take basic elements such as "secure/unsecure" variants into account, you're no really doing API design. You then just have a historically grown API, which is not necessarily bad in itself, but doesn't qualify to be called API design.
> But that removes backwards compatibility. And people don't update legacy software. So then old software becomes vulnerable to who knows how many other security problems because of it's obsolescence.
In that case, at least only old hard-to-update legacy software is affected, and not all the other (good, mostly well-written) software that is written today and in the future.
Of course, don't forget to increase the major version number, as for every API change that breaks backward compatibility.
True. I should have said the reality of the difficulty of API design.
In that case, at least only old hard-to-update legacy software is affected, and not all the other (good, mostly well-written) software that is written today and in the future.
I like the sentiment, but I can't agree with it. In an ideal world all critical systems would always be kept up the date, but that doesn't happen. Has Python 3 usage overtaken 2.x yet?
Haskellers. This is why you get lovely long and alarming names like `unsafePerformIO`
[1]: http://www.haskell.org/ghc/docs/7.4.1/html/users_guide/safe-...
https://github.com/dtao/safe_yaml
(nb: I've enable it on apps without issues)
For the others, well, we'll have to educate them, or push a safe yaml default quickly into Ruby [1].
This is no holy grail for sure, and there are plenty of other topics to be addressed (eg: secret tokens stored in SCM, pushed to third-party CI services and shared with freelancers and remote employees using non-encrypted disk, shared as well between production and non-regularly updated staging servers etc!)
So there's some piece of functionality, but security or some other obviously-important factor wasn't considered at all when it was initially "designed" and implemented. It's soon found to exhibit numerous problems, often including major security flaws.
Then a "safe" version of said functionality is offered. Yet for some reason (incompetence?) it has its own set of problems. And so the developers keep trying again and again, never seeming to get anywhere close to even a suitable solution.
An example is the mysql_escape_string() function, which was replaced by mysql_real_escape_string(), which has in turn been replaced by mysqli_escape_string(), mysqli::real_escape_string() and PDO::quote().
The end result is confusion, especially for novice users, or those coming back to PHP after some time. They don't know which of the several functions to use, or they're using older reference material that suggest the use of the faultiest of the functions.
It's better just to implement things properly the first time around, especially when security or data integrity, for example, are involved. Trying to hack on "safe" versions of functions, or even an entire "safe mode" is the wrong approach.
Changing how Rails or other libs process YAML is a behavior that can and should be performed transparently to users. And besides, most people processing user input aren't accepting YAML. It is extremely weird that Rails was accepting YAML in XML requests. It is slightly more understandable that Rails was using a YAML parser to parse JSON (although not a good idea as we've seen). But by and large, if you have an API that accepts user data, standard practice is to parse JSON, not YAML.
Your analogy is not correct.
An example of where people would be confused because YAML.load doesn't _instantiate arbitrary objects_ (which is what it really does, which results in ability to 'execute arbitrary code' as a poorly thought through side effect) -- is people using ActiveRecord::Base.serialize . Which would become broken if you were serializing any objects that weren't string, hash, integer, array.
While we've realized that allowing de-serialization of arbitrary objects ends up being incredibly likely to result in 'allowing execution of arbitrary code' -- referring to the problem simply as the latter confuses about the nature of the problem and the efficacy of various fixes.
Removing confusing stuffs in a language makes it better, but php language designers dont have a clue.
Safe_yaml is a first drop-in work-around that can be dumped into existing apps: a fix to quickly reduce the risk of exposure. If you have a proper set of tests in your app, it's fast to verify if something is broken here after starting using it.
But then, the underlying issue is being discussed actively [1], with talks about how to incorporate the safe default into coming versions of Ruby.
So I don't really see the parallel with what you describe...
So having "dangerous_load" is not going to help much with YAML: if it's exposed to untrusted input then grep for "YAML" not for "dangerous_load".