Maybe it breaks if you embed unicode strings or something. What do other languages do?
Maybe it breaks if you embed unicode strings or something. What do other languages do?
Based on the behavior of this bug, it appears that the way PHP handles this case insensitivity is that it just lowercases all class and function names before resolving them. And this bug in particular shows up for Turkish because 'i' is not the lowercase equivalent of 'I'.
Pretty much all other modern languages are case sensitive, so I'd be surprised to find this issue elsewhere.
EDIT: “trivial to fix” as in, doesn’t cause regression, not necessarily that it’s a small change to the code base.
$classname = $row_I_got_from_mysql['classname'];
$object = new $classname;
I'm sure this can still be solved though. It's not trivial, but it's not "takes over 9 years to fix" complex either.http://blogs.msdn.com/b/michkap/archive/2004/12/02/273619.as...
The mapping between uppercase and lowercase can be completely arbitrary, and as long as it's used consistently you shouldn't get these kind of bugs.
i.e. this will print "bar":
<?php
echo foo();
function foo() { return "bar"; }http://en.wikipedia.org/wiki/Letter_case#Unicode_case_foldin...
It's not PHP's fault that accurately performing case transformations across locales is difficult; it's just actually very difficult. The solution isn't to "fix" the process of transforming letter case; the solution is to simply not transform the names of your identifiers. Unfortunately that is simple only in a very isolated setting; in the real world, doing such a thing is liable to break a lot of software.
This is a really good example of the problem at hand:
>The Greek letter Σ has two different lowercase forms: "ς" in word-final position and "σ" elsewhere.
The identifiers are lowercased multiple times; first at parse time, presumably using the locale of the OS, some setting in php.ini, or some fixed locale. (it doesn't, in practice, matter where this initial locale is set; it just matters that it's set at parse time.) It's then lowercased again at runtime; if the locale was changed at runtime, such that the casing rules in the two locales produce any differences, the identifier will not be found.
I'm not saying this to defend PHP; just to shed some light on the case-folding problem. Having case-insensitive identifiers is a design mistake.
Obviously, case-insensitive identifiers are a bad idea, but PHP is stuck with them at this point.
Of course it will never happen, but I really think this would be a great idea for the next major version. Backwards compatibility is important but not at all costs. A language needs to remain agile enough to allow for the recognition of (and ultimately the fixing of) mistakes.
Breaking changes in programming languages are not that uncommon, C# and Perl spring to mind from personal experience, but also to a lesser degree such things have happened with PHP itself (and it became better for it). In this case, it's actually a change that moves the runtime's behavior closer to what's expected. It's a change that improves internal consistency while also eliminating silly bugs like the one discussed here.
[1] actually, I don't care at all since I moved on to greener pastures.