Just a nitpick...
Once more, as it is typical on HN, web programming is confused with the entire universe of software development.
There are plenty of software realms where ASCII not only is enough, but it actually MUST be enough.
Just a nitpick...
Once more, as it is typical on HN, web programming is confused with the entire universe of software development.
There are plenty of software realms where ASCII not only is enough, but it actually MUST be enough.
"Web" programmers can care all they want about Unicode, but if the backend people didn't deal properly with text encoding, then something will break no matter what.
> There are plenty of software realms where ASCII not only is enough, but it actually MUST be enough.
Name one.
You are right. It's not a frontend/backend issue. It's a "for human" vs "not for human" issues. Personal names must be treated in an international-friendly manner.
>> There are plenty of software realms where ASCII not only is enough, but it actually MUST be enough. > > Name one
Joel himself described an example:
> It would be convenient if you could put the Content-Type of the HTML file right in the HTML file itself, using some kind of special tag. Of course this drove purists crazy… how can you read the HTML file until you know what encoding it’s in?! Luckily, almost every encoding in common use does the same thing with characters between 32 and 127, so you can always get this far on the HTML page without starting to use funny letters:
The content of a webpage is required to be expressed in every supported language, but the HTTP protocol must not. And it would make no sense at all to add internationalization to intra-machines protocol, where ASCII is enough and has been enough for decades.
And if someone complains that ASCII only supports English, well... suck it up! I'm Italian and work in French, still I hate when a colleague sneaks in a comment not in English. Professional software development happens in English.
I guess no URLs with funny characters then. "GET /profile/renée" => 500 error, woohoo.
> And if someone complains that ASCII only supports English, well... suck it up! I'm Italian and work in French, still I hate when a colleague sneaks in a comment not in English. Professional software development happens in English.
Get over yourself, a lot of professional development happens in languages other than English.
That's not really a slam dunk. Lots of sites don't let you have your name in the URL at all, and the average person's experience is that their name would be taken by someone else before they signed up.
In the metaphor, it's someone telling you your aim is off.
UTF8 encoded diacritics work just fine in C++.
That's all that's needed for a backend language.
The backend does not need to understand, or even acknowledge the existence, of grapheme clusters. Because the frontend is already having to understand all of this, it should be normalising any multi-codepoint ambiguous cluster anyway.
Not as measured by clusters, no.
> search for one string in a database of other strings?
Hence I said "normalisation". The frontend already has to do all the unicode twiddling, it may as well normalise the input too.
We absolutely do not "need" to know about Unicode, outside of interest about other realms.
Not being able to support non-latin scripts sounds more like a limitation than a feature to me, although of course in many contexts it’s not in any individual organizations power to overcome it.