I couldn't debug the code because of my name
mikolaj-kaminski.com
mikolaj-kaminski.com
I also have non-latin characters in my name however I knew it was always an issue so I never used it in paths etc.
At some point, long time ago, I was tasked to do some maintance with Google Cloud service (can't remember the name of the service now) which was doable only through Python CLI utility and it failed with very similar Python error.
What I found out rather quickly is that utility took my name from Google+ profile, which did include those non-latin characters. No biggie - I thought and fired e-mail to support (yeah it was those times it was still that easy). Few hours passed and I received information that this won't be fixed anytime soon and the best course of action would be to change my name.
Of course, support person probably meant to remove the diacriticals from my Google+ profiles, but still it left unplesant aftertaste for years to come.
As someone who has been told this, for other reasons, I empathize. My reaction has always been - "Your system can't even handle names, you need to fix it".
Edit: I wish there was a library / service that helped you handle all sorts of edge cases in names, so that you don' t have to worry about it. Just use a user-id, and set / get a name from a lib / service that can actually handle it.
https://gdprhub.eu/index.php?title=Court_of_Appeal_of_Brusse...
Now that serious money is on the table, it might actually get solved and fixed once and for all.
We/us in tech had 30 years to fix our shit on our own. We didn't, that's the result.
These days everything should be stored as bare UTF-8 data (or utf8mb4 if you're MySQL) and presented without anything else. Don't parse it, don't slice-and-dice it, don't prepend or append titles or honorifics or suffixes, don't make assumptions about length or content beyond "must be > 0 as a whole" and DEFINITELY don't use it as an identifier. Treat it as a non-unique opaque token and you'll be fine greater than 99% of the time.
There are people with no last name. There are people with two or three or twelve middle names. There are people with a number for a last name. There are people with a symbol for their entire name.
Take what they give you and use it and be done with it.
Only if you are displaying them in a way that respects the computer's preferences (most websites and programs, especially American websites and programs, don't) and those preferences are set correctly. And certainly if you have text blocks that contain both Chinese and Japanese names you will always mangle at least one of them.
It looks like there's no general solution possible with Han unification. If you have any two of ZH and JA and KO and VI in a page, you will fail to display one of them correctly for certain characters unless (as in that wiki page) you add a LANG attribute for each element they are contained within.
Personally, I would use the browser's language or user locale to set the page language and give up. Then in Japan the local (Japanese) names would look fine, and same for China, Korea and Vietnam. Local consistency versus complicated perfection (tracking the input language as well as tokens and using them everywhere), and I could blame the browser for doing poorly at its impossible job.
One possible "perfect" fix would be to store the token and <span lang=...>$token</span> as well. The only place the non-wrapped version would be used is plaintext email or SMS, either of which are beyond lost causes for other reasons. Doing it with an embedded SPAN tag presents its own problems with sanitization, as well as guaranteeing it's always wrong if the input language was specified incorrectly when the token was populated, versus as above where it would be corrected to the local-optimal version if the user locale overrides it.
The ultimate source of this issue is that we are taking names and official IDs too seriously, but I doubt that problem will go away for "serious business". Funnily enough though, it already has for things like restaurant table reservations where all info provided is quite literally just a string for a human to do something with. No need to validate if the user's phone country code matches the country in which they are reserving a table...
Variation selectors are getting a good workout/testing technically in emoji at least (a lot of emoji are "just" "old" Unicode codepoints with a ZWJ and the variation selector known as the emoji variation selector to tell systems to always show it in "emoji styles"). I can't speak for how well it works in practice for CJK languages as I don't know them (more reason I appreciate emoji for letting me test compatibility with hard parts of UTF-8 in ways that I can read and most users want), but I do appreciate that there's at least the idea for/part of a fix in "recent" Unicode.
I'm also imagining it is not a fun thing to implement in practice, as Unicode at this point maintains a massive database just for it: https://www.unicode.org/ivd/
EVERY language should _try_ to handle Unicode such that if a data sequence were valid before it remains valid after. NONE should ever FORCE validation, since sometimes, like in the article's case, the correct answer is GIGO. Just pass it through and hope it continues to work. Sometimes the error is trying to enforce that validation.
For UNIX path names (and other OS data like environment variables), Python uses the "surrogateescape" error handling method, which does exactly what you ask. Any byte sequence can be converted to a string. If it decodes as valid UTF-8, it will do that. If it hits a byte that does not decode as valid UTF-8 (necessarily a byte >= 128), it will map it to code points U+DC80 through U+DCFF. These are in a reserved ranges of code points ("surrogates", which make it possible to represent code points > 0xFFFF in UTF-16), and they can't show up in actual Unicode text (i.e., there is no UTF-8 encoding of them, strictly speaking, and if you applied the UTF-8 encoding algorithm to a code point in the U+D800 to U+DFFF range, you would get bytes that aren't valid UTF-8).
On the way out, this is reversed. So you get the results you expect if your filenames are in UTF-8, but since UNIX has no requirement that filenames are indeed UTF-8 (the only constraint is they can't contain NUL or ASCII-forward-slash), the bytes are preserved in a funky-looking format in Python and you get the exact same output on the other end.
See https://www.python.org/dev/peps/pep-0383/ for more on what's going on. The tl;dr for users of Python is that if you want to interact with, say, subprocess output as mostly-normal strings (instead of bytes) but you want to be robust to non-UTF-8 bytes, you should do something like
subprocess.check_output(["some", "command"], errors="surrogateescape")
You don't need to do this for APIs that directly interact with pathnames, because they do it already. You just need to do it for things like subprocess output and file contents that Python doesn't know you want to handle in this way....
On Windows, however, path names must be valid Unicode and are stored in UTF-16. So the idea of a "ł" that doesn't decode properly shouldn't even happen! Mikołaj's home directory ought to be a very boring (and valid) 004d 0069 006b 006f 0142 0061 006a on disk.
Windows doesn't enforce that file paths are valid UTF-16 though (specifically, the surrogate code points are only supposed to show up in a certain way, but nothing enforces that and you can have random surrogates on disk), and hence Rust, which internally represents all strings in UTF-8, has a solution ("WTF-8") that's basically the inverse of surrogateescape - it uses extrapolated-UTF-8-encoding-of-surrogates to handle unpaired surrogates. http://simonsapin.github.io/wtf-8/ But it seems very odd to me that the directory C:\Users\Mikołaj would actually contain any of those, and if it doesn't, I would expect it to very easily turn into a Python Unicode string.
Maybe this is from a Python version before https://www.python.org/dev/peps/pep-0529/ , which is claimed to "fail to round-trip characters outside of the user's active code page"? Maybe this is from a Python version after that change and it's wrong?
Does it work if you set the environment variable PYTHONENCODING to cp1252?
(I suppose I should either contact the author, or try it myself...)
If you Google, "Mikołaj",
> Mikołaj is the Polish cognate of given name Nicholas
Then Google, "windows character encoding polish"
> Windows-1250 - Wikipedia
And 0xb3 is "ł" in that encoding.¹
> Does it work if you set the environment variable PYTHONENCODING (sic) to cp1252?
I don't know if setting PYTHONIOENCODING would work here; I don't think it should affect this. Really, fixing the YAML file is the fix. (And fixing the thing that generated it.)
¹it is queries like this that really make me love the search engines of today. This would have been hell in the days of Alta Vista.
And this is why you always validate your data when you slurp it in, or else you pass crap down several layers where it crashes or mostly works with the potential for security holes or catastrophic behavior, and a pain in the arse to track down since the actual bug is nowhere near where you are looking.
I'm not sure what sane behavior Python could have here besides errorring.
> EVERY language should _try_ to handle Unicode such that if a data sequence were valid before it remains valid after.
This sequence was never valid, and never will be.
> in the article's case, the correct answer is GIGO. Just pass it through and hope it continues to work.
Dear God, no; emit a diagnostic and abort. Countless decades of existing code have shown time and again that "plow forward with some hot garbage" is not a good idea. But that ignores that … that that isn't how any of this works; the YAML parse is going to want to emit strings, which the incoming data isn't.
Neither were four-byte UTF-8 characters at some point.
> and never will be.
We shall see.
$ ln -sT Mikołaj/ Miko�aj
That certainly isn't good, but it is> would have worked any better
My sibling has a name that has an accent, and just enters it with the plain letter most of the time. The name was once rare and "ethnic", but became popular a generation later so people know how to pronounce it regardless.
Our parents gave us two middle names, wanting to preserve our grandmothers' surnames, but also in the spirit of "Bobby Tables", having ambivalent feelings about the computerization of society tending towards inflexibility.
According to: https://en.wikipedia.org/wiki/Naming_customs_of_Hispanic_Ame...
...misunderstanding of naming customs in the US has actually led to significant consequences due to last names not matching on legal documents.
I remember reading a story about how there are people in China whose name incorporates a character that is obscure enough, the authorities are trying to eliminate it and get them to change their name. If I recall correctly, Chinese has a particular problem with characters that are part of names that have been around forever, but are no longer used for ordinary writing.
i don't think the authorities are actively trying to eliminate those characters but simply don't want to go through the effort to track down and have those characters added to the standard. the process takes years and in the mean time you have to live with the inconvenience. also most of the people faced with the problem would not even know how.
maybe has the benefit that he also can't receive speeding tickets?
> his approach was that he petitioned unicode to include that character.
Sounds like the right approach (although not the easiest). if unicode can include thousands of smileys, the least they can do is include actual characters used in people names.
Regarding Chinese names and uncommon characters, Japan has the same problem. It's specially problematic for place names, with some kanji used to write the name of a single place in the whole country! I used to write in a place with such an obscure kanji that I wouldn't type it on my Linux PC.
It's also difficult for people moving to the other country, with some characters existing in one country but not in the other. I have a Chinese coworker who needs to write his name in katakana because his characters don't exist in Japanese.
I could've been consistent in using two or three out of four, but when I was younger I was intimidated by forms that say you must enter your "full legal name", so I would, and they would mangle it unpredictably.
Checking account, credit card, drivers license, and property deed, each one different.
Well, in fact, my social security card and my birth certificate don't match, so I was doomed from the start.
It gives me some sympathy for places that try to regulate names to avoid parents doing something too goofy.
I sometimes wonder if there will come a day when all the databases will stop allowing discrepancies, and it won't matter to the powers that be, because it's such a tiny percentage of the population that becomes "unpersons".
That's usually easier than getting a company to fix their software.
The bug in question [0] was reported in 2001 and remains unsolved 20 years later.
[0] https://bugs.java.com/bugdatabase/view_bug.do?bug_id=4523159
https://shinesolutions.com/2018/01/08/falsehoods-programmers...
(but, in any case, this discussion here on HN https://news.ycombinator.com/item?id=18567548 provides some nuance)
https://www.kalzumeus.com/2010/06/17/falsehoods-programmers-...
Does it really surprise you that people are going to type just about anything into that box that it will allow?
https://www.cl.cam.ac.uk/~mgk25/ucs/examples/UTF-8-demo.txt
But this set of strings is specifically designed to cause edge-case errors.
Also don't forget Spolsky's seminal "The Absolute Minimum Every Software Developer Absolutely, Positively Must Know About Unicode and Character Sets (No Excuses!)".
https://www.joelonsoftware.com/2003/10/08/the-absolute-minim...
This entry, by the way, is a fantastic little easter egg in the list: https://github.com/minimaxir/big-list-of-naughty-strings/blo...
The symptom was that I could login if I used my password manager browser plugin, but not if I pasted it from my password manager.
Nitpick: if the number is higher than one, then it's at least two times too many.
The problem was that this field was used to enter a 10-digit code, and as it turns out, on default Windows10 system, the fonts are set up so that this field only fit 8 of them. Oops! :)
Around the time of AOL3 or early AOL4 someone found a user name exploit.
When making a new account, on the client side use winapi's EM_LIMITTEXT to bypass the max character limit on the input textbox. Enter one or two letters, a bunch of spaces, then some more letters.
The server side would truncate to the original length, leaving you with a one or two letter username, working around the 3+ requirement.
https://www.kalzumeus.com/2010/06/17/falsehoods-programmers-...
Certainly informative if you haven't seen it before.
My takeaway from it was that design your system to try to accommodate as much as possible, but it would basically be impossible to accommodate them all, so aim for your target audience.
There's no excuse for actively supported, paid products to have these problems today.
Not saying that this is OK, just explaining why using non-ascii characters, in this day and age, is still asking for trouble.
Windows 2000 is when the OS changed to UTF-16 by default. Before that Windows NT was UCS-2, IIRC only the DOS-based Windows versions were Windows-1252 internally, starting from Windows 1.0. So while ł wasn't supported in Windows 1, characters like ñ were. Windows has literally NEVER been an ASCII-based OS.
Basically - I agree: This shouldn't be a problem, and 7 months is a long time to wait for a basic fix. But there are a lot of footguns hanging around in windows code with respect to character encodings.
Just looking at the first result on google for "c++ get windows home directory" shows this: https://docs.microsoft.com/en-us/windows/win32/api/userenv/n...
Which takes a long pointer to tchar string (LPTSTR) - so this behavior is dependent on the unicode settings of the project at compile time, even today.
The documentation is simply wrong, GetUserProfileDirectoryA which you linked always takes a LPSTR (always "ANSI") while GetUserProfileDirectoryW always takes a LPWSTR (always WTF-16). This is reflected in the function prototype at the top. Only the define GetUserProfileDirectory switches between these two. The define is a compatibility hack and arguably was a mistake, but you can always the W-suffixed function no matter what the project settings are.
Paths are UTF-16 + unpaired surrogates, so a Windows path isn't legally representable in UTF-8.
Algol-68 supported localized sets of keywords; fortunately this language is gone.
You can #define non-ASCII stuff in modern C++. It's your best chance to "localize" a mainstream language.
Same would work for Clojure, but Lisp uses a lot of quirky abbreviations like `cdr` or `setq` that give awkward translations.
https://en.wikipedia.org/wiki/Non-English-based_programming_...
But it's a shame.
In Europe, we do have a lot of non-ascii characters everywhere. Ubuntu puts a "Vidéo" and a "Téléchargements" directory in my $HOME because I'm french. If I were to use my name as my username I would have even more troubles.
I'm careful with not using special chars in names for work, but it feels like I'm a girl trying to not dress sexy in the wrong part of town: necessary, but I shouldn't have to do this, and it's definitely the others to blame.
All in all, I thank the Gods of encoding for Python 3 unicode handling. Having a scripting language that does the right thing out of the box is wonderful on this side of the pond.
Want to offer Unicode validation? Sure having that as an OPTION is fine. Forcing it means I can't rely on that tool to handle real world data which happens to not be valid but is still a valid file-system address.
- if you want to treat paths like unicode strings, you can. Which is great for simple scripts where you don't want to deal with complexity. And 99% of the time, it's enough with modern OSes.
- if you want to threat path as bags of raw bytes, you can. Which is necessary to transparently copy and do not evaluate, as you said, for covering edge cases.
- if you need to actually deal with those as strings but don't want to loose data for edge cases, so a mix of the 2 above, you can use surrogate escapes
Python 2's unicode model was _closer_ to correct, the trivial coercion between byte[] and Unicode.
Conversion also shouldn't imply, force, or check Validation nor Normalization. Labeling a bytestream with an Encoding and validating / normalizing that encoding should be options. Operations on bytestreams with encoding related attributes should set them to either 'unknown' result or to a proper output type if they're aware the manipulations will still yield a valid encoding.
Normalization is more complex, since Unicode strings can be normalized in different ways, then combined, and still be a valid string but no longer uniformly normalized.
People have taken this to a ridiculous level. Giving practical advice about how to avoid problems is not blaming the victim. Telling my child to look both ways before crossing the street is not blaming him if he gets hit. It's not wanting to see him get hurt when the person actually at fault fucks up. Telling someone to avoid non-ASCII characters because a program can't handle it is not blaming them....
Your comment was unhelpful. Great! Not their fault! What now? It's also part of a larger trend that will lead to people being hurt.
My name appears differently in my passport, on plane tickets (not always using the same modification), and my green card. And for the latter two I left out the part of my name that can’t properly be represented at all in ASCII. And you are saying that somehow I am at fault?
Yes... The computer is at fault for not supporting proper names... That's literally what everyone has said. Nobody blamed the user....
> Your attitude reminds me of the people in the 60s who included punch cards in the utility bills, marking them “DO NOT FOLD, SPLINDLE, OR MUTILATE” (it became a meme).
The problem should be fixed, but it's not victim blaming to tell someone how to still submit their bill.
> And you are saying that somehow I am at fault?
No... I'm literally saying the opposite.... I think you need to reread your comment, then my comment.
Telling you to omit those characters so you can still travel internationally is not blaming you in any way shape or form... Your criticism of practical advice being victim blaming is harmful and unhelpful.
As much as I wish we lived in a better world where name characters were better handled, using anything outside of [a-ZA-Z]{12} as a username is a world of hurt. Some people just realize it later than others.
So yes, you shouldn't think of your handlename as your name, it's just another identifier, and choosing simple handle names is a life skill at this point.
"Asking for trouble" is a key part of testing. My suggestion would be for a QA person to have their username (and root folder of the testable project) to start with a space, and be followed by an accented letter, tab-symbol, apostrophe, an emoji, followed by an unicode RTL control character and some Arabic text.
[0] https://www.theregister.com/2021/09/29/weaponised_apple_airt...
I love to enter it and see what each vendor and website's backend does with it.
The Staples Canada website, for example, returns it as ' (HTML escaped) A couple times I've logged in, it seems to escape a new character. I'm currently up to &amp;#39;
The weirdest case I've had with that is the Six Flags mobile app. To add a season pass you need to provide your card number and last name. For the life of me I couldn't get it to validate, but I saw they showed the HTML escaped version in their e-mails to me. Turns out I had to type out "'" into their input box for my last name as that's apparently what they put in their database.
Every large company is just a conglomeration of smaller departments. Each department had individual contributors. Some individual contributor in that department wrote the code and if nobody else is their department caught it, nobody else at the large company would have caught it since they have their own work to consider and don't have time to look at other people's stuff.
NB: the method is still the same, it's a second (not accepted) answer here: https://superuser.com/questions/890812/how-to-rename-the-use... (about ProfileImagePath registry value).
Like the poor, it will be with us always.
When I think of 100 things I think of stuff like "some people spell their name in all lowercase and get really funny if you change it"
In this case, I'd guess CP-1250, since 0xb3, from the error, decodes to "ł", from the name, in that encoding. (But not in CP-1251, or '52.)
if you want to see how to arrive there: https://news.ycombinator.com/item?id=28939960
Yes, like IPv6.
If you never display the filename, the answer is to treat existing filenames as bags of bytes, but that breaks down as soon as you need to display them, or if you need to manipulate them by appending unicode to them, in which case you have to decide on an encoding.
Unicode encodings tend to mangle non-Unicode values because they're specified to replace whatever they can't understand with a particular Unicode character, usually represented as a diamond with an inverted ? inside of it.
There's some obscure solutions to this problem, like https://simonsapin.github.io/wtf-8/ (which includes discussion of the 16 bit encodings you need for Windows), but I haven't seen broad support for them. You need a deliberately "noncompliant" encoding/decoding system that doesn't replace unknown characters with replacement characters. Fortunately, compliant systems are becoming more and more popular and available. Unfortunately, that can make file name handling harder than when you had a non-Unicode-compliant handling system for your strings.
Other libraries handle this even worse than Rust. On Linux (filenames are bytes), Qt is unable to open files with invalid UTF-8 names, while GTK can open them (but shows an "invalid encoding" message instead of the original filename), which I think is a good-enough approach.
No you don't. On Windows you treat paths as a u16'\' an/or u16'/'-separated sequences of uint16_t. On Unix it's a '/'-separated sequence of bytes. If you want to display, you need to decode, but for display only - so errors should use replacement characters as a graceful failure. For appending you encode your string and then append the bytes. Never do you decode externally provided paths for the purpose of manipulation.
> There's some obscure solutions to this problem, like https://simonsapin.github.io/wtf-8/ (which includes discussion of the 16 bit encodings you need for Windows)
It's relatively new, but has wide enough adoption cosidering - e.g. it's what Rust uses for Windows paths. It's also straightforward - just encode the unmatched surrogate pairs as if they were the corresponding reserved unicode characters using the normal UTF-8 algorithm.
See for example: - https://cookieplmonster.github.io/2020/05/23/silentpatch-maf... - https://cookieplmonster.github.io/2021/02/27/silentpatch-yak...
If we are talking about ready game engines like Unity and Unreal... it is probably a naive assumption about input being 1 byte wide and things getting lost because of that in some gamedev-made script.
In the case of the Windows operating system, the worst fact is that every single part of it behaves differently. Some parts display the path with a wrong encoding, but handle it correctly. A third-party app can display it correctly, but fails while trying to access any file. From what I remember, even the built-in PATH variable editor/manager goes through some arcane steps to display the letters in a wrong way, but getting them to work sometimes.
I can only imagine how much more pain it is for someone using any of the less widely-used writing systems or those with more advanced features compared to ASCII (Hebrew’s RTL, Arabic scripts mid- and final forms, etcetera).
In English we simply shake the big bag of letters, pick a few at random and then throw them at the page until a few stick.
Nope. Neither can ź, ć, ś, ą or ę. You can, and people do write them as z, c, s, a and e when writing in a restriced character set, but that is not 'correct' and is not a bijection, ie. „półka” and „polka” mean two different things.
There's also the case of technically-same-sounding-especially-recently ż/rz and ó/u (whose replacement would let you get rid of two 'non standard' characters), but for historical reasons these are not interchangeable.
According to one of my employees (Polish) Ł sounds roughly like w as in win or water but not as in what. A quick read of this: https://en.wikipedia.org/wiki/%C5%81 doesn't help too much.
Does enforcing Ł instead of say w cause your written language to fail in some way? I don't want to cause offense, I want to understand the causes of difference.
If you wanna change that, you might as well change the entire writing system of the language, eg. to be more in line with some other, more common writing system (ie. other latin alphabets or the cyrillic alphabet which would probably make the most sense phonetically). But no-one's gonna go for that any time soon.
I think we have found the disconnect: you quite happily use a word like "wanna" which is nonsense in English. Its allowed because it is understandable. Wanna is "want to".
Ooh, "gonna": That'll be "going to".
What's gonna to you is l bar for me or vice versa or something 8)
What's the difference?
For more detailed explanation: https://en.wikipedia.org/wiki/Pronunciation_of_English_%E2%9...
I have a special character in my name, an apostrophe, and it causes trouble regularly online and with tooling. A number of years ago I decided just to never use it when it came to anything to do with technical work be it email, logins or usernames.
Unicode characters are a pain to deal with and I have suffered from it first hand trying to handle it. At the end of the day it is much easier just to not use the special characters and move on with your life rather then be battling the constant frustration.
I'm sure these tools have lots of issues opening and you would be surprised at the amount of time, effort and testing it would be required to provide fully Unicode support. Most people would see it as a very small positive and not worth the effort. I find it hard to disagree.
Unicode has been a thing since 1988. Names have included non a-z characters since forever.
Names have been spoken and hand-written since forever yet somehow computers aren't good at that so we all tolerate converting them to printed-looking text. Nobody cares, it doesn't matter.
String come up again and again as a hard issue to deal with especially once your start looking at Unicode. I think it would be very reasonable to assume only ASCII works and even then it doesn't always work!
Also, with Windows 10 users will often not even choose their username. It gets generated from their given name + surname (which is a whole different issue for people without one or t'other).
Both versions work most of the time these days, but I still run into trouble once in a while no matter which name I use.
Maybe ‘Siren’ is similar. It’s a pre-existing word that perhaps flags some sort of weird edge case.
I still find it funny that even in my home country you can't use a lot of local special characters in names. Also most airlines won't accept it so technically I'm not giving them my true name!
- websites telling me I have an invalid name
- post addressed to O'Rourke, O\\\Rourke, O&Rourke
- "my account" pages say "Welcome, Mr O\Rourke"
It's like not testing if your calculate application can handle negative numbers or decimals.
Similarly, what about tags (https://en.wikipedia.org/wiki/Tags_(Unicode_block) )? Do these require an U+E007F CANCEL TAG?
The 66 noncharacters certainly need consideration. http://www.unicode.org/faq/private_use.html says:
“Because of this complicated history and confusing changes of wording in the standard over the years regarding what are now known as noncharacters, there is still considerable disagreement about their use and whether they should be considered "illegal" or "invalid" in various contexts”
Edit: also, testing all code points likely is overkill and using code points in isolation likely isn’t enough. Most tests are better of with something like the big list of naughty strings (https://github.com/minimaxir/big-list-of-naughty-strings)
When written in the Latin alphabet, my surname is one letter.
I've had an amazing amount of problems with this not just due to technical limitations (like various forms marking the entry as invalid), but--much more aggravatingly--human limitations.
One particularly infuriating anecdote: at a past job many years ago, the email structure was lastname@company.com. I dutifully sent the IT person in charge of creating emails my desired email. The IT person wrote back an amazingly condescending email that as per the policy, emails had to be last names, not individual letters. I then had to go find a bunch of random websites which explained single-letter names and forwarded them to the IT person. They then obliged, but did not apologize for insulting me. That is not right that I had to put up with that.
Except single letter last names are less common than people not following policy and/or abbriviating the name. It could simply be an honest mistake and the email is just their standard response since they have other things to get to. Did you try simply pointing out that that the letter was in fact your last name instead of getting passive-agressive?
As for the benefits, which is completely off-topic, Windows Store is actually pretty awesome if you completely avoid search (and you need to do the Microsoft account thing for it AFAIK). Windows has needed a system to update 3rd-party software, to compete with Linux package managers, and the store is a really good effort (there are still annoying warts that Aur, Deb, RPM do not have). If you're willing to be a bit dumb, there is convenience.
If you pick up a halfway non-ancient framework in a somewhat common language with a somewhat non-terrible persistence like postgres, you just don't have problems. Just don't care, and it just works.
But it's super easy to derail that fragile correctness with something like MySQLs utf8-ish handling, or some OS's path handling, or 'efficiency', or a user or frontend dev submitting data in a wrong encoding. And then it gets mangled. And then the user is unhappy.
At that point, it becomes very hard to argue why one of the two things is wrong, and the other is not. While the user argues the other way around. Because both look correct, if you look from the right angle. And the only reason why I am right is because of some standard, while the customer is right because of money.
And yes, it is very 'surprising' why our software now functions correctly for russian or greek customers.
It would be bizarre if we were at the point where we had perfect translations for everything, but still struggled with character encodings specifically.
Why would you presume that when the problem seems to be that one tool uses the systems native 8-bit encoding while another tool expects UTF-8 - under sane systems these are the same.
The rationale is from an article someone linked here ("Falsehoods Programmer's Believe About Names"):
> Anything someone tells you is their name is—by definition—an appropriate identifier for them.
If you try to validate by checking for profanity, knowing full well that people can have names that contain profane substrings, I have a tongue-in-check message for you—you are a fucking asshole.
[1]: https://github.com/microsoft/WSL/issues/2577#issuecomment-90...
What is so curious there? Some names contain all non-latin characters, and some softwares don't work with non-ASCII symbols. I just cannot understand why is it interesting.
I got a name saying 'Hi John I just want to xyz'
I can skip this email as they used a fake name. Works better than other methods I have found.
Unfortunately, typical with Jetbrains.
They most certainly do not. E.g., a Turing machine assumes an alphabet Γ which is a set of some characters and is defined no further, as any exact definition is meaningless to the theory. (I.e., the algorithm is generic over any alphabet.) The alphabet need not even be text; e.g., for a Turing machine, the set of all octets suffices.
Even for something like Levenshtein distance, the only real requirement of the algorithm is that the abstract "characters" implement equality testing. For Unicode text, I'd start with graphemes, and then look for counter examples.
A variable width encoding can cause issues in principle, but useful algorithms already have to deal with strings that have variable-length physical represention anyway (eg "yes" vs "no"), so it tends not to be a problem in practice.
Changing all software to respect their perfectly valid name isn't something they can do.
They shouldn't need to change their name, but if they do, they can ignore all the broken software and go about their day.
This particular user is more capable than most, and found a workaround for this particular problem, which is good... But this is not likely to be the last of the problems.
I even had trouble booking flight tickets since their security system couldn't parse my name, and then had to go through some special security check due to it returning errors. After that, never again. Not sure how they managed to do it but they had some basic rules that they used to say "no real name can look like this, this is a fake person!" and just kicked it out.
Standard characters (ie english) are only used by a small subset (maybe 5-10%) of the global population.
'standard' by what measure? Ł is more standard than X or Q in the polish alphabet.
~ Sincerely, a person whose name contains „ń” and therefore had to deal with this bullshit his entire life.
Edit: You know how aircraft travel security always transforms your name into letters from the English alphabet to parse? Yeah, it transformed my name and then the resulting string looked so bad that the system rejected that. The original name doesn't look bad, but after transformations it did...
The first idea was to change the username to one that does not contain Polish characters. It turned out that Windows does not rename the user’s folder when changing the username. Manually renaming the folder was not an option. This way I could corrupt my profile in the system.
The end of the article is about how to change the directory where the temporary files go to one not under the user folder.
Polish may be close enough that an approximation is available in English, but there's an awful lot of languages that don't have a large overlap with English characters.
In the Asian case above, if someone with that name did try to "convert to English" they are ironically just as likely to end up with Akihito Abe as the ASCII, which will be just as broken!
Considering JetBrains seems unwilling to fix this bug, maybe the best solution of all is to switch to an IDE that works.
- IME On state. IME capture and interpret keypresses as engraved and generate corresponding Kana-Kanji texts.
- IME Off state. IME passes through keypresses as engraved on keytops.
- Direct Input state. IME becomes dormant.
In IME Off state, the keyboard behaves as a plain jp106(or ANSI if it is) keyboard, like I'm doing right now. The cases where you would use conversion with IME on for an English word is when you have reasons for the word to be in "full width"(usually for typesetting reasons).
I'm sure a computer savvy speaker of a fully-non-Latin language may still guess this is a good idea, but "computer savvy" doesn't cover everyone... and they shouldn't have to.
"Just use 7-bit-clean ASCII English" is not a solution to this problem.
It might be a "you're holding it wrong" situation, but everyone has already learned how to hold it "correctly " - it'll be a disruptive change to default a "natural" hold that you suggest.
The problem is the technology, not the user using it in a reasonable way. ł is older than computers and the only reason computers struggle with it is lack of foresight or choosing to make things harder for most of the world by some of the people involved early on.
BUT, there is an easy workaround to avoid all Unicode related bugs: don't use Unicode. If that's morally objectionable for you, then you can keep fighting this fight.
* yes, by and large. Many languages make do, but even the European languages that use the same script as English cannot be fully represented:
- Pretty much all mainland European languages use accents (simple example, in Spanish el and él are different words)
- French misses ç
- German/Swiss/Austrian misses ß
- Spanish misses ñ
- Dutch misses ij
I agree with you, and disagree strongly with dahfizz, who is essentially telling people their name and language are unacceptable.
It is not morally objectionable avoiding, it's just stupid.
Łódź
Fun fact: I was looking for an e-mail solution for a small company about a decade ago and found Zarafa. It seemed nice and I deployed it happily. Just to find out it only supports the Western European ISO codepage which was hardcoded. I hope they have switched to UTF-8 since then.
Nevertheless I find it absurd it's 2021 and they still have to. It's almost 30 years since the introduction of Unicode in Windows NT and NTFS, probably also close to that in Java. Pretty much every serious programming language or database supports Unicode by default today.
I believe it's a bug in some app in the toolchain as Windows file system API is perfectly capable of handling non-ASCII symbols. I always cared to avoid using non-ASCII symbols and spaces in my paths (incl. always installing almost everything to a custom directory outside "Program Files") but c'mon, how many decades do we need to develop handling these reliably?
I would also consider Windows' inability to optionally change the actual home directory name and distinguish between the user's full (display) name and their "username" (which are 2 distinct properties in Linux) "feature-bugs".
Would you also say Greek lambda has nothing to do with the English L? Or, slightly more relevant and complex example, actual Polish L with Slovak Ľ and Serbian Љ?
Conservative Ł would still be distinct.
The problem is some software just has problems with non-English alphabets because, roughly saying, all the software was meant to only process English text historically and much of it still has not been fixed. Users of non-latin-based alphabets have been accustomed to this and have no problem writing "Иван" as "Ivan" (despite it normally reads even more different in English, more correct phonetic transliteration would be "Eevan"). Heck they even spell "Семён" (~"Semyon", the Russian counterpart to English "Simon") as "Semen" X-). But the users of diacriticized Latin somehow get surprised with this.
If I could travel back in time to when ASCII was designed and give the engineers a hint I would ask them to add first-class diacritics to their design so anybody would be able to add the slash for Ł, the umlaut for Ü, Ö or whatever using an extra byte. Sadly, even today we mostly encode Ü as a letter absolutely distinct to U rather than a combination of the latter with the umlaut even though Unicode allows doing this the latter way AFAIK.
It was a long process of L-vocalization [0] that started around XVI century. The first segment of the population to be affected by it were peasants (it is also one of the reasons why the original sound quality of «ł» survived relatively long among Polish artists in early XX century as a sign of professionalism, similarly to English’s Mid-Atlantic accent [1]).
I suspect that L-vocalization’s proliferation was aided by multiple wars, partitions and occupations that followed, which caused many waves of both natural and forced internal migration and disappearance of most dialectal differences.
According to [2](PL) the pronunciation of «ł» as /w/ was codified as the standard around XIX/XX in both informal and formal settings. There are still remaining populations using the old sound quality, but they are mostly confined to the areas in proximity to other Slavic languages.
[0]: https://en.wikipedia.org/wiki/L-vocalization [1]: https://en.wikipedia.org/wiki/Mid-Atlantic_accent [2]: http://www.dialektologia.uw.edu.pl/index.php?l1=leksykon&lid...
No, "Яуап" is better, I think.
Take anything having to do with seamanship. There are many terms that date back to early modern English that simply don't make sense anymore yet are accepted and universal because the British Empire had a large and enduring influence on maritime matters and happened to be at the forefront of most modern developments until about 70 years ago.
In some cases this is actually built into laws and industry practice. Pilots speak English. That's the rules. Don't like it? Invent the time machine and beat Wilbur and Orville. For much the same reason, science speaks Latin.
This technical debt is difficult if not impossible to overcome, especially in regards to computers because we still haven't cracked general purpose AI. Software will only accommodate what it was written to accommodate.
Recognizing the problem and working to fix it is all well and good. But its wise to understand that this wont be solved any time soon so in the meantime it is pragmatic to operate in such a way to maximize compatibility.
After all, I still have to call it a Foc'sle even if I think that's dumb or isn't inclusive of my culture.
That reminds me that I was talking once to the guy who turn Mikołaj name into Mikowhy pseudonym
30 years later and i completely dropped all non-latin chars from my name in any and all forms. from airplane tickets to passport to you name it.
and you know what? no one cared about non-latin. not even the government. i loled when i actually realised.
i’ve encountered zero issues ever since.
and it’s been the same for lots of my friends. they just adopted some western name. case closed, no more issues.
it all depends on who much importance you attribute to your name. for me it’s always been a random variable. for others it’s a matter of pride. but to the “system” it will be a “random list of chars”, sometimes latin, other times utf.
Sometimes it is better to avoid being hit by the bus even if you are right.