EBCDIC is incompatible with GDPR
shkspr.mobi
shkspr.mobi
This led to problems when they instituted their trusted ID compliance. When renewing the license we were required to provide some combination of documentation to corroborate our identity, and obviously that documentation needs to match the name shown on the driver's license - and of course mine did not.
There was one way out for Christophers like myself. A birth certificate was considered the ultimate truth, so as long as I had a notarized (with the raised seal) birth certificate to prove my identity, they would allow me to renew my license.
The State of New Jersey is very awful at IT. My wife, who works in healthcare finance, told me about problems she was having with the State because - get this - their field for what amounts to "Medicaid ID#" was too narrow, so they had to recycle ID#s for new recipients! And to make that worse, they discarded old backup data so when checking the data for a patient several years ago, it's only possible to find that of the latest owner of ID# 12345.
Officially I am not who I am.
Interestingly, it looks like we're actually going backwards on this: accented characters have actually been dropped from the main passport page in recent years, to be replaced with ICAO transliterations. Which is shameful, to be honest, since it implies passports are now incomplete as a form of ID (unless the real name is recorded somewhere else). Airline lobbying clearly won the day, years ago. This seems to be a UK-only thing at this point.
[1] https://eur-lex.europa.eu/LexUriServ/LexUriServ.do?uri=CELEX...
In any case, I think it's fair to say that it's a plurality, if not a majority, and that the letters A-Z are the most natural "core" set of glyphs from which the other (upper case) Roman-derived letters are built.
[0] https://www.worldstandards.eu/other/alphabets/
[1] https://www.britannica.com/list/the-worlds-5-most-commonly-u...
Here's the (a?) specification for the machine-readable part of the passport, with transliteration and so on:
https://www.icao.int/publications/Documents/9303_p3_cons_en....
(toyg, did you originally include this link in your comment, then edit it out?)
> passports are now incomplete as a form of ID
Passports are, and always have been, tools for travelling internationally. Depending on a passport for general identity is arguably as much of a mistake as using social security numbers for identification.
In the U.S., every state-issued ID card or drivers license requires, among other things, a document proving identity; a current, valid U.S. passport is considered to meet that requirement. Here, a passport is a federally-issued identity document.
Why would I want a random user in a shop to know my full name?!?
Do they? I had three unexpired bank cards to compare.
My good bank issued me my non-contactless credit card, which is a backup and also the card my Phone "is" when I use that to pay for stuff, which is most of the time. That card (which a little worn) has my first and last name with middle initial.
My good bank also issued me a debit card very recently. This card is entirely black on the front except for the name of the bank and the logo of the card network. However on the back it has my initials and surname.
The other bank I use that does card transactions issued me a more traditional looking card with just my initials and surname on the front.
I think they should ask you how you want your name on the bank cards
If you ask people fifty questions they don't care about, and they don't need to complete the process, they just won't do it at all.
"Would you prefer our letters to you addressed you by your first name, your surname and initials or just Dear Sir/Madam?"
"If we have to call you about your account, should we text first or is it OK to just call immediately?"
"Our cards can be doused in artificial lemon scent instead of smelling of plastic when new. Would you prefer the lemon scent?"
"If you had to fight either one angry goose every Tuesday, or an enraged orangutan on a random day once per year, which would you prefer?"
"Would you prefer Contactless cards or should we disable the Contactless feature on your new card?"
"Going back to the goose question, does your answer to the previous question change if the orangutan has a sword?"
I'm sure there is good government software out there, but there are plenty of showcases of the opposite (especially since these are systems that NEED to work).
To be a US resident, you must be known to the government. To leave, the government will tax a percentage of your wealth. Try to tell the government to pound sand while remaining, and they can bring the full legal weight of their monopoly on force against you.
The difference in scale of harm between what CVS is capable of versus the government means we should have no issue with holding government to a much higher standard, And be proactive in pointing out when it falls short.
Not sure what did go through, but I'm sure that "Christofer" as a test string didn't.
The way to make a profit on government contracts:
1) Underbid
2) Find every possible use case that’s OBVIOUSLY needed, yet not in the specs
3) Leave them unimplemented
4) Charge through the nose for the extra work.
The system has already been paid for and put to production. It’s too late for the buyers to back out, lest they be out both money and face
My first son is named Christopher, and we realized right away that there is a 10 character first name limit in a ton of systems almost from day 1 - calling my insurance company, the automated system asked "Are you calling about... Christophe?"... in a French accent, which was hilarious.
My compiler started out with those, but I quickly realized that there were only two actual limits:
1. allocated memory
2. stack size
It turned out to be much less code and many fewer error messages to just detect out of memory and blowing up the stack.
Limits will be there whether you standardize minima for them or not.
For instance, no C compiler will handle a function with any number of parameters; as you try test cases with more and more parameters, eventually they will all crap out, right? If for no other reason, than memory.
So then if you don't have anything in the standard about this (like the ECMAScript standard doesn't, for instance) you have to conclude that all those compilers are nonconforming.
But if you have it in the standard that a conforming implementation must handle functions with 128 parameters, then those implementations which crap out at over 128 are spared from being declared nonconforming. Only those which crap out at 128 parameters or fewer have a conformance issue.
Limits settle questions. The user reads the spec and then has confidence they can have 4095 characters in a literal which will be portable to anything calling itself conforming. The implementor reads the spec and then has confidence that their implementation doesn't have to handle more than 4095 characters in a literal in order to handle maximally portable programs. There are goalposts firmly in the ground.
It can be useful for an implementation to optionally diagnose when a minimum limit is exceeded, even if it allows for much more. For instance, to tell the user they are using a larger character literal than 4095. This helps the user prepare a portable program that will work in other implementations.
If a compiler has to work with some object file format which allows only six character symbol names (with case being indistict), there is nothing that can be done; but there is little reason to impose such a stringent limit if the underlying linker technology has a much more lenient one.
I think C90 still made that 6 character the minimum limit; a conforming C implementation of C90 was not required to accept a program which makes two external definitions that do not differ in the first six characters of the name, ignoring case.
In those days, it would have been useful to have a warning if you're using longer identifiers, or if your identifiers are not unique in the first N characters.
C99 raised a bunch of these limits, including that one; it became 31 characters. No mention is made that case is indistinct; it must be distinct.
The limit is still there; use 32 and your program is not strictly conforming to C99.
That decision probably dates back decades and was based on the limited amount of space on a 80 column punch card. And as the system evolved past punch cards, no one bothered to update the spec because "it's always been 10 characters, if we change it now, something might break"
They thought that the next computer system would be better, and when they re-wrote it for the new machine, they'd be able to fix the problems they found, in about 3 years or so. They certainly didn't expect it to still be running in the 1970s or 1980s, let alone in 2021.
IBM broke everything when they introduced backwards compatibility. It saved a ton of time, in the short term, but everything before that point was frozen, and the technical debt it caused has never been paid.
The limits have probably been in place for 40+ years and nobody wants change it. Or maybe nobody knows how. It's much easier to tell them to abbreviate their name as "Chris" and move on.
Or the government not having enough funds to upgrade its systems since the 80s when storage space came at a significant premium (or not making modernizing IT systems a policy/budgetary priority)
Not to say that throwing money at government solves problems either. A competent administration is needed also to instil confidence in voters that they are actually committed to solving problems. Not sure what the situation is like in NJ
The 'government work' was probably done by a private contractor. Note that the OP is about a private bank. If you haven't seen results like that in private business, you've lived a charmed life!
Falsehoods Programmers Believe About Names
https://www.kalzumeus.com/2010/06/17/falsehoods-programmers-...
Nevermind that that might not be your name or that in the future having the untruncated name might be useful.
Definitely a big problem for then using that data to form the base of an identity system like trusted ID compliance.
If you have no mainframe or enterprise experience to relate to that observation, consider the effort involved to transition from python 2 to (UTF-8 clean) python 3!
That said, I am not even clear from the article which diacritical markings are missing from EBCDIC and if the lawyers arguments to "not change" were legitimate in the way the article implies... you do realize there are hundreds of EBCDIC code pages covering at least all the European languages ... since these are markets which IBM has sold into for 50+ years now, right?
I only learned about EBCDIC code pages when trying to proactively properly setup character encoding handling for data extraction from one of my employer's long running AS400s... "Which EBCDIC?" is not that different a headache from "which extended ASCII code page?"... EBCDIC is not just like 7-bit (non extended) ASCII as the article implies.
I don't think the author implied it was easy, just that it should have been done at some point in the 25 years since the system was first implemented. The last paragraph is just an exortation to use Unicode everywhere all the time, today.
Sure, but that's still a massive headache. You've probably never had a a headache like needing to switch ASCII or EBCDIC code pages. You generally can't just switch code pages per-record in a file, storing mixed code page data to disk is generally a bad idea, and in some operating systems you can barely switch code pages per application and sometimes need ROM hacks and entire mainframe restarts to switch code pages. (Modern z/OS supports something more like modern Linux locale switching with environment variables before running applications so should at least allow per-application code pages.)
Even if the lowest common denominator code page you choose to run your application in is a full bit or two more than the 7-bit ASCII lowest common denominator a single code page per application is still never going to cover the breadth of UTF-8 without nasty hacks. (That's of course assuming you don't have other problems such as intermediate tools that presume you are only using ASCII compatible EBCDIC subsets of code pages, which may be the case when you've got an eclectic evolution of code accreted around your mainframe apps.)
https://news.ontario.ca/en/release/58538/ontario-introduces-...
Going further, I wonder why for example the EU doesn't try to get schemes going that facilitate the copying of IT solutions between member states. Why does every country have to reinvent the wheel?
(sorry, couldn't resist - on topic for the thread...)
Hopefully we'll move past these eventually.
There's a law that says: "For a computer system to be purchased by the government it must work in French".
Implementation is then left to potential sellers.
https://www.microsoft.com/en-us/download/details.aspx?id=102...
[0]: https://bitbucket.org/coldacid/usintalt/src/master/
[1]: https://web.archive.org/web/20160327005949/https://chris.cha...
And plenty of people have already made "United States-International (no dead keys)" for Windows, so if you don't want to figure out the MS tool, you can just download+install a layout from GitHub.
https://en.everybodywiki.com/EBCDIC_297
So, to be unable to represent á, è, ô, ü, ç, etc, the application would have to be locked into not just Ebcdic but also a particular Ebcdic code page that seems unsuited to the locale where the program was running.
Admittedly, an Ebcdic system will have difficulty representing French, Greek and Russian names at the same time, because there's no code page that encodes all the necessary characters.
An application hard-coded to US-Ascii would also be unable to support accented characters, and an application using any one Ascii code page (as opposed to Unicode) would have the same difficulty representing French, Greek and Russian names at the same time. Which is why, in 2021, we don't do that.
Now since I'm Hongkongese where my English legal name is as legal as the Chinese one the law might be different but for Chinese people though...
If you have special letters in your name, you'll have a different name in another country without that letter. My surname is supposed to be Özcan, but it's Ozcan or Oezcan in many official documents. Don't even let me start with the "Turkish iİıI problem"...
I mean it's not totally unrecognizable but it's a different name nevertheless.
I was talking to a Romanian colleague recently and she told me that most of the country uses some US keyboard layout instead of Romanian and cannot type Romanian letters, so people have 2 names even in their home country.
Latin. Yes, a lot of textbooks add some diacritics to show pronunciation (as Latin wasn't 100% consistent between spelling and pronunciation), but the Romans themselves didn't use them.
[0] https://www.reddit.com/r/todayilearned/comments/5vcd2h/til_t...
They did sometimes: https://en.wikipedia.org/wiki/Apex_(diacritic)
Or just consider how the Japanese put up with computing in pure katakana (their writing system's equivalent of all caps) well into the 1980s.
Indonesian and Malay languages.
Also, probably thanks to Dutch colonialism, the unmodified Latin alphabet is the official writing system in Malaysia, Indonesia, Brunei, and Singapore, and is used to write the Malay and Indonesian languages.
It is also used as the base for romanization systems for languages that don't have a latin-style alphabet already. These are often designed to stick as close to plain lating as possible. Apart from academia and language teaching, a few of them are actually used by governments to render names in latin characters for passports, street signs etc.
Personally, I think that the plain latin alphabet is quite limited and that extensions are necessary. Accents, macrons, circumflexes, etc. are certainly annoying to input, but certainly not worse than inventing completely new letters or using digraphs for everything. I rather think that our educational systems don't teach well how to handle them. We don't have to pronounce them all correctly, and certainly can't be expected to, but typing them is not impossible at all!
https://www.opentaal.org/het-laatste-nieuws/171-karakterfreq...
The diacritical marks however have some familiarity and are in common use.
On a sidenote: lots of airlines also have this issue where an accent or other dimark will remove the character completely making your name different from the one in your passport. Could be quite annoying.
edit: thought it was in the Netherlands but it was in Flanders/Belgium.
Edit: I just saw it was in Belgium, but the same should apply there. Although they seem to be using a variant of code page 37 called code page 500 (also in [0]).
There is actually no good way to tell whether a name is Mandarin or Cantonese, except maybe by looking at the place of birth or residence. Ironically, the romanized form might give clues as there are many different romanization systems in use.
Edit: ah on the linked wiki article it says:
> The Court of Appeal of Brussels held that, in accordance with Article 16 GDPR, the data subject has the right for their name to be correctly spelled when processed by the computer systems of the Bank
So the plaintiff won, but no word on if/how the bank actually fixed it.
Source (Dutch): https://www.gegevensbeschermingsautoriteit.be/publications/a...
This tweet says it was ING Bank: https://twitter.com/simonhania/status/1270812210584043521
Ouch. Basically "That you're not able to write a customers name correctly in 2018 because you use a system from 1995 is not an excuse".
Iso 8859-1: 1985 Unicode: 1991 Utf-8: 1992
There is even IBM277 for an ebcdic version.
I think the point here is this is completely out of left field as far as what anyone has insisted would be non-compliant with GDPR... If you had done a compliance audit the day GDPR passed, I highly doubt this shortcoming would have even made the footnotes.
"Overly broad and interpretable law with rabid defenders is stretched to painful limits just as critics predicted" is the real story.
To me the story with GDPR has consistently looked like "IT companies unable and/or unwilling to comply with (or even read) laws when they feel they go against their established practices, no matter how bad such practices might be".
Before today, you cannot seriously tell me that (hypothetical) United Airlines being unable to print æ on your boarding pass would be a GDPR violation. No one would even have considered it. The best "GDPR auditors" that popped up to save the day with expensive consulting would have glossed right over it. And yet the overly broad language of the regulation allowed this contrived gotcha. And now any company that can't support emojis in your surname is now in the Naughty Bucket of GDPR Violators.
I'm just shocked how so many hackers are ok with this law existing in its current form, just because it sometimes achieves things that they like.
If we find ourselves asking "what else can we hit with this hammer," it's a bad law.
Correcting incorrect data sure, that's part of what the law grants you. But I believe this case is novel in that the data is as correct as possible (for intents and purpose of banking) yet the courts are requiring a cosmetic adjustment to the data. Cosmetic as in: it does not change the bank's or customer's understanding of the contract and business organization (i.e. I'm not trying to downplay ones attachment to accented letters, I'm talking about correct identification for business purposes).
> and at the same time, this doesn't mean every contrived example of a name and where a name might appear actually has to support everything)
Why not? What language in the GDPR would prevent that? It's the same violation as this case: the name is not displaying how the data subject wants it.
I also fail to see the great importance of it, but apparently some people see it more than just a cosmetic issue...
And this isn't someone "being offended", this is a legal requirement of the GDPR to accurately record someone's name.
"Zoë" is not the same as "Zoe" when you go searching for it in a DB.
That isn't necessarily true. We all have a certain set of phonemes we can enunciate, and further have limits on how they can be combined together. It is far from inconceivable that the OP could have a name which effectively _is_ unpronounceable to people speaking other languages (and you can't just put in more effort to fix this, so those people aren't "just lazy").
But they're not unpronouceable... they just take some effort to learn.
Why do you think you can't fix this? I've never encountered something that is physically unpronounceable (fictional eldritch abominations and extra terrrestials aside)
Your name may be difficult to pronounce for native English speakers, but it's not impossible.
so I guess we're two peas in a pedantic pod :)
Were they superseded by more modern solutions? Absolutely.
Were they nonfunctional? Hell no.
I worked on several systems in the early 2000s that still had a big old mainframe at the back end.
I’m pretty sure most airlines and banks still run them.
The bank's lawyers took the wrong approach, IMO. The law (as quoted in the article) says:
> The data subject shall have the right to obtain from the controller without undue delay the rectification of inaccurate personal data concerning him or her.
This doesn't have that much to do with how a certain name is displayed anywhere. Can I get an airline to change their systems if they abbreviated my name on my boarding ticket? Yeah, I don't think so. The airline could say "well we have the proper name in this database over here". And so could the bank.
I expect this is how they will eventually solve the issue - the customer-visible parts will be insulated from the old system with stuff that can handle Unicode. Chances are they currently don't have such insulation, producing documents with the wrong names, hence the complaint (bank statements are often used as proof of ID).
Btw this is not a stupid law. Accents are important parts of languages, the tech to handle them has been around for decades now, there is no excuse for willful illiteracy.
The same goes for names. Misspelling the name by not using the correct characters is frustrating and shouldn't be necessary when we have Unicode. I think everyone can imagine how annoying it would be to have a random letter in one's own name be replaced by some other letter.
Yes, I'm personally annoyed by this, since my name also has one of these letters.
Indeed it doesn't, so no thanks to you for muddling the issue by mixing that in.
> Can I get an airline to change their systems if they abbreviated my name on my boarding ticket?
Well, since they're the ones insisting that it look exactly like on your ID, if it looks different because they mangled it and then they don't let you board because of their own cock-up, then I'd imagine even you would be slightly miffed...?
So yes, you should be able to expect them to fix their systems so that doesn't happen, shouldn't you?
> The airline could say "well we have the proper name in this database over here". And so could the bank.
No they can't, because they don't: The point here was that this wasn't just some printing or other display function, but how the name was stored; as far as this bank was concerned, this customer's name was something it actually wasn't.
How one can claim this isn't in contravention of a law that states personally identifying data stored about people should be correct is frankly incomprehensible; our names are among the most personal things about us and our main identifyer.
All in all, your whole screed comes off as yet another piece of typical reflexive-but-unreflected pro-corporate (usually American) Libertarian dri^H^H^H propaganda.
I took the magnetic reel to college with me that summer and asked around. Turns out they had magnetic tape reader for reels of this size hooked up their VAX system. A friendly sysadmin tried to read the data for me, but it came back has gibberish.
I wasn't surprised. Then he said "Aha! EBCDIC!" I hadn't heard it, but as the reel spoon and the names of the dead spun off the reel, he spun his own yard about this arcane format that was an ancient as the magnetic tape reel I'd brought it.
And yes, there were some dead people voting in Kentucky.
$ echo "très intéressant" | iconv -f iso-8859-1 -t utf-8
très intéressant
$ echo très intéressant | iconv -f utf-8 -t iso-8859-1
très intéressant
Admittedly no autodetection. Luckily EU mangling is usually just one or two encodings.
Just to list the iso-8859 parts concerning EU member states:
- iso-8859-1 (Latin-1, Western European, including German umlauts, French accents, etc)
- iso-8859-2 (Latin-2, Central European, including characters to support Polish, Czech, Slovakian, Hungarian and other)
- iso-8859-3 (Latin-3, South European, including characters to support Maltese)
- iso-8859-4 (Latin-4, North European, including characters to support the Baltic states)
- iso-8859-5 (Latin/Cyrillic, including characters to support Bulgarian)
- iso-8859-7 (Latin/Greek, including characters to support Greek)
- iso-8859-10 (Latin-6, Nordic, refinement of Latin-4, popular in Baltic states)
- iso-8859-13 (Latin-7, Baltic Rim, because -10 was not enough)
- iso-8859-15 (Latin-9, basically Latin-1 with the €-sign and some commonly used characters missing in Latin-1)
- iso-8859-16 (Latin-10, South-Eastern European, "Intended for Albanian, Croatian, Hungarian, Italian, Polish, Romanian and Slovene, but also Finnish, French, German and Irish Gaelic (new orthography)")
And they are all still in use. ;)
It seems to me -15 is now more popular than -1, probably because it supports the Euro currency sign.
ISO-8559-1: très intéressant
ISO-8559-9: très intéressant
;)
Microsoft had its own Windows-125x codepages, which were not always compatible with ISO ones.
At least I believe I did, because often times you get that stuff without any hint what it is and then you can make a somewhat educated guess only.
There is still a lot of software out there that doesn't default to unicode but either some of the iso-8859's or some windows code page.
E.g. WordPad in Windows 10 uses the an ANSI code page when saving an .rtf (WordPad default format)[0] or as plain .txt[1]. But fret not if you're on macOS, my TextEdit just saved this when I created a new document and inserted a few ä-s (the TextEdit default is .rtf as well)
>Ohne Titel.rtf: Rich Text Format data, version 1, ANSI, code page 1252
At least .rtf has some embedded metadata specifying the code page of the document and more! Including mixing in some unicode[2].
Why do I mention .rtf? Because there is a lot of .rtf out there, being saved and emailed around each day.
All bets are off for plain text formats that do not come with such useful metadata (except when it's some valid UTF (with BOM)).
There is also a ton of (old) email software/webmailers deployed that do not produce unicode. Some old and/or shitty enough to produce beautiful HTML soup that does not specify any character set at all. And all kinds of tools that produce text or csv and the like in all kinds of funky encodings that aren't anything unicode and quite often are selected based on the system locale or ANSI code page (on Windows). Or just hardcoded. Or yet better: mixing different encodings in the same file.
Do you want to read the exif metadata your camera embedded into those jpegs it spit out? The standard says everything is 7-bit ASCII, but of course that won't fly, so camera vendors and image editor vendors started using all kinds of encodings. And then there are vendor-specific additional tags. Windows XP's photo editor e.g. managed to smuggle in a creator tag that embeds the name of the user account which created/edited a file and it's using UCS-2/WTF-16, of course!
[0] To be fair, the format was specified long before there was unicode. To be even fairer, it's all 7-bit ASCII with escaped 8-bit whatever-codepage.
[1] There is an "Unicode text file" format option, which, you guessed it, saves it as... maybe WTF-16, maybe UCS-2, maybe UTF-16, not entirely sure, didn't check.
[2] If you need unicode, then the .rtf format lets you embed escaped unicode sequences inside the ANSI code page document :P
And people shouldn't criticize EBCDIC too much, after all Windows still dumps a lot of crap in legacy 8-bit coding that can cause applications to break (there was a recent post on HN about someone being unable to run the IntelliJ debugger because of an accent in their username). At least EBCDIC is clear about its limitations.¹
⸻⸻⸻
1. I'd be remiss if I didn't point out one other EBCDIC weirdness: It has two vertical bars, | and ¦ which always caused complications in translations between EBCDIC and ASCII. IIRC, ¦ was the more common symbol in EBCDIC coding but some converters wanted to translate | to | instead (or maybe it was the other way around—the last time I did IBM big metal was 30 years ago).
That's not Windows, that's JVM weirdness. Using the right calls, this sort of thing has been fine in Windows for some time.
E.g., can they assume that names can be expressed as a sequence of (current) Unicode characters with some specific maximum length? Can they assume that names have no leading / trailing spaces?
If you want to be safe maybe allow for up to 1000 arbitrary codepoints? :)
I mean something like "${name} spelled with an acute accent on the e" would be technically a correct description even if it is impractical to use. The GDPR does grant you the right to correct your personal information but doesn't specify how this information is represented.
As far as I can tell the GDPR also doesn't grant the customer the right to have their name represented correctly on their bank pass (otherwise everyone with a long surname would require impractically long bank passes), the court only ruled that the inability of the bank to store the name correctly simply isn't an excuse.
If you did officially change your name to something interesting, I presume the bank would definitely have to accommodate it; but the restrictive part would be the process of actually changing your name.
So, you can make reasonable assumptions. What is reasonable will change, which is fine because the way courts figure out what's reasonable in some particular case is to either have the judge decide, or have a jury decide, and people change too.
The nice thing about reasonableness is that you are equipped to make a first pass at judging it yourself, since you are presumably a reasonable person. If you need second guessing, have a team mate consider it, and, if you're worried that your collective idea of "reasonable" might be distorted in an important way, that'll be why your organisation probably encouraged diversity to avoid that.
You might say, this seems awful because it isn't precise enough to say, implement it as a Javascript library. That's true, but intentional. Justice will necessarily involve such judgement calls, and trying to evade that by specifying everything precisely with no room for judgement is a bug not a feature.
Regarding trailing spaces etc, IMHO the standard would be "as shown in passport" i.e. trailing spaces definitely would not matter, but spaces and punctuation between words would (e.g. D'Artagnan as a name). I looked for but did not find any specific restrictions on name length. In general, the country will have regulations on what they accept as names in their official IDs, and again you may piggyback on other institutions - as long as you accept everything for which your government have issued documents, you should be fine; and if someone has an interesting case that requires changing the process, let that fight happen between them and the government first.
Maybe their IT team had other priorities than replacing EBCDIC with Unicode (or whatever they find more appropriate for their systems), but this is an indicator of poor interest in technological progress by the bank itself. It reminds me some banks that gave millions to Microsoft to keep ATMs running Windows XP after its end of life.
Edit: I elaborated a bit more and I realized that it might be more difficult than just replace the character encoding standard to a more modern one. For example, the name of the account owner likely needs to match exactly the holder name on the credit card associated with the account, and I'm not sure if diacritics can be embossed correctly on the card.
Yes, already in 1995 Unicode was an established standard (even Windows 95 started to support it). The bank should have known it would be a requirement in the future.
UTF-8 was already a thing in Plan9.
The famous "Plan 9 implemented UTF-8 first" thread's most specific date mentioned was September 1992 which only three months more lead time before the standardization notice in January 1993.
Are you suggesting the NT Kernel team should have somehow better paid attention to a not-yet-standard from a research laboratory Operating System? It still probably would have been a couple years too late in the design/architecture process even if they had, given the release data in June 1993.
Those decisions were already way beyond locked back then.
[1] https://datatracker.ietf.org/doc/html/rfc2044
[2] https://www.iso.org/standard/26223.html
[3] https://www.debian.org/releases/etch/amd64/release-notes/ch-...
Crazy shit everywhere in these old systems.
Jokes aside. I know a person named exactly like me just with a small diacritic difference. I realize they use secondary identifiers but this is identity theft waiting to happen.
The number of places that don't support diacritics this is absolutely mind-boggling.
I don't see how that really makes identity theft any worse. There's already a ton of people with exactly identical names, and differing only in diacritics isn't any worse than that is.
I unfortunately have a number of people who have the same name as me. Get emails for them sometimes...
This makes checking in for flights a nightmare, since the names do not match. Domestic flights are OK, but I cannot checkin online for any international flight, I have to go to the airport to check in.
No, I'm not entirely serious.
It used to be a given that nobody wanted their bank to run on anything, but mainframes. Now we'd rather they used cloud computing and Postgres. Mainframes have had their day. They may have a future, but they need modern databases and development tools.
I suspect you ask random Joe on the street they would say they would rather be on a mainframe.
Also, I would much rather my bank didn't host on AWS, GCP, Azure, etc.
For example, IBM z runs linux.
Backward compatibility is one of the biggest selling point mainframes provide, you can run code written 30 years ago with all limitation and bug preserved as is.
And IBM z can also run many program write for original ibm/360 unmodified.
Probably many ones decided to do so.
Mainframes have outrageous transaction processing and reliability/redundancy capabilities. If anything, mainframes and modern programming techniques for them are underrated, largely because dumb licensing on the tooling keeps people from realizing their capabilities.
Unless I'm misunderstanding, EBCDIC does have support for all of the characters used as examples. Specifically, code page 37 - the "default" US EBCDIC code page and one apparently widely used on IBM mainframes[0] - has the characters at 0x45, 0x54, 0xCB, 0xDC, and 0x48 respectively.
(Actually, code page 37 has all characters that are in ISO-8859-1, they're just in different positions.)
The problem is that they are difficult to deal with. They cause problems with data entry. They make sorting more complicated. They make searching difficult. etc etc. At my current site names are even stored in uppercase to reduce these issues.
I welcome efforts like this to force computers to be more compatible with people - many mainframe systems are stuck in the dark ages of "to hard to be flexible" and need to be kicked into the 21st century.
(Okay, privacy wise, some might have uneasiness with SWIFT but we're talking about how it can't handle (in this case) characters outside US-ASCII, unless you have negotiated it with the bank you're sending on, which if it's a US bank, is not supported: https://twitter.com/ajlobster/status/735240869859753985)
However, it's worth noting that it's not just a single exceptional person - Belgium has accented letters in two of their three official languages and names with accents are reasonably common, so if you tried that, you would have to discard many customers, and also those customers would be overwhelmingly from the french-speaking part of the country so that might be treated as explicit discrimination targeting the french-speaking minority community.
The bank is entitled to use whatever internal representation they like (although they better don't misidentify persons), but they better be able to print the correct name in forms and letters.
The embossed name on payment cards is a separate, more troublesome issue. The bank should provide appropriate documentation as evidence if the cardholder's embossed name could not be spelled correctly because of limited technical standards.
It is WAY more complicated than I thought it would be. There is so much code that manipulates strings that is not unicode aware.
I've fixed the simple things, like places where the user name is displayed, etc., but the email subsystem is a train wreck and there are still places in the database where I couldn't retroactively fix old entries. Going on 4 years fixing this!
But EBCDIC? Damn, opportunity to make lots of $$ here fixing people's code. I had friends that made bank on Y2K prep in the mid-90's.
Actually, which email system supports email addresses with non-ASCII user names? And which additionally supports IDN domains?
https://oikeusministerio-fi.translate.goog/nimilautakunta?_x...
Edit: on the downside, they would probably also reject many reasonable names.
That suggests that the bank made a 'mistake' when it recorded the name in its system as well as it could. I don't think that should count as a mistake. The information in the bank's system is as good as it could be, so there is nothing to rectify.
It feels weird to me when privacy legislation turns out to require supporting UTF-8. I think something in the legal process went wrong here.
No. It requires you to store data correctly. And in the case of a bank storing data incorrectly could have potential ramifications (think two different people, one with diacritics and without).
The law doesn't care whether you use UTF-8 or a manually written translation table, or a 15th-century printing press
Surely this situation wouldn't cause any issues though right? If they are relying on names as unique identifiers, they've got far bigger problems than a users name being spelled incorrectly.
I find this whole ordeal delightful, and applaud the intention of the GDPR and the ways the courts upheld it in this case.
In the specific case of this bank, like everyone else, they were expected to update systems unable to comply with the legislation within the grace period and yet it seems that they were unwilling or unable to update or replace a system that is incapable of achieving data integrity in a matter as basic as the name of a customer.
I suspect this will be fixed by storing the correct name in an external system and using an ID number or similar to refer to the customer on EBCDIC systems.
> Following an eight-month investigation, the Data Protection Commission (DPC) have ruled that individuals do not have an 'absolute right' to have their names spelled with fadas.
https://ireland.bloomsburyprofessional.com/blog/no-right-to-...
I find it amazing that the Data Protection Commissioner basically went against the constitution which clearly states that Irish is the first language of Ireland.
Firstly, if a government decides to "comply" with the GDPR by just seizing and revoking your passport, you might not have a case against them as the granting of passports could be considered a Royal Prerogative (or an equivalent under other systems of government) and thus non-justiciable. You might try to claim this is discrimination, but I don't think that "non-ASCII characters in name" is a legally protected class, and of course anyone could change their name to have or avoid non-ASCII characters.
Also, if the format is designed to be machine-readable, then arguably the "accuracy" of your name on the passport has to be judged by the machine, not by you as the holder of that name. Moreover, the format is agreed as a consequence of an international treaty, which again might put it beyond the jurisdiction of a domestic court, and if your passport was declared invalid by a nation you were attempting to enter because it contained non-standard characters, that is not something that a domestic court could provide a remedy to.
https://gdprhub.eu/index.php?title=Court_of_Appeal_of_Brusse...
I wonder why this fact is absent from the blog post?
This is a baffling ruling. I don't think an inability to support funny characters should be considered a GDPR violation. Anyone can put any characters they want in their name, and everyone is not breaking the law just because Unicode doesn't have made-up squiggles.
Second, if you were trying to make a point that people can put Unicode emoji in their names, well, try doing that on a birth certificate and tell me what the registration office tells you. If you successfully manage to get an actual “funny character” in your legal name, let me know.
I want a pony.