The curious case of the disappearing Polish Ś (2015)
medium.engineering
medium.engineering
"Cze" is a very informal greeting, sth like "Yo". There has been thousands of such e-mails in Polish companies sent to people who really shouldn't be greeted with "Yo" :) It happened to me too, and I'm still not sure how it happened. I think it might have been caused by windows update.
And in the past it was even worse - Windows installed both Polish-programmer and Polish-typist keyboard layouts by default and there was a very easy to hit accidentally shortcut to switch between them. If you hit it most of things still work but your "y" and "z" were swapped. So if you had "z" or "y" in your password - the next time you log in - you suddenly can't.
This caused countless problems for IT support all over the country :)
Brilliant! Growing up in Slovenia there was only the typist locale, qwertz of course. Coding was full of altgr + <key> combos and it gets real annoying real fast.
Most of us nerds switched to the US layout fairly early and simply ignored diactritics. š becomes s, č becomes c, ž becomes z. Especially when texting came around and mobile phones didn’t even have those letters. Most people from Slovenia can fluently read words without diacritics and recognize based on context. Even if you weren’t a nerd, in the 90’s and 00’s it was a gamble if the software or website you’re using will render those characters so people learned not to use them.
iOS still doesn’t have a Slovenian layout. Not sure about Android.
I hear in the commodore days people used to write ss, cc, zz but that didn’t catch on in my generation.
Maybe it was brilliant early on, but by 2000 like 99% of computer users didn't knew about the typist layout and yet it was always one accidental keypress away :) It was a shortut that you could press accidentally when you highlighted whole words with ctrl+shift+left/right.
> ignored diactritics
Yeah we wrote without diacritics too early on. On computers because in 90s there were 3 competing 8bit encodings so it was like 66% chance you got the encoding wrong. Most nerds (me included) could recognize the most common encoding mismatches visually :) Thank god for utf-8 :). We had diacritics in phones but it was a hassle to type them with numerical keyboards on dumbphones, so most people didn't. So now it looks "lazy" to avoid diacritics and most people use them AFAIK.
BTW it's funny how outdated Polish ortography (we still have digraphs sz/cz/rz/dz and even some trigraphs like dzi which most Latin Slavic scripts replaced with š etc.) became a positive early on during the computer revolution.
Android is going to depend on which keyboard you're using. SwiftKey has Slovenian in the list, but I can't comment on how well it works.
I also had a formal footer setup added to every email, so it looked really legit.
*** WAŻNA INFORMACJA/ZASTRZEŻENIE *** Niniejszy dokument powinien być czytany wyłącznie przez osoby, do których jest adresowany. Jeśli otrzymałeś tę wiadomość, była ona oczywiście zaadresowana do Ciebie i dlatego możesz ją przeczytać, nawet jeśli nie chcieliśmy jej wysłać do Ciebie. Jeśli jednak treść tego e-maila nie ma żadnego sensu, prawdopodobnie nie jesteś zamierzonym odbiorcą lub jesteś bezmyślnym kretynem; tak czy inaczej, powinieneś natychmiast się usunąć i zniszczyć swój komputer! Gdy już podejmiesz tę akcję, skontaktuj się z nami.. nie, idioto, nie możesz używać swojego komputera, właśnie go zniszczyłeś, a przy okazji, ty też zostałeś usunięty, ale odchodzimy... Nadawca tej wiadomości e-mail nie ponosi odpowiedzialności za przekazanie informacji zawartych w tej wiadomości, chyba że jest nadawcą i w takim przypadku prawdopodobnie ponosi odpowiedzialność i słusznie, biorąc pod uwagę treść powyższej wiadomości.
Jeżeli nadawca nie wysłał do Ciebie tego e-maila, zwróć go nam i załącz zeskanowane zdjęcie żony brata Twojej matki ubranej wyłącznie w majtki na ramiączkach, a my natychmiast zwrócimy Ci dokładnie połowę kwoty, którą zapłaciłeś zapłaciłeś za puszkę Pal Meaty-Bites, którą kupiłeś wczoraj, gdy byłeś w Woolies.
Nie bierzemy odpowiedzialności za nieotrzymanie tej wiadomości e-mail, ponieważ korzystamy z systemu Windows NT i wszyscy wiedzą, jakie mogą wystąpić problemy. Jeśli otrzymasz tę wiadomość, pamiętaj, że nie bierzemy za to żadnej odpowiedzialności. Nie ponosimy również żadnej odpowiedzialności, milczącej lub dorozumianej, za jakiekolwiek szkody, które możesz ponieść lub nie, w wyniku otrzymania lub nie, w zależności od przypadku, od czasu do czasu, niezależnie od wszystkich dorozumianych lub innych zobowiązań, hmmm, cholera , gdzie ja byłem..umm, nieważne co się stanie, TO NIE JEST I NIGDY NIE BĘDZIE NASZA WINA!
Komentarze i opinie wyrażone w tym dokumencie są moimi własnymi, a NIE mojego pracodawcy, który gdyby wiedział, że wysyłam e-maile i surfuję po stronach porno, odciąłby mi gonady i nakarmił mnie nimi do popołudniowej herbaty.
It feels similar to developers removing the ability to select text from not just one-word buttons, but all kinds of UI elements on a website (because it looks ugly when people accidentally select the text), which has the effect of also removing the user's ability to, well, select text, for all sorts of valid reasons (dictionary lookup of an unfamiliar word, copying a line of text to someone else (bug report: “the app says '[…]' when I do this, but it should say […]”)).
This happens in many programs for various keys, depending on which keyboard shortcuts a program implements. A solution to this is to allow remapping of keyboard shortcuts in an application.
Btw, I cannot recommend enough the other work by Marcin Wichary. For me, he's the paragon of making the most of what you can do in terms of, to put it best, tailoring things exactly to your specific needs and tastes. (My copy of Shift Happens is on its way, can't wait!)
- ł and ć: https://news.ycombinator.com/item?id=8988553
> Right Alt in Windows was internally mapped as a rarely-used combination of Ctrl and Alt pressed together.
Why "was" ? (the link is broken) You will also note that the new AZERTY doesn't use left Alt, rightly assuming that it's expected to be used for menu shortcuts.
(It's a shame though that the new AZERTY can do Greek... but not Cyrillic, a real missed opportunity for a true pan-European keyboard !)
----
> Both Windows 3.x and 95 had terrific keyboard support. The menu items and dialogs had controls that could be accessed easily by mouse… but also much quicker by pressing Alt and the underlined letter:
[...]
> Most of Microsoft Windows UI could be turned into a sequence of keyboard shortcuts, which was incredibly powerful (and something Mac could still learn from).
Is that specifically Windows though, or rather IBM's Common User Access ? (Apple, as a competitor, and being Apple, did their own thing.)
https://news.ycombinator.com/item?id=39130984
http://www.susandoreydesigns.com/software/CommonUserAccessGU...
In favour maybe of the likes of Dvorak (hey, wasn't he Czech ?) or BÉPO (any central/eastern-European equivalents ?). (The same researchers seem to have made a "new BÉPO" version, which, I guess, might be a "next step", when/if the new AZERTY becomes common ??)
But that's not the point of the new AZERTY, its point is to provide a keyboard layout that finally allows proper typing in (Latin- and Greek-based) languages other than English. (The old AZERTY still didn't allow that, for instance no capital letters with diacritics, unless if maybe you use more advanced software ?) (Also, a much easier access to the {([])} symbols than the old AZERTY, which programmers are sure to appreciate !)
All the while still keeping a familiar layout - and you'll have to concede that it should be more familiar to QWERTY / QWERTZ users than Dvorak / BÉPO, won't you ?
Microsoft was following the CUA principles, more or less, but they still get credit for actually doing it. Apple has some stuff for keyboard access, but it never felt complete; maybe I just needed to read more stuff, but it was always hard getting along on a Mac without using the keyboard, whereas Windows was always mouse optional (although it gets worse with every release).
Also, maybe less so today, but at the time when internationalization was usually even worse than it is today on typical computers available to general public, it was very common, at least for Poles outside Poland to write w/o any diacritics at all. There's not so much overlap, and usually you can figure it out from the context.
BTW, in a very similar way, Russian ё is often replaced by e. To the point that the "incorrect" variant is today acceptable even in newspapers. Alternatively, if I see Russian text that uses "ё", I get a feeling that it's about patriotism and must be some self-aggrandizing propaganda.
Countless browser extensions do the same. US-centric devs look around for unused key combinations and are happy to find them without asking themselves the question of 'why are these free?' Can't blame them too much, though.
For lower-case "ś", the pop-up mini-menu will have three options, the first being "ß".
Strase -> shtrazeh
Straße -> shtrasseh
The name of the character (eszett) means "sz" but actually represents "ss" for tortured historic reasons that I can't recall, but which I think [1] tries to explain.
More typographic trivia: The uppercase transform of "ß" is "SS", which causes bugs in code that assumes case changes won't change the character count. The language council fought for decades against recognizing an uppercase ß but eventually gave in if only to have the ability to computer-encode erroneous capitalizations, and it's been in Unicode for over a decade now: ẞ
[1] https://thelanguagecloset.com/2022/11/05/the-story-of-eszett...
I can't decide whether this is hilarious, depressing, or both at the same time. (BTW, another Pole here, who can't really comprehend the American obsession with race/skin color. I'd understand it – in a sense – 100 or 500 years ago, but it's 2024, why does anyone even care?)
Funfact: Nobody knows exactly how many people of Slovenian descent live in USA because until 1918 many of them just said "Austrian" on the immigration forms to avoid anti-slavic sentiment. And we were technically part of AustroHungary so it wasn't even lying.
Officially there's about 180,000 of us. Orders of magnitude fewer than the minorities that DEI efforts tend to think about.
America's very founding was based on the "obsession" of race/skin color. The expansion west wasn't very, uh, egalitarian.
> I'd understand it – in a sense – 100 or 500 years ago
100 years ago was 1924 - the "Tulsa race riots" had been 3 years prior, and a certain voluntary organization known for grandiose titles and its obsession with race/skin color was having a massive resurgence.
> ABCDEFGHIJKLMNOPQRSTUVWXYZ
That's not all that similar. The Latin alphabet would be:
> ABCDEFGHIKLMNOPQRSTVX
(You can browse the Pompeii graffiti and many inscriptions are just someone writing down an alphabet. It's pretty cool, if you're me.)
The five letters we've added (though Y and probably Z were already in use in Latin, representing Greek letters in Greek words) are pretty comparable to Polish adding 9 letters to ours.
> To find room for the extra letters, typewriters needed to dispense with some punctuation, most notably semicolons (comma + backspace + colon), and parentheses (replaced in common use by slashes).
...why? This would make sense if manufacturing typewriters were difficult and expensive, making it imperative to manufacture Polish typewriter chassis on all the same machinery that makes English ones. I tend to suspect that even at the time this was not the case, but maybe.
There's no excuse today; just use more keys on Polish keyboards.
In Portugal we had a similar workaround in the early days of computers not supporting our alphabet properly.
Like in Polish there are plenty of words that without diacritics get another completly unrelated meaning, e.g. caça vs caca, which you didn't want the interpretation to be left to the receiver.
So tricks got invented, like adding additional letters for the missing diatrics, é becomes eh, è becomes he or eh as in the former case, the example above would be cac,a and so on.
However it was still quite flexible, not everyone uses the same extension set.
But that isn't really helping. He had to find the code to begin with, go down the debug trail from scratch to get to the point of reading the comment. He had special knowledge to help him; most of the rest of us would have taken longer.
Big code bases need a 'business requirements' document. If e.g. you wanted to rewrite the codebase, you could read the BR doc and know that if you accomplished all that, you'd be feature-compliant in your new implementation. That is, you weren't going to get any rash of calls from customers about their feature being broken or entirely missing.
You can pore over the codebase and get all that info without the BR doc. But like my son complains at his job, shepherding a million-line behemoth ten years forward, much of what he sees in the code is opaque, badly done (and thus hard to figure out what's the feature and what's accidental side effects) or obsolete (that customer special became irrelevant five years ago or whatever).
A business requirements doc can be reviewed by actual non-programmers and judgements made of what's still relevant. So revisiting the project (reimplementing, integrating a new feature into the rats nest without disturbing previous features) becomes if not easy, at least possible in a bounded timeframe.
You young folks will have lots of tools for this kind of thing, born of Agile processes I suppose. I suspect any such database of entries is woefully sparse, badly described or never referenced. Like most 'productivity tools', if they aren't part of somebody's job description they become orphans.
Don't know the ultimate solution. People are people, and diligent leaving of breadcrumbs is just another annoying thing to do. But somebody (maybe you?) will bless you for it down the line!
But they do not help you find the code responsible for a certain business requirement. Not unless that's what you put in the comment e.g. a BR reference.
I find code comments to be at their best, when they explain how the code works or is to be used or how it relates to something elsewhere. You know, comments about the code or other code. Relevant to understanding what you're about to read.
Comments that are essentially user stories disrupt the flow and don't help interpret the actual logic. Those comments belong elsewhere.