The plague of emoji insertion in French docs
bibelo.info
bibelo.info
This has been a problem since the typewriter age. People having to get on with their jobs coped with it by using a full, breaking em-space. Unless this gets replaced automatically by the word processor, you get horrid typography and misplaced line breaks all over the place.
The Académie Française should have dealt with this years ago, if their ass wasn't stuck in the 17th century.
But honestly, rather than changing keyboards (which is hard), why doesn't Google just pick a shorthand that doesn't break typography rules, like `@` instead of `:`
If triggering it by accident wasn't intended, they could have used a regular keyboard shortcut like Alt+Shift+E or even instruct users to press Win+. to use the system-wide emoji menu.
But I agree that we need to make several alternative space characters easy to type:
- non-breaking space (for this French rule)
- wide space (for disambiguating sentence
ending periods from non-sentence-ending
periods)
- zero-width non-breaking space (for
preventing word-splitting?)> Why should we change language to appease lazy software developers?
It's been done.
For example, in Spanish it is no longer the rule that "ch" and "ll" sort as if they were distinct letters (this change was made in 2010) precisely because that was such a difficult rule to implement. And that was a 256 year-old rule per-wikipedia:
The digraphs "ch" and "ll" were
considered single letters of the
alphabet from 1754 to 2010 (and
sorted separately from "c" and "l"
from 1803 to 1994).
For another example, in Spanish capital letters were required to not carry accents, but now they are allowed to not carry accents. This was due to the use of overstriking on typewriters working to accent lower-case letters but not upper-case letters (apostrophe would collide with the glyphs for upper-case vowels). But the technology to resolve this has existed in the Spanish-speaking world for a long time now, so the rule was finally dropped. (Not accenting upper-case letter can lead to ambiguities that are annoying.)It's not just precedent. It's that the original reason for some typographic (not even orthographic) rule is simply not relevant in 2022.
And it's not unreasonable for either French people, non-French French speakers, or just non-French French-non-users to propose the ditching of hard-to-implement French rules. Now, this particular rule is decidedly not difficult to implement, but it is an annoying rule to apply as a user -- I should know, since I speak and write French (though I am not French).
Also, it doesn't matter what the French Academy says, or what the Spanish Royal Academy says, or what Webster's dictionary says, or whatever. Language evolves, even to their consternation. Moreover, developers don't have to care that much -- I18N/G11N is fun enough, and employers have to care for legal reasons, but rules like the Spanish ch/ll rule can be much too hard even for non-lazy developers, and the Royal Spanish Academy can and did have to change, and it was for the better.
This problem has primarily come about because the Catholic church enforced the latin alphabet on languages where a different alphabet might have been more appropriate. Spanish, although more closely related to Latin, still has a few sounds which there's not a good latin character for. There's no (particular) reason (as far as I'm aware) that 'ch' and 'll' became diagraphs, while ñ acquired an accent.
Why should traditions change just because it's a bit more difficult to do things the 'old way'? why do we still bother with capital letters at the beginning of sentences? or speling things with two leters when one wil do, or riting silent leters wen you cant tell the difrens? & i dont think we need apostrofees n e more.
I didn't know that Hungarian had a similar issue.
> Why should traditions change just because it's a bit more difficult to do things the 'old way'? why do we still bother with capital letters at the beginning of sentences? or speling things with two leters when one wil do, or riting silent leters wen you cant tell the difrens? & i dont think we need apostrofees n e more.
I distinguish typographic and orthographic rules. The non-breaking, thin space before punctuation rule is typographic and outdated (i.e., motivated by outdated typographic technology).
I do want some orthographic rules reformed too, but I'm more interested in the ones that are just hard. In particular I'm interested in collation reform because we do often have to collate multi-language text items but with one collation -this is especially true in databases- so having collations for Latin-script-using languages be similar is rather useful. This is also true given that I'm not going to be switching locales when I switch languages -- I speak, read, and write multiple languages, but I never ever change locales.
>A modern solution is simply to have non-breaking space easily accessible in your keyboard layout
An even better solution would be --for grown up people who have progressed beyond cave-painting and want to communicate using, you know, actual words-- to be able to disable emojis completely.Fucking moronic shite that they are. I've seen people on Twatter and FB have entire conversations in bloody emojis. Talk about reverse evolution! Why don't we just go back to grunting and gesturing and have done with it?
Certainly more useful than the requirement to put a space before a colon, non breaking or not.
My god! --you're right. This is so much better and clearer than using those boring old fashioned words.
It looks like this rule is based on old typographic considerations. Much like the Spanish Royal Academy's rule that capitalized letters carry no accents (unlike the opposite French rule that capitalized letters do carry accents!), which stems from typewriters not having accented letters, so one would type a vowel, backspace, then an apostrophe to make an accented vowel, but for capitals there's not enough space so you couldn't and wouldn't overstrike them.
Users and language academies should distinguish typographic from non-typographic language rules, and typographic rules should be context-specific (well technology-specific, since technology is the context).
No human language on Earth is in a position where it can laughs at others for their idiomatisms.
English
that person was wrong. :P
"some-chars" + <whitespace> + ":"
must be treated as a single word in French.
(I guess it's more complicated than I imagine it is, alright)
The former is easy enough, but also very annoying to multilingual people since one might run in a Spanish locale but occasionally write in French. So that's not a solution.
The latter is... hard to do, because while Unicode has language tags that you can embed in documents, those are deprecated and they were never well supported, and so there's no way to mark-up text as being in one language or another, and a document-wide setting wouldn't be enough nor sufficiently generic and standard and portable.
The best solution here is to relax the French typographic rule (since it isn't needed anymore). But that would take time to filter through to French speakers (writers, and readers) so that they learn to not put that pesky space before punctuation, but also so that they don't complain when it's missing.
Or... you know, this business of emoji pickers could be something you could turn off. Nahhh, that would never fly! (/s)
Huh? MDN doesn’t mention this … why are they deprecated?
IMO we should have tried harder to make Unicode language tags useful and used. But it didn't happen, so they're a thing of the past. Of course, they're still there, and one could attempt to resurrect them, but most likely one would fail.
Choice quotes below:
https://www.rfc-editor.org/rfc/rfc6082
> RFC 2482, "Language Tagging in Unicode Plain
> Text" [RFC2482], describes a mechanism
> for using special Unicode language tag
> characters to identify languages when needed.
> It is an idea whose time never quite came.
> It has been superseded by whole-transaction
> language identification such as the MIME
> Content-language header [RFC3282] and more
> general markup mechanisms such as those
> provided by XML. The Unicode Consortium
> has deprecated the language tag character
> facility and strongly recommends against
> its use. RFC 2482 has been moved to
> Historic status to reduce the possibility
> that Internet implementers would consider
> that tagging system an appropriate mechanism
> for identifying languages.
>
> A discussion of the status of the language tag
> characters and their applicability appears
> in Section 16.9 of The Unicode Standard
> [Unicode52].
https://www.unicode.org/versions/Unicode5.2.0/ch16.pdf (section 9 of that chapter, 16) > 16.9 Deprecated Tag Characters 519 The Unicode
> Standard, Version 5.2 Copyright © 1991–2009
> Unicode, Inc. for detailed recommendations
> on the use of U+FFFD as replacement for
> ill-formed sequences. See also Section 5.3,
> Unknown and Missing Characters for related
> topics. 16.9 Deprecated Tag Characters
> Deprecated Tag Characters: U+E0000–U+E007F
> The characters in this block provide a
> mechanism for language tagging in Unicode
> plain text. These characters are deprecated,
> and should not be used—particularly with any
> protocols that provide alternate means of
> language tagging. The Unicode Standard recom-
> mends the use of higher-level protocols, such as
> HTML or XML, which provide for language tagging
> via markup. See Unicode Technical Report #20,
> “Unicode in XML and Other Markup Languages.”
> The requirement for language information embedded
> in plain text data is often overstated, and
> markup or other rich text mechanisms constitute
> best current practice. See Section 5.10,
> Language Information in Plain Text for further
> discussion.
(Reformatting is mine.)This stuff should just stop. Operating Systems / Browser vendors should instead standardize on a hotkey to bring up an emoji selector that steals focus to filter via typing and inserts on enter key.
Finding some sequence that also isn’t a valid Perl substring is impossible, in any case.
I'd be surprised MS Word doesn't do the same. No need for a "true" typesetting solution.
And another "US English" centered thing; spacing does not really matter in English, but can have functional differences in other languages and scripts.
https://www.youtube.com/watch?v=2yWWFLI5kFU is a fun look at one of the problems with Unicode in general.
Did the 646 standard account for variable-width characters at all?
What are the actual official rules in this case (and are there links to those? Because that'd be fascinating information to read through)?
Actually before the ":" specifically there should be a regular non-breaking space, not a narrow one. Except in Switzerland. Other punctuation marks take the narrow non-breaking space.
Which I guess means we're completely free to ignore it. Pour mêmes raisons =P
I just tested with `setxkmap fr`. These are not shifted:
:!;
Only this requires Shift: ?
Also French layouts use an inverted number row (although none of those are accessed through that row).In fact, french standard body AFNOR actually updated their AZERTY layout standard three years ago to include more characters, including the narrow non-breaking space. In traditional ISO-like fashion, one must pay to access this standard, but you can find an example here:
https://commons.wikimedia.org/wiki/File:KB_-_AZERTY_-_AFNOR....
It's mapped to AltGr + Maj + Space. Now you just need to find how to install/enable this layout on the platforms you care about.
This is the best reference I could find: https://www.lalanguefrancaise.com/articles/espace-insecable
I do not want to have to close boxes that appear at arbitrary times with the esc key to continue regular typing, since the box interrupts flow and there is a reaction time between seeing the box and closing it with esc.
It would be ok to me if you have to use another less regularly used key, e.g. tab or ctrl+space, first to get into the suggestion box and then use arrows to select one.
This applies also (especially in fact) in code editors, chrome devtools, etc...
I wish others felt the same so this feature would be implemented in a less interrupting way by default. I've never seen it implemented in a non-interrupting way in anything ever.
I don't see why the Google Docs implementation wouldn't work that way. If it really must pop up a box, allow the user to continue typing in document and gradually reduce the list of candidate emojis. As soon as a character is typed that either complete the emoji or rules all of them out, the box goes away.
Of course as always no settings to have a way to enable auto completion but not use the standard typing keys :(
Keyboard layout nerd: you should use a keymap that has it. Like AFNOR's latest AZERTY or BÉPO.
Historical nerd: traditionally wordprocessors have auto-inserted non-breaking spaces before colons, why should I care now ?
Because I don't want to switch locales in order to write short bits of text in another language. I don't even want to have to switch keyboard layouts -- just use compose sequences for diacritical marks and so on. And I don't use a French locale, but I do write French text sometimes.
I would imagine that in Europe this is a big deal. So many Europeans are multi-lingual... But you don't expect a Spaniard to switch to a French locale to write in French and vice-versa -- that's too disruptive.
Microsoft is probably the biggest development-focused company in the world, but their own work communications app, Teams, doesn't allow you to paste small pieces of code in a chat because it replaces punctuation with emoji.
Regardless, if I want an emoji, I will type the emoji. If I typed ':)' I want ':)' not some stupid yellow face.
Meta's Facebook Messenger doesn't allow those forbidden strings either. Meta's WhatsApp mobile does, but WhatsApp web does not.
My phone has an emoji picker if I want an emoji. I don't want what I type or paste replaced with an emoji. Anywhere. Ever. This practice should stop.
Don't you just put it in backticks `like this`? Even for really short things, it seems clearer and it avoids the machine "helpfully" changing things.
(Your general point is correct, of course, just suggesting a solution to this very particular nuisance)
I so so wish we could standardize on GitHub markdown for all these things!
But are you sure that you didn’t really want `:slightly_smiling_face:`? (Auto-complete to passive-aggressiveness)
```
multi
line
```
So annoying in fact that I've just spent the last two hours with Microsoft Keyboard Layout Creator (yeah, I'm still a Windows person) shifting the dead key version to hide behind alt-gr.
Then I used it to mitigate some oddities on the dynabook keyboard (pipe and smaller/greater than weirdly moved to where the windows connect menu button usually resides) and continued duplicating some of the usual alt-gr suspects at locations inspired by their US KB position. Right now I'm an inappropriately happy person, let's see how I'll think about it in the future (getting too used to something non-standard cab be a terrible cost).
even messenger supports the backticks and triple backticks now
One of my pet peeves: Badly implemented "press CTRL + / to search". Entering an actual slash, which requires e.g. SHIFT+7 on German keyboards, won't open the search box, but pressing the key where American layouts have a slash will.
Edit for clarity: I type "a to get ä and 'a to get à. If I want to type actual quotes followed by a vocal I need to “confirm” the quotes with space before typing the vocal. That’s all there is to it, really.
The funny thing is that for all the getting used to different layouts, the one thing that really trips me up all the time is the different hand positioning for Control+C on Mac vs PC. Luckily, you can change the modifier key bindings in MacOS independently from the keyboard layout. But yeah, shortcuts that assume everyone has a / or whatever key are a pain too.
I used a QWERTY keyboard in France for a while, but switching back and forth between AZERTY and QWERTY each time I went on someone else computer, which happened often, I had to switch, qnd it is qnnoying.
You can learn to touch type but keep in mind that not all layouts are the same, you can bring your own keyboard, change the mapping, etc... There are solutions, but it wasn't worth it for me, so I got back to that terrible AZERTY keyboard because that's what everyone has in France.
Also, if you program or use a shell, then the Windows U.S. international keyboard layout is just unspeakably unbearable.
That sounds pretty good indeed. Any idea how could I go about setting this up under macOS?
> Also, if you program or use a shell, then the Windows U.S. international keyboard layout is just unspeakably unbearable.
I have indeed a couple of little, weird problems when working in the terminal. That's inconvenient but still beats switching between 3 layouts 50 times a day.
Sorry, no idea, but I'd expect it to be possible on OS X. And I really want this for Windows, too, as I have to use it.
> I have indeed a couple of little, weird problems when working in the terminal. That's inconvenient but still beats switching between 3 layouts 50 times a day.
Oof, I can't get used to having to type '' to make one ', "" to make one ", `` to make one `, in the shell, in $EDITOR, etc. It's too painful. So I switch layouts as needed, and I hate it. Compose keys would be so much better!
option-` is for grave (è) option-i is for circumflex (é) option-n is for eñe (ñ) option-c is for cédille (ç)
When you called QWERTZ an abomination I simply had to upvote your comment.
It's slow if you have to type in another language, but in English it's fast "enough" for words like résumé.
Thankfully I only need to write in two languages. With almost the same script.
In the Netherlands letters with diacritics don't even have their own key (Dutch computers tend to use US plus €).
When writing email/docs, autocorrector for the win.
Same thing in Italian: `e instead of è. Although, the German replacement is cleaner, I'd say.
That's… quite certainly not an opinion shared with all of your fellow inhabitants.
As someone from Espanya, I welcome you to the club.
We don't use the accents in the same way as Spanish, Portuguese, Catalan and Italian, it doesn't mark stress.
⌥u e = ë
⌥w e = ė
For example, the key sequence required to type Kayodé is: ⇧k a y o d ⌥e e
etc.Edit: and ß is just a Opt+s. It becomes quite logical after you get the hang of it.
I mapped my Caps-Lock key as Compose, so I hit compose-a-" to get ä, and so on through the set. I've made customs for Ω and µ since I use those a lot and couldn't remember their default sequences. It's handy having ® and ™ on tap as well, and I get typographical niceties like — and … vs - and ...
I use it because programming language syntax is mostly ergonomic only on US-like layouts, e.g. typing "{" causes a lot of strain if you have to do it on every other line.
After pressing the right Alt key, AltGr, all of those characters are possible to type. Here's the layout:
https://upload.wikimedia.org/wikipedia/commons/thumb/2/22/KB...
⌥s creates the ß.
(Both obviously only applicable for Macs)
right alt + q = ä
right alt + w = å
right alt + s = ß
right alt + y = ü
It took a day or two to get a hang of it but it's such a relief not being constrained by everything being developed for US layouts
Once comited to muscle memory it's very easy
> Danmark, officielt Kongeriget Danmark, er et land i Skandinavien og en suveræn stat. Det er den sydligste af de skandinaviske nationer, sydvest for Sverige og syd for Norge, og det grænser op til Tyskland mod syd. Grønland og Færøerne indgår også i den danske stat som rigsdele med selvstyre indenfor Danmarks Rige (Danmarks forhold til rigsdelene kaldes Rigsfællesskabet).
Eight Æ+Ø+Å in 382 characters.
defaults write -g ApplePressAndHoldEnabled -bool true
Might help if it got turned off.This can be one of the reasons that switching to Dvorak can be super hard, all the shortcuts become quite weird unless you remap them each by hand.
bombcar: often the shortcuts will remain the same even on keyboards where those keys are nowhere near each other
You're objecting to opposite things: you don't like that the shortcuts stayed with the same symbols instead of staying with the physical key, and 9ev the other way around.
(And also 9ev is objecting to misleading documentation)
Issues like this are why often people just give up and learn US English enough to use the computer, then at least it’s tested and consistent.
... Oh, for fuck's sake! What English word do they think "V" is short for in the first place, there?
The validation was done by checking for the AZERTY-specific key combos that result in 0-9; on AZERTY to get a <number> you need to press SHIFT+<number>.
So if you had a QWERTY keyboard, you could not enter your phone number, the input stayed blank. (If you tried with SHIFT, you could get !@#$%^&*(() to the field however.)
I believe either FedEx or UPS still can't deal with it.
In top of that Swiss mail is quite strict to how a letter is addressed and what is written on the mailbox. Luckily the mail carriers know to ignore garbled umlauts.
However, my home country postal system has since changed to an online system for sending post out of the country: You go to a website and enter the destination address, then you pay and it'll print a code which you just give to the actual post office when you bring the letter/package whatever. What's infuriating is that they don't allow Japanese addresses - I can enter it, but it just creates an error message. So I again have to come up with a crude romanized address and cross fingers that it'll arrive (which it usually does - but up to a few weeks delayed).
In short - things are not getting betters over the years.
Now the hard part is changing how to deal with it, because now you need to patch all these huge production systems that were working before, and it's a breaking change.
So if you have a big stack: UI --> business layer --> DB
You can't just change the UI, as that will cause breakage in the pieces below. You start with the DB change, then the business layer, and you do the UI last.
Moreover this is not a simple change. Getting a DB to handle unicode is a known thing, but how do you train your customer support personnel to do data entry with unicode characters? Do you expect the customer to remember the code point? There are many characters that look the same which have different code points. So what you need to do is come up with a list of officially supported extended character code points and do entry with those, but still you will have others -- users of Asian scripts, for example, whose characters will be missing.
And you will get customers still doing the old "u" approach together with the official umlaut code point, sometimes for the same name and the same package. People will call up, trying to find their package, and randomly used one convention, since even though you can train your staff (which isn't cheap) you can't train your customers. So you need some system of identifying different orthography for the same underlying name, some with "u" and some with "ü" and some with "ue". And then that's going to require changes in a lot of other systems, for example reporting systems, etc. We see this problem today, taken to the extreme, in the various different spellings of "Gadafi".
But really it's worse, as a distributed system has
UI <--> business layer <--> regional DB <--> business layer <--> regional UI
And now it's much harder to make these changes without taking systems off line, for a distributed global network.
What they probably did, as most companies do, is look at all these problems and costs, compare them to relatively small benefits, and kick the can down the road, waiting for the next major system upgrade, and sometimes these next major upgrades can take 20 years to happen. Or even longer.
In other words, the hart part is the incremental breaking change to the system, not the dealing with the umlaut. If this was a brand new code base that was just being written, they could make it handle umlaut with much less cost.
[0] https://intellij-support.jetbrains.com/hc/en-us/community/po...
Guess what, y and z are swapped on a German keyboard.
What your parent comment complains about is that the keyboard shortcut is based on the location of the key, and not the letter.
What you complain about is that the keyboard shortcut is based on the letter, and not the location.
Obviously, these complaints don't contradict each other, both make sense in different circumstances. But figuring that out requires awareness of different keyboard layouts, and of the difference between KeyboardEvent.code (location) and KeyboardEvent.key (letter).
Seem rather straight forward.
Many shortcuts are positional, like "hjkl" or "wasd" for moving , or control/meta+numbers but there's no good way to denote them without assuming you are using a qwerty keyboard. Programming against them is possible,but many times there's limitations, ignorance about the different layouts, or straight neglect.
[1]: https://developer.mozilla.org/en-US/docs/Web/API/KeyboardEve...
Google's currently demanding I settle an invoice for an account I closed... and the only way to do this, or even get support, is if you have a Google account. And zero way of replying to the email their Collections department.
On the one hand they have a massive dominant position in multiple markets. On the other hand, if you're even slightly off the happy path you don't exist.
Maybe the blunt response to complaints about the US American bias is, it's our software, our (ASCII/Unicode) language, get used to it.
I mean, seriously, what kind of antiquated statement is that? Do you actually want to keep geographical borders to persist to the internet, or work towards lower barriers for everyone..?
They haven’t and they got the money anyway.
I wrote it over my lunch break quickly, it only closes the menu the first time. I was trying to add the functionality to close it every time you enter the key combo, but it wasn't working. If someone wants to improve this feel free. MIT licensed.
Here is a discussion in the product support forum:
https://support.google.com/docs/thread/186496870/how-to-deac...
I expect that the emoji picker will become part of the OS eventually. Heck, Windows has one that you have to call up explicitly (with <Alt>-.), iOS has an emoji "keyboard" (input mode), etc, so we're headed in that direction.
var x:Pointer;
The :P changes to tongue out emoji
I'm not saying G.Docs or any editor should dictate how French is written, only being pragmatic if I have to choose between (no space): or unwanted emojis ending up in docs.
Well, that's is a bit like saying "it might be time to write your whenever you mean you're" just because Google Docs' autocorrect feature kept messing up the two when you wrote English.
How would it look like in an email to a client?
We absolutely do not have to choose between those two things. We need to stop implementing cute automatic bullshit that makes assumptions about the way people interact with UIs.
Typographically speaking English language text should have a wide space following a sentence-ending period [that isn't also a paragraph-ending period]. However, because wide spaces are not easy to type on normal keyboard layouts, the simplest thing to do is to type two spaces after sentence-ending periods and let the word processor change those to wide spaces when using a proportional font. When using a fixed-width font, however, it should always be two spaces after a sentence-ending period for two reasons: to disambiguate non-sentence-ending periods, and to make it easier to read. Coders write a lot of text in fixed-width fonts, so we should write two spaces after sentence-ending periods in, e.g., code comment blocks.
What do you mean? Every sentence in this comment has two spaces after the punctuation. That one and this one. Just because browsers and html condense it all into one space doesn't make it right.
How about getting rid somehow of the letter ù which is used for only one word (où)? Or replacing quatre-vingt by huitante (and same for 70 and 90) like in Switzerland? Or getting rid of œ and æ?
But none of it is realistic. We officially changed from oignon to ognon 32 years ago and people still don't know about it...
I don't mind "où". I especially do not want the circumflex accents removed (you didn't mention it, but there was an attempt in... the 90s?).
I wouldn't mind some spelling simplifications around all those eau/eaux/oeu/oeux type word endings.
> Or getting rid of œ and æ?
Yes.