Why isn't the external link symbol in Unicode? (2018)
dafoster.net
dafoster.net
"External link" symbol tells you, "that last bit of differently-formatted text is an active element in this application, leading to an outside resource". It's not something that makes sense in plain text, because any external link in a plain text message is both visible and obviously a link.
Elsewhere in the thread someone mentioned the play/pause/stop symbols from cassette recorders/VCRs. But those have been used culturally as symbols denoting starting, pausing and stopping for decades now, so they're an idea communication tool that makes sense in plain text, and thus pass the SMS sniff test.
(Note that I'm not sure if all symbols in Unicode actually pass the SMS sniff test. I suppose the best option for those wanting "external link" symbol to be included would be to resubmit it as the DOUBLE ARROW POINTING TOP RIGHT OUT OF SQUARE symbol, or something like that.)
But for a text mode interface that does have full blown anchor tags -- not very mainstream -- you have a point.
Modern SMS readers will linkify the URL, but will set the visible text equal to the URL, so again it wouldn't benefit from the symbol.
This explanation mirrors the rejection rationale: you don't need a symbol to be a Unicode character if the document is already rich text; the symbol can be an image instead. If you have anchor tags, you probably also have image tags or CSS.
No, that's not the rejection rationale. The rejection rationale isn't about having access to image tags or CSS. It's about hypertext (click link, go to another text) vs having text (can't click link, no other text). If that was the rationale, they wouldn't allow emojis in unicode, because you could insert them as images.
Emoji are useful (arguably I guess) in plain text scenarios because they exist to convey additional information about an author's emotion, and that author might use plain text. The external link symbol is not useful in plain text scenarios because it exists to convey additional information about an author's preceeding hypertext, and that author isn't using hypertext.
But the question is not if the icon is practical or silly, the question is if it makes sense as a Unicode character or should be handled at some other level like a GUI control or widget.
Should GUI icons, widgets and controls be covered by Unicode? That is a can of worms.
I understand the usefulness of the "external link" icon, but should it really be something you arbitrarily type into a text? Shouldn't it be part of the rendering of a link?
They already got all the Wingdings symbols https://en.wikipedia.org/wiki/Wingdings
The obvious hack is to make a custom character set called “Rejected by Unicode”.
You still see this kind of mischief when Outlook users type a smiley, and it is rendered as a "J" with the wingdings font. Readers of the mail where styling is stripped or who haven't the font installed just sees a confusing "J".
The reason: Softbank and Docomo both used proprietary emoji sets, widely used in Japan for text messages. Not including these in Unicode would have resulted in a choice: be compatible with the rest of the world, or keep using the emojis that became part of their culture. Not satisfactory. So basically, they had to put emoji in Unicode for the Japanese to use it, and excluding Japan is not really an option if you want a universal standard.
It did open an Pandora box though.
And a great idea to put them outside the 16-bit plane, so platforms would be forced to support characters beyond the BMP because of the public demand for emojis.
But I think going forward, emojis should be handled outside of Unicode. They are really small illustrations rather than characters or symbols.
The encoding of emoji was in order to unify Softbank/docomo/au under one set so that iPhones/Androids sold by the different carriers could send emoji among each other without relying on email translators (as had been the technical solution with feature phones until then)
* It has APL symbols (checking the Wiki article on APL syntax and the APL standard, it seems that APL programs can be represented entirely in Unicode - which makes sense, but is still a little surprising).
* There's a benzene code point: ⌬
* "ERASE TO THE LEFT", better known as backspace, is a thing: ⌫
...and a little more bizarro: https://en.wikipedia.org/wiki/Miscellaneous_Technical
To actually get the character included, you can't just reason that it might be used in a certain way, you will have to demonstrate it with actual and authentic use in the wild. The accepted proposal for the inclusion of the power symbol is a nice example of how this works.
If you can find a bunch of printed manuals or books or websites that actually use this icon in this manner you might be able to submit a successful proposal.
I mean, you might want to say "External links can be identified by their □ marker" or "The Windows key is marked □ and can be found to the left of the space bar"
That's simple. The justification for most characters, including the exemplary one you mentioned, is that they were already encoded in some other repertoire. The consortium's task is to collect and unify them.
If the "external link" symbol had occurred in some legacy character set, then it would have been included automatically.
But since it is a new proposal, it will be validated according to Unicodes policy for new symbols.
„Plz send me the <LINK SYMBOL>.
<RAISED HAND WITH PART BETWEEN MIDDLE AND RING FINGERS><EMOJI MODIFIER FITZPATRICK TYPE-4>“
https://www.fileformat.info/info/unicode/char/1f517/index.ht...
Can't paste it on HN, as the Unicode/Emoji filter seems to be removing it.
I can buy that the name in unicode should be something other than "external link". It could be something more generic like "external reference". Boom, now it can be accepted because it's not tied to <a>.
The problem with the rejection is that "that last bit of differently-formatted text is an active element in this application, leading to an outside resource" is that it's kind of a strawman: why be so specific just to reject it because you've decided to be so specific?
Though all of that justification is kind of unnecessary. Really, I just find it ridiculous that the unicode group is going to suddenly become prude about a ubiquitous symbol yet have 20 different check mark symbols. Sure, if everyone made the "but there are 20 different check mark symbols so why can't I have <pet symbol>?" argument, unicode would be even more ridiculous. And that's an argument to reject <pet symbol>, not one of the most ubiquitous symbols on the internet.
"Unicode has a bunch of stupid symbols, so it should at least add the most useful symbols that we use on the internet" really is a good argument.
One of them, BROKEN CIRCLE WITH NORTHWEST ARROW, could actually be used for the external link symbol if it were mirrored!
Or, to put it another way: if you think it needs to point to the right, is that because the text you read flows that way?
This then makes me wonder if the large portion of the world's population that read right-to-left find the currently widely-used external link symbol (as discussed in the article) a bit jarring.
The symbol appears mirrored in RTL languages on Wikipedia, e.g. this random Arabic article: https://ar.wikipedia.org/wiki/%D8%A7%D9%84%D9%83%D8%B3%D9%86...
Come to think of it, are tabs aligned toward the right side when the browser/OS is set to an RTL language?
I accept that heuristic, and find myself agreeing with it for the purposes of the external link character. Let me try to convince you:
This is a real conversation:
A> Are you familiar with Homebrew? You can get it from brew.sh.
B> Ok. Where is brew.sh?
If we were not limited by plain text, we could use colour and underlining to distinguish that as a domain name. A lot of websites find this poor for accessibility reasons so they use the link character.
Another way to think about this is to imagine the link-character is pronounced "https://" as in:
* Are you familiar with Homebrew? You can get it from https://brew.sh.
* Are you familiar with Homebrew? You can get it from {link character} brew.sh.
Isn't that plain text? I'm not sure how to pronounce it, but there's a lot of emoji I don't know how to pronounce. Do you think this is important?
So if you get that far, how much further is this really?
* Are you familiar with Homebrew? You can get it from brew.sh {link character}.
Phones already do it for .com, .gov, etc, but our cutesy .sh, .dev, .rocks, .xyz, etc will take a bit to catch up.
Negative. My renderer is a piece of paper.
The request is for a code point to represent the visible external link character, not for an invisible control-code to decorate some structured data (which cannot "appear" on paper).
<sarcasm>Oh yeah, that's right</sarcasm>
Which also lack a decent Unicode symbol despite being a word used in plain text. ️
That's a contingent answer. If you, say, set out to try to convince the world to use it that way, and created that demand, then by golly, that demand would exist at that point and the answer could change. But hypothesizing some possible text message that could use it isn't strong enough to argue for inclusion in the standard because every possible proposal passes that test.
(The hsts preload list helps a bit with this, but only for domains registered with it.)
I would propose updating this SMS sniff test to a Wikipedia sniff test.
I explicitly don't want it to be "Wikipedia sniff test", because that's equivalent to a "Webpage sniff test", which is equivalent to "anything goes", because icon fonts exist now and are used.
That said, other comments made me realize that Unicode is better described as "whatever was there in all the codepages around the world at the time of bootstrapping the Unicode standard" plus what's typically used in a written language, plus SMS sniff test.
Here's some random characters from the "Math Symbols" page of OSX's character viewer.
⨐⨔⊹⊰⨶⨹⩫⫸⧯⧼
Did you know that Unicode has an entire set of dominos in it?
🀱🀲🀳🀴🀵🀶🀷🀸🀹🀺🀻🀼🀽🀾🀿🁀🁁🁂🁃🁄🁅🁆🁇🁈🁉🁊🁋🁌🁍🁎🁏🁐🁑🁒🁓🁔🁕🁖🁗🁘🁙🁚🁛🁜🁝🁞🁟🁠🁡
Twice: once horizontally, once vertically.
🁣🁤🁥🁦🁧🁨🁩🁪🁫🁬🁭🁮🁰🁱🁲🁳🁴🁵🁶🁷🁸🁹🁺🁻🁼🁽🁾🁿🂀🂁🂂🂃🂄🂅🂆🂇🂈🂉🂊🂋🂌🂍🂎🂏🂐🂑🂒🂓
I know I sure use Ogham runes in my daily text communications on a constant basis.
ᚃᚄᚓᚈᚚᚙ᚛ᚘ
And Egyptian hieroglyphs.
𓁑𓀳𓂩𓃋𓇀𓈣𓇻𓇼𓇽𓀫
Oh hey and here's a "next page" and "previous page" symbol that "external link" could comfortably sit next to.
⎘⎗
And some miscellaneous arrows.
⤥↪︎↴↶↺⥀⟳⤩⤭⤱⥹⥰⥼⥺⥽⥬⇆⇼⇟⥣↬⇴
If your browser feels like displaying Wingdings codepoints, something like this: http://egypt.urnash.com
[1] https://emojipedia.org/white-right-pointing-index/
[2] https://i.ebayimg.com/images/g/t60AAOSwAPVZDOK3/s-l300.jpg
That's certainly a useful sniff test, and I think you've made a good case with it. I wonder, though, whether one should also consider a "documentation sniff test", asking: would it make sense to type that symbol into some documentation one is writing about how to use, for example, some software? I think the answer to that question is a clear yes: it might make sense to put this symbol onto such a page, and although being on inert paper (it wouldn't also be an active element) it would certainly be a useful symbol in that hypothetical text.
I don't really have a horse in this race, but... isn't the existing widespread practice of representing the sign using images attributable to the fact there is not a character available? What else do they want people to do? The point of adding the character is so that in the future people can use text instead of an image.
The rejection is based on the first standard for adding new characters.
The arguments against it seem to be based on taking the second standard as precedent. But I think as far as the Unicode committees are concerned, arguments based on the precedent of whatever random crap was grandfathered in from preexisting character sets do not apply to characters that are not from preexisting character sets. I think any argument that relies on “you already allow all this other random crap” must also argue that this symbol exists in some other character set which ought to be merged into Unicode, or that the standard for new characters should be more precedent-based/different, but I don’t see any such arguments other than some implied “common sense guess as to how one expects Unicode to work”
One has to make a better argument for the external link symbol getting a codepoint assignment. TFA, for example, makes an argument based on emojis -- certainly that's strong enough to blunt the UTC's rejection rationale, but perhaps not enough to win approval outright.
Addressing all the arguments used in the rejection is important, of course. The fact that currently images are used is hardly dispositive: that's business as usual for missing Unicode assignments!!
But there are probably stronger arguments for rejection than adoption than the UTC made that it could make the next time this comes up.
The best argument for rejection that I can think of has to do with layering. An external link character isn't very useful without the actual external link, but the link belongs a layer up: not in the text, but in the markup. Well, if the link belongs a layer up, why not also the symbol? Alternatively, more markup can move into Unicode. There has been and will continue to be some pressure to move more semantics from markup to text, but it's probably best to resist that pressure.
On the other hand, a solid argument for adoption may involve text rendering of HTML. Think of lynx/elinks and other such browsers, which can't use images. An external link character could prove useful in distinguishing the rendering of linked text from, say, underlined non-linked text.
I'm surprised these arguments didn't come up. Or maybe they did -- I've not gone down the rabbit hole on this one, and probably I won't.
Because it's not in Unicode it can't be copied. If it was in unicode symbol I would expect it.
On a lot of pages links are hidden as plain text and only show up if someone hovers over them. (Great? Confusing? Bad UX? Sure, but still a choice.)
At the same time someone else might just use underlining, but no different color. And someone might just want to use a symbol.
I've seen a right-pointing magnifying glass printed in a book, and a left-pointing one would be equivalent for Arabic or Hebrew. I'd expect to see it in a school science book, for example.
Annoy everyone enough, and we can redefine the nunicode codepoints as rationals and hand the denominator to another group.
And since when was a smiling poop an ‘element of plain text’
Neither applies to the "external link" symbol.
Since the Japanese phone vendors shoved it into the character encoding and forced everyone to deal with it if they wanted to be compatible.
Anyway I think we can avoid target="_blank" most of the time, you can have a https://developer.mozilla.org/en-US/docs/Web/API/WindowEvent... event listener if something needs to be saved before the page location changes
I suspect HN will eat the character: ⬀⃞ although they sometimes pass through if there's enough other text in the comment.
There's a useful selection of arrows here: http://xahlee.info/comp/unicode_arrows.html
Obviously enough text-mode browsers would be likely to benefit.
But it seems that to be a codepoint, it needs to be useful in plain text scenarios. Text mode browsers which render links as discussed are not plain text.
As the author pointed out, emojis seem to clearly violate this excuse. What possible justification do they give for emojis? There is no way [I originally inserted the poop emoji here, but it was stripped out. That kind of reinforces my point.] is an element of plain text.
So there's a situation of dual standards - all the weird characters that were included in any pre-Unicode text encodings for whatever arbitrary reasons are in Unicode and are always going to be there; but all the new weird characters need appropriate justification for inclusion and are likely to be denied.
Nobody would type an external link symbol. It would be added by the UI presentation layer. It is not part of plain text, because it is not typed as part of textual content. But the poop emoji is.
Nor is Pile of Poo, and that's in Unicode.
If external Link was added to Unicode, I expect it would be more used than 1000s of characters that are in it.
I really really don't understand this point. We have evidence that people use pile of poo in plaintext today in instant messaging all the time. Isn't this enough evidence that pile of poo belongs to plaintext? I'm not trying to be facetious; it's really puzzling to me. Since every single comment in this thread is about pile of poo, whereas it seems to me the worst possible example since it's such a widely used emoji.
In that way, it would be no different from 99% of unicode characters then.
a::after {
content: "↗";
}I happen to disagree with the unicode decision because there should be a section for basic commands so they can be rendered in diff fonts. But whatever — as I say don’t wait for standards if they don’t exist, make your own!
https://designnotes.blog.gov.uk/2016/11/28/removing-the-exte...
It was summarily rejected with the nonsense statement (in full):
> Thank you for your submission. This was discussed during last week's UTC meeting. I was directed to let you know UTC feels that, as submitted, the proposal does not sufficiently demonstrate a plain text need for such a symbol. The context for usage is mark-up with links by default.
The author seems to be using a very limited definition of "plain text". Emoji are clearly used in the same way as latin-character text in communication, so they are effectively plain text.
I think a more effective argument may be that the external-link symbol should be allowed for the same reasons that the power on/off/toggle and eject symbols were allowed.
Unicode aims to encode every character used for human communication in every culture, and Japan uses U+1F4A9.
They clap become clap petty clap tyrants.
https://www.reddit.com/r/justlegbeardthings/comments/6rl6mu/...
Which is something no one asked for, but Unicode gave us anyway and now everyone is annoyed by.
j/k there isnt one.
https://en.wikipedia.org/wiki/Unicode_subscripts_and_supersc...
No doubt those emoji's are more important and historically common than superscript uppercase C F Q S X Y and Z.
now you have applications where you can't find a string because it was written in Unicode bold letters instead of the letters' normal ASCII counterparts. And then you have applications that are confused about those bold letters because they are not actually surrounded by bold markup.
The worst part is that you can't criticize unicode without attracting a crowd of bullies who handwave about human languages being complicated (no, that does not justify poor design) or say it has to be this way because ugh shift-jis (or whatever nasty old encoding you can come up with) is not nice.
Well designed technology makes complex things simple. People defending poorly designed technology blame the problem for being complex.
𝔽𝕦𝕔𝕜 𝕦𝕟𝕚𝕔𝕠𝕕𝕖!
(Why don't you try search for "fuck" in Firefox?)
There doesn't exist any parser that does it 100% correct. And parsing it is becoming so complex that it's causing bugs and vulnerabilities (it's not a coincidence that so many remote exploits use some kind of unicode to trigger it).
The same link in different contexts can be as either external or local. So that is a business of UA/renderer to mark it properly.
CSS is quite adequate for that (https://davidwalsh.name/external-links-css) and the image (or whatever author/UA decided to use) can go inline in CSS itself.
I think it's a missed opportunity for unicode.
a::after {
content: url(my_external_link_symbol.gif);
}It would be nice if we could all use a standard icon or some other constant UI element--like a single underline for internal references and double underline for external? I imagine unique colors would be too difficult to standardize.
What do people do with the information that a link is external? Do people think 'I'll follow this link - oh no wait a minute it's external I won't'?
Techies will know to just open it in a new tab, but most will have to remember to browse back to the site manually. It can be a disruptive process.
Some do, sometimes. Or perhaps they'll think 'I won't follow this link - oh no wait a minute, it's external I will'.
Intranet: bespoke documentation to internal processes verses generic documents used for reference.
Government: external references can be hijacked (asking for personal information) or may not represent the government but still have relevant info.
PDF documents: jumping around a PDF document (from a table of contents) is different than going to an external website. Especially if you don't have Internet access at that moment.
I think all of this is significantly more important since browser have been hiding more URLs.
@media print {
a::after{
content: " (" attr(href) ") ";
}
}I wonder why there are actually so many people ITT who argue this new symbol would have no use and no meaning in print...
> The 'symbol fallacy’
> The 'symbol fallacy’ is to confuse the fact that "symbols have semantic content" with "in text, it is customary to use the symbol directly for communication". These are two different concepts. An example is traffic signs and the communication of traffic engineers about traffic signs. In their (hand-)written communication the engineers are much more likely to use the words "stop sign" when referring to a stop sign, than to draw the image. Mathematicians are more likely to draw an integral sign and its limits and integrands than to write an equation in words.
So where "stop sign" is in Unicode, it's a bit nuanced as per the manner.