So Long Surrogates: How We Moved to UTF-8 in Haskell
channable.com
channable.com
Because of modifier characters, control characters like for bidi, stuff like soft-hyphens and ligatures, locale-dependent semantics (upper/lowercase, collation etc.), the general discordance between glyphs and characters, and so on and so forth, Unicode is so complex, and in general always requires careful processing of code point (or code unit) sequences, that honestly the surrogate encoding doesn’t make that much of a difference. It’s just an additional wrinkle in a sea of wrinkles.
Unicode uses numeric values from 0 to 1112063. You can invent all kinds of methods to encode numbers from 0 to 1112063 (variable length, fixed length, decimal, hexadecimal, anything else). But most ways I can think of to encode these numbers, including variable length ones that would use 8 bit or 16 bit primitives, don't require me to actually reserve some of those to-be-encoded numbers themselves for a special meaning. Yet for UTF-16 they managed to do it. Imagine that all other encodings out there would also want to reserve some Unicode values for their own purpose!
Noncharacters can be represented in any Unicode encoding. Surrogates code points cannot, but can be found in unvalidated UTF-16 (which is most UTF-16).
Dealing with Unicode text semantically requires that you be aware of a great many factors such as those you name, but you don’t need to be aware of those for just storing and transferring Unicode text. But with surrogates, UTF-16 managed to break it for everyone: any part of the system that uses unvalidated UTF-16 can introduce errors that you must care about. Hence surrogates are a special kind of atrocity.
I prefer the name WTF-16 [0]
But beyond that, the difference lies in Unicode having been compromised for the sake of one particular encoding, and now there are all kinds of places that have to deal with potentially ill-formed UTF-16, more that potentially ill-formed UTF-8, I would say. Look at what Python 3 did: its string type is not a Unicode string, but rather a sequence of Unicode code points. That is, it can encode surrogates. Blech.
> Unicode having been compromised for the sake of one particular encoding
They really didn’t have much choice though, because UCS-2 existed and was in major use at the time Unicode was extended beyond 16 bits (which wasn’t part of the original Unicode design), and too many code points were already used up to turn it into an UTF-8-like encoding. Not supporting the additional planes with the 16-bit encoding would have been a nonstarter. Switching to 32-bit in-memory representation (UCS-4) even less so, given the resource constraints of the time. The reason platforms like Windows and Java didn’t add strict input validation was due to backwards compatibility, because they previously were UCS-2 where arbitrary sequences are valid.
One can lament the chronology that lead to those technical choices, but they each were reasonable and appropriate choices under the constraints present at each time.
More importantly, surrogate characters rarely create any actual problems in practice. I haven’t encountered any in the past 20+ years (other than missing support for characters beyond U+FFFF). Other parts of Unicode cause significantly more complexity.
You may not have encountered such bugs, but I’m very familiar with surrogate-related bugs, because I use a Compose key extensively. I haven’t been using Windows for the last year, but from time to time I would definitely encounter bugs that are certainly due to surrogates. On the web, I found bugs a few times, all but once in Rust WebAssembly things, such as https://github.com/Pauan/rust-dominator/issues/10. And even now I’m back on Linux, I know of one almost certainly surrogate-related bug: I can’t type astral plane characters in Zoom at all; pretty sure I had this problem back on Windows, too. Copy and paste, sure, but type, no, they become REPLACEMENT CHARACTER.
The history is unfortunate but I strongly refute that they had not much choice. UCS-2 should have been abandoned as a failed experiment. Certainly there had been significant investment into it in the last few years, but with the benefit of hindsight, switching to UTF-8 (which was invented before they decided on surrogates) would have made everyone’s life much easier, especially given its ASCII-compatibility, which would have allowed the Windows API to retire the misbegotten A/W split and return to sanity a few decades early.
Ah, BOMs. Haven’t seen one in years. Good riddance.
UTF-16 is great for lots of East Asian languages, which billions of people use. In UTF-8, most of those languages require 3 bytes to encode a 32-bit codepoint, in UTF-16 they only ever need 2. This ends up being a huge savings.
The main benefit of UTF-8 if you're say, Chinese, is interop. Everything else is worse.
You might think "but BOMs are super evil." Checking a BOM is extremely, extremely easy. Furthermore, you don't get to bail out of checking anything just by using UTF-8, you have to check to ensure you have _valid_ UTF-8. That's right, you gotta scan the whole bytestream anyway, so you may as well just check the 2-byte BOM at the beginning too.
You might also think "what about ASCII compatibility?" ASCII compatibility is an anti-feature. You should never be indexing into UTF strings (you always have to iterate, or save the results of an iteration), upper/lowercasing isn't addition/subtraction, etc. etc. You also can't just forget about encodings as a result--you can store ASCII in something expecting UTF-8, but you definitely can't store UTF-8 in something expecting ASCII. So if you're sniffing/decoding/tagging a format anyway, you may as well be agnostic.
You might also think "OK OK, you could be right, but what about HTML, which is mostly ASCII and would nearly double in size if it went from UTF-8 to UTF-16." Practically all HTML is gzipped, so the difference is pretty small, plus the majority of text isn't HTML (almost anything stored in a database, almost anything in a file on your computer, etc.)
Different encodings are good at different things. There's no one superior encoding for all uses. What we need is text encoding agnosticism.
---
In fairness, I will say I've heard that UTF-8 is pretty popular in countries with exactly the kind of languages I'm talking about, so the issue is mostly moot at this point. I just think UTF-16 gets a really bad rap, and I think we shouldn't just gloss over UTF-8 having won because it's good for European languages.
> ASCII compatibility is an anti-feature.
ASCII compatibility is extremely useful if you're working with, for instance, filenames or programming languages. You can lex UTF-8 and handle separators like `/` or quotes like `"` and `'`, because those bytes can never occur otherwise.
Reading about it, filename encoding looks like a mess, GLib apparently assumes UTF-8 unless you set it with an environment variable, which I think is wild. Qt uses the locale, which feels reasonable (ignoring that locales are themselves bananas). But assuming UTF-8 and relying on ASCII compatibility and no NULs feels risky.
This stuff just has to be OOB configurable. ASCII-compatibility lets us think we can do stuff like this, but you're always just assuming and hoping it doesn't blow up.
UTF-8 "can contain NULLs as a matter of course" in exactly the same way that ASCII can.
Any system that doesn't allow null bytes in ASCII can just as easily not allow them in UTF-8.
A system that does allow them can just as easily allow them in both.
There is no difference between the two here.
-
Edit: Oh, wait, maybe you misunderstand UTF-8. The 2+ byte encodings never use bytes below 0x80. They never cause a null byte to happen. The only way you get a null byte is if it's an ASCII-compatible null character, code point 00000.
You are absolutely right! If your use case demands storage of CJK strings, UTF-8 is probably not your best bet.
But our data (mostly stringly typed, think product descriptions or links) is >99% Basic Latin characters and we are getting to a point where memory is actually becoming an issue. So it's neat that Haskell allows us to "upgrade" to UTF-8. With the data being in memory I think compression would not be very helpful either.
Edit: I also kind of agree that UTF-16 gets to much hate :D
> ASCII compatibility is an anti-feature.
Actually, it's not. One of the main reasons it's important is...
> What we need is text encoding agnosticism
... that bit that you want. If you're supporting multiple encodings, you need a reliable way of signalling that encoding. And file metadata tends to be insufficiently reliable, so it ends up being in-band for text (e.g., HTML files). If all the encodings you need to support are ASCII-compatible, and the directive to set the charset is itself pure ASCII, that means it is perfectly safe to parse the file as ASCII to figure out how to parse the file. If the charset isn't ASCII-compatible... well, that historically has led to security vulnerabilities. Ergo, a non-ASCII-compatible charset has heightened security risks.
That said, my own experience when dealing with the mess that is charsets is that supporting multiple charsets is itself a painful process that isn't worth it.
100% this. In javascript you can accidentally cut characters in half. This lets you accidentally construct strings which have invalid encodings - and that can lead to all sorts of weird problems. In comparison, using rust's safe API, its impossible to construct a string which contains invalid unicode.
For example, if I put an emoji like this U+1F970 in javascript, I can then split it in half and put one half into a JSON string - like this: "\"\\ud83e\"". Javascript will treat this as "one character". Rust will treat this as an encoding error, and will refuse to parse it back into a string.
Another implication is that javascript's .length field looks useful, but its almost never what you actually want. How many unicode characters in this string: "\ud83e\udd70"? Trick question - even though javascript says the string's length is 2, there's only 1 unicode character in that string. How many bytes does this string (with JS length 1) "\u270C" take up over the wire? Javascript says the length is 1, but it takes 2 bytes to send in any encoding (utf8 or utf16).
What a mess.
In which language(s)?
Its quite hard to write invalid UTF-8 in Rust, because the UTF8 encoding is checked when the string is created. I think this is the right choice.
Javascript, C# and java should work like rust, since a string should always ... y'know, contain a correctly encoded string. The current situation in these languages feels like a weird middleground where String classes are actually arrays of 16 bit integers wearing a fur coat.
Oh 100%. You gotta validate every time. Again I'm not a C# or Java expert; if their stdlibs let you build invalid UTF-16 strings then that's unfortunate, but again a problem with an implementation, not an encoding.
In other words, if Rust used UTF-16 internally it would behave exactly the same as you're describing.
This is a good point. I suppose the real problem here isn't UTF-16. Its that all these languages and systems assumed unicode would never grow past 65536 characters. They weren't designed to use UTF16. They were designed to use UCS2. It was only after unicode grew past UCS2's limit that they added awful hacks to their systems in order to (badly) support UTF16. The problems I'm complaining about aren't UTF16's fault. They're problems that arise as a result of the hacks these languages added in order to support surrogate pairs.
You could build a string type in rust today which used UTF16 strings internally. It would work fine, and be just as safe as the UTF8 strings in the standard library. I just really don't see the point. The verdict is in. UTF8 has won the war to be the world's interoperable format for unicode text. UTF8 is usually smaller than UTF16 (except apparently in some parts of Asia). And it doesn't need byte-order-marks or any of that guff. Its a relief that this point is basically settled.
I hope that a future generation of programmers can grow up without needing to understand UCS2, UTF-16, byte-order-marks, surrogate pairs or any of this junk.
It's really disappointing, honestly. Go has rune, which is 32-bit, and adding a new rune-based API seems doable. But hey, what do I know.
At least I haven't seen languages claim UTF-8 support by just giving you raw byte arrays.
Before the advent of emoji (which I'm personally grateful for), the vast majority of developers aren't even aware that UCS-2 is not UTF-16 and outside the BMP Javascript's String functions are totally broken and that you actually have to check surrogates yourself to properly decode strings.
And of course, it's not only Javascript... all systems/languages that mistakenly adopted UCS-2 ended up in similar places.
IDK, again probably you can goof using UTF-8, but it does seem harder, like writing a bad byte or whatever. It would be better if string libraries actually worked though, maybe in 2023.
Totally agree these discussions circle around to "you need some kind of OOB way to know what encoding you're dealing with", which puts us back to everything being encoding agnostic, with some way to configure it. Maybe that's locales, maybe that's a type 32-byte integer at the beginning, dunno.
People have been dealing with 8-bit ASCII extensions for longer than Unicode has been around. It isn't overwhelmingly difficult to have 8-bit clean text handling without any awareness of the encoding.
An example of the latter effect in action is the system charset of Linux and most other Unixen... in practice, virtually every system is using UTF-8 for things like pathnames and the like. However, the mechanism to indicate that this is the case is frequently enough unreliable that it's often less bug-prone to assume the system is UTF-8 anyways.
Of course, there's a few more advantages to always using UTF-8:
* You eliminate the code path of "I'm sorry, I can't save the text that you wrote," since every character is usable in UTF-8, which is a feature very rare outside of UTF-* charsets. And, honestly, that's a wonderful code path to eliminate.
* One of the features of UTF-8 is that virtually any text that isn't UTF-8 will fail to parse as UTF-8. This allows you to build systems that produce errors on bad input, which means that maybe, just maybe, the people who produced the broken input will discover their problem and fix it.
I mean, *nix filenames are a pretty weird case where you have these requirements of 0x2F always being a path separator and 0x00 always meaning "end of path". In truth, we can't guarantee anything about filename encoding other than those two things. So, for example, you might be dealing with UTF-8 or EUC-JP. In practice (as a sibling comment pointed out) all you usually need to do is scan for 0x2F and stop at 0x00, so this works out OK, but if you ever need to do anything else, welcome to guessing the encoding based on locales, environment variables, or other (OOB) configurations.
> However, the mechanism to indicate that this is the case is frequently enough unreliable that it's often less bug-prone to assume the system is UTF-8 anyways.
My argument, which I guess looking at it I've not articulated at all in this thread, is that UTF-8's ASCII compatibility and *nix filename compatibility have allowed us to be slack in fixing this. It oughta be possible to configure an encoding when creating a filesystem, just like creating a database. There's nothing inherently buggy about configuring an encoding, we just haven't gotten our systems programming act together, and UTF-8's amenability to the existing situation is probably at fault here.
> since every character is usable in UTF-8
I guess you mean like, vs. EUC-JP or something? Sure but UTF-16 is fine here too.
> One of the features of UTF-8 is that virtually any text that isn't UTF-8 will fail to parse as UTF-8.
IDK the "brittleness" of an encoding feels hard to quantify. Besides, I think if you're doing the (always required) validation step, this goes for any input. Look at things like weirdo MySQL/Oracle/MSSQL database encodings for example.
---
I should say I see where you're coming from. But I think we'd be in just as good a place if we specified a filename/text encoding at filesystem creation (OOB configuration) and quit relying on magic 0x2F/0x00 bytes. I don't know why this makes people tear their hair out as some kind of impossible problem to solve.
Meanwhile in lots of other countries, UTF-8 lets you use 1 byte to store each character instead of 2; which is also a huge savings.
The question is, is it better for the world to pick one encoding and have all programs use it, or pick several encodings and switch between them depending on the context / country? Picking one encoding means we can reuse our code more easily. Picking multiple encodings means we get slightly smaller file sizes.
For my money, the best approach is to use UTF-8 everywhere. This lets us reuse code. Then if you're worried about file size on disk or over the wire, compress your text content with LZ4 or snappy or something. That'll halve the size of text for everyone and LZ4 is so fast its essentially free from a computational standpoint.
As this blog post demonstrates, quite a lot of code needs to deeply understand text encoding to work (be that UTF-8 or UTF-16 or whatever). Its a big savings in programming hours if we only have to write this code once.
Well, for HTML it turns out to be a wash.
While Han characters are typically three bytes in UTF-8, markup is ASCII and so only one byte — whereas in UTF-17, both the character data and the markup are two bytes. When you add them together the average cost per character is more or less two bytes, whether your Chinese-content HTML is encoded as UTF-8 or UTF-17.
Not really. That's just an excuse to be contrarian.
Any significant length of text will be compressed and/or have lots of latin characters.
> you definitely can't store UTF-8 in something expecting ASCII
That's not true.
A reasonable counterargument is Asian books on Project Gutenberg [0]. I guess you can say they're compressed on the wire, but they're not compressed in my RAM or cache.
>> you definitely can't store UTF-8 in something expecting ASCII
> That's not true.
Well alright, you can jam bytes pretty much into anything. And you get luckier with UTF-8 than you do with probably any other encoding. But lucky isn't robust, is all I'm saying.
You don't really need to have very much of a book decompressed at once, and in your average program handling that book there's going to be much more RAM dedicated to layout and display than the actual character bytes.
> Well alright, you can jam bytes pretty much into anything. And you get luckier with UTF-8 than you do with probably any other encoding. But lucky isn't robust, is all I'm saying.
It won't work in all systems but it'll work in most of them. And you can check beforehand.
I don't know why we treat text encodings any differently than any other binary format. We have OOB format info for all of those. People are just downright religious about UTF-8, and it's my goal in these threads to push (hopefully gently) against that. UTF-8 isn't always the best, you always have to validate, you shouldn't be indexing into strings, we should have OOB encoding info (whether in a filesystem config or whatever). This is the only possible path to robust handling of encodings: "always assume UTF-8 and hope non-UTF-8 systems get updated or fall into disuse" is a recipe for dealing with broken systems for decades (what we're doing now).
Different image formats have different use cases and substantial advantages over each other. And it's often impossible to convert without losing quality.
But everything gets to be HTML. Or the compatible variant of XHTML. Nothing else has any real support, and when it did, like flash web pages, that was not good.
For text unicode wins out easily, with UTF-8 and UTF-16 far above the other encodings, and the differences are minor enough that we should just use the better one.
> we should have OOB encoding info (whether in a filesystem config or whatever). This is the only possible path to robust handling of encoding
I don't know about that. Old files, which represent almost all the non-unicode non-ascii text we have, are unlikely to ever be properly tagged. For new files, are those tags really needed? If you have a firewall where you mandate tagging behind it, you're already mandating a format for attaching tags, so why not mandate utf-8 instead?
> You might think "but BOMs are super evil." Checking a BOM is extremely, extremely easy. Furthermore, you don't get to bail out of checking anything just by using UTF-8, you have to check to ensure you have _valid_ UTF-8. That's right, you gotta scan the whole bytestream anyway, so you may as well just check the 2-byte BOM at the beginning too.
BOM breaks substring. Everything you take a substring, you need to prepend a BOM if you want to serialize it to somewhere else. If you want to concat two strings with BOM, you need to remove one of them. All of these are unnecessary pains.
> You might also think "OK OK, you could be right, but what about HTML, which is mostly ASCII and would nearly double in size if it went from UTF-8 to UTF-16." Practically all HTML is gzipped, so the difference is pretty small, plus the majority of text isn't HTML (almost anything stored in a database, almost anything in a file on your computer, etc.)
This just contradicts your own reasoning that UTF-16 is better for East Asians due to size savings.
Re: HTML, people usually bring it up because it's a lot of ASCII that'll all double up under UTF-16. Mostly my argument is that most text isn't HTML, so size savings on other text files still matter.
I would probably do BOM, LUT, BYTES. Where LUT is a lookup table of indices for every 64 chars so that you could do o(1) random access for the vast majority of cases.
> UTF-16 is great for lots of East Asian languages, which billions of people use. In UTF-8, most of those languages require 3 bytes to encode a 32-bit codepoint, in UTF-16 they only ever need 2. This ends up being a huge savings.
That doesn't make UTF-16 not a terrible format, though. East Asian countries use better formats instead. For example, in China they use GB18030.
When your case for "UTF-16 is not an irredeemable format" is "it's better than UTF-8 at a task where nobody would use either format anyway", you're not making a strong case.
Haha this is pretty fair; I know I'm being pedantic. But I wouldn't say any GB* encoding is better than UTF-16. Another commenter pointed out that you really want the ability to--say--paste arbitrary text into an editor and for that editor to be using an encoding that can handle it.
Or even for the web, if we actually took the Accept-Language header [0] seriously, that could also be a big savings.
"Only ever" is an exaggeration. The vast majority of standard Chinese (aka. Putonghua) characters are in the BMP, but occasionally there are characters outside of it. In particular a lot of Hong Kong characters are outside of the BMP, and before emoji forced software to fix things, systems claiming to support UTF-16 (but were instead UCS-2) kept failing for say 1% of the characters.
I've had much better luck with UTF-8 systems than purported UTF-16 ones. The problem with UTF-16 is not necessarily in its inherent technical issues, but rather, it's how often systems with UCS-2 support just claim they support UTF-16, and leave users/developers with a crappy experience when working with characters outside the BMP. There's also the encoding agnosticism thing -- with UTF-16 you run into issues about byte order marks (with or without), endian problems, and generally incompatibility with ASCII.
As I said, the situation probably improved quite a bit after emoji became popular. But if I had the choice, I'd choose UTF-8 over UTF-16 any time.
Also, although it's generally true that you "can't store UTF-8 in something expecting ASCII", in practice a lot of systems are somewhat more graceful than that. For example, if I grep a UTF-8 text file, grep doesn't need to know the text encoding if all you're trying to find is an ASCII string. Similarly it's possible to edit a UTF-8 text file as an ASCII file if the editor preserves the "binary" bits.
In conclusion, "different encodings are good at different things" is only true when considered in isolation. The historical baggage of UCS-2, and separately, of ASCII-compatibility tilts the balance strongly in favor of UTF-8 IMHO. The alleged 33% saving in Chinese text doesn't really matter that much in the grand scheme of things -- which is perhaps why people use UTF-8 in spite of the "advantages" you mentioned.
But I think that 99% of the things you have to do for UTF-8 (validate, encode/decode, no indexing) you have to do for UTF-16, and if we're arguing for systems that are UTF-aware, UTF-16 does great. If we're arguing for trying to jam encoded bytes into systems, well it's a crapshoot. UTF-8 being better at winning that crapshoot doesn't feel like a recipe for robustness to me, and certainly doesn't strike me as a perfect solution. OOB configuration does, though.
> For example, if I grep a UTF-8 text file, grep doesn't need to know the text encoding if all you're trying to find is an ASCII string.
Well, this is a weird example. grep uses locales, so yeah you're probably fine if you're grepping for ASCII using an ASCII-compatible locale. But if you're not, then what you feed to grep has to be encoded in your locale's encoding (I think). So I think this is actually an example in favor of "this should be OOB configurable", and an example of "ASCII-compatibility lulls into a false sense of security".
Checking the BOM is not enough, you also have to handle it. Non-LE BOMs (like surrogates, at least before the emojipocalypse) are rare enough that many "UTF-16" based tools simply only support UTF-16LE. They are also stupid because you need to know that you are dealing with UTF-16LE or UTF-16BE in the first place - either through heuristics or because it is specidified and then you can also guess/specify the byte order.
And like surrogates, the BOM is another wasted Unicode character that is only needed for the UTF-16 mess.
> Furthermore, you don't get to bail out of checking anything just by using UTF-8, you have to check to ensure you have _valid_ UTF-8. That's right, you gotta scan the whole bytestream anyway, so you may as well just check the 2-byte BOM at the beginning too.
You don't need validation for most UTF-8 tasks, GIGO is often more reasonable for things entered by humans.
> you definitely can't store UTF-8 in something expecting ASCII.
Unix tools would disagree.
$ echo "Hello Wörld!" | tr '!' '?' Hello Wörld?
Many tools only care about finding substrings. UTF-8 gurantees if the byte sequence making up the encoding of a Unicode character (or sequence) appears in valid UTF-8 encoded text then it also decodes to that Unicode sequence.
> Practically all HTML is gzipped
Not when it is being parsed it isn't.
Gymnastics like this always occur when you don't know the encoding. If you didn't know you were getting UTF-8, welcome to heuristics or whatever.
> And like surrogates, the BOM is another wasted Unicode character that is only needed for the UTF-16 mess.
This isn't a big deal at all.
> You don't need validation for most UTF-8 tasks, GIGO is often more reasonable for things entered by humans.
You can't build robust systems this way. You always have to validate, and in order to do that, you gotta know what encoding you're getting.
> Unix tools would disagree.
Most of those use locales, or they don't expect ASCII they expect bytes. Again you can get lucky with this, but you can't build a robust system with luck.
> Many tools only care about finding substrings.
You have to scan through the string or save offsets (after an initial scan) either way, UTF-8 or 16.
> Not when it is being parsed it isn't.
Fair point!
https://docs.microsoft.com/en-us/windows/apps/design/globali...
> As of Windows Version 1903 (May 2019 Update), you can [...] use UTF-8 as the process code page.
> Until recently, Windows has emphasized "Unicode" -W variants over -A APIs. However, recent releases have used the ANSI code page and -A APIs as a means to introduce UTF-8 support to apps.
Of course if you use this, your app will only run on very recent Windows versions. But that's how it goes with OS features. We'll start reaping the benefits 10-20 years from now.
I have to say, I never thought that the benefit of Haskell having a horrible native string type would be "you can just upgrade strings like any other dependency," which is really kinda slick. You think about how much pain there was for Py2 -> Py3 where one of the big sticking factors was all of the distinctions around strings and encoding and byte arrays... this is comparatively quite nice. Makes me wonder how much of a programming language can be hotswappable.
This is very different from going from python2, which conflated bytes and ascii strings, to python3, which intentionally changed the api to propely distinguish sequences of bytes and strings.
For a research language that can make a lot of sense, not so much for a language to be used in industry.
The downside is that different libraries will have different string representations, so you can end up being forced to do a lot of conversion if you're using different libraries that have made different choices from each other, or from your own code.
There are at least 5 commonly used string types - String (linked list of Char), ByteString lazy & strict, and Text lazy and strict. The latter two have a good rationale for being different - byte strings are not necessarily text - but, for various reasons, they're often used to represent text anyway.
These five also have corresponding `readFile` functions - see https://www.snoyman.com/blog/2016/12/beware-of-readfile/ . As Snoyman recommends in that post, it's probably best to "Stick with Data.ByteString.readFile for known-small data, use a streaming package (e.g, conduit) if your choice for large data, and handle the character encoding yourself. And apply this to writeFile and other file-related functions as well."
The first comment on that post starts out with "This problem extends well beyond readFile." Having the string handling more standardized at the language level can make life quite a bit simpler for developers.
That... sounds awful for performance. Is that a real thing?
It generalizes really well to Unicode; encoding is unnecessary. :p
On the other hand, I think the name of the function that converts an integer to its own string representation, integer_to_list, could have been chosen better.
Haskell's type class capability (similar to traits or interfaces) was still new/experimental at that time, and one benefit of strings as lists is it allowed them to easily be manipulated using the same syntactic and semantic machinery - pattern matching, recursive processing etc. - as other list data.
These days real, non-trivial code uses much more optimized string representations, but the original String type still exists and is used in various standard library functions, like "error". But as the original commenter pointed out, "you can just upgrade strings like any other dependency," so if you want, you can always import some other library that uses e.g. the Text type for error message, if you care for some reason.
The performance is good enough for teaching and learning and mocking out examples... But the reason for the user libraries is that real string processing workloads need something way better as their default.
Really? Is all that necessary?
They are things that your browser happily rebroadcasts back to the server with no real UI for it outside of the shitty devtool bar made for devs, even after all this outcry about cookies.
It reminds me of the meme of the guy riding a bicycle, throwing a branch into the spokes (rebroadcasting cookies), and then roaring in pain on the ground about how evil websites/advertisers are tracking him with cookies.
That said, what a lame HN thread on a post about Haskell.
I wish there was a good way to visually differentiate when a new top comment starts except by squinting and figuring out the whitespace from the left of the mobile screen, more painful than necessary I presume.
Yeah, I think both (A) defaulting to auto-expanded threads and (B) making them annoying to collapse make HN worse than it could be.
You tend to read the top-level thread because it's already there. And then it ends up being longer than you expected, or you're trapped in a subtree that just won't end, or you just want to see what other people are saying. And there's no good way to move past it.
Would be nice to click the indentation to collapse the thread anywhere inside the tree.
Every marketing cookie generates revenue for the website in some way or another. The website wants revenue, so it asks the user agent to maintain those cookies. The user agent agrees. Then the operator of the user agent gets upset that the website asked their agent to store the cookies? Get upset that your agent agreed, not that a request was made.
Or better yet, don't get upset at all and just solve the darned problem yourself. Is this Hacker News or Complier News?
Or how do you solve this problem? Personally, the most I can be arsed to do is install some Adblock Plugin. I did that only a few months ago and I'm not even sure that it improved my experience by a lot.
Absence of cookies don't make things unstable (non-solid?), and fuck knows what 'modern' is supposed to mean, or why it's good.
> Or how do you solve this problem?
Block all cookies except for rare moments like posting on HN, which then immediately get deleted. And no JS, which means CPU is trivial (so no burn-a-core-for-every-open-tab which is so common with page-sized pointless animations). Many problems can be solved if you want them to be.
If I do need anything more, there's VMs. BTW what 'fun and inspiring information' do you refer to? Shadertoy is a loss I grant, but what else?
Deleting Cookies on exit (and/or at regular intervals) will probably not help much in terms of avoiding tracking, especially if you log back in using your reinitialized cookies.
which again you don't give.
> Anything that requires interactivity ... obviously require Javascript
jeez, no shit, I get it.
> (some defeatist blah about cookies)
Whatever.
You just persistently don't get it. These are my choices. I made them carefully. They suit me. They may not suit you. We could even compromise if you made an effort to see what I'm after but you won't/can't. Now please try to understand I'm not you, and just back off!
Also, where is this burn-a-core-for-every-open-tab stuff? Many websites are highly optimized and do not use much CPU. Not enough to be noticed without actually looking at the numbers anyway.
What sites have page size animations these days?
I don't want them to. I log back in if necessary (browser remembers id/pswd). For those few I need to stay logged in, I use a VM and save the state - I'm more concerned about controlling JS than cookies in such cases.
> And how would be have any web apps that aren't horrendous without JS?
I don't use web apps. My tradeoff.
> Also, where is this burn-a-core-for-every-open-tab stuff? Many websites are highly optimized and do not use much CPU.
Oddly, it seems to be corporate bullshit sites that are the worse offenders. Can't find one but you're right, it's not all by any means. I retract.
But I think the vast majority of people would be upset if sites didn't keep you logged in and there were no web apps.
It's even worse if you prefer FOSS and use web apps, since Chromium no longer has password sync, Brave and FF block advanced features, and if you use BitWarden it takes a few extra clicks.
Yeah. The less info the more clipart/general crap you'll get. Weird innit.
As to your other points, I can't argue. I accept a higher level of inconvenience for a higher level of security, that's just my choice. I won't inflict it on others who make different tradeoffs.
The Cisco survey(https://iapp.org/news/a/new-cisco-study-emphasizes-consumer-...) says 79% are willing to invest time or money to protect their privacy, but a lot less seem to actually do anything about it.
Almost everyone I know is on Facebook and Gmail, most seem to use Chrome, etc.
It seems to vary a lot with subculture. Programmers always seem to be more willing to sacrifice convenience, and people who watch porn seem to be more interested in privacy than most.
I suspect there's a pretty large segment that only cares in theory, at most, and only then just on principle because of the other people who have more interesting data.
Maybe not 95% but probably 90% of certain subsets at least.
Hint: it's under Tools|Preferences in firefox/palemoon
There is no setting anywhere, in any web browser, to "retain cookies that are technically necessary and reject marketing cookies" which is the desirable behaviour.
(Some possible control via Tools|Preferences|Exceptions... button allows you to customise by website, although I've never used it. Or just disallow all, which is what I do)
---
Edit: answer the question please, there may be an easy solution to what you want.
Edit2: No reply because god forbid there's an actual way you could take control, that would simply ruin everything (in a parallel universe, man complains the streets are rife with face stabbing but when presented with proof they're not, stabs self in face to prove otherwise).
Biggest problem with learned helplessness is that they like it that way. Gives them something to be angrily resentful about.
The only entity that can specify whether a cookie is marketing-related or not is the website. No one else can.
I used umatrix for years but gave up. The guessing what to enable to get a site to work got tiresome, and IIRC there was also problem with browser support.
We should all complain loudly and far more than we do about the creeping tendency of many companies to do so many obviously shitty things, instead of merely shrugging our shoulders.
The only entity with any real power to decide which cookies the website uses is the website itself.
Asking the user or the user agent to comb through cookies and decide, one by one, which ones seem marketing-related and which ones are technically required, and then block, is way too much to ask from a regular internet user.
I have tried, but fail to see good faith in your reply.
You don't even have to shift through cookies for this to work, you can just reject all by default until the user explicitly request them to be stored (or use a whitelist or wait until the users tried to login that would necessitate a cookie, etc.) Lots of possibilities.
> is way too much to ask from a regular internet user.
That's kind of the point. By making it all transparent and seamless browser makers played into the hand of marketing companies. If cookies had a cost and would degrade the user experience, they might be thinking twice before putting hundreds of them on a site.
Marketing companies are just making use of the tools they are given. And browser manufacturers gave them a lot of tools, while taking control away from the user.
> That's kind of the point. By making it all transparent and seamless browser makers played into the hand of marketing companies. If cookies had a cost and would degrade the user experience, they might be thinking twice before putting hundreds of them on a site.
Cookies do have a cost, namely the bad PR from people complaining about the unnecessary tracking cookies. If you think that's not enough, then you are free to reject cookies as well to degrade your own experience. But they aren't mutually exclusive. Complaints and bad PR can also drive users away from the site and enact change.
Simply put, browser could to a lot better job at preventing this.
As for legitimate use, I don't really see much. Login handling is the obvious one, but I'd argue that login handling itself is in dire need of a rework and should be handled by a proper Web standard, not site specific hacks and "Save password" guesswork.
As for legitimate use cases, I think shopping carts on most online marketplaces use cookies.
The website is the one who decides which cookies to send in the first place. The browser never invents a cookie out of thin air.
> you can just reject all by default until the user explicitly request them to be stored
Which cookies should the user "request to be stored" and which cookies can the user safely ignore? How does the user tell them apart? Why should the user have to bother?
> If cookies had a cost and would degrade the user experience
Cookies are already degrading my user experience; you may have noticed the cookie consent popups on many sites. Those popups exist because cookies were being abused (ie. non-consensually) for purposes that are not essential to the functioning of the website. Such uses are now banned in the EU.
> And browser manufacturers gave them [marketing companies] a lot of tools
Browser manufacturers did not build those tools for the sake of marketing companies.
I can't fault websites for making use of functions the browser offers them.
> Which cookies should the user "request to be stored"
Have a simple toggle button for "Save state for this website" and discard everything when that button isn't pressed. Most website I visit I don't care about and have no need to keep any state. The few that I need to log into, I can just press that button. Knit that together with the "Save Passwords" function and it might be pretty much automatic most of the time.
> Those popups exist because cookies were being abused
Those popups exist because browsers failed to do their job. If the users wants warning for cookies, that's something the browsers can do just fine by itself, yet few do (e.g. Lynx).
> Browser manufacturers did not build those tools for the sake of marketing companies.
I'd disagree on that. Google makes their money with ads, so of course they'll optimize both Chrome and Search for maximum ad friendliness. Meanwhile Firefox is also run on Google ad money, so they can't step to far out of line either. There aren't many browsers that are build for the user first. The "you are the product" quote applies to browsers just as much as it does to Facebook.
I have JS locked down and third-party cookies disabled. This site only managed to set one cookie for me because of my power to decide. Despite that, all content was readable.
An incredible amount of the web just breaks. Twitter, Reuters, Imgur. Like it's one thing if, when I attempt to log in, your log in fails (and usually, logins fail to handle the error & will just loop back to the start, that's at least a start) but a lot of the web will have a flash-of-text and then nothing, & JS has crashed.
The post they reference, is also very honest: ..., the fastest Haskell implementation of the Aho–Corasick string searching algorithm, which powers string search in Channable.
Basically the blog posts show that if you want to program in Haskell and still optimise, this is how you can do it. I think both posts are great resources and don't overstate their claims.
I wrote this article during a short internship at Channable. Not to be apologetic but I think these kind of articles are so prevalent because young or unpopular languages usually have worse documentation than established ones (naturally). I basically wrote down the things I learned during my internship that I found noteworthy.