Going back to the post I originally replied to, how would going down to a bytes view avoid the problems you see?
Going back to the post I originally replied to, how would going down to a bytes view avoid the problems you see?
The choice of the bytes view specifically is just that it's the most popular view from which you can achieve one specific primitive: figuring out how much space a (sub)string occupies in whatever representation you store it in. A byte length achieves this. Of course, a length in bits or in utf-32 code units also achieves this, but I've found it rather uncommon to use utf-32 as a transfer encoding. So we need at least one string type with this property.
Other than this one particular niche, a codepoint view doesn't do much worse at most tasks. But it adds a layer of complexity while also not actually solving any of the problems you'd want it to. In fact, it papers over many of them, making it less obvious that the problems are still there to a team of eurocentric developers ... up until emoji suddenly become popular.
Now, I can understand the appeal of making your immediate problems vanish and leaving it for your successors, but I hope we can agree that it's not in good taste.
For example, working with the utf-8 view does not somehow foreclose on knowing how much memory a (sub)string occupies, and it certainly does not follow that, because this involves regarding the string as a sequence of bytes, this is the only way to regard it.
For another, let's consider a point from the linked article: "One false assumption that’s often made is that code points are a single column wide. They’re not. They sometimes bunch up to form characters that fit in single “columns”. This is often dependent on the font, and if your application relies on this, you should be querying the font." How does taking a bytes view make this any less of a potential problem?
Is a team of eurocentric developers likely to do any better working with bytes? Their misconceptions would seem to be at a higher level of abstraction than either bytes or utf-8.
You are claiming that taking a utf-8 view is an additional layer of complexity, but how does it simplify things to do all your operations at the byte level? Using utf-8 is more complex than using ascii, but that is beside the point: we have left ascii behind and replaced it with other, more capable abstractions, and it is a universal principle of software engineering that we should make use of abstractions, because they simplify things. It is also quite widely acknowledged that the use of types reduces the scope for error (every high-level language uses them.)
The heart of the matter is that a Unicode codepoint sequence view of a string has no real use case.
There is no "universal principle" that we use abstractions always, regardless of whether they fit the problem; that's cargo-culting. An abstraction that does no work is, ceteris paribus, worse than not having it at all.
The quote, as you presented it, leaves open the question: more capable than what? Well, there's no doubt about it if you go back to my original post: more capable than ascii. Up until now, as far as I can tell, your thesis has not been that unicode is less capable than ascii, but if that's what your argument hangs on, go ahead - make that case.
What your thesis has been, up to this point, is that manipulating text as bytes is better, to the extent that doing it as unicode is harmful.
> It must simply do something better. If there were actually anything at all it did better...
It is amusing that you mentioned the burden of proof earlier, because what you have completely avoided doing so far is justify your position that manipulating bytes is better - for example, you have not answered any of the questions I posed in my previous post.
> The heart of the matter is that a Unicode codepoint sequence view of a string has no real use case.
Here we have another assertion presented without justification.
> There is no "universal principle" that we use abstractions always, regardless of whether they fit the problem...
It is about as close as anthing gets to a universal principle in software engineering, and if you want to disagree on that, go ahead, I'm ready to defend that point of view.
>... that's cargo-culting.
How about presenting an actual argument, instead of this bullshit?
Furthermore, you could take that statement out of my previous post, and it would do nothing to support the thesis you had been pushing up to that point. You seem to be seeking anything in my words that you think you can argue against, without regard to relevance - but in doing so, you might be digging a deeper hole.
> An abstraction that does no work is, ceteris paribus, worse than not having it at all.
Your use of a Latin phrase does not alter the fact that you are still making unsubstantiated claims.
I guarantee you there will be a quick counterexample to demonstrate that the claimed use-case is incorrect. There always is.
You may review the gish gallop in the other branch of this thread for inspiration.
A Gish gallop is not, of course, a sound way to arrive at any sort of truth. Anyone employing the technique is either unaware of that, or is being duplicitous (my guess is that Gish never really understood that it is bogus.)
Meanwhile, those questions are still waiting to be milked, so to speak...