Swift UTF-8 String
swift.org
swift.org
Can anyone explain why this is done now and not when Swift was first released? I mean, UTF-8 was already the clear winner when Swift started. Is it some obj-c compat story?
Swift was created two decades later so the only justification which seems to make sense would be compatibility with Objective C and UTF-16 APIs.
Back in the day, when the Unicode people were pushing it as an alternative to whatever the ISO standard is that solved the same problem, they argued it was better than the ISO standard because the was only one encoding, UCS2 whereas ISO defined 8, 16 and 32 bit encodings. They said they could get away with that because 16 bits was enough for all modern languages (read: languages in use). Windows, Java and Javascript among others drank the kool-aid.
But they were wrong, 16 bits was not enough. USC2 morphed into UTF-16, which is actually worse than UTF-8 because if has the byte ordering problem.
To make matters worse, Unicode decided to abandon grapheme == code point. (This means there is ä and ä. They look identical, but one is '\xe4' and the other is a combining diaeresis, '\61\u0308', and the latter is now preferred.) I don't know why that did that as it makes string handling fragile, but I notice the encoding chosen for UTF-16 drastically reduced the number of code points available, so I wonder if they were worried about running out. The end result means that makes utf-32 actually harder to parse than utf-8 (which is quite an achievement), so you may as well use utf-8 for everything.
Nonetheless, UTF-16 will most likely stay. They might add wrappers around all API functions that convert arguments first, so we'd get CreateWindow8 alongside CreateWindowW and CreateWindowA, but I wouldn't hold my breath. The duality of API functions was to enable programs to be built for NT and 9x with the same codebase. It was never really intended for people to call them directly to enable non-Unicode-aware applications on NT.
The file system cannot migrate from UTF-16 anyway without breaking things, so that will stay as is.
IMHO there's pretty much nothing to gain, except a lot of risk in trying to change Windows to UTF-8 internally and the current interface might be an inconvenience to some developers, but not as much as to justify changing the whole OS.
Thus you have to be very careful when handling filesystem paths.
Now we have surrogate pairs and the IVS, there is no advantage of using UTF-16 at all.
I will not surprise the Microsoft someday convert everything including kernel to filesystem to UTF-8 someday... or more likely they lost the market share.
There's a course currently running in University of San Francisco run by Fast.ai which is using Swift-Tensorflow taught by Chris Lattner for two lessons: https://www.usfca.edu/data-institute/certificates/deep-learn...
Not exactly what you asked, but pertinent.
import Python
var np = Python.import(“numpy”)
And can then use numpy as you would normally. Same goes for plotting.
¹It technically is, but most people don't mean reference counting when they say GC.
Software Engineering is not about what people think, rather what is technically correct.
Naturally when in some juriditions one is allowed to call themselves engineers after a 6 months bootcamp we land in such customary/practical definitions.
What exactly are you trying to accomplish here?
We're interested because it's one of a handful of statically typed functional programming languages. Haskell/Ocaml are great as ML type languages, but it's a big paradigm shift for a team to jump into. Rust still has a pretty decent learning curve. Swift seems to have landed nicely in the "anyone can pick it up" territory. Kotlin is also in that territory, but the JVM is a beast.
For now, we're heavily focused on Typescript as the statically typed functional language, and we leverage types pretty aggressively (conditional types, discriminant unions, etc). Our iOS devs bring a lot of the Swift style over to our backend Typescript systems and from what I've experienced I think Swift absolutely will have a place in backend development.
My money is on Rust for really low level stuff and Kotlin for anything where JVM overhead is acceptable.
GC'd languages have some support (JNI, etc) but they require handles, pinning, write barriers, etc. which ends up being more complicated and less flexible.
For more general purpose backend work this is not usually a major consideration though, which is one of the reasons I predict most backend work will continue to be done in a GC language.
Ref counting can beat GC if you're in a domain where GC pauses can be an issue. Things like games, for example. Although I understand state-of-the-art GCs can also be fast enough to use for these kinds of applications in some cases now. Ref counting can also be more memory efficient than GCs.
It depends on how you use it. As I understand it, if you use enums everywhere it all sort of works out. If you start mixing in structs, there are certainly a lot of minefields.
https://github.com/clojure/clojure/blob/clojure-1.9.0/src/cl...
https://github.com/clojure/clojure/blob/clojure-1.9.0/src/cl...
I would love to see more Swift on the backend. It's surprisingly hard to find a language that's concise, expressive, statically typed, and easily accessible to beginners. Swift checks all these boxes.
That being said, it's much easier to succeed on the front-end when the only competitor is Objective-C (see: JavaScript), than on the backend when the competitors are Go, Rust, Elixir, Python, Node, C++, Java and more, all with large ecosystems already in place.
Swift may succeed in a niche, and "Swift for TensorFlow" appears to be an attempt at that, but I don't think it will become ubiquitous (call it traditional HN negativity).
As for being active, no signs of life since 08th February, more than one month without updates.
My favourite podcasts have weekly and bi-weekly updates.
The main problem for me is that it's still very Apple centric. Other platforms are still very second class.
For Linux, there are only Ubuntu packages. Even on Arch which is usually front of the pack with these things, there are only brittle third party (AUR) packages.
Windows support is even worse.
That GUI drag'n'drop onto code... ugh.
SmallString actually was just a tagged cocoa string at one point but that was removed in favor of more powerful implementations. The latest implementation allocates no memory on the heap either but can store more characters, so using a tagged pointer would be a step back.
⇒ it seems tagged NSString pointers indeed map to opaque strings.
Because strings are structs in Swift, I don’t expect them to implement tagged pointers in Swift. Alternatively, you can see the SmallString struct as a tagged pointer that’s twice as large as a regular pointer.
pypy, on the other hand, made this change in the latest release (7.1.0): https://twitter.com/pypyproject/status/1095971192513708032
(In memory usage, the Python way is strictly worse: adding one 4-byte character to an ASCII string forces the entire string to be upgraded to UCS-4. The advantage of the Python way is O(1) access to codepoints, but that's an operation that rarely comes up in practice, when dealing with Unicode strings.)
I'm not sure how that would be compatible with Python, either. Either you're giving up string[index], or you're causing this one subscript operation to take O(n) time, or you need a new index type just for strings (which is what Swift does). None of these seem very 'Pythonic'.
EDIT: Apparently recent versions of pypy solve this by computing their own index into your string, which seems like it would be terrible for memory usage, but I look forward to their blog post about it.
UTF-8 is a different story, as nearly every byte of UTF-8 string is a surrogate for most languages except English/ASCII.
Now, how can a char be indexed then? Only by slow iteration from the beginning of a string?
UTF-8 is a good choice for storage, true. However, it seems to be a not so good choice for string processing from the algorithmic point of view.
Or maybe there is a catch that makes UTF-8 strings well-suited for string processing as well?
Nope. UTF-16 is not UCS-2 anymore for multiple reasons. You cannot assume you can index a UTF-16 string as if it was UCS-2 anymore. UCS-4 (32-bit) is a possible option, though not generally recommended either.
There are multiple tricks for getting great performance in UTF-8 string processing such as codepoint maps of various sorts. It's quite well suited for string processing, and it's better to use UTF-8 at this point simply because it also helps from falling into the UTF-16 is not (and has not been for some time) UCS-2 trap, and UCS-2 cannot represent a lot of Unicode today (including many emoji).
Still, many popular platforms treat it like that in terms of char indexing. They are able to get away with it since a surrogate char is a rare guest in UTF-16.
P.S. It's surprising how easily HN crowd downvotes something when it falls out of a whimsical "popular contemporary view on things". Having a lot of years of text processing behind my shoulders, I raised an important question and got an immediate down-vote attack. But never mind, I'll survive.
That the most visible mechanism of accessing parts of a string is based on code units along with UTF-16 masking bad code unless Emoji are involved (for most developers at least), is a problem, though. UTF-8 makes the failure case pop up much faster as pretty much every language doesn't only use ASCII.
Even if you are restricting "most languages" to mean "most programming languages", there's an increasing rise of emoji in comments alone for languages that support UTF-8 and some languages are picking up increasing emoji usage in places like identifier names.
Platforms that treat UTF-16 like UCS-2 and allow raw char indexing rather than codepoint-oriented traversal are wrong in 2019. (It's been wrong since the very definition of UTF-16, such as RFC 2781 in 2000.) UTF-16 is a disappointing hack because it allowed UCS-2 platforms the luxury of pretending that the UCS-2 era never ended. UTF-8 at least leaves you very aware that the UCS-2 emperor has no clothes.
It bothers me that in 2019 it can still be difficult to manage even basic things with strings, and yet, whenever we talk/argue/bike-shed about 'languages' we get caught up in intellectual meandering. Please, give us strings that make sense. Then argue about monads.
This is really, really hard.
Some of that complexity is because world languages themselves are not simple. (eg. If you were to ditch Unicode and start from scratch, it would remain hard to do bi-di text, to pick a random example.)
That's not 'ignoring complexity' and it's pragmatic.
That would help in English as it's really hard today just to deal emojis.
So even if you pick one definition and decide to live with it always... You'll still have to know what tradeoffs you (or an API designer or implementer) made.
'It's complicated' is the utterly the wrong answer as to why we 'can't have better strings'.
Strings in almost every programming language were designed before Emjojis and true internationalization. If we designed them today, they would have support for dealing with most of the problems we face.
And FYI - there are not infinite corner cases. They are limited, and we can develop constructs for dealing with all of them.
As I said in another comment, I still applaud efforts to improve the status quo.
That doesn't mean anything. Do you mean a code unit, a codepoint, a grapheme cluster, if the latter is it legacy, extended or tailored, how do you handle the locale dependency, … All of them can be useful in some context, all of them are length. And if anything, the truly useful one is the first one (because it tells you how much room you need to store something).
> That's not 'ignoring complexity' and it's pragmatic.
It's true that it's not ignoring complexity, because you're ignorant of it to start with. Which isn't really different.
> That would help in English as it's really hard today just to deal emojis.
It… really isn't.
This is totally the wrong approach.
Language is shifty, but there are definitely a series of tools that should be integrated into every 'String' that can facilitate most problems.
"That doesn't mean anything" This is false. It absolutely means something, you can't brush it aside by inventing a separate series of specific measurements. Should we have those other measurements - yes, in many cases, we probably should have them as well -> that's the whole point.
This is totally the wrong approach.
Language is complex and sometimes ambiguous, but there are definitely a series of tools that should be integrated into every 'String' that can facilitate most problems.
Not every company is trying to satisfy Hindi speakers.
"That doesn't mean anything" This is false. It absolutely means something, you can't brush it aside by inventing a separate series of related measurements. Should we have those other measurements? - yes, in many cases, we probably should have them as well - which the whole point. We desperately need a series of common tools built into String, some of them which provide more nuance.
"It's true that it's ignoring complexity, because you're ignorant to start with"
Well, since I have a background in NLP, I've developed writing recognition and word prediction algorithms, as well as language models for dozens of languages for products used by millions of people every day ... maybe I'm not as ignorant as your snide remark suggests?
"> That would help in English as it's really hard today just to deal emojis.
It… really isn't."
Uh, yes, it absolutely would help, and it's painfully obvious. In most languages, which cover the exceeding majority of markets for most apps made, simply knowing the character length (i.e. 'visible character' as to avoid confusion over messy nomenclature here) is very useful. In Latin languages, after Unicode normalization and other kinds of cleaning, Emojis present a very common case for confusion.
There are very small set of tools (or rather, a slightly better, more standardized String approach) which would solve most of the problems for most companies developing for most markets. There are even a few extra tools which would solve, pragmatically all of them.
Swift is meant to be retro compatible with Obj-C and all. As someone pointed out windows in the meantime have basically no announced battle plan to transition out of UTF16.
I this game latecomer were rewarded because they started with UTF8 or switched from ASCI. But innovators actually tried UTF16 in the mean time an now have to do the job twice in a retro-compatible manner.