If so, then node can take exactly the same strategy.
If so, then node can take exactly the same strategy.
We spend many hours reading JavaScriptCore (the engine)’s code and when we wrote the code for Buffer and related APIs we spent a lot of time benchmarking and iterating on different approaches. A lot of performance work looks like this
That's bun in a nutshell. Node could be faster, but they are not. Bun is essentially using the same architecture as Node, just implemented in a more efficient way.
Therefore, the hopes that Bun would herald a breakthrough in JS performance are ill informed. However, best case-scenario Node gets a good nudge to become more efficient.
[1] https://ziglang.org/documentation/master/#String-Literals-an...
Codepoint-level string abstractions in particular are complete nonsense that only serve to give you the illusion of making things easier before learn the hard way that Unicode is more complicated than that. This also goes for UTF-16 which is only a cope extension uf UCS-2 for those that already made this mistake before additionally realizing that 2 bytes are not enough to encode all human languages.
Now you might think that declaing all your strings are UTF-8 wouldn't have any of these problems and is the way to go .. until you find out that there are strings you can't represent as (valid) UTF-8 including things that are almost UTF-8 like filenames and other OS-provided data under most POSIX operating systems. This also applies to UTF-16 under Windows btw.
And yes, UTF-16, used by Java and C#/.NET is a pain-point as it forces conversion from sockets/files from UTF-8 into UTF-16 so they can be used, then another one when writing back. But that's beside the point when talking about Zig.
You actually cannot implement substring replacement at the byte level with Unicode, just think about what happens if there is a modifier right after the substring in the original text. You cannot just avoid the fact that Unicode (and human writing in general) is a mess.
Zig also has a concept of "sentinel terminated arrays/slices" which allows easy interop with C APIs for string data, but the details and implications go a bit too far for a comment :)
Why do you say that std::string_view is more awkward to use as a result of being part of the stdlib?
As far as I'm aware, the type of string literals in C++ is still a "raw" char pointer and not a string view. Also few libraries actually make use of std::string_view, while in Zig everything is built around strings as slices, from the language to the stdlib to 3rd party libs (easy to do of course in a new language ecosystem).
You can define a constexpr std::string_view with a literal if you so choose. And if you pass a string literal to a function that accepts a string_view, then the compiler has enough information to (and typically will) construct the string_view in constant time. (A sibling commenter also points out that you can use the sv suffix: https://en.cppreference.com/w/cpp/string/basic_string_view/o...)
(Meanwhile, if you pass a string literal to a function that accepts const std::string&, then that will be a linear-time operation, as it copies that data since std::string owns its data. But with any amount of indirection, you'll end up with an implicit strlen call. So this is absolutely a pitfall.)
> few libraries actually make use of std::string_view
This is a fair point, as it's a relatively new addition to the language and by no means mandatory.