I've been musing for a while now: what would it look like if we were to discard the C library and design a new one, leaving the language itself intact?
I've been musing for a while now: what would it look like if we were to discard the C library and design a new one, leaving the language itself intact?
Building blocks for memory were also very different from stdlib, notably the use of Handles, which were pointers of pointers, so that the OS could move a block of data around to defragment the heap behind your back without breaking the memory addressing.
C++ string_view is closer to the Right Thing™ - a slice, but C++ doesn't (yet) define anywhere what the encoding is, so... that's not what it could be. Rust's str is a slice and it's defined as UTF-8 encoded.
Back then it wasn't clear which encoding method would turn out to be dominant, so we did all three. (Java was built on UTF-16.)
As it eventually became clear, UTF-8 is da winnah, and the other formats are sideshows. Windows, which uses UTF-16, is handled by converting UTF-8 to -16 just before calling a Windows function, and converting anything coming back to UTF-8.
D doesn't distinguish between a string and a string view.
Imagine you go to a library and insist on borrowing "My Cousin Rachel", but they don't have it. "Oh I don't care whether you have the book, I just want to borrow it" is clearly nonsense. If they don't have it, you can't borrow it.
> D doesn't distinguish between a string and a string view.
In C++ std::string owns the buffer and std::string_view borrows it. If there is no difference between the two in D, then how is this difference bridged?
They added a setting in Windows 10 to switch the code page over to utf-8 and then in Windows 11 they made it on by default. Individual applications can turn it on for themselves so they don't need to rely on the system setting being checked.
With that you can, in theory, just use the -A variants of the winapi with utf-8 strings. I haven't tried it out yet as we still support prior Windows releases but it's nice that Microsoft has found a way out from the utf-16 mess.
I don't mind seeing UTF-16 fade away. We've been considering scaling back the D support for UTF-16/32 in the runtime library, in favor of just using converters as necessary. We recommend using UTF-8 as much as practical.
And the wheels fall off with the first string longer than 255 characters.
However, Free Pascal has the worst documentation of any major project I've ever encountered (The exact opposite of Turbo Pascal), so I can't link to a good reference. Their Wiki is a black hole of nuance and sucks all useful stuff off the internet.
You often end up with some kind of structure, or variations of structures, for strings:
struct string {
size_t length;
char data[];
};
struct string {
size_t length;
size_t alloc;
char *data;
};
Those are just examples. The tricky part is figuring out the different ownership use cases you want to solve. Because C gives you so much freedom and very little in the standard library, you end up with a lot of variations. You might use reference-counted strings, owned buffers, or string slices, etc. You might want certain types to be distinguished at compile-time and other types to be distinguished at run-time.An example can be found in the Git source code.
https://github.com/git/git/blob/master/strbuf.h
The history of changes to this file is interesting as well. This is a relatively nice general-purpose string type—you can easily append to it or truncate it.
I've seen many libs using this style of strings, not convinced by the practicality.
If you’re not convinced of the practicality, it sounds like you are simply not convinced of the practicality of doing string processing in C at all, which is a fair view point. String processing in C is somewhat a minefield. Libraries like Git’s strbuf are very effective relative to other solutions in C, but lack safety relative to other languages.
The trick is to pass an allocator (or container) to string handling functions.
If/when I want to get rid of all the garbage I reset the container/allocator.
I’ve seen similar approaches, e.g. with APR pools, and if your application can work within those restrictions, it’s very convenient.
[1] https://www.digitalmars.com/articles/C-biggest-mistake.html
float m[10][10];
it not a an array of pointers, but a 2D dimensional array with 2D memory layout.
int[][] jagged; // an array of `int[]` (i.e. each element is a pointer to a `int[]`)
int[,] multidimensional; // a "true" 2D array laid out in memory sequentially
// allocate the jagged array; each `int[]` will be null until allocated separately
jagged = new int[][10];
Debug.Assert(jagged.All(elem => elem == null));
for (int i = 0; i < 10; i++)
jagged[i] = new double[10]; // allocate the internal arrays
Debug.Assert(jagged[i][j] == 0);
// allocate the multidimensional array; each `int` will be `default` which is 0
// element [i,j] will be at offset `10*i + j`
multiDimensional = new double[10, 10];
Debug.Assert(multiDimensional[i, j] == 0); int N = 10;
char buf[N] = { };
auto x = &buf;
and 'x' has a slice type that automatically remebers the size. This works today with GCC / clang (with extensions or C2X language mode: https://godbolt.org/z/cMbM57r46 ).We simply can not name it without referring to N and we can also not use it in structs (ouch).
How is this not a quality of implementation issue? Any implementation is free to track all sizes as much as they want with the current standard.
Either a implementation is forced to issue an error at run time if there is an out of bounds read/write and in that case its a very different language than C, or its feature as-if lets any implementation ignore.
https://godbolt.org/z/qh7P93Tcd
And I agree that this is a misuse of auto. I only used it here to show that the type we miss already exists inside the C compiler, we simply can name it only by constructing it again:
char (buf)[N] = ...
but we could simply allow
char (buf)[:] =
and be done (as suggested by Dennis Richtie: https://www.bell-labs.com/usr/dmr/www/vararray.pdf)
I believe this is used by Redis.
my_function(my_var, 3.6, "bzarflo", my_other_var, false);
The string handling functions are part of the story, but the null-terminated char * is produced when the compiler reaches a string literal, and writing code without being allowed to just use string literals when it's convenient tends to feel like coding with oven mitts on. my_function(my_var, 3.6, $("bzarflo"), my_other_var, false);
Isn't that much more of a mouthful, and as long as 'my_function' knows to free it, then you're A-OK! The only trouble is '$()' isn't legal in standard C, so a real solution would have to be something like 'str()'.C is not perfect, there are some parts of the syntax that I strongly dislike, like casting or function pointers declaration...
But it is overall a good enough syntax, much simpler than C++.
A Freudian slip, methinks.