Maybe option 1 is feasible but not really practical, i can only see it being used in extremely low level stuff with a non standard compiler and going through tons of hoops like pre-creating loop-variables used later in the function and disabling optimizer.
And yes, that's specifically written to be awful and unsafe, but there are circumstances where you need to be close to the metal and carefully resort to more complicated variations of such things. That's what C is fairly uniquely appropriate for.
char foo[9]; foo[] = "abc"; foo[3..3+5] = "1234";
No unchecked buffer overflows, and no calls to strlen. The +5 puts the terminating 0 on.But what's the mention of "terminating 0"? The article says that terminating zeros should not be needed under its proposal; and that's what I was saying didn't make sense. [Added:] So, if you didn't just happen to know that the global string variable foo contains a string that was 3 characters long, how would you concatenate "1234" to it? I don't see any way without either double-fat pointers, or terminating NUL.
No, but strcpy does need to check each char for NUL byte as it copies the string. And strcat will need to redo same check on the original strcpy'ied string plus the new one.
So "abc" length is effectively checked twice and "1234" once.
So there's some truth to the matter, even though they're not "true" strlen calls.
strcat(s1,s2) does a strlen on s1 and s2.
Now, two of the strlen's can be replaced with byte-by-byte copies checking for 0 for each, but that tends to lose the efficiency that a memcpy would bring, so you're pretty much suffering from it anyway.
BTW, here's the strcat I wrote eons ago:
https://github.com/DigitalMars/dmc/blob/master/src/CORE32/ST...
It does do two strlen's (the repne scasb instructions). With the improvements in CPUs since there are probably better ways to write it, but that was pretty good for its day.
Here's strcpy:
https://github.com/DigitalMars/dmc/blob/master/src/CORE32/ST...
which does the test-every-byte method. I think Steve Russell wrote it, but I'm not sure.
If there's anything efficiently implemented in a C compiler, it's memcpy. Being able to implement string processing in terms of memcpy leverages that very nicely. strcat() and strcpy() don't leverage it.
Which do you think is faster (s2 is 1024 bytes long)?
strcpy(s1, s2);
memcpy(s1, s2, 1024);
I've dramatically speeded up a lot of my code and other peoples' by replacing the strxxx functions with memcpy. It's low hanging fruit and one of the first things I look for.Am I missing something?
[0] https://github.com/lattera/glibc/blob/master/string/strcpy.c
[1] https://github.com/lattera/glibc/blob/master/sysdeps/generic...
1. This implementation tests every byte, as discussed in other posts here. That makes it slow.
2. This implementation is likely not used - the gcc compiler probably has an internal code sequence it emits for a strcpy.
The function still checks every byte for \0.
And if i'm reading it correctly, only checks the upper bound after copying the data. And checking it before copying would require a call to strlen.
mov EAX,[ESI]
mov EBX,4[ESI]
mov ECX,8[ESI]
mov EDX,12[ESI]
mov [EDI],EAX
mov 4[EDI],EBX
mov 8[EDI],ECX
mov 12[EDI],EDX
add ESI,16
add EDI,16
Modern processors can likely do the 4 loads and 4 stores in parallel. That can't be done with 0 terminated strings, as you have to check every byte for 0. Even worse, you have to take care not to seg fault by reading too far past the 0, as there may not be any valid memory there.I do agree with the general idea that null terminated strings are a mistake though.
Non-dynamic arrays of char is just supposed to be a simplistic representation of a sort of thing that one has occasion to want to do in C that doesn't seem to fit into the proposed model without going into "old C NUL-termination" mode, or a "keep track of the string's length yourself" scheme, either of which would seem to ruin the whole thing. Thus my claim that this single feature would be hard to graft onto C in a useful, upward-compatible way. It's fine to have a language where all strings are dynamically allocated on the heap, or have an immutable known length, but that's a non-starter in the existing C universe.
The point isn't that doing things D's way isn't great; the point is that there's no reasonable way to put this feature into C. Every reasonable approach to string (and pointer) safety ends up being a new language: C#, Swift, Java, etc.
You don't seem to have followed the discussion or the paper. The array length is known; a slice is a "fat pointer", but not "double-fat".
"The point isn't that doing things D's way isn't great; the point is that there's no reasonable way to put this feature into C."
You're plainly wrong; the proposal in the article does exactly that.
With dynamic arrays, just return a slice.
This doesn't happen 3 times, though:
> char foo[9]; strcpy(foo,"abc"); strcat(foo,"1234");
Strcpy has to iterate through all the elements because it copies them. This would happen regardless. It doesn't do strlen.
Strcat has to find the end of the destination string, so it has to iterate (or call strlen). Then it's just strcpy again.
Instead of 3 strlens, there is 1.
Do you not understand how C string/arrays work, or why do you insist on 3 strlens?
Iterate all the elements and copy a fixed size is two very different things.
strcpy has to read, byte for byte, and check it for null (which is exactly what strlen does). A "real" copy would just blindly copy a chunk of memory with no other processing on it. The speed difference is huge.
Yes, but that's much more expensive than a memcpy (of the two source strings) or just knowing the length (of the string in the target buffer).
> Strcpy has to iterate through all the elements because it copies them. This would happen regardless.
No, it wouldn't; memcpy is generally a lot faster than strcpy.
> Do you not understand how C string/arrays work
Do you not understand that he's written a few C compilers, and designed and implemented a language that is known for its runtime compatibility with C?
> or why do you insist on 3 strlens?
Um, you already noted that he really means "code has to iterate all the string elements". What he should have said is that you have to find 3 NULs. That's an expensive operation even when you're copying the string while finding it.
Safe types doesn't stop you from doing this. You do need another length-field though, one for the allocated size and one for the used size. In c++, std::string already has this feature with the reserve/capacity functions, STL is also heap based but it is possible with some effort to pass it a stack allocator. Now c++ isn't exactly the best reference when it comes to these things either but just saying conceptually fat pointers doesn't stop you from doing these things, see WalterBrights reply for a better example.
You can still do that in a C dialect with fat pointers. You can have strings with 5 byte chars and use 0x1337 as string terminator if that's what your metal needs.
The point is, you don't have to, the compiler provides you with a sane array implementation that is adequate for 99% higher level algorithmic tasks.
#include <stdio.h>
#include <string.h>
int main() {
const char* s1="abc";
const char* s2="1234";
char str[strlen (s1) + strlen (s2) + 1];
strcpy (str, s1);
strcat (str, s2);
puts(str);
return 0;
}
$ gcc -std=c99 -Wall -Werror ./tst.c -o tst && ./tst
abc1234 char[3] s = "abc";
char* p = cast(char*)malloc(s.length + 4);
assert(p != null);
char[] a = p[0 .. s.length + 4];
a[0 .. 3] = s[];
a[3 .. 3+4] = "1234";
It's more verbose than necessary, but I wanted to illustrate the idea. Note how the allocation is turned into a dynamic array.Note that my proposal is not for a new memory allocation scheme for C, just a way to map data onto arrays.
Please check who you are replying to.
char foo[9]; foo[] = "abc"; foo[3..3+5] = "1234";
`foo` is an array. `foo[3..8]` is a slice, which is an object that does its own bounds-checking. I don't think the heap is used here.Another explicit example:
char foo[5] = "abcd"; // still NUL-terminated, carries length aswell
char[] bar = foo[2..3]; // a fat pointer with length 1
bar[0] = 'C';
printf(foo); // abCd
bar[1] = 'D'// ERROR!!! bar has length 1
Note that the array `foo` is now bounds-checked, which may affect backwards-compatibility. Also, `bar` is no longer null-terminated, which means you can't do printf on it.Some JIT optimization may allocate an instance in the stack, I'm not counting that.