Conflating pointers with arrays: C's biggest mistake? (2009)
digitalmars.com
digitalmars.com
On the other hand, UBE (undefined, or unspecified behavior) are probably the nastiest stuff that can bite you in C.
I have been programming in C for a very, very long time, and I am still getting hit by UBE time to time, because, eh, you tend to forget "this case".
Last time, it took me a while to realize the bug in the following code snippet from a colleague (not the actual code, but the idea is there):
struct ip_weight { in_addr_t ip; uint64_t weight; };
const struct ip_weight ipw1 = {0x7F000001, 1}; const struct ip_weight ipw2 = {0x7F000001, 1};
const uint32_t hash1 = hash_function(&ipw1, sizeof(ipw1)); const uint32_t hash2 = hash_function(&ipw2, sizeof(ipw2));
The bug: hash1 and hash2 are not the same. For those who are fluent in C UBE, this is obvious, and you'll probably smile. But even for veterans, you tend to miss that after a long day of work.
This, my friends, is part of the real mistakes in C: leaving too many UBE. The result is coding in a minefield.
[You probably found the bug, right ? If not: the obvious issue is that 'struct ip_weight' needs padding for the second field. And while all omitted fields are by the standard initialized to 0 when you declare a structure on the stack, padding value is undefined; and gcc typically leave padding with stack dirty content.]
No amount of syntax sugar on top will prevent you from writing unsafe code, unless that basic model is thrown away.
Seriously, just don't cast object of one type to another type and you can forget about alignment until you need to optimize.
You need to keep this in mind basically any time you're going from byte buffers to more complicated types. Basically anything touching allocation or IO needs to be aware of it, IMO, even before optimization.
Quote:
Exact-width integer types
The typedef name intN_t designates a signed integer type with width N, no padding bits, and a two's-complement representation.
Thus, int8_t denotes a signed integer type with a width of exactly 8 bits.
The typedef name uintN_t designates an unsigned integer type with width N. Thus, uint24_t denotes an unsigned integer type with a
width of exactly 24 bits.I've stopped programming in C if I can help it. It's a crazy language with support for crazy machines.
No, the bug is thinking that hashing random bytes in your memory is correct. Why wouldn't you make a correct hash function for your struct ?!
The idiom here is to take the address of the struct and read the width of it's whole footprint in memory not just the field data. It's a weak idiom breaking under a common use case.
Most "UB bugs" stem from users who think they know that a struct or data type will be laid out in a certain sequence in memory.
If anything, I'd blame compilers here -- IMO, they should automatically throw at least a warning any time they need to pad/rearrange a struct to make it explicitly clear to developers what's happening.
Even the old language spec from 1989 talks about an "abstract machine with the expressions being evaluated as specified by the semantics". One of the first chapters 1.2 Scope makes this very clear and then 2.1.2.3 Program execution clarifies this even further. Other things to note is that the spec doesn't mention variables being ever stored on the stack, instead it's called automatic storage, see "Storage durations of objects", the word stack is not even mentioned once throughout the whole spec.
Today with modern CPUs all this is even more important, if you think you are operating on a global memory array you are doing it wrong.
The idea of using bytes is error-prone, but now that you mentioned it, pretty typical of the C mindset.
Of course C++ also has some of these cultural biases. I think they're an important reason why unsafe code continues to be written.
The safe programming tribe, which I include myself, usually refugees from Wirth languages, makes use of C++ abstractions and type safety to deal with unsafety, including type driven development.
Meaning heavy use of templates, type wrappers, STL (or standard library if you prefer) data types, pre-processor only for #include and yes some meta-programming as well.
Then there is the tribe of C refugees, whose C++ compiler was forced on them due to a platform SDK, usually eschew anything standard. Might write some C++ like code due to interop with the SDK APIs and that's it.
Naturally there are a few tribes in-between, but these are the two major groups.
IMO the right solution would be a special annotation on a struct that says “I want the logical value of this struct to uniquely determine the bytes of the struct’s in-memory representation.”
Of course, adding such an attribute without nasty edge cases may be tricky.
Because it's not conformant. Those possible C++ programs are broken.
Conformance is a property of the code, the programmer's decision process is the really interesting thing here to me and I was arguing that such a solution is not unusual for C, but it would be for C++ , where a programmer would avoid it not because it's not conformant, but because it looks odd. pjmpl further clarified that it depends on the type of C++ programmer - I agree.
> I was arguing that such a solution is not unusual for C
You may have asserted that (funny how people call their unsupported assertions "arguments"), but you're wrong if so. Again, it's broken, in C just as much as C++. You wrote
> instead one would use hash functions for each type and in the case of a struct the individual members would be accessed to build the hash value.
But this is just as much true in C as C++, because hashing pad bytes is wrong. In fact, C programmers are probably more aware of this than C++ programmers.
Buffer overflows are UBE, too. But the way I proposed the fix is a pretty much cost-free solution, and it's optional.
Redefining C so that struct padding is always 0'd is an expensive solution, and rarely needed.
Using std::vector/array/string::at would literally eliminate buffer overflows and yet programmers aren't generally using this style.
I would love it if I could prove to my colleagues that mandatory bounds-checking would not result in a noticeable performance loss, but my gut feeling is that it's not so. Interestingly Rust (and I guess D) does just that and seems to be getting away with it.
However, in the C UB thread the author of a Rust crate mentioned that a case for using unsafe is exactly this: avoiding the performance loss of bounds checking.
It's still up to you, the programmer. With C, though, you have no choice. No checking for you!
I meant cost-free in the sense that one way or another in C code you wind up passing the length anyway, or in the case of 0 terminated strings, you wind up recomputing it if you don't pass it.
"Many years later we asked our customers whether they wished us to provide an option to switch off these checks in the interests of efficiency on production runs. Unanimously, they urged us not to--they already knew how frequently subscript errors occur on production runs where failure to detect them could be disastrous. I note with fear and horror that even in 1980, language designers and users have not learned this lesson. In any respectable branch of engineering, failure to observe such elementary precautions would have long been against the law."
> Fortunately, the problem of program correctness has turned out to be far less serious than predicted. A recent analysis by Mackenzie has shown that of several thousand deaths so far reliably attributed to dependence on computers, only ten or so can be explained by errors in the software: most of these were due to a couple of instances of incorrect dosage calculations in the treatment of cancer by radiation.
In real life, we use safety gloves, seatbelts, elbow and knee protection, helmets, chainsaw cover, gun lock, ....
Naturally there are those that think such protections are only for children and an accident will never occur to them, until the day they happen to be part of the statistic.
I long for the day that lawsuits for buggy software become a regular activity, only then the industry will actually care to change.
Pure strawman. The statement was that program correctness is far less serious than predicted, not totally benign.
"and that quote in no way talks about bounds checking, which is the issue being discussed here"
Um, programs that exceed array bounds are not correct, so yes, it does talk about them.
[further strawmen not worth addressing]
What's the cost? It's an optional feature.
"Using std::vector/array/string::at would literally eliminate buffer overflows and yet programmers aren't generally using this style."
C programmers don't use that because it doesn't exist in C.
"I would love it if I could prove to my colleagues that mandatory bounds-checking would not result in a noticeable performance loss"
Did you read his article? There's nothing in it about mandatory bounds-checking.
But like you said, you have to be aware of this effect at least, and it is not possible to eliminate that by simply clearing all variables to zero. See,
p = (struct ip_weight *)some_used_mem;
p->ip = ip;
p->weight = weight;
At which point should imaginary p->garbage be set to zero? At cast? But that may again do unexpected thing, since casts usually do not modify data. The entire struct abstraction seems leaky as hell, but that’s the price of not dealing with asm directly.These examples show that not C itself is hard, but low-level semantics are. You have to somehow deal with struct {x, y} and the fact that y has to be aligned at the same time. And have different platforms in mind. Maybe it is platforms that should be fixed? Maybe, but these are hardware with other issues that may be even harder to get right.
I think C is okay, really (apart from compilers that push ub to the limit). Type systems in successors try to hide its rough edges, but in the end of the day you end up with semantics of the compiler (c++, rust), that a regular guy has to understand anyway; it’s trading one complexity for another. C++ folks often seen treating it as magic, simply not doing what they’re not sure about. Good part is some languages force you to write correct code, no matter how much knowledge you have. But NewC could e.g. instead force one to create that ‘auto garbage’ to make it clear (why not, safety measures are inconvenient irl too).
I have no strong conclusion, but at least let’s think of all non-cs people who make their 64KB arduinos drive around and blink leds.
The reason for undefined behaviors is to avoid over engineering. In a capable engineer's eye, it is beautiful.
That C has UBE is not a mistake, it's fundamental to the language design, which allows for unrestricted access to the bare metal. If you want a different sort of language, use Java.
Similar questions apply to general arrays, as well. Also: Am I able to take the address of an element of an array? Will that be a fat pointer too? How about a pointer to a sequence of elements? Can I do arithmetic on these pointers? If not, am I forced to pass around fat array pointers as well as index values when I want to call functions to operate on pieces of the array? How would you write Quicksort? Heapsort? And this doesn't even start to address questions like "how can I write an arena-allocation scheme when I need one"?
In short, the reason that this sort of thing hasn't appeared in C is not because nobody has thought about it, nor because the C folks are too hide-bound to accept a good idea, but rather because it's not clear that there's a real, workable, limited, concise, solution that doesn't warp the language far off into Java/C#-land. It would be great if there were, but this isn't it.
> If nul-termination of strings is gone, does that mean that the fat pointers need to be three words long, so they have a "capacity" as well as a "current length"?
No. You'll still have the same issues with how memory is allocated and resized. But, once the memory is allocated, you have a safe and reliable way to access the memory without buffer overflows.
> If not, how do you manage to get string variable on the stack if its length might change? Or in a struct? How does concatenation work such that you can avoid horrible performance (think Java's String vs. StringBuffer)?
As I mentioned, it does not address allocating memory. However, it does offer one performance advantage in not having to call strlen to determine the size of the data.
> On the other hand, if the fat pointers have a length and capacity, how do I get a fat pointer to a substring that's in the middle of a given string?
In D, we call those slices. They look like this:
T[] array = ...
T[] slice = array[lower .. upper];
The compiler can insert checks that the slice[] lies within the bounds of array[].> Am I able to take the address of an element of an array?
Yes: `T* p = &array[3];`
> Will that be a fat pointer too?
No, it'll be regular pointer. To get a fat pointer, i.e. a slice:
slice = array[lower .. upper];
> How about a pointer to a sequence of elements?Not sure what you mean. You can get a pointer or a slice of a dynamic array.
> Can I do arithmetic on these pointers?
Yes, via the slice method outlined above.
> If not, am I forced to pass around fat array pointers as well as index values when I want to call functions to operate on pieces of the array?
No, just the slice.
> How would you write Quicksort? Heapsort?
Show me your pointer version and I'll show you an array version.
> And this doesn't even start to address questions like "how can I write an arena-allocation scheme when I need one"?
The arena will likely be an array, right? Then return slices of it.
Maybe option 1 is feasible but not really practical, i can only see it being used in extremely low level stuff with a non standard compiler and going through tons of hoops like pre-creating loop-variables used later in the function and disabling optimizer.
And yes, that's specifically written to be awful and unsafe, but there are circumstances where you need to be close to the metal and carefully resort to more complicated variations of such things. That's what C is fairly uniquely appropriate for.
char foo[9]; foo[] = "abc"; foo[3..3+5] = "1234";
No unchecked buffer overflows, and no calls to strlen. The +5 puts the terminating 0 on.But what's the mention of "terminating 0"? The article says that terminating zeros should not be needed under its proposal; and that's what I was saying didn't make sense. [Added:] So, if you didn't just happen to know that the global string variable foo contains a string that was 3 characters long, how would you concatenate "1234" to it? I don't see any way without either double-fat pointers, or terminating NUL.
No, but strcpy does need to check each char for NUL byte as it copies the string. And strcat will need to redo same check on the original strcpy'ied string plus the new one.
So "abc" length is effectively checked twice and "1234" once.
So there's some truth to the matter, even though they're not "true" strlen calls.
strcat(s1,s2) does a strlen on s1 and s2.
Now, two of the strlen's can be replaced with byte-by-byte copies checking for 0 for each, but that tends to lose the efficiency that a memcpy would bring, so you're pretty much suffering from it anyway.
BTW, here's the strcat I wrote eons ago:
https://github.com/DigitalMars/dmc/blob/master/src/CORE32/ST...
It does do two strlen's (the repne scasb instructions). With the improvements in CPUs since there are probably better ways to write it, but that was pretty good for its day.
Here's strcpy:
https://github.com/DigitalMars/dmc/blob/master/src/CORE32/ST...
which does the test-every-byte method. I think Steve Russell wrote it, but I'm not sure.
If there's anything efficiently implemented in a C compiler, it's memcpy. Being able to implement string processing in terms of memcpy leverages that very nicely. strcat() and strcpy() don't leverage it.
Which do you think is faster (s2 is 1024 bytes long)?
strcpy(s1, s2);
memcpy(s1, s2, 1024);
I've dramatically speeded up a lot of my code and other peoples' by replacing the strxxx functions with memcpy. It's low hanging fruit and one of the first things I look for.Am I missing something?
[0] https://github.com/lattera/glibc/blob/master/string/strcpy.c
[1] https://github.com/lattera/glibc/blob/master/sysdeps/generic...
1. This implementation tests every byte, as discussed in other posts here. That makes it slow.
2. This implementation is likely not used - the gcc compiler probably has an internal code sequence it emits for a strcpy.
The function still checks every byte for \0.
And if i'm reading it correctly, only checks the upper bound after copying the data. And checking it before copying would require a call to strlen.
mov EAX,[ESI]
mov EBX,4[ESI]
mov ECX,8[ESI]
mov EDX,12[ESI]
mov [EDI],EAX
mov 4[EDI],EBX
mov 8[EDI],ECX
mov 12[EDI],EDX
add ESI,16
add EDI,16
Modern processors can likely do the 4 loads and 4 stores in parallel. That can't be done with 0 terminated strings, as you have to check every byte for 0. Even worse, you have to take care not to seg fault by reading too far past the 0, as there may not be any valid memory there.I do agree with the general idea that null terminated strings are a mistake though.
Non-dynamic arrays of char is just supposed to be a simplistic representation of a sort of thing that one has occasion to want to do in C that doesn't seem to fit into the proposed model without going into "old C NUL-termination" mode, or a "keep track of the string's length yourself" scheme, either of which would seem to ruin the whole thing. Thus my claim that this single feature would be hard to graft onto C in a useful, upward-compatible way. It's fine to have a language where all strings are dynamically allocated on the heap, or have an immutable known length, but that's a non-starter in the existing C universe.
The point isn't that doing things D's way isn't great; the point is that there's no reasonable way to put this feature into C. Every reasonable approach to string (and pointer) safety ends up being a new language: C#, Swift, Java, etc.
You don't seem to have followed the discussion or the paper. The array length is known; a slice is a "fat pointer", but not "double-fat".
"The point isn't that doing things D's way isn't great; the point is that there's no reasonable way to put this feature into C."
You're plainly wrong; the proposal in the article does exactly that.
With dynamic arrays, just return a slice.
This doesn't happen 3 times, though:
> char foo[9]; strcpy(foo,"abc"); strcat(foo,"1234");
Strcpy has to iterate through all the elements because it copies them. This would happen regardless. It doesn't do strlen.
Strcat has to find the end of the destination string, so it has to iterate (or call strlen). Then it's just strcpy again.
Instead of 3 strlens, there is 1.
Do you not understand how C string/arrays work, or why do you insist on 3 strlens?
Iterate all the elements and copy a fixed size is two very different things.
strcpy has to read, byte for byte, and check it for null (which is exactly what strlen does). A "real" copy would just blindly copy a chunk of memory with no other processing on it. The speed difference is huge.
Yes, but that's much more expensive than a memcpy (of the two source strings) or just knowing the length (of the string in the target buffer).
> Strcpy has to iterate through all the elements because it copies them. This would happen regardless.
No, it wouldn't; memcpy is generally a lot faster than strcpy.
> Do you not understand how C string/arrays work
Do you not understand that he's written a few C compilers, and designed and implemented a language that is known for its runtime compatibility with C?
> or why do you insist on 3 strlens?
Um, you already noted that he really means "code has to iterate all the string elements". What he should have said is that you have to find 3 NULs. That's an expensive operation even when you're copying the string while finding it.
Safe types doesn't stop you from doing this. You do need another length-field though, one for the allocated size and one for the used size. In c++, std::string already has this feature with the reserve/capacity functions, STL is also heap based but it is possible with some effort to pass it a stack allocator. Now c++ isn't exactly the best reference when it comes to these things either but just saying conceptually fat pointers doesn't stop you from doing these things, see WalterBrights reply for a better example.
You can still do that in a C dialect with fat pointers. You can have strings with 5 byte chars and use 0x1337 as string terminator if that's what your metal needs.
The point is, you don't have to, the compiler provides you with a sane array implementation that is adequate for 99% higher level algorithmic tasks.
#include <stdio.h>
#include <string.h>
int main() {
const char* s1="abc";
const char* s2="1234";
char str[strlen (s1) + strlen (s2) + 1];
strcpy (str, s1);
strcat (str, s2);
puts(str);
return 0;
}
$ gcc -std=c99 -Wall -Werror ./tst.c -o tst && ./tst
abc1234 char[3] s = "abc";
char* p = cast(char*)malloc(s.length + 4);
assert(p != null);
char[] a = p[0 .. s.length + 4];
a[0 .. 3] = s[];
a[3 .. 3+4] = "1234";
It's more verbose than necessary, but I wanted to illustrate the idea. Note how the allocation is turned into a dynamic array.Note that my proposal is not for a new memory allocation scheme for C, just a way to map data onto arrays.
Please check who you are replying to.
char foo[9]; foo[] = "abc"; foo[3..3+5] = "1234";
`foo` is an array. `foo[3..8]` is a slice, which is an object that does its own bounds-checking. I don't think the heap is used here.Another explicit example:
char foo[5] = "abcd"; // still NUL-terminated, carries length aswell
char[] bar = foo[2..3]; // a fat pointer with length 1
bar[0] = 'C';
printf(foo); // abCd
bar[1] = 'D'// ERROR!!! bar has length 1
Note that the array `foo` is now bounds-checked, which may affect backwards-compatibility. Also, `bar` is no longer null-terminated, which means you can't do printf on it.Some JIT optimization may allocate an instance in the stack, I'm not counting that.
It was created based upon real world experience having designed and implemented it in D. Where all of your concerns have not been discussed in the years following that article (it was already in the language for about 8 years at that point aka the start and has been solidly proven to work in the exact same context as it would have done in C).
In D at least, you can grab the pointer by a simple .ptr and for length .length. To get a specific element, it is as you would expect &f[i] all nice and straight forward. But what if you want to create an array from malloc? In D that is easy, just slice it! malloc(len)[0 .. len]. And free is just as you would expect from above, free(array.ptr);
In the end, I think the Go slices are the consequential implementation of safe fat pointers, having both the capacity and length and allowing efficient and still safe reslicing. The overhead of having 24 vs 8 bytes per pointer on a 64 bit machine should be worth it in modern times.
Don't see why you shouldn't be able to make a fat pointer point into a range inside the original array either? Just point it to an element and make the length-field shorter than the original? This is usally called array_view, span or slice in other languages.
No, it's not, those ideas were implemented in practice in D and Rust, and there are no real issues with those. This feature could be easily implemented in C, there are no dependencies on features that C doesn't have.
> If nul-termination of strings is gone, does that mean that the fat pointers need to be three words long
No need to store the capacity. This is a slice, not a buffer. Go conflates those two for user's convenience, but this is not necessary, and in fact is waste of RAM - not an issue for Go, but it is an issue for C. For instance, `&str` in Rust is a pair of pointer to a string and its length and it works really well.
> If not, how do you manage to get string variable on the stack if its length might change? How does concatenation work such that you can avoid horrible performance (think Java's String vs. StringBuffer)?
Use your own slice buffer abstraction for that purpose. It can be implemented as a struct storing a slice and its capacity. Pass a pointer to slice buffer abstraction, if you want a function to be able to add elements to it. This is also how it works in Go, for that matter.
Slices don't define concatenation. This is C, not a high level programming language.
> Am I able to take the address of an element of an array?
Yes. `&a[3]`. It's still an array, it just knows its size.
> Will that be a fat pointer too?
No.
> How about a pointer to a sequence of elements?
Probably you could add some sort of a range access syntax. Say, something like `&a[1:3]`.
> Can I do arithmetic on these pointers?
I don't know whether pointer arithmetic should be allowed or not, but even if it shouldn't be, there is nothing to stop you from doing `&a[4]` as a replacement for `a + 4`.
> How would you write Quicksort? Heapsort?
The same way you would with a regular array. Think of it as a struct storing an array pointer and its length. If you prefer to working with pair of start/end pointers instead of pair of start and array size, then note that `end - start` is array length, so getting an end pointer is trivial.
Even older than that, those features already existed in NEWP, Mesa and Modula-2, just to pick some examples back when C was being designed still.
You don't need a fat pointer. It can be part of the memory layout on the heap. How do you think `free` knows the length of the memory you are deallocating? Because the length is on the heap snuggled in right before the actual pointer malloc returned.
I once had a proposal on this. See [1]. Enough people looked it over to find errors; this is version 3. The consensus is that it would work technically but not politically.
The basic idea is that the programmer knows how big the array is; they just don't have a way to tell the compiler what expression defines the length of the array. Instead of
int read(int fd, char buf[], size_t n);
you write int read(int n; int fd, char (&buf)[n], size_t n);
It generates the same calling sequence. Arrays are still passed as plain pointers. But the compiler now knows how big "buf" is, both on the caller and callee side, and can check.I also proposed adding slice syntax to C, so, when you want to talk about part of an array, you do it as a slice, not via pointer arithmetic.
The key idea here is that you can call old code from new ("strict") code, and strict code from old code. When you get to all strict code, subscript errors should be all checkable.
[1] http://www.animats.com/papers/languages/safearraysforc43.pdf
The reason I'm fairly confident of that assessment is I've had similar experiences with D when the syntax for something was too complex. Early on, the syntax for lambdas was rather clunkly. Everyone either hated it, or insisted that D didn't even have lambdas. Greatly simplifying the syntax was a revelation, suddenly D had lambdas and they became used everywhere.
Syntax matters a great deal.
int read(int n; int fd, char (&buf)[n], size_t n);
is a bit bulky. The initial "int n;" is a little used GCC extension. Allowing int read(int fd, char (&buf)[n], size_t n);
is an option. "n" is used before it is declared, which is strange for C. This is only a problem because of the UNIX idiom that buffer pointer comes before size in most system calls. char (&buf)[n]
is also a bit bulky, but that, too, is forced by C/C++ tradition. char &buf[n]
would be an array of refs, and char buf[n]
would be an array passed by copy.There have been many, many attempts to "fix" C in a non-backwards compatible way. The result is always a new language. It's the backwards compatibility that's hard.
int read(int fd, char buf[n], size_t n);
The following guarantees your array is not modified (but it's still passed by pointer). int write(int fd, const char buf[n], size_t n);
The following emulates passing by value: void foo(const int buf[n], size_t n) {
int tmp[n];
memcpy(buf, tmp, n);
}
Alternatively: void foo(const int buf[n], size_t n) {
int tmp[n];
arrcpy(buf, tmp); // may or may not check bounds
}
Maybe this would render the proposition less useful, but it would already help. Here's for instance authenticated encryption from Monocypher, my crypto library: void crypto_lock_aead(uint8_t mac[16],
uint8_t *cipher_text,
const uint8_t key[32],
const uint8_t nonce[24],
const uint8_t *ad,
size_t ad_size,
const uint8_t *plain_text,
size_t text_size);
It is not crystal clear that `text_size` is referring to the size of both the `plaintext` and the `cipher_text`. With something like your proposition, I could write this instead: void crypto_lock_aead(uint8_t mac [16],
uint8_t cipher_text[text_size],
const uint8_t key [32],
const uint8_t nonce [24],
const uint8_t ad [ad_size],
size_t ad_size,
const uint8_t plain_text [text_size],
size_t text_size);
That way, the size of each buffer is crystal clear. Bonus: a sanitizer can check that I don't overflow my bounds (and I love sanitisers for stuff as sensitive as a crypto library).Pascal's compiler was smaller and it worked in 16bit machines no problem.
Maybe C's base library was bigger? I'm not sure
I highly disagree with this. One of the advantages of conflating pointers with arrays is an obvious and very consistent way of indexing and slicing on the entire language that has minimal syntactic baggage.
Edit: That is more useful if you have function overloading, or templates, to avoid touchy ambiguities. It's still a slightly useful distinction to have in C, just for human readability.
There's nothing stopping you from simply doing it. With a couple of macros the whole thing can just be a header file.
True, it doesn't take you all the way there (you'll still need to manually check array access to make sure they don't go over), but it's a start. And those manual checks can be a macro as well, to make it easy to add them where needed.
[1]: http://man7.org/linux/man-pages/man3/malloc_usable_size.3.ht...
Probably the most damning problem is none of them are able to interoperate.
However, I'd recommend people who deal heavily with multidimensional arrays but couldn't sacrifice the low-level C environment for a dynamic language to consider using the ISO_C_BINDING of Fortran 2003. It provides fully C compatible native types, and can be compiled together with C (you get gfortran from GCC anyway).
C arrays know their length, it's always `sizeof(arr) / sizeof(*arr)`. It's just that arrays become pointers when passed between functions, and dynamically-sized regions (what is an array in most other languages) are always accessed via a pointer.
"It's just that arrays become pointers when passed between functions"
Oh, is that all?
Did you read the article, or the comment you're responding to? They point out the cost of "just" doing that.
int fred(int a[10]) {
return a[11];
}
It compiles without error with gcc and clang, even with -Wall. The code generated by clang is: mov EAX,02Ch[RDI]
ret
i.e. buffer overflow, even though the array size is given. Compile the equivalent DasBetterC program: int fred(ref int[10] a) {
return a[11];
}
fred.d(2): Error: array index 11 is out of bounds a[0 .. 10]
And the 32 bit code generated (when using 9 instead of 11 so it will compile): mov EAX,024h[EAX]
ret void foo(int arr[static 10]);
It cannot check whether a passed pointer will point to enough space, but the compiler can warn you if you pass a fixed-size array of a smaller size.When the size of such arrays is computed at compile-time via macro definitions, that feature is quite handy.
I pretty much never use static arrays because if I do I always without fail get a bug report when some user exceeds it.
Not among good programmers (unless the array is immutable).
Nice to see it get such a nice response!
Also, nobody I know constantly looks at a cheat sheet. The concepts motivating the various slice transformations get ingrained pretty quickly.
It's a bit like the first language you learn. For someone from the Latin family of languages, Mandarin's verbal & written structure might seem hard, but for native Chinese, it's second nature.
The main source of bugs in C to me would be pointer arithmetics.
You're commenting without reading the article?
But serious question, why even bother with this one fix?
The only reason for the fix is so to make it more difficult to make errors.
Fix arrays, then you would fix null pointer, then you might add objects, templating/generics to support a good collections library, rtti, and before you know it you are creating another one of c++, D, go, java. And we already have those.
C paved the way. Why not let it be the end of it?
P.S. thank you for everything you have done with D. I read in another HN thread about Better C, and it convinced me that D is the language I should be investing my time in learning and using.
A good tool to check out, but which hasn't been promoted much because it's new, is dpp[1]. You can directly reference C header files in your D code. With that, betterC mode becomes a viable option for adding to an existing C project.
There were OSes being written in better languages outside Bell Labs, had it been allowed to sell UNIX instead of giving it away for a symbolic price to universities, and the historical outcome would have been completely different.
It can even be funkier, like 12 bits in a char
https://stackoverflow.com/questions/2098149/what-platforms-h...
It's a mess
Whenever an implementation says so. There are now few machines where the addressable unit is not 8 bits though, which is why languages like D and Java can get away with not supporting them.
> I presume it'd still have a sizeof() of 1 though.
The language standard requires that.
No. There were problems when 64-bit CPUs came along, but they have been pretty much ironed out, and don't nearly compare to the pervasive bugs that Walter mentions in his article.
Arrays losing dimensionality when passed through functions is a pain every now and then.
The real 'mistake', is programmers not stating their intention explicitly.
Thought of this as I was reading the article.
> There's probably an advantage in the two types being the same.
Not really.
They are still as insecure as any traditional C string and memory function.
Yes, they sorted out the issues about then a string always gets its null terminator.
However given that buffer and size are still two different parameters, the issue of mixing up the values is still present.
So...30+ years later, we decide Pascal was right. Just saying. Shoutout to FreePascal/Lazarus!
The only thing I really liked from Pascal were the nested functions, which I put in D.
Is it really just about the issues with begin..end and verbosity ???
Object Pascal was originally developed by Apple and adopted by Borland into Turbo Pascal 5.5, which then started to adopt ideas from C++.
In the PC world, Turbo Pascal was the king of Pascal dialects, for a Pascal compiler it was more relevant to be Turbo Pascal compatible than ISO Extended Pascal (the standard revision that fixed the issues with ISO Pascal).
Borland switched focus to the enterprise, leaving the hobby developers behind, increasing the prices of their compilers to enterprise tools range, and then went through an identity crisis with Inprise and Codegear.
The Kylix attempt to bring Delphi and C++ Builder into Linux was never that serious.
They lost key people like Anders to Microsoft, nice story of events why he left in this interview.
https://behindthetech.libsynpro.com/001-anders-hejlsberg-a-c...
So most of us moved away, on the mid-90's C++ was an welcoming home for Object Pascal refugees.
Had the mix of OOP and procedural programming, thanks to classes and overloading it was possible to write type safe abstractions, RAII better than Object Pascal had, and even if the standard was a few years away, every compiler had a nice framework that would relive us from the pain of dealing with plain old C arrays and strings.
And for those moments that we were forced to deal with C APIs, being almost copy-paste compatible with C helped. Which incidentally is one of the pain points in modern C++.
This in the PC world.
On the Mac, Apple decided to cater to the UNIX crowd and started to move away from Object Pascal.
http://basalgangster.macgui.com/RetroMacComputing/The_Long_V...
http://basalgangster.macgui.com/RetroMacComputing/The_Long_V...
http://basalgangster.macgui.com/RetroMacComputing/The_Long_V...
Outside PC and Mac worlds Object Pascal was hardly used.
Max Weinreich said "A language is a dialect with an army and navy".
On the context of systems programming languages, "A systems programming language is a language with an OS".
If it isn't tied to an OS SDK there will be always attrition why use it at all.
But, my original point was about the pros and cons of the language, itself. Object Pascal seems to solve some major pain points experienced with other languages (specifically strings and dynamic arrays, but also others), but doesn't get copied or adopted as much as one would think. Instead, newer languages keep copying the same bad ideas that kill performance and/or limit the versatility of the language.
When it came out I was disappointed that they adopted and interpreter, followed by JIT with 1.2, leaving to commercial third parties the AOT compiler toolchain.
I was then double disappointed with .NET, due to the NGEN/JIT mix, because NGEN was no match for a proper AOT compilation, just for faster startups.
And it took them Singularity, Midori, to finally arrive at CoreRT and .NET Native, and still it only applies to certain deployment scenarios.
Back then it wasn't only Delphi, there was Oberon, Component Pascal, Eiffel.
But they were all commercial and then around the same time FOSS started to pick up steam, Kylix was very badly managed, and due to its UNIX roots everyone was mostly writing GNU tools in C, which wasn't actually that much used in the PC world where we were already quite happily using OWL, VCL and MFC.
At least Pascal style syntax is fashionable again.
As for Kylix, I simply think that there was no way that it was going to work on Linux. IOW, trying to do a GUI-based development tool on Linux was a bad idea from the start. They couldn't even nail down a few distributions very well - it was a constant moving target...
Assembler programs are very tedious to write, and so they tend to be rather small. You don't get any help from the non-existent compiler for even simple HLL features like static type checking.
For stuff like SIMD we have compiler intrisics, loop vectorization and compute shaders.
Even when writing straight Assembly, unless we are talking about a PIC class processor, the amount of opcodes and their behaviors across a CPU family are so broad no human manages to fit the instruction manuals on their head, several thousand pages long.
Modern CPUs, aren't a Z80 on a Speccy where we could fit the opcodes and memory map on our head.
So it is constrained to things like a few hundred KB PICs, people writing compiler backends, software decoding for video/audio codecs or kernel level drivers, a very specialized set of tasks, the very definition of niche.
- there are only two CPU's nowadays to code for, intel and ARM, so if you don't need portability across operating systems, that is manageable; and with cpp(1) macros, one might even be able to write code which would assemble across operating systems on the same processor (for example illumos and GNU/Linux on intel);
- even if you use the simplest of instructions, assembler code written by a human will always beat the compiler - ALWAYS!; don't take my word for it - try it out for yourself.
On slower / older systems, assembler is the only way to get the required speed. And it's fun, really lots and lots of fun to code in assembler. Not to mention that it's easy. Lots of people do it, just look at the demo / cracking scene.
You're just hanging around in the wrong circles, with the wrong crowd if you think assembler is niche.
You can keep your Assembler, my crowd rather uses C++ with intrisics when we need performance.
Being part of the Demoscene was cool, like when Amiga 500 actually mattered.
As someone who is part of the C++ experts group and maintains the GCC compilers at my organization, I will “keep my assembler” over C++ any day of the week.
But in one thing I’m starting to think that you actually might be right: it is I who seem to be hanging around with the wrong crowd by being here on “Hacker News”, where actual hackers (in the MIT sense of the word) are in very short supply. The more I read what people write and how they think around here, the more n-gate.com is right, critique by critique, point for point. As it stands right now, this site is a gross misnomer.
Ignorance is not a virtue.
> The solution is to learn assembler first, then move on to C
So if you learn assembler first, then suddenly C has fat pointers, strings aren't NUL-terminated, and the massive code base written by millions of programmers doesn't contain any buffer overflows?
It's not ignorance, but knowledge: since I know assembler, pointers are no big deal in C. For one who does not understand how the machine functions, they are a big deal. I don't understand why understanding assembler is so hard for so many people since machine code is so simple.
So if you learn assembler first, then suddenly C has fat pointers, strings aren't NUL-terminated, and the massive code base written by millions of programmers doesn't contain any buffer overflows?
No, suddenly your code doesn't have those problems any more because you actually understand what's going on and how it works. It's not magic. Except apparently on "Hacker News", where hackers seem to be in very short supply.
"I don't understand" is clearly a statement of ignorance and not knowledge. The ignorance you expressed was about why people make an issue of something. That ignorance could be dispelled if you actually read and attempted to understand their points, but that requires qualities like humility and intellectual honesty.
As a top class programmer who wrote his first ASM program in 1967, was on the C Standards committee, and has programmed at every other level, I will simply smh at the naivety and point missing of your comments, and avoid engaging you further. Ta ta.
I’m interested about getting to the bottom of fighting to master assembler because that’s the real issue here, everything else is overhead. Ta ta!
There's nothing whatsoever wrong with C. The problem are programmers who grew up completely and utterly disconnected from the machine.
I am from that generation that actually did useful things with machine language. I said "machine language" not "assembler". Yes, I am one of those guys who actually programmed IMSAI era machines using toggle switches. Thankfully not for long.
There is no such thing as an "array". That's a human construct. All you have is some registers and a pile of memory with addresses to go store and retrieve things from it. That's it. That is the entire reality of computing.
And so, you can choose to be a knowledgeable software developer and be keenly aware of what the words you type on your screen actually do or you can live in ignorance of this and perennially think things are broken.
In C you are responsible for understanding that you are not typing magical words that solve all your problems. You are in charge. An array, as such, is just the address of the starting point of some bunch of numbers you are storing in a chunk of memory. Done. Period.
Past that, one can choose to understand and work with this or saddle a language with all kinds of additional code that removes the programmer from the responsibility of knowing what's going on at the expense of having to execute TONS of UNNECESSARY code every single time one wants to do anything at all. An array ceases to be a chunk-o-data and becomes that plus a bunch of other stuff in memory which, in turn, relies on a pile of code that wraps it into something that a programmer can use without much thought given.
This is how, for example, coding something like a Genetic Algorithm in Objective-C can be hundreds of times slower than re-coding it in C (or C++), where you actually have to mind what you are doing.
To me that's just laziness. Or lack of education. Or both. I have never, ever, had any issues with magical things happening in C because, well, I understand what it is and what it is not. Sure, yeah, I program and have programmed in dozens of languages far more advanced than C, from C++ to APL, LISP, Python, Objective-C and others. And I have found that C --or the language-- is never the problem, it's the programmer that's the problem.
I wonder how much energy the world wastes because of the overhead of "advanced" languages? There's a real cost to this in time, energy and resources.
This reminds me of something completely unrelated to programming. On a visit to windmills in The Netherlands we noted that there were no safety barriers to the spinning gears within the windmill. In the US you would likely have lexan shields protecting people and kids from sticking their hands into a gear. In other parts of the world people are expected to be intelligent and responsible enough to understand the danger, not do stupid things and teach their children the same. Only one of those is a formula for breeding people who will not do dumb things.
Stop trying to fix it. There's nothing wrong with it. Fix the software developer.
Oh yeah; social construct, I would say, like gender.
> I am from that generation that actually did useful things with machine language.
Unfortunately, most of them are undefined behavior in C.
> You are in charge.
Less so than you may imagine. You're in charge as long as you follow the ISO C standard to the letter, and deviate from it only in ways granted by the compiler documentation (or else, careful object code inspection and testing).
Despite what many might believe the universe didn't come to a halt when all we had was C and other "primitive" languages. The world ran and runs on massive amounts of code written in C. And any issues were due to programmers, not the language.
In the end it all reduces down to data and code in memory. It doesn't matter what language it is created with. Languages that are closer to the metal require the programmer to be highly skilled and also carefully plan and understand the code down to the machine level.
Higher level languages --say, APL, which I used professionally for about ten years-- disconnect you from all of that. They pad the heck out of data structures and use costly (time and space) code to access these data structures.
Object oriented languages add yet another layer of code on top of it all.
In the end a programmer can do absolutely everything done with advanced OO languages in assembler, or more conveniently, C. The cost is in the initial planning and the fact that a much more knowledgeable and skilled programmer is required in order to get close to the machine.
As an example, someone who thinks of the machine as something that can evaluate list comprehensions in Python and use OO to access data elements has no clue whatsoever about what and how might be happening at the memory level with their creations. Hence code bloat and slow code.
I am not, even for a second, proposing that the world must switch to pure C. There is justification for being lazy and using languages that operate at a much higher level of abstraction. Like I said above, I used APL for about ten years and it was fantastic.
My point is that blaming C for a lack of understanding or awareness of what happens at low levels isn't very honest at all. The processor does exactly what you, the programmer, tell it do to. Save failures (whether by design or such things as radiation triggered) I don't know of any processor that creatively misinterprets or modifies instructions loaded from memory, instructions put there by a programmer through one method or another.
Stop blaming languages and become better software developers.
Sure.
Only problem is, all you have to do is change some code generation option on the compiler command line and millions of lines of code now produce different instructions. Or, keep those options the same, but use a different version of that compiler: same thing.
> The processor does exactly what you, the programmer, tell it do to.
Well, yes; and when you're doing that through C, you're telling the processor what to do via sort of autistic middleman.
C is not the low level; you can understand your processor on a very detailed level and that expertise won't mean a thing if you don't understand the ways in which you can be screwed by the C language that have nothing to do with that processor.
I suspect that you don't know some important things about C if you think it's just a straightforward way to instruct the processor at the low level.
> Languages that are closer to the metal require the programmer to be highly skilled and also carefully plan and understand the code down to the machine level.
C isn't one of these languages. (At least not any more!) It's considerably far from the metal, and requires a somewhat different set of skills than what the assembly language coder brings to the table, yet without entirely rendering useless what that coder does bring to the table.
It is the responsibility of a capable software engineer to KNOW these things and NOT break code in this manner.
You are trying to blame compilers and languages for the failure of modern software engineers to truly understand what they are doing and the machine they are doing it on.
If you truly understand the chosen language, the compiler, the machine and take the time to plan, guess what happens? You write excellent code that has few, if any bugs, and everyone walks away happy.
And you sure as heck are not confused or challenged in any way by pointers. I mean, for Picard's sake, they are just memory addresses. I'll never understand why people get wrapped around an axle with the concept.
I wonder, when people program in, say Python, do they take the time to know --and I mean really know-- how various data types are stored, represented and managed in memory? My guess is that 99.999% of Python programmers have no clue. And I might be short by a few zeros.
We've reached a moment in software engineering were people call themselves "software engineers" and yet have no clue what the very technologies they are using might be doing under the hood. And then, when things go wrong, they blame the language, the compiler, the platform and the phase of the moon. They never stop to think that it is their professional duty to KNOW these things and KNOW how to use the tools correctly in the context of the hardware they might be addressing.
I've also been working with programmable logic and FPGA's, well, ever since the stuff was invented. Hardware is far less forgiving than software --and costly. It forces one to be far more aware of, quite literally, what ever single bit is doing and how it is being handled. One has to understand what the funny words one types translate into at the hardware level. You have to think hardware as you type what looks like software. You see flip-flops and shift registers in your statements.
This is very much the way a skilled software developer used to function before people started to pull farther and farther away from the machine. It is undeniable that today's software is bloated and slow. Horribly so. And 100% of that is because we've gotten lazy. Not more productive, lazy.
Nobody is saying that it's a acceptable for an engineer to screw up and then blame it on the tools (compiler, slide rule, calculator, ...).
However, if something goes wrong in your work, it's foolish not to recognize the role of the tools, even though it's not acceptable to blame them as a public position.
As objective observers of a situation gone wrong in engineering, we do have the privilege of assigning blame between people and tools. Tools are the work of people also. The choice of tools is also susceptible to criticism. We have to be able to take an objective look at our own work.
Read the C Standard. (Do you even understand that it defines an abstract machine? Do you have any idea what an abstraction is?)
For example, reading the processor data book to understand it, the instruction set and how it works could be crucially important in certain contexts. I would not expect someone doing Javascript to do this but how many have studied the virtual machine in depth?
Don't confuse being lazy with problems with languages and compilers.
I once worked with on a project that needed specialized timing in relation to high speed (well, 38.4k) RS422 communications. I don't remember all of the details, it's been decades. I remember one of the engineers came up with a super clever way to trigger the time measurement and actually measure it. Rather than using a UART he bit-banged the communications and actually used the serial stream for timing (meaning the one's and zero's). It worked amazingly well. If I remember correctly that was a Z80 processor with limited resources.
This is the least intelligent and least intellectually honest hackneyed phrase on the internet. In this case it's a complete non sequitur. It would tell me a lot about you if you hadn't already made it evident. Over and out, forever.
But there's no arguing with extreme ignorance coupled with extreme unwarranted arrogance.
And if you (plural) are an ENGINEER, it is your JOB to KNOW these things and prevent them from happening.
I get the sense that the term "software engineer" has been extended so far that we grant it to absolute hacks who know nothing about what they are doing and what their responsibilities might be. Blaming a language, compiler and machine are perfect examples of this.
True engineering isn't about HOPING things will work. It is about KNOWING things will work. And testing to ensure success.
I've been involved in aerospace for quite some time. People can die. This isn't a game. And it requires real engineering not "oh, shit!" engineering that finds problems by pure chance. Sadly, though, we are not perfect and things do happen. It isn't for lack of trying though.
That's nice; not all engineering is aerospace and not all aerospace processes are always appropriate everywhere else.
Even in aerospace, still I don't want to write code that depends on knowing exactly how the compiler works. I will write code mostly to the language spec. Then treat the compiler as a black box: obtain the object code, and verify that it implements the source code (whose own correctness is separately validated).
Safety is not treated the same way regardless of project. For instance, an electronic device that has a maximum potential difference of 12V inside the chassis is not designed the same way, from a safety point of view, as one that deals with 1200V.
Your parent comment is utterly irrelevant. The conversation is about the C language and the perception some seem to have that it has problems. My only argument here is that a capable software engineer knows the language and tools he or she uses and has no such problems, particularly with a language as simple as C. Things like pointer "surprises" are 100% pilot error, not a deficiency of the language itself.
>As an example, someone who thinks of the machine as something that can evaluate list comprehensions in Python and use OO to access data elements has no clue whatsoever about what and how might be happening at the memory level with their creations. Hence code bloat and slow code.
Not having to care about details that aren't contextually important is a good thing. When someone is constrained more by development time than by computational resources, working in a high level language means you're explicitly shunting low level concerns so you can spend more time dealing with domain logic.
There are many situations where finishing something faster, which will run 10x slower and use more memory, is a worthwhile tradeoff.
You might be reading far more into my comments than what they were intended to address. Namely that blaming languages for the failings of software engineers is dishonest. A true software engineer will know the chosen tools and languages and use them appropriately. Blaming C for pointer issues is dishonest and misguided. There's nothing wrong with the language if used correctly.
There's also no such thing as a computer, or memory, or operating systems ... they're all just a bunch of molecules.
I too am from the generation before people understood the power of abstraction ... but I'm intellectually honest and so I managed to learn.
> Fix the software developer.
Which one?
OK. Prove it. And you have to do it without laying out a set of rules and conventions that might allow us to interpret a list of bytes as an array.
An array is a fabrication by convention. At the simplest level it is a list of numbers in memory. Adding complexity you can store additional numbers that indicate type size and shape. Adding yet more complexity you can extend that to be lists of memory addresses to other lists of numbers, thereby supporting the concept of each array element storing more than just a byte or a word. And, yet another layer removed you can create a pile of subroutines that allow you to do a bunch of standard stuff with these data structures (sort, print, search, add, subtract, trim, reshape, etc.).
Nowhere in this description does an array exist. There were experimental architectures ages ago that actually defined the concept of arrays in hardware and attempted to build array processors. These lost out to simpler machines where multidimensional arrays could be represented and utilized via convention and software.
Arrays do not exist. If you land in the middle of a bunch of memory and read the data at that location without having access to the conventions used for that processor or language nothing whatsoever tells you that byte or word is part of an n-dimensional array. The best you can say is "The number at location 1234 is 23". No clue about what that might mean at all.
See also discussion from 9 years ago: https://news.ycombinator.com/item?id=1014533 (47 comments)
They are implemented in the C libraries for OpenBSD, FreeBSD, NetBSD, Solaris, OS X, and QNX.
They have not been included in the GNU C library used by Linux.
Since then I thought, not knowing any amount of C, that strl* was part of the language and available to all.
Imagine if you will my confusion everytime I read people complaining about C being insecure.
Your comment corrected my perception of this. Thank you.
I'm in DevOps field, and this 'make it right from the start' resonates strongly with me. Richard Feynman: Disregard others :)
When it doesn't (work), it is NOT because of a failure in the language; It is because C has (and always will have) the "basic philosophy that programmers know what they are doing;".
Criticising C, is like criticising assembly. What's the point?
If people want to criticise a programming language, then they should always start with C++, not C.
C++ was designed to allow us to develop bigger and more complex programs, and yet, C++ inherited from C?
How stupid was that! But people are happy to give out various awards and medals to the person who made one of the dumbest decisions ever made, in the whole history of computing!
Leave C alone. It's fine. It's C++ that is the problem.
Trying to make C 'foolproof' however, is an excercise in futility, and in any case, can only come about by morphing it into a fundamentally different language.
An argument in this thread, is that you shouldn't be able to pass an array without the argument being passed having some implicit 'size' element associated with it. That is NOT C.
Conflating pointers with arrays, that is C.
Again I feel the need to quote this:
"C retains the basic philosophy that programmers know what they are doing; it only requires that they state their intentions explicitly."
If you don't know what you're doing, don't use C.
C should be considered a 'specialist' language - much like doing brain surgery - if you're doing it, you better know what you're doing, else go be a GP or something.
And, if you're project doesn't absolutely require that you use C, don't use it. Instead, use something that is more 'foolproof'. (and I don't mean C++!!!)
D should focus less on being a better C, and more on being a replacement for C++. Then, I might take D more seriously.
No attempt to morph C (i.e. the language, not the library) into something else will ever succeed.
Leave C alone!
So we would like to improve our foundations, to move UNIX derived OSes to some kind of safer C, instead of having it be the backdoor of the whole security infrastructure.
What's the downside?
> by changing it into D?
That comment is tendentious, hyperbolic, and generally credibility-damaging.
> Concentrate on that D language.
You think he doesn't?
I would remove go from that list, and add rust and zig.
EDIT: besides, you should start new projects in Rust anyway, because it takes security to whole other level. C did a great job, but it's a bit old. :)
Thanks, I almost forgot what website I was on for a second.
That is literally the meaning of "fat pointer", and the linked article even explains it.