But that's the problem with C. It looks like a simple language with a small number of straightforward constructs, when in fact it has very poor predictability.
Real-world results are implementation and platform-dependent, which is a polite way of saying they're unhelpfully unpredictable.
Coding standards will help, but they only get you so far because some ambiguities will always be there.
Basically C is a half-assed wrapper around assembler - useful in its day, but machine-dependent and outdated now.
You can keep the syntax and all the useful parts of the standard library in your head.
A lot of modern languages sacrificed simplicity for features no one (with enough skill) asked. Even with Python, I can only keep some part of built-in functions and methods together with syntax in my head.
For a large part of standard library, I absolutely need IDE features. Especially since Python is an object oriented language where every data instance has a bunch of methods attached.
you drug in "offshored code delivered by sweep shop consultancies", got a response, and then move the goalposts.
go away.
You were the one asserting that it only happens in sweetshops, hence my remark wasn't valid, thus I decided to broaden the scope of people failing to write C.
Those mythical C developers that are yet to be found after 50 years searching for them.
But, it is, in fact, possible for good developers to write incorrect C code. Noone's claiming that's not the case. It's just a lot less likely for someone who actually knows the language and the codebase.
I'm in the process of moving an old codebase from a compiler with a very loose adherence to an old specification to an ARM chip that is using GCC for the compiler. Enabling -Wall to catch all issues that the compiler highlighted, 16-bit ints to 32-bit ints, and quite a few ternary operator issues.
None of these were issues on a system that's been running in production for many years but porting to a new architecture will mean there’s a lot of extra testing to be done just to make sure we catch all the edge cases in the code.
I like the C language and have been using it for close to 30 years now. The main part of the language can be kept in your head but there is so much more to working with it than just the base language.
Not to mention the chances of so many things to go wrong, in writing C code. Its like climbing a mountain without a safety gear.
If you're implementing custom data structures, you'll likely be able to make them optimal for your use-case, both from ease-of-use and performance. (except hashmaps. those do suck to write)
Membership testing? Use sets. Data storage? Use classes with slots or tuples/namedtuples. Custom switch statement? Use the new match/case keywords or just fall back to plain if/elif towers.
It's so much simpler.
Like in JavaScript, you can understand and remember how double equals work, if you try hard enough... Or you can just decide to always use triple equals and voila, your problem is solved once and forever.
But I don't write C, so take this with a grain of salt.
We aren't computers, and that's not what we mean by having things in our head.
"The patient needs more dozing/dosing."
There’s no need to have any subset of the language in memory, this is the whole standard.
No, you're wrong. "Undefined Behavior" (UB) in the context of C has a very different meaning from the colloquial meaning "not specified". The meaning of UB is explicitly defined in the C standard, and the C standard explicitly lists certain program execution patterns as UB.
> If a ‘‘shall’’ or ‘‘shall not’’ requirement that appears outside of a constraint is violated, the behavior is undefined. Undefined behavior is otherwise indicated in this International Standard by the words ‘‘undefined behavior’’ or by the omission of any explicit definition of behavior. There is no difference in emphasis among these three; they all describe ‘‘behavior that is undefined’’
And I would never excrete such a line of code without automatically thinking "hmmm.. does this really do what I think it does?"
This is an undefined operation in almost all programming languages.
There is no intrinsic reason why a[i] = i++ has to result in undefined behavior.
I think the second example may contain UB on a <=32-bit architecture (right shift by a value greater or equal to the number of bits), or at least this is UB in C++. On a 64-bit architecture it would be fine (but the result would not be 0).
if i is signed, this is undefined. You'd think it's just a loop though.
while (1) for(;;)int this_is_a_fairly_long_variable_name_1;
int this_is_a_fairly_long_variable_name_2;
Why? Because section 5.2.4.1 requires implementations to support at least 32 significant characters in external identifiers and section 6.4.2.1 says "If two identifiers differ only in nonsignificant characters, the behavior is undefined."
These ones may be more "obvious", but:
Given e.g. int a[4][5] then accessing a[1][7] is undefined even though one might assume it's like accessing a[2][2]
x << -3 is undefined behaviour though logically one might assume it should be equivalent to x >> 3
And the classic, very common newbie use of "fflush(stdin);" is also undefined.
Not saying anything about the other things, but this (in my opinion) is something that should raise eyebrows. At least to me personally that looks rather weird.
(It's also not unique to C)
It's not obvious which actions are atomic.
How can you identify legal C without understanding undefined behavior? Or rather, what is the list of UB that nay professional C developer should be expected to know?
> If someone has some examples that a reasonable person would write, specifically examples that look like the simplest way to do something, then I’d love to hear them.
Basically all OOP-style C. For instance, the Windows macro
#define CONTAINING_RECORD(address, type, field) ((type *)( \
(PCHAR)(address) - \
(ULONG_PTR)(&((type *)0)->field)))
The Linux kernel is also an excellent source of these.Also, serialization/deserialization libraries are an excellent source for UB. And huge amounts of code depends on UB by using unsigned char arrays for storage of other types.
Furthermore, in practice, C code assumes most of the following:
- null pointers are represented as 0, e.g. if(ptr).
- pointer provenance does not exist
- char and unsigned char are 8-bit bytes
- two's-complement arithmetic
- page-based memory protection
- no types have trap representations
- the values of base types, besides floating point, do not have multiple representations.
- etc
Off the top of my head, as a non professional C programmer:
- can’t read from uninitialized memory
- one malloc = one free
- signed overflow is undefined
- pointers to stack variables can’t outlive the stack variable itself
- can’t dereference NULL pointers
(Obviously take this with a massive grain of salt, there are almost certainly important ones that I’m forgetting.)
> null pointers are represented as 0, e.g. if(ptr).
If(ptr) is guaranteed by the standard to work. Memset(ptr, 0, len) is not (and is very broken in other ways as well).
> char and unsigned char are 8-bit bytes
They’re guaranteed to be one byte, where a byte is at least 8 bits. So unless you’re trying to simulate %256 with an implicit conversion, or relying on unsigned overflow to turn 0xFF into 0x00, then you’ll be fine.
Much of the stuff programmers are just used to being "part of C" is actually from a library. For example NULL is just a macro that some library provided, C's "null pointers" are just the literal zero for a pointer, even if your hardware does not represent "null pointers" by an all-zeroes value.
Unless we are speaking about the very niche use case of allocating a big blob of static memory on the global segment.
Most of the stuff done with ISO C + language extensions/external Assembly, can be as easily done in safer languages.
Language extensions and external Assembly aren't a C priviledge.
But C doesn't come with functions like memcmp() and memcpy() or memmove() either and those are pretty fundamental even if you're on a tiny constrained system that can't aspire to stuff like "strings" and "allocation" you've probably heard of data structures and copying.
Or do you mean you can't write your own malloc that complies with the OS's memory management in C?
Because I'm pretty sure you can write a malloc in ISO C when you aren't dealing with an OS (which is usually when you need to write a malloc).
Same deal as it being impossible to implement BSD sockets in ISO C because of the sockaddr funny business.