Some things every C programmer should know about C (2002)
web.archive.org
web.archive.org
You can certainly have a pointer to a struct member, so this isn't the reason why you can't have pointers to bitfields. The reason is that bits are not addressable, fullstop.
The restriction in C can only be explained as a limitation of the language itself -- although probably motivated by the implementation complexity it would require, for a niche use case.
Nowadays a byte is conventionally eight bits, especially for measures like "megabyte", but the term octet is often used to avoid ambiguity. Commonly they're used for pointers, yet often only words are addressable by machine instructions (e.g. many ARM instructions take a byte address yet raise a hardware exception on use of unaligned addresses).
"byte: addressable unit of data storage large enough to hold any member of the basic character set of the execution environment"
Hence why the type that corresponds to it is "char"! Beyond that, the only thing that kinda sorta implies that it's the smallest addressable unit is the definition of CHAR_BIT:
"number of bits for smallest object that is not a bit-field (byte)"
This might be why the code space alphabet is defined by the standard so it will at least put an emphasis on 8 bits == 1 byte.
The question, though, was about whether it's the minimum addressable unit of memory. In the C memory model, it is, but by implication - you can't have two pointers that compare non-equal, but differ by less than 1, so a type with sizeof==1 is by definition the smallest you can uniquely address. However, the C memory model doesn't have to reflect the underlying hardware architecture.
The CPU itself used 32-bit addresses to access machine words, the size of which was determined by what was being accessed. External memory was limited to 32-bit. Internal memory had regions that could be 32-bit, 40-bit, or 48-bit. An address increment of 1 would thus move by that many bits.
Mercury Computer Systems shipped a byte-oriented port of gcc. Pointers to char and short were rotated and XORed as needed to reduce incompatibility. Pointers to larger objects were in the hardware format. This allowed a high degree of compatibility with ordinary software while still running efficiently when working with the larger objects. There was also a 64-bit double, unlike the 32-bit one in the other compiler. Data structures were all compatible with PowerPC and i860, allowing heterogeneous shared memory multiprocessor systems.
Even if the ISA only allows word- or dword-aligned loads from memory, the addresses still typically enumerate bytes, not words or dwords.
Based on a quick summary of the MCS-51 that I googled up, it looks like its memory addressing scheme still assigns addresses to bytes, and has special operations that allow you to further specify a bit offset within that memory address.
There are also instructions which use an addressing scheme which takes an 8-bit bit address, with the 0x00 - 0x7f corresponding to lower memory, and 0x80 - 0xff corresponding to 16 specific registers in the Special Function Register set.
For example, there are machines (some DSPs) that individual octects are not efficiently addressable and usually a C byte in these machines is 16 or 32 bita.
https://www.ralfj.de/blog/2018/07/24/pointers-and-bytes.html
I also happen to very much enjoy this piece on how the C abstract machine has very little in common with modern architecture.
$ grep -rE ':[[:digit:]][[:digit:]]*;' /usr/include/ | wc -l
702
$ uname -a
OpenBSD orville.25thandClement.com 6.6 GENERIC.MP#3 amd64
$ grep -rE ':[[:digit:]][[:digit:]]*;' /usr/include/ | wc -l
276
$ uname -a
Linux alpine-3-10 4.9.65-1-hardened #2-Alpine SMP Mon Nov 27 15:36:10 GMT 2017 x86_64 Linux
$ grep -rE ':[[:digit:]][[:digit:]]*;' /usr/include/ | wc -l
532
$ uname -a
Linux splunk0 5.0.0-36-generic #39-Ubuntu SMP Tue Nov 12 09:46:06 UTC 2019 x86_64 x86_64 x86_64 GNU/LinuxIn fact, many internal coding standards forbid its use because it is poorly defined and compiler-dependent.
Moreover, it does not guarantee atomicity so if atomicity is critical then you want to access the actual bitfield 'manually'.
They are indeed trouble for super-portable interfaces. They also tempt some people into trouble with hardware drivers, but then that would already be trouble anyway due to instruction reordering in CPU or memory operation reordering beyond the CPU. (volatile is neither sufficient nor necessary)
In a normal emulator, none of that is a problem. Reasonable uses of bitfields are compatible between the x86_64 ELF ABI (Linux, etc.) and the Windows ABI. What you gain is C code that is simultaneously concise and readable.
For example, a CPU instruction can be represented in a header file by a union of anonymous structs, each of which contains a padding bitfield and a bitfield for part of the instruction. Access in the C file is then just as easy as for normal struct members. It works for non-CPU hardware too, like DMA descriptors and motherboard registers.
As I said, C's bit fields are often banned in coding guidelines.
ethdev->phy.foo = val; // foo is a bitfield in struct phy
cpu->status.irqlevel = newlevel; // irqlevel is a bitfield in the CPU's status word
No, I don't want a 2-argument macro (or worse, with an implied variable) called something like SET_FOO or SET_IRQLEVEL.
Crummy coding guidelines are not a useful argument against the value of bitfields in the C programming language. Coding guidelines can also ban floating-point types, recursion, or symbol names longer than 6 characters. The problem is not C. The problem is the coding guidelines.
struct a {
struct *a next;
};
struct b {
struct *b next;
};
was illegal because the second declaration of "next" is a redeclaration.I find code dealing with structures that have their members prefixed with a unique prefix much easier to read, usually.
Given the pervasiveness of APIs that contain struct definitions nowadays, it's also understandable that the unique-constraint was removed.
I have not seen a case of it creating problems where there weren't already plenty of other problems with the code and metaprograming had to be abandoned anyway. But it does create problems.
Then again, OCaml still has module-scoped scoped record field names today, even though it also uses the dot syntax.
#define st_mtime st_mtimensec.tv_sec
POSIX preserves many prefixes for this reason. C99 brought anonymous unions (a Plan 9 invention), which makes many of these API tricks possible without using macros.I've known for a long time that Microsoft headers make use of anonymous unions, which makes it a rare case where they have implemented a recent ISO C feature ahead of time. I seem to recall GCC in the same timeframe would accept this with a warning that it's nonstandard. [I guess not anymore.]
struct foo { union { int i; double d; }; }
Visual Studio also supports anonymous tagged members, union bar { int i; double d; };
struct foo { union bar; };
Is the latter more common in Microsoft land? I've always thought the latter were more useful and wished they were standardized. Anonymous tagged fields make it possible to reuse plain struct and union definitions, permitting a simple form of inheritance; whereas with untagged fields you have to rely on macros for sharing definitions.GCC (https://gcc.gnu.org/onlinedocs/gcc/Unnamed-Fields.html) and clang (confirmed macOS clang-1001.0.46.4) support anonymous tagged members with -fms-extensions, but I can't bring myself to make use of it as I still try to keep many of my projects working with Sun Studio and otherwise try to avoid non-standard features.
I don't know if I have an authoritative sample on that. I saw it exactly once when I worked at MS and I thought it was weird. But I guess it aids in the "psuedo-inheritance" game that people play with structs and their first member being a "base class". Like GObject or gtk+ do. Or struct sockaddr, which duplicates the first few fields.
struct a *next;
or was putting the asterisk before the type name legal back then? > There is an implicit "x != 0" in: if (x), while (x), ... etc.
> An explicit "x != 0" in these contexts serves no semantic purpose.
> And "x == 0" in these contexts might be better written as "!x".
I disagree with this advice because the idea is not portable across languages.For example, if we let the variable x hold the integer value 0, then:
Java 'if (x)' is a compile-time type error.
Python 'if x:' is falsy, similar to C.
Ruby 'if x' is truthy!So either no idioms are correct, or it is correct to apply the idioms pertaining to the language of choice.
Consider how you would read it: "if grass is green ..." is effectively shortened to "if grass ...".
This makes no sense.
I always use explicit conditions and consider "if(x)" and antipattern and have done since I was learning C back in the mid eighties.
if (grass_is_green) { ... }
vs if (grass_is_green != 0) { ... }
or if (grass_is_green == 1) { ... }
The one with the implit comparison to zero is more readableif (bool_name == TRUE)
There is no ambiguity, and it forces you to handle exactly equal to 1 and not equal to zero separately.
if (bool_name == TRUE)
and
if (bit_field_value > 0)
are logically different even if they can be simplified to the same thing.
The worst part is that sometimes these redundancies grow and combine to create even more confounding expressions, frequently resulting in monsters resembling this:
if ( ((((!(var)!= FALSE)) != TRUE)) )
{
return FALSE;
}
else
{
return TRUE;
}The better approach is to always compare against FALSE, because that avoids the ambiguity of == TRUE only catching a single truthy value out of the millions of possible truthy values, contrary to the well-known semantics of C.
If you really need to distinguish between 1 and other truthy values, then just compare against 1 explicitly. Don't muddle the concept of truthiness.
The reason is for writing embedded code that doesn't depend on a specific compiler or architecture, which is important when code that was written in the 80s is still getting used everyday in new safety critical systems, where any kind of standard libraries are forbidden.
I agree it's not the best but it's definitely not uncommon.
The comp.lang.c FAQ touches upon this:
if (isupper(c) != 0) if (isupper(c) != 0)
This code is illegible to me. How do you read it aloud? "if c is upper is not zero" ? It makes no sense. Compare it with the normal way of writing "if(isupper(c))" which is pronounced "if c is upper", which is perfectly clear, grammatical English.The other two leave the reader with the lingering suspicion that the variable takes on two or more values. The latter more so than the former.
Insisting if takes a bool like java does is perfectly acceptable. But just taking a 0 value and making it "true" makes sense only in Bash?
E.g. in Common Lisp, the only falsey value is NIL (the empty list). This is easy to understand, easy to remember, and straightforward in practice because idiomatic Lisp does a lot of list operations. It wouldn't make sense for 0 to be truthy, because that's a legit value, and NIL represents the absence of a value.
I don't know what values are falsey in Ruby, so I can't comment on it.
If I return a number or nil, I want to check `if ret`, not `if type(ret) == "number"`. The latter suggests that it might be some other type‡, the former that it might be missing entirely.
A language which provides a proper Boolean has no need for the association between zero and falseness, it is a conflation of levels which can lead to subtle bugs.
‡ yes, nil is also a type. I mean some other type of type.
timeout = options[:timeout] || :infinity
if zero is falsy, you can be royally screwed with that code, if timeout should be settable to zero (definitely a thing).Jokes aside, polyglot source files are pretty cool, and generally things like abuse of the c++ preprocessor are really neat to see, if completely unwise to do in production.
So if(!strcmp(str1,str2)) is a valid way to check for equality, but I think the negation makes it confusing; it is tempting to intuitively read it as checking if the strings are not equal. if(strcmp(str1,str2) == 0) is the better choice.
Does anyone know for sure whether or not the Arduino compiler deviates from this standard, or ever had, in previous versions?
Does multiplication qualify as a binary operator?
IIRC, I think I was compound multiplying the product of an int and a char into a long (see below), and I had to explicitly cast the char as int to get the right results, much to my surprise, as I had expected the char to get promoted to an int. But I could also have made some mistakes, it's been awhile...
longType *= intType * (int)charType;
EDIT: Found my old code snippet, and it kinda invalidates my original question, what I was actually doing was longType *= charType * charType * charType
which leads the right side to evaluate first without promotion, thus modulating the end result by 256, which was not intended. I then cast all of them as long, but probably (certainly!...?) only one would've been enough.EDIT2: Or at least, two of them have to be casted, else it might evaluate two chars first, modulating the result, before actually promoting to long to multiply with another long. (Or not, because it promotes to the most significant data type on the right side? Ugh... this is why I prefer to be extra verbose.)
I am not exactly certain, but what I think happens is that when you define your charType variable that the upper bytes in the word are still the same random data, so when you multiply without casting, it is treating the char as an int and multiplying the whole 4 bytes. When you cast, it probably just multiplies the lowest byte and carries if necessary. I am also not certain if the compiler would optimize to using a shift operation if one operand is a multiple of 2.
Again, not exactly certain if this is the case, but it's similar to what I've noticed.
char foo;
may be signed or unsigned depending on the architecture and compiler, which may lead to surprising results.
The compiler is configured to follow the -std=gnu++11 variant. Here are the flags: https://github.com/arduino/ArduinoCore-avr/blob/master/platf...
If you use other Arduino-compatible boards (e.g. ARM chips, ESP8266 or ESP32), they will use a different C++ compiler and different settings. See for example:
* SAMD uses 'arm-none-eabi-g++' (https://github.com/arduino/ArduinoCore-sam/blob/master/platf...)
* ESP8266 uses xtensa-lx106-elf-gcc (https://github.com/arduino/esp8266/blob/master/platform.txt)
* ESP32 uses xtensa-esp32-elf-gcc (https://github.com/espressif/arduino-esp32/blob/master/platf...)
So Arduino programming is "just" C++ with additional libraries and an IDE. Though I wrote "just" in quotes because they've done a lot of work to make things easier and approachable for beginners, so I don't want to diminish their work.
I think the answers I'm looking for are buried in those specifications. Gotta check them out, one of these days...
> So Arduino programming is "just" C++ with additional libraries and an IDE.
Considering that flag, I think it's something more (or less) than just plain C++...
It is an 8-bit chip, and that would take a number of cycles more, than to just leave them...?
And depending on the exact expression, it doesn't have to actually do anything differently - these integer conversions describe the expected result of the operation, not how it's performed in assembly. So when you write something like this:
char a, b, c;
...
a = b + c
the spec requires that b+c is treated as an int addition, but the resulting int is then assigned to a char variable, which basically truncates it. Furthermore, if b+c is larger than can fit in a char, then we have signed overflow (assuming that char is signed, which is usually the case), which is undefined behavior. Altogether, this means that the compiler can actually do an 8-bit addition here while remaining within the bounds of the spec.OTOH if the result is assigned to an int variable, then yeah, it'd have to promote them - but that would also be far less surprising to you than if it didn't, no?
Where this can potentially lead to more overhead than you'd expect is use of the temporary as input to another subexpresion. E.g. suppose we have this:
a = (b + c) / 2
Suppose b is 200 and c is 100. The spec requires that the result of this operation is 150, because b+c is performed on ints, and thus won't overflow. If it were an 8-bit addition of signed chars, assuming the typical wraparound behavior on overflow, you'd get 22.Indeed, though, on the Arduino platform, an int is usually 16 bit, but not because that's the chip's native archticture, but because that's the range required. Though I can't say for sure, I highly doubt (and now am actually almost completely certain in my doubts) that it would promote two chars to ints to perform operations on them, and last night also realized why, namely, there are counts for the number of cycles required to do operations on various data types [0].
If what you say about integer promotion is true, there would be no difference in clock cycles between an int and a byte, but there is.
So in essence, no, the Arduino is definitely somehow deviating from spec when it comes to promotion.
[0]: https://forum.arduino.cc/index.php?topic=92684.msg696420#msg...
EDIT: I just realized, I think I need to actually look into the (Atmel's) AVR specs, as the Arduino IDE is basically just a wrapper for that.
But, again, this would only manifest in expressions where you mix operators, or try to assign the result to a variable of a broader type. So (a+b+c) could still be evaluated entirely in 8 bits without becoming non-conforming, if target is a char, and ditto for (a-b-c) etc. So long as you know that only the last 8 bits of the result is all that matters, the promotion can be disregarded. It's only when you apply some other operator - multiplication, division, bit shift, or comparison - to the intermediate result that you can observe the increased width. But that doesn't actually happen all that often.
(Conversely, it might also mean that it is not conforming, but there are relatively few practical cases where it actually manifests.)
There's also the operator(7) manpage for quick reference.
I assumed much of the steps listed happened during parsing, such as processing escape characters and converting newlines, but those are done beforehand.
Perhaps it is because C uses a preprocessor and some macros would not be possible if all the steps were performed while parsing?
I can't remember the case in C, but in C++ the initialisation of a function local static variable is guaranteed to be thread-safe.
In C, all statics can only be initialized with compile-time constants. Thus, they can all be initialized before anything in the program starts running (indeed, there's no code initializing them in most implementations - they are just bytes in the data segment).
In C++, initializers can be arbitrary expressions, which means that an initializer for a global variable can spawn a thread, and that thread might still be running when another global is being initialized with a different expression that is not a compile-time constant. There are no thread safety guarantees wrt those kinds of conflicts, for either globals or locals. But for static locals with runtime initializers, the evaluation of initializer is deferred until execution actually reaches the definition of that variable - and thus C++ has a special provision for thread-safe synchronization if that happens on two different threads concurrently.
you forget the dlopen case :)
Only in C++11 onwards.