How SerenityOS declares ssize_t
awesomekling.github.io
awesomekling.github.io
#if SIZEOF_SIZE_T == 8
typedef int64_t ssize_t
#elif SIZEOF_SIZE_T == 4
typedef int32_t ssize_t
#elif SIZEOF_SIZE_T == 2 // LOL
typedef int16_t ssize_t
#else
#error port me!
#endif
SIZEOF_SIZE_T can be obtained using a script which compiles a test program without executing it.Over they years I used more than one approach, settling on this one:
https://www.kylheku.com/cgit/txr/tree/configure?id=1f902ca63...
Here, in the test program, a DEC macro has been defined which given a constant expression, produces two decimal digits as the initializer for a two-character array. For instance:
DEC(42) -> { '4', '2' } // not exactly: character constants are not used
With this trick we can use DEC(sizeof (size_t)) to get a value like { ' ', '8' } into the portion of some character data, which we can prefix with an identifying string we can look for, like: SIZEOF_SIZE_T= 8
That we can basically grep out. In my configure script, this data is extracted, spaces are removed from it, and it's evaluated directly as shell assignments, so then the values are available in shell variables. #define SIZE_T_BITS 64
#define PASTE3(a,b,c) a##b##c
#define TYPEDEF_UINT(bits,name) typedef PASTE3(uint,bits,_t) name
#define TYPEDEF_INT(bits,name) typedef PASTE3(int,bits,_t) name
TYPEDEF_UINT(SIZE_T_BITS,size_t);
TYPEDEF_INT(SIZE_T_BITS,ssize_t);
Should work for any value of SIZE_T_BITS, even something weird like 128 or 36 (PDP-10 port?), so long as you already have [u]intN_t defined.It does require SIZE_T_BITS to be in bits rather than bytes, as your SIZEOF_SIZE_T is. But surely if your script can calculate SIZEOF_SIZE_T, it can multiply the answer by 8? (Or by CHAR_BIT, if we want to support weird platforms without 8-bit bytes – Lars Brinkhoff's PDP-10 port of GCC 3.2 has CHAR_BIT==9.)
The simple, dumb #if blocks have the virtue is that they are not hostile to simple tooling, like Exuberant Ctags, Cscope, and whatnot. When we ask the editor to jump to the definition of ssize_t, it knows the three possible places where it is defined and serves them up.
This alternative is possible:
#define PASTE3(a,b,c) a##b##c
#define UINT_TYPE(bits) PASTE3(uint,bits,_t)
#define INT_TYPE(bits,name) PASTE3(int,bits,_t)
typedef UINT_TYPE(SIZE_T_BITS) usize_t;
typedef INT_TYPE(SIZE_T_BITS) ssize_t;
I think in this form, there is a good chance the tools will grok the typedefs and index them, because we have not disguised the basic phrase structure.Fun fact - for a while, this was true for the third most popular processor architecture in the world, one that most people haven't heard of.
CSR (the bluetooth chip maker) was originally a spin-out from Cambridge Consultants, and their chips were for years based on a Cambridge Consultants processer design called XAP. As CSR had made billions of devices, XAP was likely right up there in terms of number of processors made.
XAP was unusual in a couple of ways (at least unusual compared to the fairly universal modern design). Here's a fun set of statements that all evaluate to true:
sizeof(uint8_t) == 1 // So far so normal, this is required by the C standard
sizeof(uint16_t) == 1 // Surprise! BTW, uint16_t is defined as "unsigned int"
sizeof(void *) == 1 // Weird
sizeof(void (*)(void)) == 2 // Yes, function pointers are not the same size as data pointers
This is all completely legal according to the C standard. Fun fact, the size of a byte isn't "8 bits", it's technically "the minimum addressable size on your platform"[1]. On every modern processor design I know of, that is an 8 bit quantity, but the XAP only allowed you to address 16 bits at a time. Presumably this was for efficiency - the XAP was very low gate count for a fairly capable processor (at the time). 16 bits was also the size that the instruction set operated on, the "natural size for calculation" as described in the C spec, so that's the size of an int. This means that sizeof(int) == 1, which is a surprise to most people. It also means that you can't do tricks like "cast a uint16_t to an array of 2 uint8_ts" (which is non-portable also for endian reasons anyway).The other unusual aspect is that it's a Harvard architecture, which separates the address space for code and data. Code size being large, this meant function pointers were 24 bit, whereas data pointers were 16. Code that assumed it could cast a function pointer to "void *" to put it in callback context could get a nasty surprise (usually a subtle one, because it only broke if the function was late in the address space).
This was what I worked on early in my career, it was a great education in writing C that was actually portable. If it ran on PC and XAP, then it would likely run on anything!
[1] This is why picky/anal engineers will refer to "octets" in a network protocol rather than bytes. Bytes don't have a meaning except in the context of execution, which obviously doesn't apply to a network protocol.
I don't know if the compiler actually does this as different types and how it internally handles it. Maybe someone can elaborate on that.
> This is all completely legal according to the C standard.
Is it? I don't have access to the standard, but from secondary sources[1] it seems not?
unsigned integer type with width of exactly 8, 16, 32 and 64 bits respectively (provided if and only if the implementation directly supports the type)
If it is, it kinda defeats the whole purpose of uint16_t and friends.
So rather, a byte is 16 bits.
The "minimum addressable thing" is called a byte, and on XAP that is a 16-bit value.
Storing a uint8_t wastes a 8 bits, because there isn't anything smaller than an int to put it in.
I realise know when I've previously written the above examples I've done it was basic types, so sizeof(char), sizeof(int) etc.. Sounds like it would be more correct as well!
(I've worked with multiple processors with MAU>8; I recall at least one with 24-bit ‘bytes’, though I've forgotten what domain drove that.)
sizeof(char) == 1
sizeof(short) == 1
sizeof(void *) == 1
all being true with CHAR_BIT equal to 16, but it seems pointless to support uint8_t if it's not truly 8 bits.Too far from my language-lawyer mood now to dig deeper. :)
Serious throughput DSP chips don't sully themselves with a mere 8 bits, they're designed to do 32 bit and more FFT pipelines that modular index multiple vectors, fetch, multiply, add, and store every clock cycle.
Eg: the Texas Instruments TMS320C54x has CHAR_BIT 16 ( Table 7-1 page 192 [1] )
Other modern DSP family chips have CHAR_BIT 32 .. they're for numerics not ASCII text processing.
The Kalimba was equally weird with sizeof(int) = 1 being 24 bit and sizeof (long) was 48bit. I ported an ECDSA library to it for Pixel Buds 1 because the XAP was too slow. It was really really hard to get that to work properly and writing a simulated environment to make sure I emulated the math correctly with masking and whatnot.
It was such an annoying architecture to work with that all the senior SW leads for Pixel Buds fought extra hard to avoid Qualcomm’s solution. Their evolution for the next set of chips to compete with Apple’s W1 saw them go down a weird path where they doubled down on getting rid of GCC and instead using their home grown compiler (none of their shit ran natively on Linux not Mac) and similarly weird architecture decisions (forget all the details now). Our job was made easier in that they couldn’t actually deliver a W1 competitor. Their best was exposing each bud as a separate device which would have been a terrible experience and they could only improve that experience for Android if we mainlined their weird decisions into Android (sorry - no). By comparison BESTech delivered a proper competitor design to the W1 (transparent sniffing and hand off) and their SW architecture was totally sane (ARM chip + gcc for sure + FreeRTOS if I recall correctly). Much better partners than Qualcomm.
I was writing code natively for the processor - it was a XAP5 and natively 16 bit. Possibly CSR later moved to XAP6 (which was 32-bit) and kept application code portable using the VM.
Kalimba was my first experience writing hand-coded assembly - I was implementing sample rate conversion for a hearing aid manufacturer. It was good fun - a dedicated multiply accumulate that meant you could do filtering in a single instruction per sample.
For the kalimba I just used the Qualcomm C compiler. There may have been some assembly for audio related things although I can’t recall. It was mostly C code though I think.
That’s actually what they tried to do for the new chips if I recall correctly - they put Kalimba everywhere. I was like - wtf are you doing Qualcomm.
If we had to be a bit more portable to include such systems, we could test on bit instead of byte sizes, e.g.
#if SIZEOF_SIZE_T * CHAR_BIT == 16This guy is awesome, and his positivity is outstanding. Just watch a few of his YouTube videos and you'll understand what I mean (:
> Nor shall such a translation unit define macros for names lexically identical to keywords.
https://stackoverflow.com/questions/9109377/is-it-legal-to-r...
For some projects and people, fun is more important than readability, speed, or any other dimension.
(edit: not to imply that hack isn't readable or fast or such! I think it's cute and very readable)
Counterpoint: I'm a C developer and a big fan of Andreas and I prefer mildly clever code over longer explicit code. But even then I was still surprised by his choice to keep this specific hack. It feels very brittle to me.
The story of my life . . .
I have doubts about the legality of this solution though. The user might have #defined unsigned and this would break that. So far none of my users was mad enough to do it but I think they would be in their rights if they did.
If you only support "standard" platforms you can just typedef signed long ssize_t However some platforms (looking at you here, Windows!) will define long as 32-bit even on 64-bit and for those that will break. Not sure if __SIZE_TYPE__ is intrinsically declared on Windows in the first place. The C standard allows platforms where pointers have more bits than integers, in which case long would not work.
Hey, I just had an epiphany. You could use __PTRDIFF_SIZE__!
Here the trivial, unthinking preprocessor allows to do a pretty crazy (though very understandable and predictable) thing, which allows to acceptably solve a problem which would take years and a ton of effort to be solved "properly" (with standardization and compiler support).
When programming in C++, where limited compile and type metaprogramming exists, one is constantly hitting the limits and it causes endless frustration. I go through a mini cycle of grief until I give in and use a macro, or an otherwise less elegant implementation.
You're right in that it has taken years (decades, even) to standardise some better alternatives to macros. But even now C++ lacks some of their power.
I have to wonder how much faster alternatives would have been implemented, if macros weren't "good enough" for so many use cases.
With the addition of “constexpr” and “consteval” compile time programming is the same as runtime for many cases. Templates are obtuse for meta programming but usually can get the job done.
The need for macros much less common in modern code.
Existing c++ reflection has mostly been done with macros, which you sacrifice readability for by declaring your class with macros instead, and I believe is often a runtime thing anyway. Complex type metaprogramming is possible, sure, but often so obtuse and illegible I'd dare say the preprocessor is a better alternative if it works.
A really cursed implementation could hack up compiler support for an asymmetric integer type that treats all-bits-set as -1 and everything else as a positive number, allowing ssize_t to hold all but the largest size_t values. While perfectly standards compliant, I guess this might break a bunch of implicit assumptions across various existing programs.
It's easiest to explain in terms of how to interpret a particular bitpattern. As the first step, interpret the bitpattern as an unsigned int, u.
If u <= T, for some threshold T, then we are done, the final value is u. If u > T, then we interpret it as the negative number u - UINT_MAX.
T = UINT_MAX gives you the unsigned numbers, and T = INT_MAX gives you the two's complement.
Because of modulus intense handwaving it's also easy to do addition, subtraction and multiplication using this representation.
If an abstraction is placed into a standard, its answer to "how many people are benefiting from this headache we're giving everybody" really ought to be noticeably greater than zero.
I'm glad you brought this up because is kind of my point. C++ realized this was useless baggage and finally left it behind. [1] I don't see why ptrdiff_t is much different here. C just doesn't want to let things to, I guess. Literally any feature you put into a language will end up being used (or abused) by someone for something. "It's nice" that at some point in the future someone can pick up any random shiny thing once in a while and twirl it around doesn't seem like a reason to keep it into the standard for decades and burden everyone else with it the whole time. (Not to mention there are much nicer things that C and C++ lack, and that would make people's lives easier rather than harder...)
[1] https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2018/p09...
If the future is still flat, then... we've paid the cost of an extra few paragraphs in the standard? Folks interested in writing non-portable code can ignore this and use implementation-defined behaviors; those who want to be fully portable to future machines can be more careful (although ptrdiff_t is basically a cursed type anyway, so I don't see this particular overhead mattering). Yes, getting rid of ones complement makes sense these days -- but if you were worried about the compatibility issues around supporting it properly and avoiding undefined behavior, you're probably spending exactly as much effort today dealing with the fact that INT_MIN and friends are cursed mathematically on twos complement machines.
Meanwhile, suppose that we actually break out of this local minima. That's an interesting world, and an even more interesting one if we can carry most of our software forward with us.
We're not even a century into designing computers yet. I can't even begin to predict what architectures will look like in another five or ten centuries -- even assuming that transistor budgets continue to taper off. But I'll say that of the languages I use daily, C is one of the few that I'd still expect to be around and functional in that future, even if only as an archeological curiosity; it allows a decently high fidelity description of how an algorithm should be implemented across more than half a century of hardware.
If people are looking for examples, I'm wondering about C compilers for Burroughs Large Systems. Or C compilers for Lisp machines (Symbolics had one). Those are the kind of weird architectures on which you'd do this, if anyone ever did. Indeed, it is rather obvious that the C standards committee gave compiler developers these unusual options with those weird architectures in mind. But it can't force them to make use of them, even if they are on a platform in which they might make sense.
C2x applies a proposal to permit pre-C99 limits:
N2808 Allow 16-bit ptrdiff_t
N2808: https://www.open-std.org/jtc1/sc22/wg14/www/docs/n2808.htmDraft C2x: https://www.open-std.org/jtc1/sc22/wg14/www/docs/n3047.pdf
But yes, the "cute" "hack" aspect is the primary endearing factor :^)
I understand doing hacks when there's no other way around something, or when a hack is much cheaper than the proper solution, but in this case, I think those guys have chosen a more complicated, more expensive, less functional hack than a cheap and more complete proper solution (in "how others are doing it").
>Other C libraries typically use more careful techniques, such as wrapping the declarations in architecture-specific #ifdefs
They don't have to define it multiple times for different architectures. This is theoretically platform agnostic and saves a few lines. Not really that significant, but then again it's not like he's recommending people do it.
I've always parsed it as "the previous statement was intentionally wrong or irritating", sort of like the /s sarcasm tag except it can also denote trolling those who aren't in on the joke. I'm unsure whether it's used as a normal smiley here or whether there's something I'm missing.
"__SIZE_TYPE__" is a builtin symbol in the preprocessor, seriously?
echo __SIZE_TYPE__ | cpp -
# 1 "<stdin>"
# 1 "<built-in>"
# 1 "<command-line>"
# 31 "<command-line>"
# 1 "/usr/include/stdc-predef.h" 1 3 4
# 32 "<command-line>" 2
# 1 "<stdin>"
long unsigned int
Okay, so it's defined in a file apparently. But why is that file automatically included in every program, and not the more usual type definitions? And is there really no better way to define "an integer the size of a pointer" than by going through this rat's nest of text substitutions?The C language seems almost designed on purpose to maximize uglyness. It is called "portable assembly language", but even in this day and age when there are only two or three relevant processor architectures, which have been largely designed around being able to run C code efficiently (to the detriment of everything else), it falls short of that.
It is only portable at all because of include files that are included from other files, containing __MACROS__ calling on __OTHER__MACRO_S___ to the n-th degree.
The most advanced compiler algorithms are then applied to the problem of transforming the resulting ((void *)(__pile_of(cr*p))()) back into machine code that is at least not completely terrible. Follow every obscure rule of the standard, and they might be so nice not to remove your carefully written null pointer check in the process!
And people worship this uglyness and needless complexity, even as it strangles the life out of every other technology like a cancer that has been growing for 50 years. Professionals and hobbyists both, they celebrate how clever they are, being able to work around its deficiencies, think that they are dealing with the fundamentals of computer science rather than the grotesque evolution of an operating system originally written to typeset documents and play SpaceWar.
Just once I'd like to see a new operating system - better yet, a new CPU! - that is not based on C and UNIX, one that is outright hostile to them at every level of abstraction. Not a single line of C code anywhere, different calling convention, different filesystem, user interface, networking etc.
You may, or may not, have to pay it back.
So some technical debt is good in many cases as long as it is properly managed.
If you have zero debt it means you are not efficient (the same in finance or if you purchase a home cash while you could get a low and fixed interest rate loan for example).
i used to do "cute" tricks like this. then i learned better. always write the least mysterious code unless you have a good reason...
Why not doing the same?
Declarations can be definitions but a typedef is not a definition in particular. https://en.cppreference.com/w/c/language/typedef