HNHacker News
TopNewBestAskShowJobs

sparkie

2,309 karma · joined February 15, 2012

submissionscomments
sparkie··on C's Flexible Integer Sizes Were Not a Design Mistake
That's an implementation choice. The standard makes no such guarantees.

> An integer may be converted to any pointer type. Except as previously specified, the result is implementation-defined, may not be correctly aligned, may not point to an entity of the referenced type, and can produce an indeterminate representation when stored into an object

The compiler may specify otherwise. GCC specifies that the pointer cast to integer back to pointer must point to the original object, else the behavior is undefined.

sparkie··on C's Flexible Integer Sizes Were Not a Design Mistake
Partial register stalls occurred when you assigned something to say `eax`, then read `ax`, or vice versa - if you wrote `ax` then read `eax`. If you wrote to `ax` then read `ax`, there was no stall. This was due to the register renaming implementation.

AFAIK, this was only an issue in some older CPUs and isn't a problem today.

32-bit `int` is still cheaper than 16-bits though, because 16-bit instructions require an operand size override prefix (0x66), or address size override prefix (0x67), or both. Technically, these aren't "16 bit prefixes" - if the machine was running in 16-bit protected mode, then you would need those prefixes to use 32-bit instructions and the non-prefixed ones would be 16-bit and thus cheaper, so `int` would be better as 16-bits in 16-bit protected mode - though this mode is essentially unused today, so for all intents and purposes the prefixes are used to issue 16-bit instructions and 16-bits is more expensive (in code size, i-cache usage, which may impact performance).

sparkie··on C's Flexible Integer Sizes Were Not a Design Mistake
If you want to know how bad it can be, only need to take a look at GCC's implementation of `stddef.h` for `size_t` and `ptrdiff_t` etc.

https://github.com/gcc-mirror/gcc/blob/master/gcc/ginclude/s...

There's a simpler way to define these types in newer versions of C which have `typeof`, we can take advantage of the fact that `sizeof()` always returns a `size_t`, `ptr - ptr` always returns a `ptrdiff_t`, and a literal L'c' has type `wchar_t`, etc.

    typedef typeof(sizeof(0)) size_t;
    typedef typeof(nullptr) nullptr_t;
    typedef typeof((void*)0-(void*)0) ptrdiff_t;
    typedef typeof(L'\0') wchar_t;
    typedef typeof(u8'\0') char8_t;
    typedef typeof(u'\0') char16_t;
    typedef typeof(U'\0') char32_t;
Don't need ifdef soup to test what the sizes of things are on different architectures - we just utilize the compiler's implementation of those types.
sparkie··on C's Flexible Integer Sizes Were Not a Design Mistake
It's safe to dereference in the same thread.

A `thread_local`'s actual virtual address is not merely what the pointer contains - it's an offset from some other virtual address stored in the FS or GS register (on x86-64), which is swapped when you change thread. Casting the pointer to `intptr_t` does not retain the segment base address - only the offset. The pointer is not the absolute address.

If you dereference in another thread, it's the same offset, but from a different base address.

There may be other things besides segment registers on other architectures that also make it unsafe. The C standard makes no guarantee that you can safely dereference a pointer cast from `intptr_t`.

sparkie··on C's Flexible Integer Sizes Were Not a Design Mistake
Maybe you are talking past grandparent, but I think they meant "it's safe to assume float and double are IEEE-754," which is not mandated by the C standard.
sparkie··on C's Flexible Integer Sizes Were Not a Design Mistake
`int` being 32-bits on amd64 was the correct decision, even in hindsight. If you are familiar with the ISA you will understand this. Existing 32-bit code just worked on the 64-bit chip, because the instruction encodings for 32-bit are unchanged. The 64-bit instructions are basically "opt-in", by placing a REX prefix on them - that makes them more expensive to encode and uses more instruction cache. Even today, compilers will emit 32-bit versions of instructions when the upper 32-bits are not needed, because it's cheaper.

When amd64 was released, x86 was almost ubiquitous on desktops and ran the majority of servers - most of the software used by the world could continue being used. If AMD had not gone through this effort to make it backward compatible, it's likely IA64 would've won and we wouldn't have this debate. Hard to understate the importance of not breaking things.

If you were designing a greenfield 64-bit ISA, then yes, it might make sense to have `int` be 64-bits, but it was definitely not, and still is not a mistake that it's 32-bits on amd64.

On RISC-V for example, it's questionable. The RV32 ecosystem is tiny and almost irrelevant - if they decided to break things for RV64 it wouldn't be a big problem - probably better to fix any problems early rather than hold baggage to run software that never existed - though it's much easier to port software to RV64 if `int` is still 32-bits.

sparkie··on C's Flexible Integer Sizes Were Not a Design Mistake
Yeah, in GCC for example, we can define exact width types without including `<stdint.h>`.

    typedef unsigned __attribute__((mode(HI))) uint16_t;
    typedef signed __attribute__((mode(SI))) int32_t;
Of course, the mode needs to be supported by the compiler for the target arch, but we don't need to include anything.
sparkie··on C's Flexible Integer Sizes Were Not a Design Mistake
Yeah, but my point is that two pointers can compare equal but point to different addresses, thus it's not necessarily a safe operation to dereference a pointer cast from `intptr_r`.
sparkie··on C's Flexible Integer Sizes Were Not a Design Mistake
I think backward compatibility was the main aim. Intel tried redesigning the architecture as 64-bit native (Itanium), but AMD done a better job at backward compatibility - and intel eventually adopted it as x86-64.

The amd64 design could run most 32-bit code with minimal changes. All the 32-bit instructions had the same encoding, besides push/pop which instead acted on 64-bits. The 64-bit instructions were basically opt-in, though a few opcodes (0x40..0x4F) had to be deprecated for the REX prefix.

sparkie··on C's Flexible Integer Sizes Were Not a Design Mistake
> int is basically the signed version of size_t aka a word sized data type. It's not meant to have a fixed size.

It isn't. `int` is at least as large as `short` and at least 16-bits. On modern systems `int` is still typically 32-bits whereas `size_t` is typically 64-bits. There's a `ssize_t` in POSIX for signed sizes.

> When people want the classic 4 byte data type they should choose long instead.

`long` is only at least 4 bytes, and at least as large as `int`. On MSVC (LLP64 data model) it's 4 bytes, but on SYSV (LP64 data model) it's 8 bytes. `long` should almost never be used if you actually want portable code today.

`int` is 32-bits and `long long` is 64-bits on both LP64 and LLP64. If you want portable code using the native integer types, these are the ones you should use, definitely not `long`.

sparkie··on C's Flexible Integer Sizes Were Not a Design Mistake
What is the size of a pointer though?

On Intel 286 we had a 16-bit machine word and 24-bit addresses. A pointer wasn't just two machine words concatenated - the upper 8 bits were stored somewhere else - a segment register.

On modern machines we don't (usually) need to consider this because we have a single linear virtual address space, though the size is architecture dependant - usually above 40 bits and below 64. Most common size is 48-bits, but also up to 57-bits with 5 level paging enabled.

Either way we round up to 64-bits to store the pointer as one integer. C optionally provides types `intptr_t` and `uintptr_t`, which are integers large enough to hold the value of a pointer. Converting a pointer to `intptr_t` and back to the pointer type results in a pointer that compares equal to the original.

However, there is no guarantee that a pointer converted to `intptr_t` and back to a pointer can be dereferenced! It works most of the time because of our linear address space and non-use of segmentation, but segmentation can still be used - the FS and GS segment registers are still available on x86_64 and are commonly used for thread local storage. If you take a `thread_local T*`, convert it to `intptr_t`, and then convert it back to a `thread_local T*` on another thread and attempt to dereference it, then despite the pointers comparing equal, they dereference to different virtual addresses.

Integers tied to the size of a pointer would have been misguided. Pointers are not integers! (They just happen to use an integer in their representation).

Another one, `size_t` is supposed to represent the maximum size of any object. However, that's also not well-defined. The maximum object size on the Intel 256 would have been 16-bits, because that is all you can fit in a single segment.

On a modern machine, a `size_t` should really be 48-bits (4LP) or 57-bits (5LP), because we can't have an object larger than our maximum virtual address size - but `size_t` is typically 64-bits.

sparkie··on Parsing Expression Grammar vs. Regexes: Building Org Parser in Lisp, Export HTML
PEGs aren't contained within the context-free languages, so it's not really intuitive that they're weaker, particularly as nobody has yet come up with a context-free language that a PEG cannot parse. They're capable of parsing things context-free grammars cannot.
sparkie··on Parsing Expression Grammar vs. Regexes: Building Org Parser in Lisp, Export HTML
Worth looking into nested words aka visibly pushdown languages.

They're a proper superset of regular languages and a proper subset of deterministic context-free languages, but they retain many of the nice properties of regular languages that DCFLs don't - they're closed under intersection, union, concatenation, Kleene Star and reversal.

They can parse more languages that Regular Expressions (non PCRE), but fewer than deterministic CFG subsets like LL/LR. They're expressive enough to parse languages which have a regular tree structure like S-expressions, JSON, XHTML.

sparkie··on Zig: Pointer Stability for ArrayLists
I made a mini example in C.

It's awkward to do get right because you need an indirect pointer whose address remains fixed, but points to another pointer which can change (and is volatile).

While it might be possible to make something like this lockless - it's much simpler to stick a mutex in the array header. When we access the array_segment we can take a lock to prevent some other thread reallocating mid-way through accessing.

There's probably a few improvements that could be made. In particular it doesn't handle use-after-free, so it's not thread safe w.r.t cleanup.

https://godbolt.org/z/rYzn5KGre

sparkie··on Zig: Pointer Stability for ArrayLists
I think parent was after base+offset+index rather than just base+index.

Examples would be eg, `string_view` or `ArraySegment`. They hold some offset relative to a base allocation, and when we index the string_view or ArraySegment we're indexing relative to that offset.

sparkie··on Zig: Pointer Stability for ArrayLists
That's Fat pointers, not Far pointers. A fat pointer is a pointer with some other associated data which is stored in the pointer itself - typically by widening the number of bits used to hold a pointer value. The addressable bits usually remain unchanged - the added bits contain the auxiliary data.

Segmentation isn't used. There's no separate registers to hold the bounds information in CHERI - the bounds are held in the pointer value, unlike for example, the now obsolete Intel MPX, which held bounds information in separate registers.

There's some similarity to segmentation because the CHERI pointer restricts which addresses can be accessed, but I wouldn't compare them to far pointers.

Most modern processors have a single linear virtual address space and don't use segmentation, and even where segment registers exist (eg, FS and GS on x86-64), they're only superficial "address spaces" - allocated sections of the process's linear virtual address space which could be accessed without segmentation registers if you knew the base address held in FS or GS.

sparkie··on Zig: Pointer Stability for ArrayLists
Far pointers are for accessing memory in different segments. They're basically obsolete now. They were necessary in older machines with limited sized pointers or address spaces.

GCC still supports `__seg_fs` and `__seg_gs`, which behave similar to `far` in the example on the wiki page, as the FS and GS segment registers are still valid in x86-64 and used for TLS. Clang uses attributes `address_space(257)` and `address_space(256)` for the same thing.

The `__based` pointer in MSVC exploits the addressing modes by pinning the base in eg: `[base+index*scale+displacement]`. It's unrelated to segmentation.

sparkie··on The world may have less time than it thinks on climate change
The poorest in society are the lowest CO2 emitters - they don't fly often, many don't own a car and use public transport - they can't afford meat in every meal - can't afford air con in summer or heating in winter.

Conversely, the wealthiest are frequent flyers, are chauffeured about in SUVs and rarely walk. They have everything they'll ever need and don't concern themselves about running the AC or heating.

They're preaching to the poor guy about what he needs to do to stop climate change.

Even if the poor guy wanted to care, he can't do much, but the people who claim to care could do a lot more, if they really did care.

The media circus tries to put the blame on the poor, while ignoring the emissions of the celebrities and treating them as royalty. Is it any wonder most people simply don't care?

If the media cared, they would be making villains of the celebrities and idols of the poor.

sparkie··on Zig: Pointer Stability for ArrayLists
The FS and GS segment selectors are still used in x86-64, typically for `thread_local` storage, but they can be repurposed.

`thread_local` is an example of a "relative pointer" though. Instructions to access the thread local are prefixed with `fs:` or `gs:`, and point relative to the address in the respective segment register.

sparkie··on Tail-call optimization in C is relatively recent (2025)
That's not entirely true, but it's a valid reason to prefer using musttail.

`register` is a hint if you don't specify which register you want to use - however, if you specify the register it will clobber it.

    noinline void bar() 
    {
        register void *parent __asm__("r10");
        ...
    }
You can also use GCCs extended asm syntax to clobber a register for specific portions of code - such as the start of a function where you expect a register to have been given a value from the caller just before the call. Use `volatile` to prevent the compiler from making certain assumptions that might remove or reorder the instruction - as long as it is at the top it should execute immediately after the function prelude and before any of the function body.

    noinline void bar() 
    {
        void *volatile parent;
        // set parent = %r10 before anything else.
        asm volatile ("mov{q}\t{%%r10, %0|%0, r10}" : "=r"(parent) : : "r10");
        ...
    }
Note that this will probably be less efficient than the former example, but maybe useful where you want to limit the scope in which `r10` is clobbered.

In both cases you would set the register immediately before making the call, again using `volatile`. Since `r10` is not used by a typical call in SYSV - it's the static chain pointer in the SYSV convention, but otherwise usable as a GP register, a call will not overwrite it.

    void foo()
    {
        struct foo_frame {
            int x;
        } locals = { 
            .x = 999
        };

        // Set `r10` to our function's local frame
        asm volatile("mov{q}\t{%0, %%r10|r10, %0}" : : "r"(&locals) : "r10")
        
        bar();
    }
That's pretty ugly but we can write a few macros to implement it more tersely - we can use this to have efficient closures in C without requiring an executable stack. (There's also `__builtin_call_with_static_chain`, but I've found it more troublesome to use than the manual way).

Demo: https://godbolt.org/z/cM9d8e1r5

For other registers which are part of the regular calling convention, we might be able to clobber them if they wouldn't normally be used for the call. Eg, if our function takes regular 2 arguments, they would be in `rdi` and `rsi` - so we could use `rdx`, `rcx`, `r8`, `r9` like the above, but if our function took 6 or more regular arguments we wouldn't be able to use any of these in this way. If we wanted a custom calling convention we could just make all functions have zero-arguments and perform all of the setting and capturing ourself - which gives us more control than using [[musttail]] - though less portable, and may prevent optimizations the compiler could otherwise make.

sparkie··on Tail-call optimization in C is relatively recent (2025)
> What practical patterns are enabled by TCO in C?

Continuation Passing Style - an important construction for interpreters, but which is also useful for compilers as it's a nice way to do control flow analysis, data flow analysis and more.

The missing feature is closures - functions which capture values from their static environment, which are basically needed to make CPS useful. GCC has nested functions, but they cannot capture without making the stack executable, which is terrible. There's a proposal[1] to get closures into C, but at present you need to simulate the capturing yourself, which is cumbersome, but can be done efficiently.

[1]:https://thephd.dev/_vendor/future_cxx/papers/C%20-%20Functio...

sparkie··on A generic dynamic array in C that stores no capacity and needs no struct
That may be true, but it may also mean you utilize more memory than you need to. If you aren't shrinking the array when you no longer need previously allocated capacity then you're wasting memory. You could end up with an array of 10 elements and an allocation of 2^10.

The capacity as bit_ceil(len) ensures that at most, half of the allocated space is wasted - excess space is O(n).

sparkie··on A generic dynamic array in C that stores no capacity and needs no struct
I don't think it's that rare. I've been using the technique for years, and I've seen it done in other work. Bagwell's VList[1] for example uses the equivalent of `bit_ceil` to determine the size of each block without having to store it - and there are earlier works based on the same trick. RAOTS, which is referenced by the VList, mentions using the technique, but itself uses a slightly more complex trick where we can calculate the size of a block based on the approx square root of the length.

You can use the trick if the array can shrink as long as you always shrink the allocation when length goes below the next power of 2 not greater than len (which may make use of stdc_bit_floor).

[1]:https://cl-pdx.com/static/techlists.pdf

sparkie··on A generic dynamic array in C that stores no capacity and needs no struct
If you use a struct with a `void*`, you also need to specify the type on usage, where here it's done with `typeof`.
sparkie··on A generic dynamic array in C that stores no capacity and needs no struct
In C23 this approach is nice, but in older versions of C we end up with awful macros where we need to define the structure before we use it.

    #define Array(T) struct array_##T
    #define DEFINE_ARRAY(T) struct array_##T { size_t len; T *elems; }
    
    DEFINE_ARRAY(int);
    Array(int) foo;
    Array(int) bar;
C23 has relaxed rules for redefining the same struct, so we can avoid having to create the struct up front.

    #define Array(T) struct array_##T { size_t len; T *elems; }

    Array(int) foo;
    Array(int) bar;
sparkie··on A generic dynamic array in C that stores no capacity and needs no struct
The reason the struct is avoided here is so the array can be typed to its element type (rather than casting to and from `void*`).

With a struct we would need one struct for each element type - at least prior to C23 which provides a better approach where we can declare the same struct multiple times in a translation unit.

    #define Array(T) struct array_##T { size_t len; T *elems; }
We can use `Array(int)` in multiple places in the same TU - but in C11 or earlier, this is an error.
sparkie··on A generic dynamic array in C that stores no capacity and needs no struct
The concept of not storing capacity isn't silly. If you need to reserve space then it's not the appropriate structure, but it's otherwise fine.

However, using an 2-element array to avoid using a struct is silly.

sparkie··on A generic dynamic array in C that stores no capacity and needs no struct
Really? It's been done plenty and I thought was quite common knowledge. Some of the <stdbit.h> provided functions are basically for this purpose.

stdc_bit_ceil(len) gets the smallest power of 2 not less than len, which is our capacity. This is usually implemented with a clz instruction.

stdc_has_single_bit(len) determines if it's a power of 2 - typically implemented with a popcount instruction (popcount(len)==1).

The approach isn't used in older (90s and earlier) texts because hardware support for popcount/clz wasn't commonplace and the cost to do it in software wasn't worth it, but it is mentioned in some texts.

sparkie··on Bijou64: A variable-length integer encoding
Bitcoin has a variable width encoding (`CompactSize`), but it doesn't prevent overlong encodings - however there are various canonicalization rules in the Bitcoin protocol to require minimal encoding.
sparkie··on Lib0xc: A set of C standard library-adjacent APIs for safer systems programming
The C charter has a rule of "no invention".

Anything needs to be demonstrated and used in practice before being included in the standard. The standard is only meant to codify existing practices, not introduce new ideas.

It's up to compiler developers to ship first, standardize later.

Page 1 of 34Next →