Fat pointers in C using libcello
libcello.org
libcello.org
- slices become a total non-issue since a pair of (start, end) already is a slice and you can just move start and end.
- comparing against an end pointer is generally easier than adding up a length value first, particularly if you're slicing at the same time.
- the end pointer value is independent of the array element type, so if you e.g. cast to uint8_t * (which arguably you shouldn't in most cases) it stays exactly the same. If you store a count you need to adjust a multiplier. If you store a byte length, you need to do a lot of divides or casts to deal with pointer arithmetics.
Also, this is a huge red flag to me:
https://github.com/orangeduck/Cello/blob/master/include/Cell...
#define is ==
#define isnt !=
#define not !
#define and &&
#define or ||
#define in ,
P.S.: This also is a "try to invent a new programming language without inventing a new programming language" thing. Have your cake and eat it... either it's C or it isn't, and this library is leaving the space of "normal" C.I know he's doing a little hackiness by placing the size before the start of the object, but they also take that approach here (https://www.piumarta.com/software/cola/objmodel2.pdf) so maybe it's not that uncommon. How would you do the dual pointer setup? Is there any overhead from that, or is it small enough to not worry about?
Placing a length "before" a pointer is perfectly fine on a _technical_ level. It's also how glibc's malloc works, it has its own data before any allocation it returns. However, hackiness is not the question here - it's whether it's the "best" approach. I simply believe, based on my own experiences, that twin pointers cover/win out in a much larger subset of applicable scenarios.
As for how to implement it - declare a struct with 2 pointers in it. Or just pass 2 pointers around.
Puke! Who would ever want to create that kind of macro abomination.
And I'm not even joking :)
That said, my opinion about those macros is unchanged.
(C is my "native" language, so it doesn't really happen the other way around, but that's just me)
But honestly, I'm pretty sure we can call iso646.h "fringe" ;)
Certainly, speaking as someone brought up on imperative and later OO code, it took a fair bit of unlearning to understand functional programming, and I still have no clue about type theory.
> I have seen people coming to C++ from Python use "and", "or", and "not" in their if-statements :)
Technically, I believe this is actually UB, just like any other #define for a keyword or an identifier from the standard library.
When I do #include <stdbool.h> it also brings the #define bool _Bool, a keyword from the standard library.
How is it UB?
"Notwithstanding the provisions of 7.1.3, a program may undefine and perhaps then redefine the macros bool, true, and false"
I guess it's because defining these was so common even before C99. But there's no similar verbiage for <iso646.h>, so...
It’s been a while since I watched the talk but I believe his intention was to do just that, to push the bounds of what could be done in a header file purely for the fun of it. The second half of the talk specifically addresses the “why are you doing this?” in quite a charming way.
type slice struct {
zerothElement *type
len int
cap int
}
Though I suspect under the hood the len and cap are ordered first.If it's good enough for C greybeard Ken Thompson and unix hacker Rob Pike, its good enough for me.
In fact I've looked for a port of go-style slices to C and haven't found one. Maybe people think sds is good enough?
A slice is stored in memory as a `reflect.SliceHeader` https://golang.org/pkg/reflect/#SliceHeader ; the pointer does come first.
https://github.com/jgbaldwinbrown/slice
It's just a proof-of-concept, though, and would need a lot of work to be used in anything serious.
There’s a PARC paper from the Cedar team benchmarking base-and bounds vs marker-terminated (e.g. null terminated) strings and base-and-bounds won hands down on a variety of use cases. Won on speed, not just the safety grounds. In those days some people actually worried about the space taken up by the extra bounds variable.
But most people these days have a hard time realizing what it was like to program a machine with only a few K of memory. PL/1 (the Multics implementation language) had bounded strings. C did not as it was believed the PDP-7 couldn't afford it. Of course by today's standards the GE 645 is an absurdly tiny machine.
Just to think that what is now on a ESP32 die took over a complete desk at the high school computer club and was still weaker than what ESP32 is capable of, yet that mentality also persists for what 80's hardware was already capable of.
typedef void* var;
struct Header {
var type;
};
// ...
#define alloc_stack(T) header_init( \
(char[sizeof(struct Header) + sizeof(struct T)]){0}, T)
var header_init(var head, var type) {
struct Header* self = head;
self->type = type;
return ((char*)self) + sizeof(struct Header);
}
The section "struct Header* self = head" is UB.
The alignement requirement of the local char array is 1 but the alignment
requirement of struct Header is that of void* which is probably 8.var is a typedef for void* and no & appears in the function.
It's an array, so you don't need & to take its address, it decays into a pointer without &.
Imagine:
char buf[sizeof(struct Header) + sizeof(struct T)];
char *p = buf;
Then take away the names so that you are effectively passing p as a parameter... Then returning p.As in, an anonymous temporary being given to the function, and the function returns its address back.
It's assuming that this temporary parameter buffer will exist after it is used and the function has returned. I'm not sure what the standard says for that but it is crazy sketchy. [Edit: Googling around, it seems like maybe this is illegal in C99 but possibly legal in C11? Or that C11 changed the rules for this. Does not seem like a great thing to rely upon.]
Many standard library functions return a pointer which they got as a parameter (or another pointer offset from it, as is the case here). The compound literal is no more "temporary" than a variable that was introduced right before the function call. Lifetimes of compound literals are specified in section 6.5.2.5 of both C99 and C11.
func((int[256]){ 0 });
// is mostly equivalent to
int __a__[256] = { 0 };
func(__a__);Obviously. But this does not extend the lifetime of the buffer they are passed. Namely you can't use this as a technique to extend the life of automatic storage falling out of scope.
> Lifetimes of compound literals are specified in section 6.5.2.5 of both C99 and C11.
This is what I was missing. So it is valid by the standard. Which is good if you have looked it up. It remains not obvious when reading source without a copy of the standard on hand, or prior knowledge of that section. Passing an expression of that sort and keeping a pointer to it visually looks like the intent is to retain a pointer of more limited scope. Intuitively it would make just as much sense if the lifetime were shorter. If you are seeking clarity of intent this is not a great thing to rely upon.
A C standard library has a header file for complex math but it doesn't define a simple fixed size array struct? Why is that? Is it because they become pointless when there is no generics to deal with the stride?
I've read that it's because there used to be binary interface issues with structures. They can be returned from functions and passed as parameters but it isn't immediately clear how that happens: is it on the stack, in one register or in several registers? Even today there are compiler options that affect the generated code in those cases:
-fpcc-struct-return
Return “short” struct and union values in memory like
longer ones, rather than in registers.
-freg-struct-return
Return struct and union values in registers when possible.
https://gcc.gnu.org/onlinedocs/gcc/Code-Gen-Options.html#ind...https://gcc.gnu.org/onlinedocs/gcc/Code-Gen-Options.html#ind...
https://gcc.gnu.org/onlinedocs/gcc/Incompatibilities.html#in...
Why does it have to be clear? It can be unspecified and the compiler will do what it thinks is best given the struct, e.g. return `struct {int x,y};` in registers, return `struct int[80] x}` as pointer to memory or write in-place to the caller's stack, via RVO.
Because it's part of the binary interface. Changing the binary interface prevents existing programs from interoperating. Everything breaks until the software is recompiled and that can be a major nuisance at best and impossible in the worst case.
Simple, well-defined and stable binary interfaces are a major reason why C is still widely used. Uncertainty in this area is never a good thing so people will actively avoid language features that introduce it. Looks like enums aren't favored for similar reasons: what is the underlying type?
To answer the question though, a number of people do define structures containing a buffer and a length (and potentially capacity), there just isn't such a structure standardized so everybody who wants to do this has to bring their own.
Some examples from Unix: iovec, sendmsg/recvmsg. Surely there are others I'm just not thinking of right now.
In the Windows world you have UNICODE_STRING and similar structures. SChannel has "PSecBufferDesc". Again, surely there are others.
And prominent libraries might also have their own.
I don't think ABI is a concern these days, though, even if things were different 30 years ago. The specs that we have today do cover how to pass and return structs in the standard manner.
A ptr to an array of unknown length can’t be used without the length being passed around next to it whereever it’s passed.
You cant deref ptr+i without knowing that it’s under length etc.
Two bits of data that belong together seems like they would be convenient to pass (or return!) as one argument or return value.
It’s especially terrible with methods that e.g take two lists and return a third. That should be two arguments and a return value, not six arguments.
I had discussions with people who claimed that null terminated strings are the only idiomatic thing to do in C - because that is how C does strings. They assumed that since the standard library only provided methods which acted on those kinds of strings it was a preferred way to do things. Even though that is a lot less efficient than the string/array types that other languages use as defaults.
For one thing, as I recall in the original K&R version of C (before ANSI C89), the language didn't support passing a struct as a function argument or return value.
That means if you did make a struct, then every time you wanted to pass one of these pairs around, you'd have to pass a pointer to the struct, and you'd have to dereference that on every use. Which is arguably just as cumbersome as just passing two arguments, at least in terms of how much code you have to type. Plus it was probably slower.
From there it's no surprise if using separate arguments becomes the normal, idiomatic way to do it.
Given that years ago we added a mandatory piece of hardware to most systems to implement virtual memory, I'm now starting to wonder what security and/or performance benefits could be achieved by delegating memory allocation to (or through) hardware.
Since years, Oracle has shipped SPARC Solaris with ADI turned on.
Since iPhone X iOS makes use of memory tagging for pointers.
Starting with Android 11, hardware memory tagging is a required feature on ARM platforms on the CPUs that support it, while on other CPUs the kernel will randomly attach GWP-ASan to user processes and is enabled by default on all system processes during the ongoing preview releases.
The summary version (from Walter Bright's article) is:
> C can still be fixed. All it needs is a little new syntax:
> void foo(char a[..])
> meaning an array is passed as a so-called "fat pointer", i.e. a pair consisting of a pointer to the start of the array, and a size_t of the array dimension.
The paper mentions fat pointers in passing, not putting the term in quotes, not defining it, and not giving a citation -- which makes it clear that the term was already well established at the time.
edit: Pascal pointers were just a location and size, however, not a slice-type fat pointer; however, I have always heard of any pointer containing more information than a memory address referred to as a fat pointer (except tagged pointers). YMMV.
They even link to it.
The difference is massive. Bright's idea would allow me to use existing interfaces with fixed structure: as long as I know the declared length of an array, I could pass a fat pointer to a function that accepts one. The Cello approach would require me to modify the interface to accommodate the new size header that is packed with the array. ABI break, not compatible with existing C code and libraries.
One could justify the name of fat pointer, the other is really just arrays with headers.
Isn't the entire point of the article that this is not the case?
My understanding is that a Cello array is laid out like this:
<metadata header>.<actual payload data in standard C format>
^
|
this pointer is what you pass to existing C code
The existing C code gets a pointer to data exactly as it expects. It does not get a pointer to the metadata. It could access it by subtracting an offset from the pointer it's given, but it will not do that since it does not expect anything meaningful to be there. "This pointer is fully compatible with normal pointers".If I pick up some library (or system call interface..), and it uses this structure
struct foo {
struct bar a;
char b[64];
};
The metadata field that cello wants is not there. I cannot use cello's "fat pointers" because they are not fat pointers at all. I would have to modify the struct to have array-with-length-header (which is what cello calls a fat pointer.. talk about confusing pointers and arrays!) , which might be impossible if it's someone else's binary interface that I am using.If I had actual fat pointers, then the size would be simply passed together with the pointer and I wouldn't have to have modify structures to accommodate an extra field. Of course, I can't pass these pointers to functions that expect normal pointers, but it wouldn't make sense to do that anyway, and if implemented right, doing that would require me to misdeclare a function. You would pass normal pointers to old C code that expects and deals with normal pointers.
It's true that the internal array inside a struct cannot magically become a Cello fat pointer. So yes, you cannot pass a pointer to that array into a Cello function that expects a fat pointer, but that's the opposite of "pass normal pointers to old C code".
Is it actually? Here's a quote: "The trick is to place the value representing the number of items in the array in memory just before the pointer we actually pass to functions. This pointer is fully compatible with normal pointers,"
So suppose I pass this "fully compatible with normal pointers" pointer to a function.. and it ends up being stored in a register. Now where is this location "just before the pointer"? In the previous register?
I don't think this post is describing fat pointers at all. I think it is describing arrays with a prepended length header. There is no data in or before the pointer, there is metadata before the pointee assuming you're pointing at something with metadata prepended to it.. No extra information is passed with a pointer, it's just assumed to be there at the pointee. Call it a fat pointee if you want a fancy name; at least that acknowledges the onus is on the pointee to have the right metadata in the right place. There's nothing in the pointer here.
Yes, the extra data being the length that is located in memory before the pointer.
> So suppose I pass this "fully compatible with normal pointers" pointer to a function.. and it ends up being stored in a register. Now where is this location "just before the pointer"? In the previous register?
You’re at the wrong level of indirection; the header in in the memory before what the pointer points to, as in p - 1 rather &p - 1.
> I think it is describing arrays with a prepended length header.
Correct.
> Call it a fat pointee if you want a fancy name; at least that acknowledges the onus is on the pointee to have the right metadata in the right place.
I would prefer that this was not called a “fat pointer”, but the claim made above is that a pointer with any implicit data associated with it is “fat”.
> There's nothing in the pointer here.
No data, but there is an implicit guarantee that it points to the data portion of a length-prepended array.
> You’re at the wrong level of indirection; the header in in the memory before what the pointer points to, as in p - 1 rather &p - 1.
You are contradicting yourself. &p - 1 is before the pointer. p - 1 is before the pointee.
> I would prefer that this was not called a “fat pointer”, but the claim made above is that a pointer with any implicit data associated with it is “fat”.
> No data, but there is an implicit guarantee that it points to the data portion of a length-prepended array.
I think that's a borderline useless definition. In my C, there's an implicit guarantee that any pointer, in a context where it may be dereferenced, points to a valid object (as long as it's not NULL and there's no programming error causing it to point at nothing well defined). Usually it points at the start of a struct, sometimes it points at list node (or whatever) embedded in a struct and I might have to work my way back with offsetof. Either virtually all of my pointers are fat or none of them are. I go by that none of them are, because whatever data I have (implicit or explicit) is not a property of the pointers I use but of the data I point them to.
Just locate anything declared as an array in a particular linker section so the pointer manipulation can be done with two (or one if it's at the top of memory) comparison, possibly even to a constant.
If you do this you can even forbid pointer arithmetic except in actual []-declared memory, and can do transparent bounds checking (&array-1 can hold the array length or, possibly faster, the address of the location after the end of the array).
An advantage of this over the library route is you can prevent pointer/array punning but otherwise allow any C program to work fine. And apart from a few corner cases (there are legit non-array uses of pointer arithmetic, though very few) and noncompliant program can be changed to use [] and still work perfectly fine without this option being used.
Walter often shows up on HN, so I'll ask: was this proposal merely on the Dr Dobbs article or did it actually go to a committee for review? If the latter, why wasn't it accepted?
Should C reconsider this? Especially now that C++ has std::span<> and std::string_view<>?
Contrary to common HN wisdom, most C and C++ related surveys show that only up to 50% actually use some kind of analysis tooling.
(Fixed to actually point at discussion of fat pointers; thanks, ‘dgellow.)
typedef struct MyStr {
size_t mystr_length;
char [1];
} MyStr;
typedef char *PMyStr;
size_t MyStrLen(const PMyStr p) {
const MyStr *pmystr = ?;
const size_t *psize_t = ?;
return *psize_t;
}
How do you make a legal cast to get either pmystr or psize_t?Definitely don't use this in production code.
> Can it be used in Production?
> It might be better to try Cello out on a hobby project first. Cello does aim to be production ready, but because it is a hack it has its fair share of oddities and pitfalls, and if you are working in a team, or to a deadline, there is much better tooling, support and community for languages such as C++.
I've seen it on HN before, whenever this project gets mentioned. People who don't know much about C confuse it for a really cool thing you can do with C, as if it's just another legit library that you can pick up and use. It's a lot of undefined behavior. People have enough problems writing safe C as it is, and on top of this complaint about alignment and misuse of temporaries, this thing makes the problem worse in other ways too, removing the few safeguards that exist by treating everything as void* for instance.
The new code on GitHub makes it clear that this library is not about fat pointers, but about writing unsafe C in another language, powered by undefined behaviour and the glory of the C preprocessor.
If you want to have safe arrays in C, just use a proper library for that instead of going the route of abusing undefined behaviour. There are plenty of libraries to choose from: https://github.com/search?l=C&q=vector+array&type=Repositori... (not all of them are serious projects, so you should be careful when you pick one. I personally have been using this one: https://github.com/iscgar/cvec2 because I like the easy interface it provides and especially the fact that type safety is a top priority, but other may have different tastes and priorities).