C Macro Magic
sagartewari01.com
sagartewari01.com
But I do not think this trick or list implementation is newly introduced in Wayland. It existed at least as far back as kernel 2.6.11 [2].
[1] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
[2] https://elixir.bootlin.com/linux/v2.6.11/source/include/linu...
It's used pretty much in every OS kernel, Windows included.
Note that the one that ships in linux is pretty simple, the one that is in BSD is a lot more complete.
But you can't do it in C. __typeof__ is a GNU extension.
> Note that wl_list is not a very safe data structure, in the sense that a programmer need to know exactly how to use it.
You can do it in C++ with template, and it would be type safe.
You can do it in C, you just can't do it in the part of C the standards committee has decided to standardize.
I'm not sure what the value is of defining C as standard C when loads of software has shipped for decades using this pattern.
(The C standard explicitly calls this out as a valid way of handling UB.)
But extensions like typeof are different from UB. They have documented behavior, it is just documented by GNU instead of the C standards committee.
Yes you can. You just have to provide the type every time you use the macro yourself. You can even create a macro that creates helper “methods” for you.
With that said, I really dislike how common it is for C programmers to just use whatever their compiler's authors felt like adding and still calling it “portable C”. At this point, if you dislike “plain” C so much, you might as well just switch to a subset of C++ or D's “Better C” mode.
#define wl_container_of(ptr, sample, member) \
(void*)((char *)(ptr) + ((char*)(sample) - (char*)(&(sample)->member))) const std = @import("std");
const wl_list = struct {
prev: *wl_list,
next: *wl_list,
fn init(list: *wl_list) void {
list.prev = list;
list.next = list;
}
fn insert(list: *wl_list, elem: *wl_list) void {
elem.prev = list;
elem.next = list.next;
list.next = elem;
elem.next.prev = elem;
}
};
const Data = struct {
data: i32,
link: wl_list,
fn init(data: i32) Data {
return Data{
.data = data,
.link = undefined,
};
}
};
pub fn main() void {
var foo_list: wl_list = undefined;
foo_list.init();
var e1 = Data.init(1);
var e2 = Data.init(2);
var e3 = Data.init(3);
foo_list.insert(&e1.link);
foo_list.insert(&e2.link);
e2.link.insert(&e3.link);
var entry = @fieldParentPtr(Data, "link", foo_list.next);
while (&entry.link != &foo_list) : (entry = @fieldParentPtr(Data, "link", entry.link.next)) {
std.debug.warn("{}\n", entry.data);
}
}
[1] https://ziglang.org/documentation/master/#fieldParentPtr1) In C++ you can use struct inheritance to get the same functionality without macros or templates. ie: struct ThingList : wl_list { /* data elements */ }
2) Like in the C++ case above, the next and prev elements should probably be at the top if you plan on iterating a lot since the next/prev pointers will be in the first cacheline that you load when you dereference previous list elem. In this case, it's also worth allocating list nodes with posix_memalign.
OTOH, there might be reasons why you want the list buried deep in the struct, like if the other data element access times are more important.
I think the most common reason for this generic-location offset-of style is so that the struct data can have multiple list nodes and be on multiple lists (without any extra allocation). This is pretty common in the kernel.
No way to do this cleanly in C++. At best, it's going to be an iterator galore, at worst some boost-style frankencode.
https://stackoverflow.com/questions/15832301/understanding-c...
EDIT: check out some examples to see for yourself: https://elixir.bootlin.com/linux/latest/ident/hlist_add_head
struct foo { uint8_t data[3]; struct wl_list link; };
For an application link member may be at offset 3 while for another it would be at offset 4.
That may be the reason link should be put as the first member, and a simple cast would let you walk the list.
Unless you want to write them somewhere or put it as part of API and ABI. Which you probably shouldn't.
http://www.catb.org/esr/structure-packing/#_structure_alignm...
>It’s a simple doubly linked list. Here’s it’s definition
The first one's right, the second should be its.
Though almost most native speakers seem to do this, so I guess it won't be 'wrong' for much longer...
The gorilla guards it's bananas.
It makes sense now, thanks.
Oh right, thank you, I didn't know the history. Short version: https://www.merriam-webster.com/words-at-play/the-tangled-hi...
Did you mean, some people now alive were taught that "it's" is right (in "it's definition"), or just ownership generally (Harry's). Well, it's better than using apostrophes for plural's, which seems very common too.
Hence I recommend kernel refugees start with CCAN's list: http://ccodearchive.net/info/list.html
No but really, there isn’t anything to it. Just think about how a CPU sees memory and registers and it becomes intuitive. But that doesn’t make it easy to use because it doesn’t give you any guard rails to show you the one true way to accomplish a task.
The issues with C are around some of the weirder bits of the parsing process, namely the context-sensitive parser and the "interesting" corner cases of how the C preprocessor works.
Edit:
Here's a fun example - try it with and without the typedef line commented out:
typedef int T1;
int main() {
const T1 (*T1);
printf("%d\n", _Generic((T1), const int *: 1, const int (*)(int *): 2));
}Note that the type of a variable defaults to int in C, so you can write 'const a' instead of 'const int a'. The return type of a function also defaults to int.
In a different time, that something was left undefined by the standards might have meant that the result might be some artefact of the target machine architecture, but with aggressively optimizing C compilers the undefined behavior is more likely to be an artefact of the optimizer.
A thin light wrapper wouldn't compile this:
static int doub(i) { return i + i; }
int main(void)
{
int i;
int sum;
sum = 0;
i = 5;
while (i--)
sum += doub(i);
return sum;
}
into this: movl $20, %eax
ret
I'm not programming memory or registers, I'm instructing an optimizer.There's no undefined behavior in that example, as far as I can see. The point is that the compiler optimizes away all of the actual computation, since it's able to determine at compile time that the result will always be 0x20.
There is no undefined behavior in the example. It's meant to further illustrate my point that C is not a light wrapper around memory and registers. In the example, it's just a deceptively imperative looking way of describing the return value to the optimizer.
The compiler in this case was gcc with -O3.
For an example of behavior that is straight forward in the machine but is undefined in C, try signed integer overflow. Intuitively on x64 it should behave just like ADD, possibly adjust the carry and overflow flags and wrap around. In C, the behavior is entirely undefined. Some operation that results in UB is low hanging fruit for the optimizer. If it can deduce that the operation will result in UB, in this case signed integer overflow, it might omit the operation altogether.
Have you seen that claim backed up by any evidence produced within the last 25 years?
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
The main reason a language like SML or Haskell is difficult to optimize is because it is incredibly difficult to canonicalize for the purpose of making peephole optimizations and abstract interpretation effective. https://sunfishcode.github.io/blog/2018/10/22/Canonicalizati...
C macros on the other hand, should die. I use C a lot but never ever use c macros, they are evil. Extremely hard to debug, confusing and very weak.
When we want macros, we use Lisp, that have proper macros.
Except languages with parametric polymorphism, such as ML and Haskell.
The compiler doing my offsetting for me using type information is much safer and more reliable.
https://en.wikipedia.org/wiki/Mixin
Also, is this macro magic, or simply memory layout magic? The macros look like appreviations for readability, not magic required to make the structure work.
I don't see what this extra complexity of sizeof() offset calculations gains you.
Ref:
https://en.wikipedia.org/wiki/Offsetof
https://stackoverflow.com/questions/26906621/does-struct-nam...
I prefer not to use it due to the loss of type information and the potential for screwups, but you can write ANSI portable C using it.
Using __builtin_offsetof is undefined, because it is not standardised. The fact that __builtin_offsetof is used to implement offsetof by GCC does not affect the fact that offsetof itself is standardised (in the same way that a given internal heap operation is not standardised, but might be used by GCC to implement malloc which is).
Check e.g. https://docs.microsoft.com/en-us/windows/desktop/api/ntdef/n... and https://git.reactos.org/?p=reactos.git&a=search&h=HEAD&st=gr...
C++ Preprocessor is pretty much the same as C (they might go formally out of sync from o e standard release to another but in practice compilers implement the same features as far as I know).
http://man7.org/linux/man-pages/man3/queue.3.html
https://github.com/freebsd/freebsd/blob/master/sys/sys/queue...
ifndef TTD_SWITCH_STRING
#define TTD_SWITCH_STRING
#define TTD_CASE(str) if ([__s__ isEqualToString:(str)])
#define TTD_SWITCH(s) for (NSString *__s__ = (s); ; )
#define TTD_DEFAULT
#endifWhile you're at it, just make the feature generic. Handle arbitrary arrays, structs, unions, and floating-point types.
The ... feature supported by gcc is also a huge usability improvement. Don't allow weird sorting order issues with strings. Don't allow it for structs and unions.
const char * s const = "bar";
switch (s, strcmp)
{
case "foo":
printf("I got foo\n");
break;
case "bar":
printf("nopes, I got bar\n");
break;
default:
printf("neither foo nor bar :/\n");
}
Which would be fine/acceptable, but not exactly how C tends to roll.https://github.com/glouw/andvaranaut/blob/master/src/Theme.c
The problem: You have a string. (pointer to char with NUL termination) Depending on what that string is, you want to run different code. You want to do this with high performance and a minimum of fuss. No, it isn't an option to say "but what if I used an enum instead?". You have a string. A string is what you have.
Currently in C, you must choose between high performance and a minimum of fuss. Pick one.
The high-performance solution is to use bsearch or a perfect hash, mapping the string to an index of some sort. With purely standard C you would then use a normal switch. Using gcc extensions, you could get slightly better performance with a computed goto. Writing and maintaining this code is a pain, so most people don't bother.
The minimum-fuss solution is a whole bunch of "else if ... else if ... else if ..." that tries each possibility in turn. The performance is terrible, but good enough for stuff like command line option parsing.
For strings, the simplest syntax is to have the new behavior triggered by passing a pointer to any type which can be evaluated as an integer. That last requirement makes scanning for a terminator work in an obvious way. This does seem to lose the possibility of comparing arrays that are not NUL terminated, but that isn't the use case that people are dying for.
Since the vast majority of the desire to switch on strings is involving ASCII keywords, it would be reasonable to prohibit gcc-style ranges (like "case 1 ... 4:") for non-ASCII strings. This means that the strings can be treated like giant integers stored in big-endian form. (after the NUL, treat it as more NUL going on forever) Clearly it is well-understood how to switch on integer values, and thus we can switch on strings.
Right now the ways to switch on strings are terrible. If they are short, they can be memcpy into long long. With gcc one can use the computed goto extension, plucking values out of a table via bsearch or via a perfect hash function generated by something like gperf. It's all really miserable, leading low-performance programmers to just use cascading "if".