The trouble with struct sockaddr's fake flexible array
lwn.net
lwn.net
Probably, sockaddr_storage was introduced rather than making sockaddr's array bigger because by that time the definition of sockaddr had become a mystical sacred cow, not to be touched. Plus that size of 14 is the same on every system. It is a one size fits all approach which is not a way to define the maximum size structure over all the address types if you want every platform to be able to make it a tight fit. The size is going to vary with implementation due to different alignment and padding requirements.
int inet_connect(AF_TYPE, destination, port, options, &error);
Also: int inet_listen(AF_TYPE, port, options, &error);
So for example: int sock;
inet_err err = NULL;
sock = inet_connect(SOCK_STREAM, “google.com”, 443, NULL, &err);
if ( sock < 0 )
{
printf(“Failed to connect: %s\n”, err->message);
return -1;
}
Basically one level more abstraction in the API. The old API could stick around for people who are doing something weird, but this would cover most use cases I think. Added bonus: programs using this API are trivial to add support for IPv6, it is just a recompile.You just added a DNS lookup to connect(2).
Also, you linked socket creation to all of this. This would make it hard to do an asynchronous version. Today you can create a socket and do a non-blocking connect(2) (after doing your asynchronous DNS lookup)
Lastly, I think if you tasked someone with creating that API, they would probably use the old APIs to build it ...
If you're doing that, though, you'd better dang well give me secure key storage in the Kernel, too, or full TLS capability. Allow my program to n load it's RSA/ECC key at startup, then drop privileges to access that file, and have a kernel module handle the cryptography.
If there's enough volume of requests that a kernel level cache meaningfully improves real world scenarios, something is broken in those scenarios. :p (but I think you agree with me)
Certainly you could implement these functions using the existing BSD socket libraries, but that misses out on having them used everywhere so you get "free" IPv6 transitions and whatnot.
Your example would fit perfectly in a libinet library, but I doubt the C standard library is going to include improvements upon this. If you want good internet support, pick a library or, better yet, another language because plain C is terrible for managing arbitrary network inputs anyway.
There is also some stuff your code doesn't handle (do you retry when a connection fails but an SRV record lists multiple options? who is supposed to free the inet_err struct, because your example leaks memory? how do you handle attempts to connect over IPv6 on hosts still stuck on IPv4 only?) that I would expect a low-level C interface to be more explicit about.
> As long as this usage remains, the checking tools built into both compilers must treat any trailing array in a structure as if it were flexible; that can disable overflow checking on that array entirely.
No, they don't? Why couldn't they just have a mechanism to suppress the check for this particular struct (like a whitelist)?
I suppose C was made for Unix at Bell Labs and GCC is also inextricably tied, you could say that C is first and foremost the language of Linux, fair enough.
What I have not seen is this happening in the middle of the struct, for obvious reasons; it's always the last element in the struct.
Have you seen that in something that's ABI-critical, though? i.e. whose code simply cannot be changed due to backward compatibility, like is the case with sockaddr? Because otherwise I'd consider it a non-issue.
And each one can behave differently: https://lwn.net/Articles/908817/
It won't make existing uses any clearer or safer though, requiring rewrites to take advantage of them.
For defining or allocating an object that can hold any address, sa_storage should be used, mentioned in the article.
A function that takes a sockaddr * can dereference it to get at that field, and then know exactly which address type it's dealing with.
The same bind() API has to work for any kind of socket: AF_UNIX, AF_INET, AF_INET6, AF_X25 or what have you. That any kind of socket pairs with the matching address type, whose type is erased at the API level down to struct sockaddr *. But from the socket type, it is inferred what type it must be and the appropriate network stack that is called can check the type of the address matches.
I don't see how you'd get rid of this with a new bind_ex function, or why you would want to.
Of course we could have a dedicated API for every address family: bind_inet, bind_inet6, bind_unix, ... which is bletcherous.
Strictly speaking, I think we could drop the address family field from addresses, and then just assume they are the right type. The system API's all have a socket argument from which the type can be assumed. Having the type field in the address structure lets there be generic functions that just work with addresses. E.g. an address to text function that works with any sockaddr.
This surprised me. What kind of mechanism is there in place to prevent the compiler from formatting /dev/sda when you write code like this?
C is not a safe language. The more abrasive among fans of the language would say "skill issue" and "git gud" if you want to avoid footguns.
That's nothing to do with the language, that's always on the compiler. Nothing stops the compiler from taking even correctly-specified code and doing whatever it wants.
The thing stopping the compiler from doing dangerous behaviours in response to commonly abused UB is obviously that people wouldn't use the compiler if it did that. Just like how the thing stopping the compiler from doing dangerous behaviours in response to spec-legal code is that people wouldn't use it if it did that.
See also: "What Every C Programmer Should Know about Undefined Behavior": https://blog.llvm.org/2011/05/what-every-c-programmer-should...
Beyond that, I believe Linux and other kernels technically use a slight variant of C by taking advantage of compiler extensions/flags to better fit their use cases. For example, Linux compiles (compiled?) with -fno-strict-aliasing and -fwrapv and uses a GCC extensions that allows type punning via unions [0], so that they can compile what the standard calls "incorrect" C code without worry.
I'm not sure whether recent versions of C have changed their stance on this particular UB, but since it'll probably be a while until they're adopted (if ever) kernels will be making do with their workarounds for a while longer.
Within the body of a single function, the compiler can see where addresses came from and figure out what pointers may or may not alias. But when pointers are passed in as function parameters, the compiler has no idea. Thus, the C standard allows compilers to assume that pointers to different types never refer to the same object; an assumption that's true for almost all code anyone would want to write. But this adds surprising undefined behavior to some things you might try, like casting between layout-compatible structs, or the fast inverse square root trick from Quake. (Casting a pointer to a different type can still be tricky to get right even if it wasn't UB though. Alignment requirements are another footgun: for example, it's not legal to cast a uint8_t pointer to a uint16_t pointer unless you're sure its address is even.)
The blessed-by-the-standard way to do type punning is to either use a union or a memcpy, not a pointer cast.
For comparison, Rust prohibits mutable pointer aliasing for safe reference types, so the compiler knows when references may or may not alias. (Raw pointers are assumed to always be aliasable unless the compiler can prove that they're unique, e.g. by being derived from a safe reference). This leads to more efficient codegen in many cases, but it also means that type punning through pointer casting is fully legal in unsafe code (provided validity and alignment requirements are met), since the compiler does not need the type information to figure out aliasing.
"restrict" was added to give the compiler an idea.
And the API's should follow. For the folks who need more bytes, invent your own type. Or set socket attributes as everyone else.
You can determine the size by looking at sa_family.
Some platforms also have an sa_len field.
It's still unsafe if you allocate for the "base" type of struct sockaddr then try to use it for something larger, but that generally is not done. People usually allocate for the exact type they want, and only pointers to the structure are passed around, often opaquely.