Nolibc: A minimal C-library replacement shipped with the kernel
lwn.net
lwn.net
Nolibc: A minimal C-library replacement shipped with the kernel - https://news.ycombinator.com/item?id=34479284 - Jan 2023 (13 comments)
But on reflection, the reason why open source has been successful is all the duplicate projects. You pick the one closest to your needs and get your thing done. The ones that survive were simply meant to survive, for all sorts of reasons, but the least of them being cooperation for its own sake. Most successful open source is just somebody scratching their own itch.
So I'm wondering now about how you could encourage more of these one-off pet projects. What if all open source projects required a one-time lifetime license fee? The first time you use some open source project, you pay $1, and you have a license for life. You can adjust it for inflation, salary, cost of living, etc. But just imagine if every project were implicitly funded by everyone that used it. That would encourage more projects to appear, and incentivize them to be easier or better for the user. Plus the projects that are really popular yet have no funding, would finally have some.
Imagine if that idea grew to the point that a majority of software in the world were written by people who just got paid to write open source all day. A global economy of open source software. More software would become implicitly compatible, in order to deal with the deluge of competing software.
And so less expertise would be needed to combine different software into "a product". Companies would no longer need to spend as much time or money building up world-class engineering organizations just to kick out a CRUD app and take payments. Instead, companies could hire "integrators" to basically connect the unlimited array of "new apps" via "new pipes". A lack of expertise needed to make new products a reality would enable faster building of products, more reliably, with less expense.
Obviously, the more software that is used, the more we'd be paying for licensing. To deal with that, programmers would begin to merge their source trees together, to build more complete solutions, to lower licensing fees for users. More turn-key software, supported by more devs, that does more, for less money. Making it easier still to adopt the software and turn it into new products. In the end, programmers join forces to improve one product. But maybe that's just me scratching my own itch (˙ᵕ˙)
Also, at $1 per user, your projects would need 15,000 new users per year just to match USA minimum wage. There's a reason commercial software has high price points (e.g. $50 for video games) or operates on a subscription model. You'd need marketing for that kind of userbase growth, which implies (1) overhead and (2) annihilation of forums like HN under a tidal wave of spam.
Finally, as a developer, there's no way I'd pay $1 per open-source project because nearly all open-source code is crap. I'm willing to pay high prices for good software, which is why the niche of commercial software is saturated and any developer with a pulse can get $100,000 per year at a cubicle farm in San Jose.
Honestly compilers should straight up offer support for a Linux system call builtin that just generates the code in those stubs. They have support for insane calling conventions, there's no reason not to support something simple like that.
I think all we need is low barriers to entry; existing code bases virtually never manage to be all things to all people, so we want it to be easy to say, "the existing options don't do what I want; time to write something better!" (And then to do it)
struct auxiliary {
long type;
union {
char *c_string;
void *pointer;
long integer;
} as;
};
long main(long argc, char **argv, char **envp, struct auxiliary *auxv);
This is the complete set of parameters the kernel passes to the process. /* this function is only used with arguments that are not constants or when
* it's not known because optimizations are disabled. Note that gcc 12
* recognizes an strlen() pattern and replaces it with a jump to strlen(),
* thus itself, hence the asm() statement below that's meant to disable this
* confusing practice.
*/
static __attribute__((unused))
size_t strlen(const char *str)
{
size_t len;
for (len = 0; str[len]; len++)
asm("");
return len;
}
/* do not trust __builtin_constant_p() at -O0, as clang will emit a test and
* the two branches, then will rely on an external definition of strlen().
*/
#if defined(__OPTIMIZE__)
#define nolibc_strlen(x) strlen(x)
#define strlen(str) ({ \
__builtin_constant_p((str)) ? \
__builtin_strlen((str)) : \
nolibc_strlen((str)); \
})
#endifWouldn't it be easier to just use a simpler POSIX-like OS to begin with? Or better yet, a so-called "library kernel" that effectively turns a userspace program into its own OS?
What's the upside of using the (relatively large and complex) Linux kernel for this? I doubt it's reliability, since the kernel in this particular configuration is essentially untested.
More generally, the Linux kernel has a lot of code in it that is of generally reasonable quality and has received a lot of benchmarking/testing from well-resourced users. Sure, I could run some unikernel with a third-party network stack and SCSI drivers and ext4 implementation, but Linux already has all that stuff and it's ubiquitous. Why would I care about the extra ~30 MiB of RAM or whatever that it takes?
And that's before we even get to the topic of sharing code between environments. I can run the same unmodified binary on my desktop and on a minimal kernel-only Linux, which is not generally true of most alternative kernels.
It would be simpler, but harder.
The point of the linux kernel is that it's familiar. Many developers know it from inside, many wrote device drivers or file systems, its network stack is well-understood. There's plenty of reference and expertise.
POSIX is an ancient standard, which, while still useful, does not give you enough to be able to just switch to another compliant OS. Say, io_uring is not POSIX, and EBPF is not POSIX, and both can be hugely important for your embedded system.
With that, you can still run multiple threads with that highly minimal nolibc (not pthreads, but clone() is supported), AFAICT, so you can use a complex architecture. I didn't notice if cgroups are surfaced; with them, you'd be able to add defense in depth.
This is on top of Linux's superior tooling, compilers, a collection of drivers for nearly everything, and other creature comforts, usually FOSS.
The appeal is you can eliminate every single dependency. Linux is the only kernel with a stable user space interface. It's the only system where you can do this.
With a single 6-ary system call function, you can do anything you want on Linux. There's no need for anything else.
> What's the upside of using the (relatively large and complex) Linux kernel for this?
The kernel's massive amount of features and drivers for one.
> I doubt it's reliability, since the kernel in this particular configuration is essentially untested.
This is not a kernel configuration though. User space is completely separate from the Linux kernel.
RTEMS, Azure RTOS, AWS RTOS, NuttX, Mbed...
The system call numbers are stable on most unixes (including XNU/Darwin). It's just Windows that doesn't have stable syscall numbers.
Years ago I wrote a rationale for such a thing in my liblinux project:
https://github.com/matheusmoreira/liblinux/blob/master/READM...
I don't maintain it anymore because nolibc is better but there's a lot of references in that text if you'd like to read them.