Supporting Linux kernel development in Rust
lwn.net
lwn.net
People underestimate the role of copyleft licenses in preserving long-running FOSS products like GCC, Linux, etc.
Although this is probably not the best example because this is sold by the company that largely funds the open source project. Although that is a big conflict of interest.
https://blog.cloudflare.com/nginx-structural-enhancements-fo...
No, some bits are open, and some bits track what's actually being shipped. Much of what is open sourced lags behind what is running on end user machines by several years. iOS modifications to XNU and Darwin were never made open source.
Apple is also dropping GPL internals in favor of open source but not copyleft licenses so that they don't have to release source code changes to the software they release.
Furthermore, many non-GPL OSS projects have eventually gone proprietary (e.g. MongoDB's switch from AGPL to SSPL).
Half the linux foundation's funding comes from 6 companies. Not to mention the full time employees paid to work on it. And I find that there's similar patterns in a lot of other widely used gpl software. Similarly, MIT/Apache software gets tons of contributions from non-corp people too.
https://www.linuxfoundation.org/membership/members/
and that:
https://www.linuxfoundation.org/blog/2016/08/the-top-10-deve...
Every time it happens people are upset and then forget everything a day later. To me non-GPL + corporate involved is a huge flag, especially if it's a more complex or niche project that'd be difficult to fork.
As someone who works within LLVM professionally, I don't think this is particularly likely -- compilers are massive and complicated beasts, and the "moat" for proprietary improvements is small: they're either niche and therefore not of interest to the majority of programmers, or they're sufficiently general and easiest to maintain by releasing back upstream (where Apple, Google, and Microsoft will pay a small army of compiler engineers to keep them working).
Your concern is the one that kept GCC from stabilizing its various intermediate representations for decades, which is why virtually all program analysis research happens in LLVM these days.
Edit: To elaborate on the above: neither Apple, nor Google, nor Microsoft wants to individually maintain LLVM. Microsoft appears (to this outsider) to be actively looking to replace (parts of) MSVC/cl with the LLVM ecosystem, because they're tired on maintaining their own optimizing compiler.
(Rust does maintain a fork of LLVM, but it's just a couple minor patches that typically get upstreamed eventually)
...keep them working for their use-cases.
Are you under the impression that C and C++ code within Apple, Microsoft, and Google is fundamentally different from C and C++ code elsewhere? Because it isn’t. Google’s engineers maintain LLVM’s ASan for their purposes, but those purposes happen to be everybody else’s as well.
The necessity of open source for industry has both negative and positive implications for those of us who care about open source / free software as an ideal and not simply a tool of capitalism. The negative one (which GCC's leadership failed to really internalize) is that the number of engineer-hours at the command of for-profit companies is much higher than the number of engineer-hours at the command of community-driven projects. If you deliberately build a worse product to prevent corporations from using it for profit, given enough time, the corporations will replace it. The positive one, though, is that those engineer-hours are generally more valuable when pointed at some common cross-company codebase unless it is the specific thing that makes you money, and FOSS as an ideal provides a well-accepted model under which they can organize cross-company work (doubly so when they employ idealists like us as implementors). A compiler makes very few people money directly. It's generally a tool that you want to work well, and it's helpful to have other people run into the majority of problems and fix them before they cause trouble for the things that do make money.
So it's not surprising that LLVM is catching up with GCC, nor is it surprising that LLVM is and remains open source. If you are concerned about the LLVM monoculture, build a competitor such that it is cheaper / more profitable for companies to work on your competitor than to either work on LLVM or build their own compiler. Figure out what will make companies want to contribute and encourage it. (GCC did not do this, but it is perhaps slowly relaxing here.) If you are concerned about LLVM becoming proprietary, make it so that is cheaper / more profitable for companies to release their changes instead of holding onto them; that is, figure out what will make companies feel like they must contribute and encourage it. (One common strategy, used by Linux itself, is to maintain a high level of internal API churn coupled with genuinely good improvements in that churn and a policy that people must update in-tree callers; at that point, the more API surface you use in your private fork, the farther behind you get, and you'll watch competing companies outpace you.)
Interesting angle that I had never thought of as deliberate. As someone who works on a (bespoke) integration/embedding of Chromium, I could say exactly the same thing about it too.
So that the ppl at Google would keep your code working? And you didn't need to spend time resolving git merge conflicts, and compilation errors?
But you cannot, because some of those bespoke changes are "secret" and what you make money from? (And maybe some changes are off topic to Google)
I wonder how much time does it take to merge a new chromium version into your repo? Like, hours? Days? Weeks
Also, the fact that Chromium is a high profile security critical software with occasional emergency security updates for bugs exploited in the wild, doesn't help at all when you want to maintain your own fork.
This is not to say that I am not appreciative of all the work Google/Apple/etc engineers do in LLVM (I will be eternally grateful for the Clang targeting MSVC work).
Personally I don't care, but I bet many FOSS anti-GPL advocates will eventually care.
As mentioned I don't care, after university my main UNIX platforms were HP-UX, Aix and Solaris with their respective system compilers anyway.
Linux is already getting replacement candidates in IoT space via Zephyr, NuttX, mbed, RTOS, Azure RTOS, and who knows if Fuchsia will eventually get out of the lab, so be it.
No, for practical purposes, there is not. You can only enforce delivery of the source when a product based on GPL software gets delivered. But of course companies, which create software, know about this. Any usage of GPL software for delivered products only happens after the decision to publish the created software has been made. In doubt, companies tend to not use GPL software as part of deliveries.
GPL advocates try to claim otherwise of course.
I keep wondering if Microsoft opens up the Windows Kernel ( And Kernel only ), would it change the landscape much.
On Android Linux is an implementation detail, most of it isn't exposed to userspace, not even on the NDK as Linux APIs aren't part of the stable interface specification.
And ART is being ported to run on top of Fuchsia as well (https://android-review.googlesource.com/q/fuchsia), so...
_ph_ meant to write "This way, they can at least upstream everything which doesn't fall under these restrictions." (note the deleted comma before "which"), i.e., they can at least upstream something instead of nothing.
So the current options are: these companies don't open-source anything (which is what you get with GPL), or they open-source something (which is what you get with BSD, and in practice they open source a lot).
The claim that with the GPL they would open source _everything_ isn't true; they would just not open source anything instead. I also don't see any current legal framework that would allow them to opens ource everything without loosing a significant competitive advantage, and am very skeptical that such a framework can be conceived.
Since I am a commercial software user for most part, this is more of a philosophical question than anything else.
Lets see how long Linux will hold out against the new generation of IoT OSes all being MIT/BSD based, and what the bazaar will get out of them.
Or how long GCC will hold, when all major OSes use clang as their default compiler, and then how the baazar will go from there on.
I'm not sure this is true. See LLVM as an example. GCC existing and being GPL only meant that, e.g., Apple couldn't properly use it. Apple could have bought LLVM and kept it private, or develop their own proprietary solution and not make it open source, or fork the open source project into a private project and never contribute anything back, etc. There were many options.
Open source software has always existed, if anything, the GPL demonstrated that a particular open source model does not work well for a big part of the industry, while it works well for other parts, and non industrial usage.
Linux being successful seems incidental to it being GPL'ed, at least to me. Other open source OSes, like BSD 4.x, have also been quite successfull (powering the whole MacOS and iOS ecosystems, after a significant frankenstransform into Mach). Maybe Linux would have been even more succesfull with a BSD license, or maybe it would be dead.
I wouldn't consider BSD layer from NeXTSTEP an example of success for BSD's market adoption, given that not everything goes upstream, by now it hardly represents the current state of affairs.
If anything it represents what would have happened without Linux, all major UNIX vendors would continue to take pieces of BSD and not necessarily contribute anything back.
Its essentially what happens if (1) only one company wants to use the open source product, and (2) the open source product has a tiny community where no development happens.
In that scenario, there is no benefit from anybody forking the open source project (a private company or an individual) for upstreaming anything. It just costs time, but adds no value for them.
Linux and GCC never was like this (not even early days), and I think this is independent of its license. LLVM never was like this either.
In fact, there are many private companies that maintain a GCC fork, like arm, and due to the GPL need to provide its sources with a copy, which they do. But they are not required to reintegrate anything upstream, which they very often don't, and ARM support in gcc-arm from ARM is much better than on GCC upstream. A volunteer can't merge (review, rebase, ...) a 200k LOC patch on their free time, so once these forks diverge, its essentially game over. You'll need a team of volunteers equal in man power to what the private company provides.
So while the license affects which companies can use an open source project in practice, and what they can or cannot contribute. The GPL2 and GPL3 licenses are not a good tool to actually allow everybody to benefit from those contributions.
Maybe a GPL4 could require companies to upstream and get their contributions accepted, but IMO that would just get even less companies to use those projects.
BSD is Older than Linux:
https://en.wikipedia.org/wiki/History_of_the_Berkeley_Softwa...
Ah and there was that small legal issue back in the early 90's.
BTW: Market-share means nothing, you know Android is NOT Gnu/Linux
EDIT: Maybe your too young..but do you remember that SCO/Microsoft thingy with Linux
I surely remember it, except BSD did not ever had anyone like IBM jumping into the battle. :)
As for being young, thanks for the compliment
> ...as anyone coding since mid-80's....
And as information, during the early days, the Internet ran on commercial UNIXes.
They did not...the opposite is true, they shat their pants because shortly before they said "yes" to Linux, that's why SCO came...first time you had really big money (IBM) behind Linux, please don't change the timeline...it's kind of important.
>And as information, during the early days, the Internet ran on commercial UNIXes.
What do you wanna say with that? During the early days of Smartphones they ran on commercial OS's??
RIOT OS is LGPLv2.1.
Exactly; maintaining a fork is a big pain. The ongoing cost of keeping it up-to-date is a pretty big incentive to merge it, even if there's not a legal requirement.
with GPL you need a legal team to decide and plan how much to use, with BSD/MIT it can be an afterthought.
Of course none of it goes back to LLVM because updating your production copy of LLVM is just as messy and broken as the non-stable GCC IR you complain about and the fact that OSS development is still fundamentally incompatible with 1-year industry release schedules that shift all the time.
And you think if they didn't have the option of using LLVM they would have released an open source driver instead? That makes no sense to me.
But that's the exact choice Apple faced in 2005 and they did not choose gcc. They paid Chris Lattner and his team to develop an alternative compiler. Excerpt from wikipedia[1]:
>"Finally, GCC is licensed under the terms of GNU General Public License (GPL) version 3, which requires developers who distribute extensions for, or modified versions of, GCC to make their source code available, whereas LLVM has a BSD-like license which does not have such a requirement.
>Apple chose to develop a new compiler front end from scratch, supporting C, Objective-C and C++. This "clang" project was open-sourced in July 2007."
For some reason, the enthusiastic focus on the benefits of GPL principles seems to ignore the actual game theoretic behavior of actors in the real world to not choose GPL at all.
Of course the answer then reveals that this was only possible because there was comfy GCC to fall back on all along. They started in 2005, Clang became default in XCode with the 4.2 release in October 2011.
Ask yourself if you can sell your manager on 6 years of effort for no functional difference, likely even inferior.
LLVM is just a proprietary-able version of GCC. And thus compiler authors mostly focus on benchmarks while not caring about other metrics such as compilation speed, correctness and codebase quality / ease of contribution. There appears to be some tribal knowledge requirement to add a new target.
With all the hype and apple funding LLVM gets, it could have been better.
It isn't just that. It's also a compiler framework, somewhat usable as a library without being part of its codebase. And I really do mean somewhat usable; the LLVM experience for use as a library is not great, but it's wildly better than GCC's. I think that, not just the license, is a big part of what made LLVM successful.
The license was the biggest deal.
LLVM's license was certainly a factor, and I'd never suggest otherwise, but there were many other reasons that GCC couldn't help kick off a wave of language design and experimentation the way LLVM did.
GNU Pascal, Modula-2, Modula-3, BASIC and plenty of OEM derived ones were never part of GCC source tree.
I had to study GIMPLE and GCC integration as part of my compiler design studies back in the 90's, using those nice Walnut Creek CD-ROMs.
Sometimes in life it's better to be a bit heterodox, but have influence and leverage, than being alone on your high castle. GCC can be as open as it gets, but pushing people away in the name of freedom has actually only given reasons to the industry to move away from it, which is sad and could be avoided.
If he could just stop being a zealot for a split second and actually tried to understand the situation, he could have maybe been able to foresee that Apple had the people, resources and will to reimplement a whole compiler infrastructure that could threaten GCC's dominance, but he neglected it (it famously didn't care about Apple's offer to merge LLVM into GCC under the GPL).
Sometimes even the best of intents are shadowed by people's inability to compromise.
Sometimes the hate for a license prevents people to grasp what is ahead of them.
https://www.phoronix.com/scan.php?page=news_item&px=Sony-Tha...
or that:
There are some talks from Sony at LLVM meetings regarding those.
That PR from Apple has nothing to do with the bitcode reference I made.
I don't know, maybe other AMD users would like to get them?
>might provide clues how to bipass PS 4 security
>AMD users would like to get them
They got em:
https://www.phoronix.com/scan.php?page=news_item&px=Sony-Tha...
https://www.phoronix.com/scan.php?page=news_item&px=LLVM-10-...
https://www.phoronix.com/scan.php?page=news_item&px=Sony-LLV...
Why do i want ANY PS4 specific security feautures in the Compiler??
The FSF is shooting themselves in the foot with the way they handle many of their projects today, GCC is no exception. Free software purity is a great goal and all, but what use is it when nobody uses it or it falls behind more permissively licenses projects.
Sometimes in order to achieve your goals you must compromise or risk loosing the footing you already have, because you have close to zero chances of succeeding.
If RMS had compromised on GCC in the '00s, i.e. if GNU had spun off the C/C++/ObjC front end as a separate GPLv2-or-later library upon which something like clangd could have been made, Clang would have never existed. Yes, LLVM would have still been present, but as a special-case backend that used GCC as its frontend instead than its own, as they were planning to do since the beginning. All these improvements you talk about would have been done under the GPL, and not BSD licenses or proprietary. IDEs like Xcode and such would still be proprietary like they are nowadays, because there was a 0% chance of getting Apple or whomever to release them under the GPL.
It doesn't matter how much strong your moral principles are, or how much you value integrity. The world is definitely more pragmatic about software and values different things than RMS; while this might or might not be beneficial to our overall society, that's the way it is. Ignoring it is myopic, if it does not outright amount to shooting yourself in the foot.
As commercial software user, I don't have much issue with it, after all I have been coding since the early 80's.
And in spite of it and my occasional Linux rants, I am thankful for GPL, because without it there wouldn't be a UNIX to install at home to get my university work from DG/UX done without having to spend one hour traveling into the campus and fighting for a terminal.
Without Linux + GNU based userspace, the commercial UNIXes would all be around.
Can you prove that? EMC, Netflix and Sony thinks otherwise.
EDIT: And please stop with that LMAO (Sound like a childish Child from reddit)
?? What?
>guess why
Because protecting your games is kind of important for a Gaming-Console.
Your answered your own question:
>No, something like the PS4 CPU features that don't get upstreamed because they might provide clues how to bipass PS 4 security.
Would it be possible for a hypothetical new kernel, presumably written in Rust, to run an existing kernel such as Linux in a VM, just to tap into its drivers and emulate its syscalls? As more drivers are ported to the base kernel, the reliance on the donor kernel would shrink over time.
Edit: This was an aside. Of course I realize that the conference was about adding Rust code to the Linux kernel, and not about rewriting any large, important, and functional codebase in language-of-the-day™. Hence my speculative language: "hypothetical" new kernel, "presumably" written in Rust.
Interesting idea, though!
It would probably shake up the OS ecosystem to no small degree, but I imagine the fresh ideas that could be experimented with by non-experts would be hugely beneficial. Isn't that the big argument against so many new projects? "It looks cool, but the hardware support isn't there."
FreeBSD might be a better target though, since that's specifically designed to be compatible with everything.
NetBSD, perhaps, would be a better choice for "tries to run everywhere"?
There are kernels written in Rust, such as Redox, but those are separate projects, and that's not what this conference session or article are talking about.
Standing reminder: the Rust project is not a fan of rabid Rust over-evangelism (e.g. "Rewrite It In Rust", "Rust Evangelism Strike Force"), and sees it as damaging and unhelpful. We discourage that kind of thing wherever we see it. We're much more in favor of a measured, cautious approach.
I was really caught off guard by this stuff recently. I was around the rust subreddit a lot back in 2013-2015 and never saw stuff like that. Stepped out while working on other stuff and now it seems common. Super strange.
Examples: https://gitlab.com/fdroid/fdroidclient/-/issues/1049 and https://github.com/rapid7/metasploit-framework/issues/9092
Re-reading my comment I realize "you don't see this kind of behavior from the rust community" was unclear. I meant that people who are part of rust community don't see it, because it happens outside.
Agreed. If you want to rewrite everything in Rust, then please help out with a project such as Redox where the goal is indeed to write everything in Rust. But don't bother people by talking about that goal, just write some code!
BTW, I'm not a Rust dev, and it would appear the OP isn't a Rust developer either. They don't even seem aware of Redox. Seems odd to accuse them of "Rust evangelism".
So reading your last paragraph makes me feel welcome as someone that says "Rust is great - but rewriting everything in Rust will not solve security". Thanks. <3
The idea of that project was to turn Linux into a user mode application run by a microkernel; rewriting that microkernel in Rust would essentially give you what you're suggesting.
This reminds me of an interview with Linus Torvalds from many many years ago, when he was still very young. The interviewer asked him if he was afraid of getting replaced by someone young and hungry. Torvalds shrugged it of with something along the lines of: Nah, no one likes to do driver development.
That's why there exists crates like https://docs.rs/cpp/ which allow to embed C++ (and thus C) directly into the rust source code, simplifying the manual work and reducing the duplicated code.
C++ is a richer language than C and has more information about ownership (e.g., if you return a std::string you know how ownership works; if you have two arguments that are a char * and a length it's much less clear), but additional annotations in the kernel's C headers to convey this level of information in a machine-parseable way would be awesome.
The cpp crate is about convenience, but require the use of unsafe within the rust code.
The cxx crate do extra checks making the rust code safer but needs more boilerplate.
In the case of C however, these type checking are not really possible, as you say.
Two examples: My understanding is that there's a ton of code and added complexity in Swift itself to support Obj-C interoperation. And while you can call C and Obj-C APIs, as-is, there was/is basically a company-wide effort to write API wrappers (aka. overlays) for existing system APIs.
For the second part, maybe if Rust for Linux kernel programming gets to the level of popularity as, say, TypeScript in the JavaScript community, the Linux community will end up creating the needed API wrappers as an organic, group effort.
One can even completely ignore the RCW/CCW support in .Net entirely and just hack together some struct types that match the vtable layout.
Also you ignored that .NET was made to fully support C++ via Managed C++, replaced with C++/CLI on version 2.0.
C++/CLI ultimately still uses the same features the runtime uses for p/invoke calls, though it uses slightly different IL to transition to native code since it doesn’t have to worry about an external ABI within a mixed mode assembly. Sure, there’s a whole compiler that supports this weird mixed mode world from a language perspective - but it’s really boring at runtime.
Something that no UNIX does, and we only find similar ideas in Swift/Objective-C++, IBM language environments, and the old OS/2 SOM (Smalltalk/C++).
While everyone else just writes glue code with C like interfaces, as if the world has hardly changed from UNIX V6.
I agree that C as the lowest common denominator sucks.
Project Reunion seems to have rebooted everything, which is most likely the reason for it to have stalled.
https://github.com/solana-labs/rust
If this kind of seems interesting please send us a CV.
It’s typical GNU/Pendantry.
The kernel will not be shipping a rust compiler, patched or otherwise.
I can see why people would be opposed to that.
I'm thinking about LLVM (and, obviously, rustc)
Like GCC?
Your point is certainly valid, but I'm not sure it's a new problem.
That said, it would certainly be uncomfortable to move from "it compiles with Clang, though it's a pain, or GCC, which works fine" to "it compiles with GCC, if you don't use any of this functionality, or LLVM only, if you want any of these features or anything that depends on them".
So, we're currently expecting that supporting Rust will not place any requirements on the C compiler you use, unless you want to use cross-language LTO (which we'll eventually want to do, but it won't be a hard requirement for a working kernel).
Does Linux even work with LTO right now? It seems to do a lot of weird-linking-magic things that I wouldn't expect to work well with LTO.
https://linuxplumbersconf.org/event/7/contributions/798/ (slide link at the bottom)
It's not fully complete. But there is no reason why it couldn't be.
But no, there are other reasons outside of trademarks, some technical ones are in submisson’s article.
We do also, you know, give permission.
https://lwn.net/Articles/118279/
Debian wanted to allow anyone to patch their firefox. Mozilla said OK, but not with our name attached to it.
And this makes sense on some level : No project wants to provide support for someone else's bugs.
1. https://wiki.hyperbola.info/doku.php?id=en:main:rusts_freedo...
GNR: Gnr Not Rust
TRust: Treated Rust
...
Some attempts: GRR: Gnu Renames Rust, GAR: Gnu assimilates rust, GER: Gnu non Est Rust (gnu is not rust, but in latin to get an E), GIR: Gnu Isn't Rust, GOR: Gnu Oxidizes Rust, GUR: Gnu Unseats Rust.
You can also go for GunsNRoses, but they won't like it and complain. Which might be good for visibility at the beginning.
GNU can either cut a deal with the Rust Foundation or use fair-use previsions afforded by trademark law. I think there is some leeway for cases like these.
It's a rant by an obscure one-developer distribution that formerly used the Linux kernel. Some people are never happy no matter what you do, and if you make the mistake of treating their rants as useful feedback, you get a list of 30 more demands before they'd consider gracing your little project with the honor of their all-important usage again.
> Linux is not "forcing adoption of HDCP" (it's just something Linux has a driver for)
What they actually said in the article:
> Historically, some features began as optional ones until they reached total functionality. Then they became forced and difficult to patch out. Even if this does not happen in the case of HDCP, we remain cautious about such implementations.
> Linux kernel forcing adaption of DRM, including HDCP.
(Which linked to a patch adding driver support for handling HDCP, with absolutely no possibility of it somehow being forced.)
Regarding the quote from that interview: sure, and there's also no guarantee that the Linux kernel won't drop support for every architecture except SPARC, except of course for all the people involved not putting up with it. In what possible universe would Linux developers decide it was a good idea to mandate HDCP? It's completely unfounded and unsupported hyperbole and fearmongering, without even a hint of potential truth.
It's also roughly consistent with what I'd expect from someone who thinks that gettext daring to have support for Java and C# format strings is a horrific problem that should be ripped out by the roots.
There were specific reasons for that: https://yoric.github.io/post/why-did-mozilla-remove-xul-addo...
Mozilla markets Firefox as ‘the only browser made for people, not profit’. But they’ve done the exact opposite time and time again. Virtually the single most prominent UI element, the omnibar, has been sold out to the single largest threat to the open web, the gatekeeper to the entire internet to the vast majority of its users. I don’t want to spend all day railing against Mozilla, as easily as I could, but I can easily find a couple other examples off the top of my head. The Mr. Robot scandal is another example of selling out and (not necessarily directly harming but very significantly) alienating users. And it’s been a few years since I’ve used FF as my main browser but I still remember every time I installed it on a new device going into settings and changing default after default to revert it back from a profit-seeking product to a personal tool, disabling telemetry, etc.
1: Although it does make a difference to me. I’d use pre-Chromium Edge if it was FLOSS, on my OS, and let me use Pentadactyl. The excuses for XUL’s removal — security, stability, and speed — don’t concern me: https://yoric.github.io/post/why-did-mozilla-remove-xul-addo... seems to claim it would allow someone to somehow steal my passwords. I’ve been using Palemoon for years and I don’t have any weird activity on any of my accounts, and I keep my browser sandboxed (something Firefox should do better by default, one way it’s behind Chromium) so I don’t have any other security worries. I have no stability or speed issues, and maintainability seems like a non-issue given that an apparently one-man main team keeps Palemoon running.
Where? I was curious as there has been no upstream improvements. And there are no new code in what I assume is their upstream repository?
This is debatable. The default workflow requires access to crates.io, and it is pretty hard to not use crates.io.
Even with using crates from crates.io.
As someone who uses a source-based distro I hate language package managers. They always end up inferior to a distro package manager and create extra work for everyone, including the language developers, who tolerate it because it is in their favorite language.
> Some users have correctly mentioned that many other software packages have trademarks, do we plan to remove them all? No. We are not against all trademarks, only those which explicitly prohibit normal use, patching, and modification.
> As an example, neither Python PSF nor Perl Trademarks currently prohibit patching the code without prior approval. They do prohibit abuse of their trademarks, e.g. you cannot create a company called “Python”, but this does not affect your ability to modify their free software and/or apply patches.
> Due to the anti-modification clause, Rust is a non-permissive trademark that violates user freedom.
[0] https://wiki.hyperbola.info/doku.php?id=en:main:rusts_freedo...
> Licensed redistributors of Perl code are permitted by their licenses to use the Perl logo in connection with their distribution services, on product packaging, and in promotional materials.
IANAL, but that text would especially suggest to me that you would have no right to use the trademark even to distribute an unmodified binary of Perl without approval.
I suspect the real reason is Debian and Mozilla had a very public, well-known spat over licensing, and Python and Perl have not had very public, well-known spats with anybody.
Shame really.
But it seems, from an RESF standpoint, that the fact that any C code is out there running at all is a crisis-level problem.
It's very Chrome and C++ specific, but the problems will likely be similar in practice.
And for what it's worth, for almost all langauges C <-> X interoperability is a lot cleaner than C++ <-> X. Since almost all languages (including rust) "speak C", both in terms of compiler support and and being able to map every important C concept to a concept in their language. The same isn't true with C++.
One thing to worry about with adding any language (including rust) to the kernel, is that the next good looking language will probably interoperate well with C and not interoperate well with Rust. So there is a much higher cost the second time you try to add a language.
This is something that Rust is acutely aware of. There have been many discussions about the idea of a higher-level "safe ABI", such that languages with notions like counted strings or bounded buffers could interoperate with each other without having to go by way of an unsafe C interface.
I'd say the biggest questions for Linux that may be blocking here relate to LLVM not being available on nearly as many platforms as GCC (many of which will likely never be supported cause they're legacy systems that are only supported in the sense that they aren't purposefully broken) which limits it's use in any core-ish kernel systems.
The safety guarantees that Rust provides are neither unique nor complete, and if we discuss the amount of effort necessary for bringing new interfaces like the one this article mentions ("one can define a kmalloc_for_rust() symbol containing an un-inlined version"), we should compare it to other existing solutions, like ATS [1], that are designed to be seamlessly interoperable with C codebases without giving up on the safety side of the argument.
Just a week ago there was a link in ATS reddit channel that advertised ATS Linux [2][3] initiative, that may be of great interest for the same people who are interested in bringing more memory sefety to Kernel development, without giving up on existing C interfaces and the toolchain[4].
[1] http://www.ats-lang.org/Documents.html#INT2PROGINATS
[2] https://www.reddit.com/r/ATS/comments/ibyczp/ats_linux/
[4] http://ats-lang.sourceforge.net/DOCUMENT/INT2PROGINATS/HTML/...
There can be multiple great C replacements.
Not being "perfect" in terms of the guarantees you provide doesn't mean you aren't a great improvement. Nor would not providing any guarantees at all mean that a language is necessarily not a great improvement.
That said, I would be very interested in seeing a comparison between C, Rust, and ATS in practical code. This is the first I'm hearing of ATS and it sounds interesting.
I didn't claim that Rust isn't a great replacement overall, I was specifically mentioning that it's not great for existing C codebases, by the fact that it requires significant work to replace internal interfaces that may be important for performance-, backward-compatibility- and conventional tooling reasons. It doesn't interoperate with C as well as other alternatives are capable of, whilst not being objectively better at safety guarantees either.
[1] http://www.ats-lang.org/MYDATA/SPPSV-padl05.pdf
[2] http://ats-lang.github.io/DOCUMENT/ATS2TUTORIAL/HTML/c1267.h...
I’ve never used ATS to do what the GP is suggesting of incrementally wrapping C with proofs, but I can’t imagine this being simpler than just rewriting the C in Rust. You wouldn’t get the same proofs, but you would get memory and thread safety, which for many apps would be an incremental improvement.
I haven’t taught ATS, but I found it harder to learn than Idris, and I can’t imagine C programmers which have a hard time with Rust learning it quicker than they would learn Idris or Rust.
sometimes it's impossible without falling back to unsafe Rust, which kind of undermines the initial incentive. C codebases heavily utilise pointer arithmetic programming.
Simplicity is a good trait though, and there are a few promising initiatives in that regard in ATS3 - https://github.com/githwxi/ATS-Xanadu#project-description
In reality most Rust code hardly ever needs to be unsafe and the existence of unsafe code in libraries you use hardly ever has any impact on the security and stability of what you ship, because there just isn't very much of it compared to the safe code. The "trophy cases" for Rust fuzzing bear witness to this.
> C codebases heavily utilise pointer arithmetic programming.
You don't write unsafe Rust code everywhere C code would use pointer arithmetic. You use safe Rust idioms and APIs instead.
If it turns out that writing Linux drivers in Rust requires writing a lot of unsafe Rust code in each driver, then that would certainly be a failure. I don't see any reason to believe that will be the case.
I haven't tried F*, but from quick googling of "aliasing", it seems that a similar kind of checks can be achieved with view-changes in ATS (but please correct me if I'm missing the point) - http://ats-lang.sourceforge.net/DOCUMENT/INT2PROGINATS/HTML/...
I'm not sure what you mean by "FFI" (I would apply the term "FFI" to the way ATS interoperates with C, too), but you can certainly get cross-language LTO between C and Rust, since they both compile to LLVM IR.
> you can iterate indefinitely at the desired pace to bring new layers of formally verified API in places of previously unverified C calls without breaking backward compatibility and expected runtime profiles
You can do this in Rust, too.
> You can do this in Rust, too.
not always within boundaries of safe Rust though (assuming each iteration is allowed to brings zero unsafe/unverified code), which undermines the initial safety argument.
Rust's Box<T> and Option<Box<T>> have the same representation as a C pointer to the same type, so this is true of Rust too.
You do have to wrap the type in source code to indicate what the ownership guarantees/contracts are. (Is this not also true of ATS? That is, if ATS code calls a C function that returns a pointer, how do you know how long that pointer is valid for and whether you're expected to free it?) But there is no runtime overhead/conversion.
> which undermines the initial safety argument
No, adding unsafe Rust does not undermine the safety argument (at least, no more than having any C somewhere in your address space does). See my reply in the comment section of TFA: https://lwn.net/Articles/830158/
(It's unfortunate that the Rust language used the "unsafe" keyword to mark blocks where you can access raw pointers, because it causes people to think that using such code is unsafe. Something like "manually_verified_safe" would have been better - the whole point of putting an "unsafe" block on a safe function is to write code that is safe to use that cannot benefit from the Rust compiler's checks. But at the boundary with C, it is impossible to have automated checks, because the relevant typing/safety information doesn't exist in C in the first place. It's no different from ATS code using $UN where you've manually verified that you're using a C function or C-created data in the way it's intended to be used.)
But it does, see my answer in the thread below regarding ring buffers - https://news.ycombinator.com/item?id=24337184
> Is this not also true of ATS? That is, if ATS code calls a C function that returns a pointer, how do you know how long that pointer is valid for and whether you're expected to free it?
the C implementation needs to be known, then the signature (interface) of the extern-annotated C function would define a pointer that will be associated with a proof object. Every ATS program that compiles must consume all instantiated proofs (with a proper proof-consuming function), and the proof itself can indicate whether a pointer is stored/taken care of inside the C function or it should be taken care of outside the function (the latter is indicated by a prefixed "!" to the pointer type). An example could be found here - https://bluishcoder.co.nz/2012/08/30/safer-handling-of-c-mem...
First, you're talking about two different things here. Originally you were talking about incrementally adding Rust bindings to C code and how that obligated you to write "unsafe". Now you are talking about pure-Rust implementations of data structures that use "unsafe" for optimization purposes. The second one is much more dangerous.
Second, it does not undermine the safety argument - quite the opposite, as demonstrated by the fact that code that used VecDeque remained unchanged after the bug fix. Everything that the Rust compiler proved about code that used VecDeque, on the assumption that VecDeque was sound, remained just as proven after the bug was fixed and VecDeque was changed to be actually sound. There was no ex-falso-quodlibet problem. There was a clear delineation in the safety argument: on one side, a human was obligated to check that VecDeque was implemented soundly, and on the other side, the Rust compiler checked that every user of VecDeque did so soundly. The first part was done incorrectly and then fixed; the second part remained done correctly.
> the proof itself can indicate whether a pointer is stored/taken care of inside the C function or it should be taken care of outside the function (the latter is indicated by a prefixed "!" to the pointer type).
Then this is an axiom as input to the proof, not a proof itself. That is, ATS is no more capable than Rust of magically determining what the unstated invariants of C code are; it can only prove that certain things logically follow from certain assumptions. This is precisely what Rust does, too, except it is clearer than ATS because any use of such unchecked assumptions are clearly marked with the "unsafe" keyword.
In the linked ATS code, I see no indication, as a human reviewer, about which parts I should carefully audit (e.g., the annotation of the "extern fun") and which parts ATS has checked for me (e.g., the implementation of "string_to_base64"). In Rust, you can very quickly search a codebase for "unsafe" and see what needs auditing. (And in fact this technique empirically works well for finding unsoundness in production Rust code, and I can point to multiple examples of it. I wonder what people do when reviewing production ATS code.)
Again, I'm not saying ATS isn't cool. I'm not saying it wouldn't be great to have ATS support in the Linux kernel. Please work on this. There are, in fact, lots of things that ATS can do that Rust cannot do, and they are helpful to have. But I think the specific things you are claiming are things that Rust can do just fine or that ATS in fact cannot do.
I've never stated that there's such a capability. I was saying that unlike Rust, ATS can be used to gradually rewrite every C call into a safe and formally-verified equivalent function that has exactly the same performance and memory charactersitics as the original C function.
compared to what?
Maybe I should've highlighted it in my initial comment, because I definitely agree that this list is non-empty for Rust. I was mentioning specifically existing C codebases that are heavily used in production around the world (last time a similar topic was related to QUEMU project).
In any case, if one is discussing modern and safe languages, there's a dramatic difference between using a C library and using a library native to the language. Even ignoring safety, the ergonomics/developer experience is dramatically different and the impedance mismatch can be very frustrating (for a "best case" example of this, Swift puts a lot of effort into exposing Objective-C interfaces as "swiftier" APIs automatically, but this benefits significantly from a relatively opinionated set of idioms in the source language, something arbitrary C libraries do not have).
You're right that being able to import a C library directly is a very nice feature.
> And Rust libraries regularly suffer from the same lack of formally-verified and proven-to-be-safe APIs.
Right... but having a larger community means there's a much higher chance of a safe or easier-to-use library existing for any particular task. In particular, I think it means there's almost certainly more libraries in total, with more safe libraries and more unsafe libraries overall, so simply comparing number (or proportion) of low-safety libraries is misleading.
(Formally verified is shifting the goal posts here: the bar is just "has a native library".)
The dramatic nature of that difference is whether that library exists or not. Odds are, if exists then it's written in C. Thus this point is mute with regards to Rust because at best it's relegated to a nice-to-have, in the sense you can enjoy the same features that are already available in C but with language-specific assertions.
It is, but it is for a good reason - ATS formally verifies your safe functions, so you only need to implement your proof once and make a compiler agree with you, and then you can save on a community-driven code-review process. Rust, on the other hand, provides safety guarantees only for a subset of safe guarantees that ATS provides. Pointer manipulations have to reside in "unsafe" blocks, and that leads to CVEs - https://gts3.org/2019/cve-2018-1000657.html
Focusing on pointer manipulations is also misleading: it's entirely true that it's dangerous, but most Rust code does not need to do any sort of raw pointer manipulation. For instance, there's extensive work on operating systems and other low level code that involves both inherent unsafety due to hardware specifics (that is, ATS almost certainly does not model it natively) and safe Rust wrappers (i.e. proofs of safety) for the unsafety at a surprisingly low level:
- a recent series: https://www.ecorax.net/as-above-so-below-1/ https://www.ecorax.net/as-above-so-below-2/
- a long-standing operating system: https://os.phil-opp.com/
Finally, if you are happy to cherry-pick specific examples, https://bluishcoder.co.nz/2017/02/22/borrowing-internal-poin... discusses a case where Rust is able to prove more things safe than ATS 2 (at the time of writing).
Plus... this is still ignoring all of the other factors why popularity and momentum are useful reasons to choose a language.
You cannot say that while this kind of CVE is possible https://gts3.org/2019/cve-2018-1000657.html and as long as Rust is not capable of performing safe pointer manipulations. The above issue may be solved for VecDeque in stdlib, but what about other data structures and algorithms in the wild?
> Focusing on pointer manipulations is also misleading
Remember that the discussion is happening in the topic about brining Rust into well-established C codebase, and C uses pointer-based arithmetics all the time, sometimes the whole algorithms are implemented this way because it brings efficiency to lower-level interfaces that are supposed to be very fast. If the motivation to bring a new tool is pronounced as "let's make it safe", why should we allow the argument to fallback into the "unsafe Rust" territory?
> For instance, there's extensive work on operating systems and other low level code that involves both inherent unsafety due to hardware specifics
I'm not going argue against that, because it's not related to the current topic related to Linux Kernel.
> Finally, if you are happy to cherry-pick specific examples, https://bluishcoder.co.nz/2017/02/22/borrowing-internal-poin.... discusses a case where Rust is able to prove more things safe than ATS 2 (at the time of writing).
it's not more, it's one specific example. Shall we now enumerate all the things that both languages are able to prove? I think we should, to make it clear what the differences are and what is possible to prove safe. That's exactly my point about comparing brining Rust into existing C codebases with alternative tooling available specifically for C codebases.
> If the motivation to bring a new tool is pronounced as "let's make it safe", why should we allow the argument to fallback into the "unsafe Rust" territory?
Because going from 0% of code proven safe by the compiler to 99% is very valuable. Verifying the remaining code is desirable but not as valuable as what Rust already provides.
Having said that, it would indeed be great to have a proof system for verifying properties of unsafe Rust code. Much work has been done in that area: https://alastairreid.github.io/rust-verification-tools/ Something to look forward to.
one has to do it all the time if cyclic mutable graphs or specific buffers/caches are involved.
> Verifying the remaining code is desirable but not as valuable as what Rust already provides.
How do we know that? What are the criteria and the thresholds that lead us to that conclusion? Is it true for all fields where the language can be used? What should we do about inability to express more precise constraints at compile time? Shall we stop on Rust, or try to embrace more powerful tools that already support Dependent Types in low-level systems programming? These checks enable a whole new world of expressive powers and correctness guarantees, even compared to the cool borrow-checker. ATS supports them today [1]
[1] http://ats-lang.sourceforge.net/DOCUMENT/INT2PROGINATS/HTML/...
"Specific buffers/caches" is ambiguous. For embedded systems there are crates that provide safe interfaces to memory-mapped hardware.
You will likely argue that using unsafe code in a library is just as bad as writing unsafe code, even if that library is used and tested by a lot of people and the unsafety is corralled behind a safe API. You would be wrong.
> How do we know that? What are the criteria and the thresholds that lead us to that conclusion?
My current project is 170K lines of Rust code, and has 225 uses of unsafe. That's about 1.3 uses of 'unsafe' per 1000 lines of code. If I could write C++ code and introduce less than 2 exploitable vulnerabilities per 1000 lines of code I'd have an even higher opinion of myself than I already do.
> Is it true for all fields where the language can be used?
I'm sure we're both imaginative enough to dream up some "field" narrow enough to disprove any universally quantified proposition.
> What should we do about inability to express more precise constraints at compile time?
We should adopt proof systems that let us verify safety properties for those little bits of unsafe Rust code. We should not, however, make that a precondition for writing that vast majority of code that can be written in safe Rust in safe Rust.
> Shall we stop on Rust, or try to embrace more powerful tools that already support Dependent Types in low-level systems programming?
That is a false dichotomy.
Safe Rust is a sweet spot where the compiler and tools can verify a strong set of safety properties without the developer having to deal with proof systems and dependent types, with lots of engineering to produce helpful messages when things go wrong. Plus a large library ecosystem that, among other things, provides lots of safe abstractions over unsafe code. Trying to put the brakes on Rust and get everyone to buy into ATS instead is putting the needs of the few over the needs of the many.
You can, because the libraries that you use fallback to unsafe blocks. How about those who want to write these libraries without having to use unsafe?
> You will likely argue that using unsafe code in a library is just as bad as writing unsafe code, even if that library is used and tested by a lot of people and the unsafety is corralled behind a safe API. You would be wrong.
just because you think I'd be wrong doesn't prove me being wrong. Why should I rely on unsafe if I can implement the same algorithms and data structures with proper safety guarantees, but not in Rust? Why should I stick to Rust specifically for that matter?
> My current project is 170K lines of Rust code, and has 225 uses of unsafe. That's about 1.3 uses of 'unsafe' per 1000 lines of code. If I could write C++ code and introduce less than 2 exploitable vulnerabilities per 1000 lines of code I'd have an even higher opinion of myself than I already do.
yeah, it shows your personal commitment to Rust toolchain. This is good. But what if you've already acquired some knowledge of more powerful tooling out there, and you're given a task to estimate applicability of these two toolchains to the development of Linux Kernel?
> We should adopt proof systems that let us verify safety properties for those little bits of unsafe Rust code. We should not, however, make that a precondition for writing that vast majority of code that can be written in safe Rust in safe Rust.
Why shouldn't we do it specifically for Linux Kernel? Why shouldn't we use compile-time checks for element incusion, length-based non-emptiness of containers instead of implementing another set of runtime validators?
> Safe Rust is a sweet spot where the compiler and tools can verify a strong set of safety properties without the developer having to deal with proof systems and dependent types
How do you define a sweet spot? Is it indeed a sweet spot when it comes to kernel development? Why does this spot happen to be at the level of Rust type system, and not somewhere else?
> Trying to put the brakes on Rust and get everyone to buy into ATS instead is putting the needs of the few over the needs of the many.
I'm actually arguing here from a position of technical merit of two tools in the context of Linux Kernel development, but for some reason you bring vague and non-technical definitions of "sweet spot", "lots of people using something", "needs of the few vs needs of many", which I've not mentioned anywhere in the thread.
(The main difference seems to be ATS allows inline C, since it looks to be tied to C as a compilation target.)
Choosing an obscure language with little community support imposes real-world development costs: it's harder to find or ramp up new developers, there are fewer eyes identifying bugs in the implementation, tooling support can be subpar, documentation and blog posts are harder to find, etc.
(btw I know nothing about ATS so I'm not saying this is a good description of that language in particular.)
If PR == popularity, yes.
Otherwise we would all be writing our stuff in Haskell, Lisp, Elm, F# and similar, and Algol-derived languages would be a footnote at this point.
That's not what this is about, so even though you aren't tired of posting this, I'm not sure it's warranted.
This is about the kernel supporting other languages in select places where it makes sense, such as kernel modules, where there's already an API in place to conform to. It's also not limited to Rust, it's about providing support for multiple other languages that could be beneficial in those locations.
> I won't get tired of repeating the same comment in every topic that suggests Rust to be a great replacement of C in existing projects, that it isn't.
Nobody is giving up on existing C interfaces. They are simply making sure those interfaces are not actively hostile to languages that aren't C by nature of how they are implemented in C.
To be clear, the "kmalloc_for_rust" was mentioned as one possible way forward. Another would be a more advanced Rust bindgen tool that can determine how to deal with it automatically. Neither affect ATS, based on your description (although a "kmalloc_for_extern" might generally help any non-C language, which would probably be beneficial).
If ATS doesn't require any changes to C to interoperate with it, this whole announcement should have zero impact on ATS, other than possibly making it easier to get an ATS module in-kernel since they're trying to make it easier in general for non-C modules.
You can treat this as a zero-sum gain, where Rust or Go's gain is ATS' loss, in which case you might as well pack up your bags now, since the writing is on the wall based on publicity, or you can treat it as a rising tide raises all boats type situation, where more general acceptance of alternatives to C is still beneficial to ATS actually being allowed in-kernel, even if it doesn't play exactly to ATS' strengths. I know which one I would put my money behind as being a better strategy in the end.
But what are these benefits that other languages bring on the table? From the article it appears that proponents of the new toolchain mention memory-safety as the primary concern:
> They focused on security concerns, citing work showing that around two-thirds of the kernel vulnerabilities that were assigned CVEs in both Android and Ubuntu stem from memory-safety issues. Rust, in principle, can completely avoid this error class via safer APIs enabled by its type system and borrow checker.
So the natural follow-up question to that concern, it seems to me, would be whether there's something in the existing C/GCC toolchain (or around it) that can resolve the raised concerns without adopting a whole new toolchain (Rust/LLVM) and without having to extend the existing interface with language-specific suffixes like "kmalloc_for_rust".
The answer to this question appears to be: no.
Memory errors sit somewhere around 2/3 to 3/4 of errors irrespective of codebase when C is involved (this has been consistent in things like the Linux kernel as well as the Windows codebase). C++ hasn't seemed to have changed that, either.
Remember, Rust IS Mozilla's answer to this question.
> without having to extend the existing interface with language-specific suffixes like "kmalloc_for_rust".
The requirement to do that is pointing out the that kernel API is poor and is too intimately tied to the build chain. We have seen this before--when Linux got ported to Alpha, for example.
The point is to fix the API so that more languages than just Rust can be used.
I suspect that more "system programming" languages may appear once Rust does the difficult groundwork of breaking this kind of intimate toolchain dependence. As things currently stand, why develop a system programming language given that no one will ever use it?
C++ is not like C. Putting some C++ objects in a C codebase would, indeed, not change it much, but if you write in modern C++, some memory errors will literally disappear, and some become less likely.
Caveat: This is based on language mechanics, standard library facilities and recommended idioms; of course people can write unsafe C++ if they want to.
The problem is knowing the safe subset of the language requires continuous study and deep understanding. It also has changed with time.
I used to have a bookshelf in order to understand C++. I suspect that has not changed.
I do not need a bookshelf to understand the "safe" subset of Rust.
It's true that writing safe code requires some study and some understanding. But you don't need to "know what's safe". That _is_ complicated - there are a lot of complex depths to the language. The thing is, the non-expert user can avoid them.
For example, a green developer can start with "Avoid using new and delete directly; use vectors, other containers, or unique_ptr's instead. Maybe shared_ptr's in some cases." If at some point they really need to write their own allocating class, they'll be told to be very careful with the allocation, and take into account the possibility of exceptions, concurrent execution by other threads etc.
Also, when you're not sure of yourself, there's a good enough reference in the form of the C++ Core Guidelines [1]. (That's a long document, but you don't need to read through it.)
Finally, a whole lot of memory error hazards / unsafe code can be statically identified and discouraged. Gradually, IDEs and other tools have started doing so.
[1] - isocpp.github.io/CppCoreGuidelines/CppCoreGuidelines
I'm not optimistic about the safety of modern C++. Most people who really care about safety long ago left the C++ building.
So modern C++ will have to do for the few cases that require going out of them, e.g. runtime plugins for native agents, COM/UWP with access to OS APIs, IDE mixed language debugging, GUI components.
Also there is the whole issue that GPGPU shaders and Apple/Google/MS driver stacks are mostly about C++.
So one either does managed language + C++, or managed language + C++ + additional glue effort + Rust, not much to win currently.
I can give examples with other managed languages that are more at home with C++ infrastructure for the time being, like JavaScript, Swift, Julia.
* Dangling references from lambda captures are no different than dangling references or pointers passed without lambdas, and not more likely to be used. * spans are no worse, and usually better, than using a raw pointer and a size; they don't make memory errors more likely. * string_view - indeed, it potentially introduces a dangling pointer. But - that only happens if you store it. Which is why I qualified my claim to "using recommended idioms". If you replace your `const char*` or `const std::string&` with `std::string_view` you have not introduced a potential memory error.
More generally - new language/library construct which do not manage their own memory can introduce memory errors. But, luckily, C++ now has a better "harness" of protecting you from many of these without costing you and performance nor introducing now memory issues.
So I stand by my claim. Having said that - C++ is not a safety-guaranteed language, and if that's what you need then it is indeed not the language to use.
Of the venn diagram of languages that provide more safety than C (and that is, either require it or make it hard to bypass), and languages that people know know or are likely to know in the semi-near future, Rust and Go are probably the big contenders.
ATS may be easier to interface with the kernel in, but is random peripheral company writing a kernel module driver likely to know it and use it? Theoretically, everyone can write most or all of their modules in Coq, right? So why don't they? Why is ATS different?
> So the natural follow-up question to that concern, it seems to me, would be whether there's something in the existing C/GCC toolchain (or around it) that can resolve the raised concerns without adopting a whole new toolchain (Rust/LLVM)
You're mistaking the aims of a project to specifically adapt Rust to the linux kernel and what they focused on. This note is frm a Rust initiative, not a kernel initiative, why would they focus on doing something in a different language?
> and without having to extend the existing interface with language-specific suffixes like "kmalloc_for_rust".
I addressed this in the last comment. Did you have something new to bring to this aspect of the discussion? It feels like you're ignoring how I noted that this is but one option people are discussing, and using it as a reason why the whole thing is a bad idea. That doesn't make sense.
In fact it would be nice to have eBPF backends for them as well, naturally only possible for specific language subsets.
I.e. not actively small and fast, like an OS should be.
2. Linux is already moving in the direction of having easier-to-use internal interface over small, fast, and tricky ones. For kmalloc in particular, the memory management docs https://www.kernel.org/doc/html/latest/core-api/memory-alloc... recommends kzalloc (which zeroes data) and suggests a function called kvmalloc which calls either kmalloc or vmalloc depending on the size of the allocation. It unconditionally recommends kvfree (which handles alloations from either kmalloc or vmalloc) instead of asking you to use kfree or vfree as appropriate.
3. In this particular case, kmalloc is inlined because it calls krealloc(NULL, ...). This could almost certainly be equivalently handled by inlining the other side of it. (Again, benchmarking.)
4. LTO is a thing.
ATS is a very interesting language! But there are a wide range of options for formal verification of systems programs, including the techniques used for the formally verified seL4 kernel, ATS, recent Ada work, TLA+, etc. Some of these tools have been very successful for certain projects.
But the more powerful proof systems involve tradeoffs. Frequently, you need to invest significant amounts of developer effort in exchange for very rigorously proven code. I've played with a number of these systems over the years, and they're great. But the costs would be hard to justify for many projects.
Rust occupies a slightly different niche. It focuses on preventing several classes of memory errors and data races. It's certainly harder to learn than Python, but probably easier than some common subsets of C++.
So the right tool for the job depends on what you're trying to do. If you want to formally verify your kernel's security model, or your distributed protocol, Rust would be a poor choice. If you want a language that makes it easy to work within sight of the metal, and that takes pains to make expensive operations visible, Rust can be a reasonable choice.
Python's gradual typing with MyPy demonstrated a great utility of static type checking that could be brought to existing project through many small and non-desruptive iterations. From my observation an learning of ATS, I expect even greater utility of gradual formal proofs that ATS enables for C. There are tons of successful C codebases, they are used in production around the world and contain a great amount of battle-tested knowledge. They just miss formal proofs that now could be gradually written for them.
Well no language can ever have "complete" safety guarantees. But I would like to know what other languages offer the same set of guarantees without GC overhead. This would include:
* the usual memory safety * no nulls * no undefined behavior * no data races
Does ATS guarantee those things?
Hopefully it will get better in ATS3 - https://github.com/githwxi/ATS-Xanadu#project-description
One of the problems with repeating comments is that the previous replies are not incorporated into the discussion; for example, mine:
>Back when it was on the PL shootout, ATS was consistently in the top five for performance. However, it was also consistently the worst in program size
I looked into learning ATS some years ago before Rust existed, and shied away due to the complexity. It's been around much longer than Rust, and I think there's a reason it hasn't caught on: it's just too onerous. ATS might be appropriate for formally verified systems, but language ergonomics does matter, especially for a project with as many contributors as the Linux kernel.
That sounds like "LEDs shouldn't be used for traffic lights because they don't melt snow", while ignoring all the benefits. Just add a heater where it is required.
Rust does add benefits. Are these worth the extra effort? It seems that some people think it's worth it.
It looks like the "ATS Linux" work is happening on a Git fork at http://git.bejocama.org/ats-linux-5.7.2 . However, I don't see any commits there when I clone it - it looks like it's just the v5.7.2 tag from upstream. Do you know if there's any sample code that's been written demonstrating kernel code in ATS?
I think that, empirically, we have built something that works. You can check it out (https://github.com/fishinabarrel/linux-kernel-module-rust - see tests/* in particular). I'd be delighted to see folks who are excited about other languages build something in their language, too.
I'd like to correct two specific things about your comment:
> if we discuss the amount of effort necessary for bringing new interfaces like the one this article mentions ("one can define a kmalloc_for_rust() symbol containing an un-inlined version"), we should compare it to other existing solutions, like ATS
Here is the amount of effort necessary to work around that:
https://github.com/fishinabarrel/linux-kernel-module-rust/bl...
I took a semester-long MIT class on a proof assistant (Coq) and I thought it was great but I'm still not fully comfortable with it. I realize that Coq and ATS are different languages, but I would humbly submit that if you compare the amount of effort spent calling krealloc instead of kmalloc vs. learning ATS, the former is probably lower.
> without giving up on existing C interfaces and the toolchain
Looking at Chapter 8 "Interaction with C" from the ATS book http://ats-lang.github.io/DOCUMENT/INT2PROGINATS/HTML/c2016.... , it seems like ATS's interoperability with C is pretty comparable to Rust's https://doc.rust-lang.org/book/ch19-01-unsafe-rust.html#usin... - I don't think either of them require giving up on existing C interfaces or the C toolchain. One of the specific reasons to use Rust in this context instead of Go (which is a great language too) is that Rust is designed to fit in very closely with the C toolchain. The first paragraph of that chapter, which discusses how ATS's data structures can be zero-overhead bridged with C and you can view ATS as a better-typed frontend for C, seems to equally well apply to Rust.
(In fact, as mentioned in the article, we'd love to automatically generate bindings from existing C interfaces. It turns out that most interfaces lack documentation - certainly machine-parseable docs, but often human-readable docs too - that indicates their ownership properties and locking constraints and so forth, but definitely the best approach is to get this sort of information into C annotations instead of having manual bindings in Rust. As a bonus, if it turns out that another language like ATS or Sing# or whatever is a better fit than Rust, the information is right there.)
Edit: I think what I’m trying to say is, let’s embrace what excites people.
Easier said than done of course, and definitely a huge effort especially for such a huge project as the Linux kernel, but there's also a significant amount of resources that could be allocated to that task by several large institutions.
I don't know if it's a good idea, or if ATS or something else could be better, but I think you're overly pessimistic.
Also "the safety guarantees that Rust provides are neither unique nor complete", while technically true, is also borderline FUD-y. At least you should expand on what you have in mind here, because I doubt that there exists a programing language that deals with low level hardware configuration that could ever provide "complete guarantees" about anything.
And beyond that, keep in mind that Rust provides a lot more than just safety guarantees, its syntax is lot more advanced that good old C in particular for anything dealing with types (including destructuring, matching, inference etc...). Even if you don't get "full" safety right away it can still be a much nicer language to code in (and I say that as somebody who knows C well and quite like it overall).
I meant inability to implement within the boundaries of safe Rust certain data structures and algorithms that ATS can prove to be safe. To name a few: stack-allocated closures, and safe pointer arithmetic programming [1]. For instance, ring buffers in Rust use "unsafe" and at some point there was a logical CVE in vec_deque that broke one invariant of the unsafe block [2]. Safe implementation of the ring buffer in ATS could be found in [1]
Rust routinely proves these to be safe.
> safe pointer arithmetic programming
You can't directly prove the pointer arithmetic safe with the compiler, but we routinely write a small amount of unsafe code which the programmer is confident in (and depending on the project that can mean an informal proof), and then re-use the data-structure that unsafe code created millions of times.
A ring buffer is a great example. Vecdeque is implemented once in the standard library, everyone who uses Vecdeque can do so in completely safe code provided the small amount of code in Vecdeque is safe.
This is orders of magnitude better than C, where every user of the ring buffer might make a mistake and cause a bug (or at least the API has no way of communicating that that is not the case, and in general the type system isn't expressive enough to make APIs that force that to be the case). It's not clear that taking the next step to proving the ring buffer implementation itself is correct to is actually worth the effort (it's not clear that it's not, but you need to make that cost benefit analysis, and I highly doubt it would be anywhere close to uncontroversially the case).
it is, what about other similar cases of using pointers in other data-structures and algorithms that are not part of the Rust standard library?
As it turns out most code isn't datastructures with unique memory requirements (i.e. that can't be implemented in terms of other datastructures), most code just re-uses existing datastructures.
this doesn't automatically prevent the code from containing CVE, does it?
> As it turns out most code isn't datastructures with unique memory requirements
most code that has been written in Rust so far. But when you enter a territory of OS kernels, unique memory requirements and data structures such as ring buffers and dynamic mutable trees are the things that have to be implemented according to the requirements, and we want them to be implemented as efficiently as C is capable of, and correctly. Falling back to "unsafe" is an option, but it's hardly an argument that supports the effort of bringing safety to such a kernel.
But let's say that we've solved pointer issues. What about dependent types and totality checking? ATS has that [1] [2]
[1] http://ats-lang.github.io/DOCUMENT/INT2PROGINATS/HTML/INT2PR...
[2] http://ats-lang.github.io/DOCUMENT/INT2PROGINATS/HTML/INT2PR...
There are certainly tools to detect any known CVEs, and tools to help you audit the unsafe code that you are using. They don't provide rigorous proofs automatically of course, the same way the standard library is generally not rigorously proved to be correct.
> But when you enter a territory of OS kernels,
Nothing really changes, see redox as an example of a fairly "complete" kernel written in rust with a reasonably low amount of unsafe.
> we want them to be implemented as efficiently as C is capable of,
This is already the case with Rust's typical library system. Generic libraries in rust are not less efficient.
In the real world rust datastructures are typically more efficient in my experience because the clean separation makes it easy to optimize the hell out of them.
> What about dependent types
That's a tool not a problem - you have yet to provide an argument that they are necessary and real world experience says that Rust achieves a meaningful amount of safety and convenience without them.
> totality checking
That's also a tool not a problem, and you could argue that rust's ADTs (enums) give you a weak version of it.
so they don't automatically prevent the code from containing new CVEs.
> That's a tool not a problem - you have yet to provide an argument that they are necessary and real world experience says that Rust achieves a meaningful amount of safety and convenience without them.
I can argue along the same line that Rust is a tool not a problem, and that you need to prove that Rust has a "meaningful" amount of safety. When you say that something has "meaningful amount of sth" you should add to whom it is meaningful and based on what criteria. Or, alternatively, you can stay on a pure technical aspect of the matter and conclude that the safety part of the language is not complete.
> That's also a tool not a problem, and you could argue that rust's ADTs (enums) give you a weak version of it.
totality checking is not exhaustiveness checking, there are two more properties that enums are unable to express in any "weak" form - termination and productiveness.
[1] and really, I'm an OCaml developer and I can write rust; yet ATS' syntax looks terrible and very complicated. Its website and tooling make it appear like a research language made by one person. I don't know if that's really the case.
(For example, say we could one day make a bunch of basic Linux drivers somewhat kernel agnostic? That would be amazing.)
Now, ATS is I hear a fine language, but I think the ability to write good, and especially abstract, ATS is sadly too far outside of most kernel dev's skill set. Even among PL researchers, there is a tendency to write monolithic Coq---to wit, is there a Hackage or crates.io for Coq?---because abstraction are hard and the academic paper economy doesn't really reward it.
Rust is just easy enough, and with enough crates.io momentum, that I hope the relative cost of reusable vs unreusable code will be less.
And finally, it takes a lot of political will to introduce a new language in any existing project, let alone one as big and storied as Linux. Like it or not, but general popularity of the language absolutely does help with that.
In my experience, learning the fundamental concepts behind theorem proving is a lot harder of a problem than writing an FFI interface to a library. In fact, that problem is so easily solved that there are tools that do 99% of the work for you.
And sure, maybe rust will never be able to inline C code, but inlining function calls is such a trivially small performance gap that most applications would never know the difference.
I'm not convinced that there are better alternatives for contributing to or incrementally rewriting an existing code base. FFI is for most purposes a trivial impediment compared to the alternatives.
There's an extremely high cost to introduce and maintain another language in a stack such as the Linux kernel. Everyone involved would be permanently impacted regardless of their opinion on the matter. It's not just a flag you can toggle.
Otherwise Rust community would have already forked the Linux kernel instead of bugging their maintainers to RIIR.
On top of that Rust compilation is very costly. It would render some devices that currently can compile the kernel unable to do so.
Last I checked Rust fails to compile itself with 4GB of memory. Here's another user commenting about it: https://news.ycombinator.com/item?id=23059869
The article is not about that, though, it's specifically about the various issues surrounding the use of Rust for Linux kernel development.
So no proposal of rewriting by any Rust developers
https://news.ycombinator.com/item?id=24133292
http://blog.vmsplice.net/2020/08/why-qemu-should-move-from-c...
Perhaps the most long-lasting harm of the Rust Evangelism Strike Force has been the rise of the Anything But Rust Counterinsurgency, who sees that someone, somewhere, is thinking about Rust, and needs to convince them that they're making a mistake.
I haven't found a source for that, but I don't believe it to be true and I'm curious why you do.
I haven't seen the previous comments of the OP so for me this comment was useful.
Even when we mark stories as dupes there are often users who say it was new to them and they appreciated the post. That makes perfect sense. No one sees everything, and if you haven't encountered the earlier elements of a repetitive sequence, then for you there is no repetition. I'm always reminded of a second-hand clothing store in my home town called "New to You".
But at a site-wide level it's not hard to understand that this is a stochastic process and we have to manage it globally. The alternative would be not to moderate repetition at all, and the community would definitely not prefer that.
To be honest, it was a modest improvement, because compilation has been decently parallelized already (compiler uses incremental builds, parallel codegen, and ThinLTO even within a single crate)
The kernel is not written in C++ and is unlikely to use C++, for many reasons - Rust is a more natural C-but-better for the kernel than C is.
Cyclone hasn't been maintained for over a decade; it was a research language, and it largely served its purpose as a research language in inspiring production-supported languages like Rust.
More worryingly, the article also says that only some platforms will be "Rusted" in this manner, splitting the codebase into "mainline Rust" and "obscure platform C" as far as these sections are concerned; this seems like it will fracture the development community between people who know Rust and those who don't, and those who can work on mainstream platforms and those who can't. Fewer eyes on this code seems like the wrong way to go.
(The further point is that it's impossible to have this discussion here.)