From: Ingo Molnar
Subject: [PATCH 0000/2297] Eliminate the Linux kernel's "Dependency Hell"
...
25,288 files changed, 178,024 insertions(+), 74,720 deletions(-) From: Ingo Molnar
Subject: [PATCH 0000/2297] Eliminate the Linux kernel's "Dependency Hell"
...
25,288 files changed, 178,024 insertions(+), 74,720 deletions(-)https://git.kernel.org/pub/scm/linux/kernel/git/mingo/tip.gi...
When will Linux move to a site like github / gitlab / something similar self hosted that supports people proposing changes without having to send every single person interested in the development of Linux thousands of emails.
And if you need email to authenticate on a website, why not just use email anyway?
But this is an interesting idea, and I'm on the look-out for ideas to replace email. I'll keep "LKML" in the back of my head as a use case.
It does support making backups of repositories at least.
>and if the host goes rogue they can make it inaccessible and make everyone lose efficiency while they swap back to the old method.
Isn't this also a problem if the person managing the email list goes rogue? You have to trust someone to host the infrastructure.
>And if you need email to authenticate on a website, why not just use email anyway?
Because Email may not be the optimal user interface for handling issues, pull requests, code review, etc.
It depends on what kind of threat you are worried about. The beautiful thing is even if LKML was abused and ruined tomorrow, nothing is lost or damaged. The emails have already been sent, and moving to a new list is more of a nuisance at this point.
If you're worried about authenticating a patch series, committers can sign their commits a number of ways.
Lastly, a large chunk of Linux development happens off-list, with subsystem maintainers building branches for Linus or Greg to pull from.
Maybe that's your experience.
Others have used mailing lists for this purpose for decades and a lot of people prefer it.
For one, it's accessible. Every machine can do plaintext email. Every text editor can work with plaintext. It's simple. Integrate it into whatever shell/editor/scripting language you prefer.
As opposed to login into to a bespoke web interface, mainly primarily designed to be friendly to novice users.
If this in true then why do so few projects use a mailing list nowadays instead of something like github.
>Every machine can do plaintext email.
Every machine can do web browsing too. More people have used a web browser on their computer than a dedicated email client (as opposed to a web app like gmail).
>As opposed to login into to a bespoke web interface
You have to log in to email too.
>mainly primarily designed to be friendly to novice users.
What's wrong with that? Having good UX is a win. The UX for creating a new repo on github is a million times better than creating a mailing list (yes I have set up a mailing list for a project I made and no one e ended up using it except for me. Meanwhile I had much more success with getting people to join and discuss the project via Discord)
Tell that to most of my browsers which refuse to display any Gitlab content (they only ever show the side bar).
Lot's of projects use mailing lists. Lot's of projects use web-based git hosting services. Lot's of projects mix.
Why do people just use GitHub? Because it requires zero configuration and it's convenient when you're developing something yourself.
The Linux Kernel is by far the worlds largest open source project. Consider that it might have different needs.
>> Every machine can do plaintext email.
> Every machine can do web browsing too.
You're missing the point. A web interface is more complex than plaintext emails. To integrate into your shell environment is a lot more complicated.
> More people have used a web browser on their computer than a dedicated email client (as opposed to a web app like gmail).
I don't understand your point.
>> As opposed to login into to a bespoke web interface
> You have to log in to email too.
The emphasis was on bespoke web interface, not logging in.
>>mainly primarily designed to be friendly to novice users.
>What's wrong with that?
Fast and flexible are sacrificed. Ie. it's more convenient for the people who don't use the system all the time (people filing bug-reports) vs. the people who actually use it all the time (maintainers)
> yes I have set up a mailing list for a project I made and no one e ended up using it except for me
a tad different use-case to the development of the linux kernel it seems
> Meanwhile I had much more success with getting people to join and discuss the project via Discord
Whatever works for your project bro.
> What's wrong with that? Having good UX is a win.
What's "good UX for novice users" isn't necessarily good UX for more advanced ones.
Then I came to realize that my problem was not with email itself, but the way I had been using it. All GitHub, GitLab, etc do is take a decentralized platform and add centralization. I love the idea of using email now instead of a centralized issue tracker. Mailing lists become the issue tracker and those are then published on a website for others to use that want some centralized features.
A mailing list is centralized, but much less so than GitHub and is secondary in this instance.
Please actually attempt the ideas you are suggesting before suggesting them.
This set of changes (and other large ones like it) is partially why the kernel needs to use email instead of <insert bloated Ruby application here>.
This sort of implies that the Kernel would use something centralized if it worked well, which I don't believe to be true. Using email is very much an intentional choice and a desirable solution.
The mailing list that they use now is centralized. It works by everyone sending a message to a single email address. The owner of that address then sends the email out to all of the subscrbers of the mailing list.
Having everyone who wants to upstream their patch directly to the kernel submit their work to a single website is not much different than them sending it to a single email.
This here for example is sent to Linus directly and that part does _not_ depend on the mailing list. The mailing list is just a convenience feature for people that _want_ to be aware but may be looked over/forgotten by the submitter (which still happens, and people regularly say stuff like "also CC'ing additional relevant people").
Yes you could, at least in principle, also do that via GitHub/GitLab/whatever. But a) you wouldn't gain anything (the kernel already has all the tooling it needs and it's large enough to just create missing tools, heck git was literally written for it), b) you would introduce a massive single point of failure that either i) is _not_ under control of the kernel.org team or ii) needs a massive amount of resources to host (see the fun people poke at displaying 2k commits in GitLab). Neither seems like a good use of the kernel development resources. And if you want to contribute to the kernel... using a mailing list is an absolutely trivial problem.
So’s github, anything of a non-trivial size it just won’t show by default and you need to re-request every time.
It starts struggling around the kloc scale, and the browser itself soon follows as the dom is not the neatest and you start getting tab memory above half a gig, which I assume also makes the JS start killing itself or something.
Good times.
it's plain-text email
the display ad on your local news website is probably wasting more resources
Ingo did not send 2298 patch emails, he sent only the 0000 one, which contains the location of his public git repo for all this. People can clone this, and can even push it to a github/gitlab/etc hosted repo if they like.
I think it's natural for workflows to change overtime.
Your question is fundamentally wrong; totally bass-ackwards. The correct question is:
When will all the zillion other projects move off of GitHub / GitLab / all the other usurper sites?
Think how many lines are now duplicate (and therefore need to be updated in twice as many places, and bugs introduced when a copy is missed).
Think how much extra stuff someone needs to skim through looking for the relevant file or part of the file.
If the 100k lines were in their own subdirectory and added a major new feature, it would be worth it. But spread across the whole codebase, and introducing more duplication and chances for bugs I think outweighs a 'neater' header file system.
The 'dependency hell' is only an issue for people who only want to build part of the kernel anyway. If you're building the whole lot, you might as well just include everything and it'll work great.
That's a lot more than "shaving the build time."
As far as I understand it, the goal of this patchset isn't to improve the build time; that's just a nice consequence. The goal was to refactor the header-file hierarchy to make it more maintainable and less "brittle." Sometimes, increasing maintainability requires more code. (Almost always, if the current version is a terse mess of mixed concerns.)
Think of it this way: take an IOCCC entry, and de-obfuscate it. You're "increasing the size of the codebase." You might even be "duplicating" some things (e.g. magic constants that were forcefully squashed together because they happened to share a value, which are now separate constants per semantic meaning.) But doing this obviously increases the maintainability of the code.
I wonder what the memory savings for compilation look like? Because that's also potentially more workers in automated testing farms for the same cost.
In Ingo's post, he points out that the main speedup is coming from the fact that the expansion after the C preprocessor step is a LOT smaller.
That's a lot of decoupling. As someone who had to go rattling over the USB gadget subsystem, I can tell you that running "grep" with "find" was the standard way to find some data structure buried in a file included 8 layers deep. Having those data structures in files actually mentioned by the file you're working on would be a huge cognitive load improvement as well as make tool assistance much more plausible.
Even if this particular patch doesn't land, it lights the path. What types of changes need to be made are now clear. How many changes are required before you see the improvement is now clear. With those, the changes required can be driven down into the maintainers and rolled out incrementally, if desired.
As someone who tried poking around: Oh good; I assumed I was just missing something. Alternatively, oh no; I had assumed I was missing something and there was a more elegant tool out there.
It has always amazed me how finding something which originally seemed a trivial little thing, usually meant going through a chain of #defines and typedefs across many header files. It's the same with GLibc, by the way. It' a bit like when you hike to a summit by following a crest path: you always think the next hump in sight is the right one, your destination, the promise land; and when you reach it, dammit, it wasn't, your goal is actually the next one. Or perhaps the next after the next. Or...
When you're talking a project of half a million lines, sure.
The Linux kernel has around 27.8 million lines of code. An increase of .35%
> Think how much extra stuff someone needs to skim through looking for the relevant file or part of the file.
Why add features at all? Code has a purpose. Sometimes bringing code into a static context is a net good. It was going to be generated at runtime anyway.
> If you're building the whole lot, you might as well just include everything and it'll work great.
That's not strictly true, but it's true for these features, which is a stated reasoning.
A library I'm working on is 7000 LOC which seems pretty sizable, but 0.35% of that is 25 LOC.
The Linux patch set is actually tiny
This is horribly misleading; most of these lines of code are drivers, which this patchset doesn't even concern.
It's still a massive change that only a handful of developers will ever be able to review in entirety - a fact to which the size of the project is completely irrelevant - if anything, actually, it urges even more caution, given the implied complexity. Which I believe was (at least in part) parent comment's point - given the importance and ubiquity of the Linux kernel, this may be concerning.
That said, I am very confident in the structures put in place by the kernel devs, their competence and the necessity for such a change - but trivializing a 100k LoC patchset because the project it's intended to land in is even more colossally complex isn't how I'd choose my approach.
That's not true at all, a big part of those added lines are added includes in drivers.
E.g. this commit I picked at random adds 1500 lines, of which just a few procent are in the core kernel: https://git.kernel.org/pub/scm/linux/kernel/git/mingo/tip.gi...
Having to recompile the entire tree every time you change a seemingly unrelated header gets old fast.
yes. and this is very, very good reason. As another poster said, you at 0.35% lines to make it compile almost twice as fast? And you're not happy about that?
> Think how many lines are now duplicate (and therefore need to be updated in twice ...
OK, how many? None! That's how many. Adding proper header dependencies to the .c module doesn't duplicate anything. Unless you think adding #include <stdio.h> in every module somehow creates unmaintainable duplication.
> Think how much extra stuff someone needs to skim through looking for the relevant file or part of the file.
OK. Hrm... I think a lot, lot less is how much. That's the whole point of a major cleanup like this. Proper decoupling. Headers that you use are obvious where they belong not being brought in with some action-at-a-distance accident.