Linus says Linux has become 'bloated and huge'
theregister.co.uk
theregister.co.uk
Does anyone know how much redundant code is in Linux right now? Would finding, isolating, generalising and merging it be a good way to start? I know of examples of this happening, e.g. libata taking over a lot of the old 'ide' drivers; the many attempts at unifying the various WLAN subsystems and drivers.
It sounds like the problem is that everyone is writing new features that probably not many people actually need (since they did fine with out it before - I'm not talking about drivers here) but it's slowing the whole kernel down and people need to go back and optimise that stuff.
After all, Linux is already designed to make most of its code optional via the module systems. The only problem is that its core is becoming not-so-sleek, and they're probably already getting as much mileage out of the module system as they can hope for. One interesting possibility is that Linux could avoid the performance hit of run-time dynamism by introducing a new kind of compile-time feature selection. It's obvious to me that C isn't up to the task, but it might not be obvious to someone much smarter than me. Or maybe they could extend C. These days we think of the Linux kernel community as old, conservative, and unlikely to surprise, but who knows? I'll keep my fingers crossed.
Also you should put it in perspective, the Linux kernel itself is really not so bloated, that any normal user would notice, like they do with Gnome/KDE/Firefox's bloat. Even with that type of bloat, Linux distros still run well on a 1GB (probably less) and an old P4 CPU which people don't say about recent releases of OSX or Windows.
Now the theory behind Open Source is that developers around the world can see bloat and fix it. But with very few exceptions the corporate developers aren't going to make their code more efficient. So you have 75% of developers creating bloated code and at best 25% left to fix it.
Meaning the amount of bloated code easily overwhelms those volunteers who are willing and able to fix it.
It'd be nice to know if Linus had anything to say to that effect - it might lead to a more productive discussion than waxing lyrical about the wonders of microkernels.
I think that's more or less where things are right now. You can't have tons of features without paying a price. They do a pretty good job of balancing things, but maybe someone will come along and improve things in the future.
I understand software bloat to describe the situation that arises when unnecessary features added to a program increases it's size and/or decreases performance, without providing significant user benefits.
In the article, Linus says "Acceptable and avoidable are two different things. It's unacceptable but it's also probably unavoidable."
If it's unavoidable, then the features that were added aren't really bloat, because it implies that these features are necessary, and that they provide significant benefits. Necessary additions that provide significant benefits could be anything from security features to IO.
IMO, Linus should either state what he thinks should be removed from the kernel, or he shouldn't even bother mentioning it at all.
If on the other hand he's lamenting the fact that there's a cost associated with adding features (necessary or unnecessary), well then I think most people understand that you can't have your cake and eat it, so even from that perspective his point seems redundant.
Actually, that brings up another point. Who, besides Slackware maybe, uses a stock kernel? Don't most distros customize it anyway? Am I being unfair in saying most Linux users don't really encounter the kernel at all?
IMHO it doesn't really jive with UNIX principles to have such a monolithic kernel.
In reality, the core Linux kernel is just the stuff that compiles to vmlinux. You could read it all in a few weeks and understand it.
filesystems dont belong in the kernel.
Unfortunately native disk file systems like NTFS-3G is lot more slower than their native counterparts.
there's your problem. it's not designed for it
Unfortunately userspace file systems like NTFS-3G are lot slower than their native counterparts.
Which is understandable, regarding kernel calls are still quite expensive.
user/developer/management/customer/congress/judiciaryUnix is NOT the be-all, end-all definitive operating system. Linux is NOT even the be-all, end-all Unix (or Unix-like). There is A LOT of research left to be done in operating systems / kernels and many many years of evolution to come.
It's a shame that it is SO HARD to build the ecosystem required to launch a kernel. New ideas are hard pressed to gain traction.
reference - http://www.cs.vu.nl/~ast/reliable-os/
It would be interesting to see how the original Mach microkernel runs on modern hardware. At the time, the context switching of the Mach kernel was making performance poor. But this was on 300/400 MHz machines.
They were writing a sort of microkernel, that only multiplexes/protects users from each other, but exported the raw hardware interface to userland. The abstracting part of the OS would be implemented as (possibly shared) shared libraries.
Making people want to use it is hard, as most apps people want are boring to write and/or trapped in the poor design decisions of the past (e.g. POSIX).
But those programs have to talk to each other, and this message passing overhead increases with increased functionality (more programs). So it seems likely that there would be a definite performance drop as the functionality of a micro kernel increased.
There are a lot of other side benefits. As stable as Linux is, it still has a great deal of code executing in kernel mode, which is error-prone; microkernels can restart crashed processes on-the-fly without compromising the stability of the system rather easily. Also, some of the newer microkernels have formal proofs of their security properties, which is pretty damn fascinating! (albeit they assume the underlying proof system itself is error-free, which might not be true)
> I'm sure the majority of that Linux kernel logic is never executed; if that code were instead placed in user-level processes, they would have no performance hit.
Such code doesn't matter. Most of it is in a dynloaded module, so it doesn't even eat ram, and even the stuff baked in the kernel hardly matters for performance -- the < 5 megabytes of ram doesn't even hurt in cellphones anymore, and the thing that hurts most is L1i cache use, where only code that is actually executed goes.
The bloat we are talking about is not the increase in the total size of the kernel, but the increase of the size of the commonly executed code paths. Linux being a microkernel would not help there, in fact it would hurt.
Or it could mean that these cases do in fact have to be considered on that code path every time and that any "outsourcing" to user mode processes would in fact increase the size of the code path.
If it's the former, then the argument that a micro kernel design encourages more fine grained modularisation and leads to fewer special cases is not completely spurious in my opinion.
But all of this is impossible to judge without looking at the code itself.
This assumes that someone needs and actively maintains every feature in the code.
I take offense at the statement "Linux is designed" :)
Of course it is not optimal. Nothing else is either. How can we make it better?
Yet in the interview, Linus says the kernel is getting bigger and slower because of new features, which presumably not everybody needs. "[W]e are definitely not the streamlined, small, hyper-efficient kernel that I envisioned 15 years ago...The kernel is huge and bloated, and our icache footprint is scary. I mean, there is no question about that. And whenever we add a new feature, it only gets worse." He isn't talking about code that can be compiled out or left unloaded if you don't need it. He's talking about code that ends up in your instruction cache whether you need the features or not.
-Monolithic kernels do not enforce modular implementations, inevitably leading to unnecessary dependencies
-Monolithic kernels need to be developed by a single party, who could not possibly manage such a large code base. The result is poorly maintained code.
-Monolithic kernels need to be recompiled to turn off features (and I have better things to do)
-Microkernels can tolerate more bugs (may be good or bad)
-probably a couple more
Nor does the Linux kernel need to be recompiled to turn on/off features - most features can be turned on/off via /proc or sysctl or by loading / unloading modules.
As for the unnecessary dependencies, is this a problem in real life for Linux? No. Modules tend to depend on well defined kernel interfaces.
Remember the Tanenbaum vs. Torvalds debate?
I wonder what Linus thinks about it now.
btw, here they are together.
How the kernel is organized has no meaning on how bloated it is. A microkernel can be just as bloated, in it the bloat is just not in the kernel itself, but in the processes providing the services a monolithic kernel would provide. If anything, a microkernel is always more bloated because of the extra code needed to do all the process sync.
Linux being a microkernel would make the situation worse.
(What we need to do is to start spending more time on cutting features.)
and tell me which one is bloated.
Yeah, you have a piece of code named "the kernel" which you can point at and tell that it isn't bloated. But you can do this in Linux too, if you compile everything as loadable modules. That is just creative accounting of the bloat.
It's only marketing spin if you will always need all of the features in order to do anything meaningful.
Are you still having this debate with a purely academic definition of "monolithic" vs. "micro" kernels? Because neither of them exist anymore, and haven't for a long time; it's like CISC vs. RISC, in that arguing about it today is really missing the point since most everything is a mix. You need to address what Linux actually is, and what real microkernels exist, and what real performance they have, not regurgitate boring arguments from the 1980s that have simply been superceded (rather than one side "winning").
In fact, this goes for all you other commenters still having this argument. "Compare and contrast a microkernel running on CISC vs. a monolithic kernel running on RISC. Which will run best on a top-of-the-line Amiga, and how can this be best leveraged into flaming those who disagree with you?"
Then there is the support issue. In Linux, all the modules must be maintained by the kernel team. This may constitute bloat as the team may have an unnecessary amount of code to maintain. Some modules may only be used by a few people. The issue isn't the features, but rather Linus's philosophy.
From his posts, it seems he like to keep changing kernel APIs as he sees fit. If you have many modules, this creates problems. Without a strict API, you will frequently break modules and thus create a lot of maintenance work. Fixed APIs are both very important and very good for large scale collaboration. Without standards like POSIX or the Windows APIs, we wouldn't be where we are today.
The second aspect to this is security. Say you built a very stable API and let others make modules as they need them. In a monolithic model, to maintain security you must audit every module. Thus for Linux to stay secure, the kernel team would need to either declare many less used modules as "tainting the kernel" or audit them all: a lot of work. In a microkernel model, the kernel team, so long as they designed it right, can let users build modules without audit.
Saying something has too much bloat is ambiguous to the point of being useless. Many people use that term simply as meaning slow, or a lot of code, or even poorly designed code. It's always better to describe things precisely. If it has too many features, say so. If it has bad design, then state that.
There are similar problems in Firefox. Blame the extensions, not the browser. API or none, bad programmers can get their code to affect a lot of people.
Of course, if you want to reduce bloat, compile your own kernel. Get rid of the modules you don't need. Is that a lot different that what you propose?
The interviewer asked Linus about benchmarks that show the kernel getting slower and slower each year. Linus acknowledged that the kernel is in fact slowing down. I don't think they would be worrying about badly configured kernels, or kernels with unnecessary modules loaded. I don't know what features Linus is talking about that make a properly configured Linux kernel bigger and slower than it was ten years ago. Nor do I see any explanation or examples cited on this page. I wish someone who understands would give some examples of the feature creep they're talking about.
there are more than two types.
We have fretted recently because the plan9 kernel is struggling to fit on a floppy disk when loaded up as an installer with most things turned on.
The whole of plan9, kernel, userspace, 386 binaries & sources for 6 architectures fit into 200Mb uncompressed (and the whole shooting match compiles in 15 minutes).
Linux is a re-implementation of an OS that was already considered dead by it's maintainers.
"Not only is UNIX dead, it's starting to smell really bad." Rob Pike circa 1991.
see also "Sometimes when you fill a vacuum, it still sucks."
The beauty of plan9 comes from the interface to the userland, not from implementation details and features. If you wanted to add al those features to plan9 I seriously doubt you could get it as lean and quick as Linux.
WiFi stack? yes, and bluetooth
Support for different power modes? yes, I think so, can't say I've ever used it. does "echo blank > /dev/vgactl" count?
Framebuffers? no, it is a 21st century os
Kernel probes, high resolution timers, extended inbuilt security APIs? yes
all in 13 syscalls
Software organization can cause bloat if the feature can be implemented in another way. For example, there may be multiple ways to establish the same network connection.
I'm not defending microkernels, but some kernel functions in userspace is not bad. Dismissing microkernel concepts immediately is just as bad as questioning whether Linus was wrong in not selecting that architecture.
> "Uh, I'd love to say we have a plan," Torvalds replied to applause and chuckles from the audience. "I mean, sometimes it's a bit sad that we are definitely not the streamlined, small, hyper-efficient kernel that I envisioned 15 years ago...The kernel is huge and bloated, and our icache footprint is scary.
At the very least, badly edited.