Mach's designers simply assumed that systems would be rebooted often enough
clozure.com
clozure.com
Root cause seems to be "server up too long."
You begin to see why a microkernel that needs to be bounced more often than, say, Linux, which is not a microkernel, begins to lose some of the appeal of having a microkernel in the first place.
(Aside: Mac OS X is based on Mach, but practically all of the stuff applications see is provided by a single server which (to the best of my knowledge) can't be restarted without taking down the whole system. It's a hybrid design.)
Also, just to confirm a stereotype, do you use Windows for your servers?
Related tangent: There was a time when the the US military's COTS (Civilian Off The Shelf) initiative resulted in AEGIS missile cruisers running largely on machines running Windows NT. Yes, the ships responsible for screening carrier task forces against attack -- had to be rebooted on a weekly schedule.
To be fair, the unmolested NT kernel is rock solid. I know people used to run RAID arrays with it, but for the desktop, Microsoft allowed the crossing of certain architectural lines.
which, if memory serves me right, directly translates to "OS/2 kernel is rock solid"
"Given these issues [with OS/2], Microsoft started to work in parallel on a version of Windows which was more future-oriented and more portable. The hiring of Dave Cutler, former VMS architect, in 1988 created an immediate competition with the OS/2 team, as Cutler did not think much of the OS/2 technology and wanted to build on his work at Digital rather than creating a "DOS plus". His "NT OS/2," was a completely new architecture." http://en.wikipedia.org/wiki/OS/2
This is unsurprising: They're all designed by the same guy, Dave Cutler.
Too bad it resents running programs on it so much...
When that happens, it doesn't resent those programs so much as that it's been violated and turned into a warped version of itself. Sort of like Jeff Goldblum's character in The Fly.
So 3.0 was adopted, with success, but not as a microkernal hosting user-space servers.
That's what you get when you assign the management of Solaris servers to MCSAs...
Bouncing a server may even make the problem go away, but it does little to enlighten you on why it happened in the first place.
And yes, If your MCSAs tend to bounce servers all the time, that's because their experiences have shown this works for Windows.
You have a human resources problem, not a technical one.
I'd argue that a 5 why's analysis of "server up too long", leads to "server wasn't written well." YMMV
It's basically a scheduled sanity check for your infrastructure. It boils down to robustness and assurance. If parts of your system won't survive a reboot this is an unknown and therefore a risk.
If a system can't survive a reboot it's better to find that out when it's during a scheduled reboot and personnel is on hand to deal with it rather than relying on on-call personnel responding to alarms from monitoring at some point.
Apart from the repair time being longer when relying on on-call technicians, there is also the general observation that productive systems tend to become more valuable or have more exposure over time as their usage increases. A failure at a later point therefore also becomes more costly as time goes by.
Most of the time you don't need to do complete reboots. Restarting services usually takes care of the majority of the sanity checks.
Sure, that usually takes care of most checks, but the point is to eliminate as much as possible unusual scenarios. A service-level restart will catch service config mismatches, but it won't catch, for example, manual interface or routing table updates. It also won't catch the 'fragile root' problem where the device that the kernel is loaded from has become corrupt.
Bad things happen in the strangest places if you run enough systems.
But I agree - the objective is to detect those situations when things have gone wrong a long time ago and we just don't know it yet. ;-)
A common pattern is that each logical partition gets rebooted once per month, on a rotating basis. Say you have 4 LPARs: LPAR 1 is rebooted on the first week, LPAR 2 on the second and so on.
The goal is to demonstrate that in case of catastrophic failure that the system will come up. If it doesn't come back up during the scheduled reboot, you have the advantage that you are rebooting outside production hours and have some time to fix the problem minus frantic phonecalls.
Meanwhile, the other LPARs will keep running as if nothing is going on.
I guess bootstrapping is more common in the Lisp world.
Its a business decision. Not worth getting emotionally involved, just weight cost of fixing vs cost of rebooting.
Here's a hypothetical, but not ridiculous, scenario. Which is a better way to spend resources? Track down and stamp out an extremely difficult resource leakage bug? Or simply bounce the server? The costs associated with the former may be HUGE. The costs associated with the latter may be small in comparison (cost of potential down time, minimal labor cost of bouncing the server(s), and costs associated with user reliability).
I wouldn't always chalk it up to whatever it is that your comment implies (laziness? willful ignorance?).
It also looks really JV to see those alerts day in/day out. Like, come on. We can't just fix the problem?
Sure, we would all love to fix the bugs, but if there are no easy clues and a regular reboot fixes it - who really cares when there are features to build?
Seems to indicate that it has ;-)
I reboot my macbook pro (running leopard) as often as I reboot my windows 7 PC which is not much. Except I use the Win7 machine 90% of the time and for much more "intensive" stuff.
I did experience memory leaks with pretty much every Bittorrent client, but I haven't torrented any uh, Linux distros in forever.
Granted, I don't know much about what's going on under the hood (I'm a designer, not a programmer) but Activity Monitor shows that each tab uses a gig of virtual memory...
Also, most developers tend to buy machines with the maximum RAM they can. I've got 4GB in my macbook pro and the 20 chrome tabs can't be taking up more than 500MB all together. Comparatively, Safari seems to take up a lot more memory than Chrome.
That sounds like paging in effect, which is decades old. Maybe you just notice it more now!
This is what used to happen to Chrome on my MacBook when I had "only" 2GB of RAM:
1. Navigate to HN and click on 10 interesting stories + related comment threads.
2. Chrome tries to launch 20 processes at once. Processor usage goes through the roof and knocks several weather satellites out of their orbits.
3. The entire OS X UI freezes. I can't even Cmd+Tab to another application.
4. Some Chrome processes inexplicably crash (why?!). Nearly every other app gets swapped to disk.
5. Some old tabs get swapped to disk (which means switching to them later takes 3 to 5 seconds).
6. Processor usage becomes normal after 2 minutes. The OS becomes responsive again.
7. All my frustrated Cmd+Tabs and mouse clicks are finally registered and the UI does a halting, jitery dance as my inputs are de-queued and routed to several different apps.
8. After figuring out what buttons I accidentally clicked and in which apps, I can read HN!
Note: I know I've done a fair bit of Chrome-bashing on HN lately. If that bugs you, I apologize. It's just that I can't stand terribly engineered software grabbing a large share of the market. Chrome makes some interesting trade-offs re efficiency and the result is really not that bad. Chrome works fine for most non-geek, non-tab-happy people. For people who pretty much live in their browsers, though, Chrome is a terrible choice. Unless, of course, you have powerful machines. I believe Chrome's multiprocess model will be less of a bottleneck as processors evolve, but it's a pain in the ass on most contemporary machines.
The Firefox "BarTab" extension does something like this:
Indeed; that is the point of the article. =)
I have to reboot the dang thing every day or two. The development tools hang and screw up. The network drops and needs to be unplugged/plugged back in or rebooted. The debugger totally fails to breakpoint and only a reboot will fix it.
I alternate between banging my head on the wall, and white-hot fury at the thing. Fortunately I only occasionally have to verify the code still runs there, and don't have to live in that environment.
I think we have to assume that the port leak is still there in the most recent version (10.6.6).
I do hope that the poster had misunderstood Avie Tevanian in his conversations with him. Tevanian was one of Apple's chief architects and was one of the big wheels responsible for NextStep's transformation to OS X. It would be kind of outrageous for an OS architect to believe that it's at all acceptable to allow user-mode processes to cause kernel resource leaks.
OS X was Mach once in the same way that orcs were elves once.
Using Mach ports correctly without leaking is hard, similarly to how using C's malloc/free without leaking is hard. The former no more implies a current flaw in Mach than the latter implies a current flaw in C stdlib.
It just means that there are cases that the kernel developers didn't think of, just like the people who made Apache didn't think of the Slowloris attack. Heck, a few months ago someone exhausted Linux kernel fds in just by shuffling socket pairs around (just 40 lines of C with three syscalls).
These are problems, ranging from remote DOS attacks, local DOS attacks, to accidental DOS attacks.
https://bugzilla.redhat.com/show_bug.cgi?id=97373 (System UPTIME reported incorrectly):
"Steps to Reproduce: 1. Boot Linux system; 2. Go away for 497 days; 3. check uptime"
Microreboot – A Technique for Cheap Recovery
"A significant fraction of software failures in large-scale Internet systems are cured by rebooting, even when the exact failure causes are unknown. However, rebooting can be expensive, causing nontrivial service disruption or downtime even when clusters and failover are employed. In this work we use separation of process recovery from data recovery to enable microrebooting – a fine-grain technique for surgically recovering faulty application components, without disturbing the rest of the application."