Why Windows 95 and 98 would crash after 49.7 days of uptime
sites.google.com
sites.google.com
However, timers wrapping around is a very current problem every programmer should be aware of. Not too long ago we had a problem with corrupted video files. After weeks of investigation it turned out that some code used an 32 bit unsigned to count microseconds. Well, that wraps around after approx. 1 h 10 min... Obviously most test cases were videos shorter than that.
Not all programming is the same; there are many lessons to learn about software engineering (mostly in the "what not to do" sense) which, for practical reasons, have only ever showed up in (time- and budget-constrained) closed-source codebases.
Well, security by obscurity might happen a bit less. And hopefully many bugs in contributions are spotted instead of being merged because the author thinks it should work. But all major open source projects have had amazingly "stupid" programming bugs that went undiscovered for many years. Or the same feature has been patched again and again, but still causes more trouble.
Knowing that many company-internal software projects work with little review and very low testing coverage, it should be natural that their code quality is even worse. But that certain kinds of problems would only occur in closed source does not sound evident to me.
How could I have that, or the opposite of that, without access to closed-source code-bases?
That was my whole point: that you can't know what sort of selection bias you're implicitly accepting by only looking at a certain easily-accessible slice of the data.
In order to know that it's safe to assume that FOSS is representative of all software, you first have to get access to enough non-FOSS codebases to be representative.
This is a running theme on HN:
Why X does Y?
Well we don't know but let me talk abut Z in relation to X.
Did you seriously expect to see Windows source code published? AFAIK it is still under strict NDA regardless of how old it is or did I miss some headlines?
"Microsoft server crash nearly causes 800-plane pile-up"
https://www.techworld.com/news/tech-innovation/microsoft-ser...
I created a bat file with:
net stop wuauserv
sc config wuauserv start= disabled
and then scheduled it to run every 10 minutes. The problem with that is you now have to remember to disable it and check for updates when you have time, which I'm obviously not doing.edit: And I'm trying to help other people. I know how infuriating it is to have your computer restart in the middle of something, and the above fix is pretty straightforward.
As a result of this complete ignorance to ergonomics, the windows ecosystem is returning back to WinXP years, and windows users are now getting pwned by years old exploits again.
I think microsoft best achievement in win7 was that it was working "out of the box" without having to hack it. Now, win10 again requires to be hacked right out of the box just to work.
People now install stuff like https://www.novirusthanks.org/products/win-update-stop/ or https://securityxploded.com/windows-autoupdate-disable.php and countless others.
But when they finally need to update, the half broken windows install may give up the ghost, because the update script can't account for hundreds of ways how windows internals may have been altered, and so they learn to not to update at all, and switch to antiviruses, qihoo and other "crapware clearners" in the end
Or, to rephrase that (from Microsoft's—and most IT departments'—perspective): Windows now works, as once again users with broken workflows are now forced into fiddly sub-optimal patterns for using the OS, which are more work to maintain than just correcting their workflow. (And IT departments are empowered to not even allow these fiddly sub-optimal patterns through GPOs, allowing them to force users to immediately rectify their workflows.)
Microsoft is trying to be the teacher that slaps your hand with a ruler when you look down at the keyboard instead of touch-typing. Except, instead of touch-typing, they're trying to force people to use (and build) software in a way that makes restarting their computer irrelevant, by e.g. making long-running jobs checkpoint their work and resume from the last checkpoint after a restart.
That has not happened even once to me, and what company allows users to decide?
I can tell you that the majority of calls to my desk from mac users are a direct result of not updating. The Windows machines all have their updates controlled by WSUS, a standard component on an AD network. I have very close to zero problems with Windows 10 updates.
That last happened to me a couple of years ago and I used Windows 10 every day since.
How? UPS?
If it's a short enough move I can see doing that just for the hell of it.
I was referring to longevity in general. Back in the day, DEC computers would just about never go down. They easily had up times of longer than 5 years.
Allow Windows 10 auto update but for a chance for it to delete all your profile / data files PERMANENTLY :
https://www.bleepingcomputer.com/news/microsoft/windows-10-k... (February 19, 2020)
Or
Turn off auto update with user control on when is best time to update:
I don't object at all to regular reboots for security updates, but I do object to forced reboots. Windows not being safe to leave alone for 24 hours without risk of it rebooting itself makes it no better than a toy OS in my mind. And it isn't as if it is just the one monthly reboot around patch Tuesday plus the occasional emergency patch: I think I can count on one hand the number of months in the last ~two years when my main Windows machine hasn't needed to reboot twice or more in a month.
Actually I do object to regular reboots if I'm more honest, at least as many as Windows seems to need. The fact that just about any security update requires a full system reboot instead of many of them being applicable with just restarting component services smacks of bad design.
I assume that a lot of the packages I update either don't need to be updated or aren't getting updated on a production server, and therefore less restarts are needed?
APT::Periodic::Update-Package-Lists "1";
APT::Periodic::Download-Upgradeable-Packages "1";
APT::Periodic::AutocleanInterval "7";
APT::Periodic::Unattended-Upgrade "1";
And 50unattended-upgrades: Unattended-Upgrade::DevRelease "false";
Unattended-Upgrade::Remove-Unused-Kernel-Packages "true";
Unattended-Upgrade::Remove-Unused-Dependencies "false";
Unattended-Upgrade::Automatic-Reboot "true";
Unattended-Upgrade::Automatic-Reboot-Time "02:00";
No more worrying about updates :)Jokes aside, around 2001 I was asked to do some maintenance on a machine running Windows 3.11 that was running an app to control all access within the building. When I asked about potential upgrade the guys balked at me saying the machine has couple of years of uptime and they are not very keen to have to restart it on a schedule due to windows 95/98 bugs.
For exact this reason. They are working and doing what they meant to be doing.
This sounds crazy to modern ears, but at the time the idea of a desktop system actually staying booted for a whole month was weird and alien. They just didn't do that. Commercial unixes were better. It was routine to find servers that had been running for three digit (!!) uptimes, and people would brag about stuff like this. Linux, when it arrived, was in this category too and we'd all puff to our friends about how we "never turn off" the machine under our desk, because it would always be able to run for a month or so before it crashed.
It was just a different world, and things, despite some of the community emotion, have gotten vastly better. We're much better at writing software than we used to be, which tends to make some of old folks a little confused when people hold up "new" ideas as being paradigm shifts in software quality.
We already had the shift! We're just picking the higher hanging fruit now.
[1] Do the mid-90's even count as early?
The problem is not people writing buggy software or them crashing. At least for me problem is we keep on reinventing the wheel, without corresponding increase in productivity. I don't see people cranking s/w better or faster then they did with Delphi in 1998. Instead of standing on shoulders of the giants we disrespect them, our whole industry is built on abusing all the hardware and technology gains for frivolous wars and gains.
</rant>
DOS (before) was very, very stable (but of course was not multitasking) and widely used in industrial setups.
Windows NT (before Win9x) and Windows 2000 (right after) were both, very, very stable, I had machines running 24/7/365 that were rebooted once a year or so and for other reasons (maintenance/update, hardware replacement, black outs longer than UPS capability, etc.).
And a big issue (at least on the desktop) is that in many cases we keep throwing away old stuff and reinvent new stuff with their new bugs and issues, instead of fixing that old stuff (and i do not think that it is a coincidence that most of the stable foundation is actually old stuff that had a lot of people trying to fix its issues instead of throwing everything away).
Even more insidious, we never hit the issue until months into production because we had a pretty consistent two week release process. So the issue only came to light when development slowed down as the project became more stable.
Luckily deployments were staggered across datacenters so when the bug hit it didn’t happen to every server at once, but of course as always happens, when it did, the majority of the servers started hanging over the weekend while I was on call.
Very different than something which only happens after 50 days because obviously in small intervals it's not a noticeable bug.
But more to the point - the consumers of GetTickTime should have also had tests to verify that the consuming code appropriately handles time rollovers.
Granted, this was Win95 so I'll cut them some slack, but that should be standard procedure for any code that consumes an incrementing counter API. First question should always be - 'how will this code handle a rollover value'?
Software will always have bugs, we've chosen to approach it as "find the bugs before the customer" rather than "stomp out every possible bug". Because you are right, we don't get 2 months to pause and fix bugs, and also customers will always do some crazy configuration or workload you didn't have a test case for.
For me, it failed at around 25 day (49.7/2?). I believe the documentation at the time [since fixed] might have had GetTickCount() returning an integer instead of an unsigned integer. I had a devil of a time tracking it down!
I thought the anime really glossed over that story terribly.
https://web.archive.org/web/20041109020858/http://support.mi...
It is dated August 2004 (it is revision 3.2) but the patch (for Windows 98) has a date:
Release Date: Jun-04-1999