Leap second causing Linux server crashes?
serverfault.com
serverfault.com
So, kernels between 2.6.26 and 3.3 (inclusive) are vulnerable.
[1] https://git.kernel.org/?p=linux/kernel/git/torvalds/linux-2....
[2] https://git.kernel.org/?p=linux/kernel/git/torvalds/linux-2....
[3] https://git.kernel.org/?p=linux/kernel/git/torvalds/linux-2....
Spent the last two hours recovering servers, tomorrow will be another interesting day.
Whoever figured it'd be a good idea to INSERT[1] the leap-second instead of just slowing/accelerating time... <censored>
[1] Clock: inserting leap second 23:59:60 UTC
That would be the IERS organisation. There's going to be a vote in 2015 to abolish them entirely.
Almost all of my machines run the Debian stable kernel and were still affected.
1. https://bugzilla.redhat.com/show_bug.cgi?id=479765
2. http://it.slashdot.org/story/12/06/30/2123248/the-leap-secon...
3. http://serverfault.com/questions/402087/does-centos-5-4-prop...
Even if I had googled (which I didn't) then I'd probably have assumed the fixes for bugs from 2009 to have long made it into the current distro kernels.
I just didn't expect something so basic to be still (or again) broken.
And still, something in your app stack could crash on this as well, leaving the kernel patching pointless.
[1] http://googleblog.blogspot.com/2011/09/time-technology-and-l...
Any chance Google could just make a GPT NTP server available as a public service anyway, just as 8.8.8.8 is their public ping responder. ;-)
Google does provide time servers, although I'm not sure whether they are officially supported. The addresses are:
time1.google.com
time2.google.com
time3.google.com
time4.google.comGood news about the time{1..4} NTP servers, I'll give them a try, thanks.
Because nothing well known blew up, many people wrongly assumed that Y2K was never a real problem to begin with.
[1] I moved a Fortune 100 manufacturing company's database off an ancient mainframe that would've been disastrous come Y2K. It went smoothly and was thus a thankless job. They paid well though (mid six figures - those were the days).
http://news.ycombinator.com/item?id=2261637
It means 500k, or there abouts; 400-600k would probably qualify, with anything higher or lower being mid to high or low to mid six figures, respectively.
Personally, I fixed 3 Y2K bugs back then, 2 of them would have brought down a rather critical business support to simply crash every time new data arrived.
Did you read the parent's post?
From the outsider's perspective, it is indistinguishable from any number of other putative disasters that required lots of money to fix, yet didn't come to pass... in some cases including putative disasters in which the money wasn't spent and the disaster didn't happen anyway.
I have the insider's perspective and I agree that it is the more accurate, that Y2K was, if not necessarily going to end the world, certainly a bad thing and was largely averted through effective engineering. But I can still see how from the outside it sure doesn't look that way.
Many, if not a majority, of my non-technical friends and acquaintances have expressed at one time or another a reference to the "Y2K disaster" and rolled their eyes to suggest it was somehow not an issue. I was the 'Y2K compliance officer' at my startup at the time (we even got certified, and that actually may have been a scam (the certifying part)) but we identified and fixed a number of issues our box would have suffered had we not done the work.
P.S. for people wanting to know more this video is simple to understand but really amazing http://www.youtube.com/watch?v=xX96xng7sAE
It seems to be the unique class of bug that not only is it easy to forget to test, and won't ever show up until a particular date... but then affects everyone!
I can't think of any other kind of bug that never shows up ever, but then affects everyone. Rare bugs tend to stay rare, common bugs tend to get caught before they affect everyone... this is the exception.
(a) it's really not at all difficult to handle leap seconds, but
(b) the POSIX standard specifically disallows them, by specifying that a day must contain exactly 86400 seconds. (Analogously, imagine if leap days occurred as normal, but a "year" by definition contained exactly 365 days.)
The existence of leap seconds means that it's not possible to simultaneously have (1) system time representing the number of seconds since the epoch, and (2) system time equal to (86400 * number_of_days_since_epoch) + seconds_elapsed_today, and all the proposed methods of dealing with the problem involve preserving (2), which seems worthless to me, and throwing away (1), which I would have thought was a better model.
edit: actual system times may be in units other than seconds, but the point remains
The problem of leap seconds is therefore closer to that of time zone definitions -- which are a total mess, because they depend on keeping rapidly changing system tables up to date. I can see why people don't relish the idea of requiring similar tables just to keep system time accurate.
It seems like we already have a much bigger lead time for notification than we could possibly need.
> I can see why people don't relish the idea of requiring similar tables just to keep system time accurate.
But the 'solution' we're using now is to make system time less accurate, not more accurate. Accurate would be if leap seconds incremented the system clock like normal seconds do. If the accuracy you're worried about is displaying a clock time rather than time since the epoch, you already need a time zone to do that.
I am not an expert, but as far as I know the most automated solutions are doing it via NTP, which just resets the second, then relies on clock drift to bring everything back into synch. Otherwise, I think your only option is to keep the timezone packages up-to-date (which is a non-trivial task for large deployments). A quick search found this:
http://www.novell.com/support/kb/doc.php?id=7001865
"But the 'solution' we're using now is to make system time less accurate, not more accurate."
Yeah, I'm not disputing this. I'm just saying that preserving the assumption that "day == 86400 seconds" probably breaks less code than the alternative. NTP messes with the notion of seconds-since-epoch anyway, so we know that single-second variations in unix time aren't automatically deadly to most unix software.
That is time in the same way that leaves moving on a tree is wind.
The instance of the radioactive decay indicates the defined period of time has passed within the local frame of reference. That tells us very little useful about time itself.
What you had in mind is the SI-definition of the second, which reads (http://www.bipm.org/en/si/si_brochure/chapter2/2-1/second.ht...):
"""
The second is the duration of 9 192 631 770 periods
of the radiation corresponding to the transition
between the two hyperfine levels of the ground state
of the caesium 133 atom.
"""
The details of this effect are a little more complicated, but it boils down to the fact that you can measure the "angular momentum" of your nuclei when you have them pass through a non-uniform magnetic field. Particles in the one state are deflected differently than particles in the other state. And when you irradiate a particle beam in the F=0 state with the right frequency (9.2 GHz) you can very efficiently swap many particles over to the F=1 state. By adjusting your frequency sufficiently to find the maximum rate of flips, you can can tune for the exact 9'192'631'770 Hz.The cesium particles have not decayed, in principle you could run forever on a certain supply of atoms... even though in practice a Cs-beam is produced on one side, at a hot filament, and dumped to the other side of the clock after passed through the apparatus that performs the steps described above. They will be disposed of when the lifetime of the beam-tube is reached (typically 10 years or so with a few grams of cesium inside).
Edit: In fact there is strong indication that they may be abolished: http://en.wikipedia.org/wiki/Leap_second#Proposal_to_abolish...
We have 25 years to get ready. I still think we'll be patching at the last minute.
(Yeah, lots of systems will be 64-bit by then, but there will still be a lot of embedded crackerbox systems running 32-bit timestamps. It's all the embedded stuff I'm worried about).
Less than a year ago there were already people thinking about your job security. (It's a better explanation than "the glibc maintainers are insane".)
It's gonna be a fun one.
The problem is that 32-bit time is embedded in filesystem representations and related protocols. (eg the POSIX specification for file times in tar.) Therefore even if your machine is 64-bit, it still needs to use 32-bit time for many purposes.
To name a random example, the POSIX specification for times in the tar format is 32-bit. GNU tar has a non-standard extension that already takes care of it. But will everything else that expects to read/write tar files that a GNU tar program implement the same non-standard extension to the format in the same non-standard way? Almost certainly not. And there will be no sign of disaster until the second that we need to start relying on that more precise representation.
At first, you can confirm the status flag like this.
$ ./adjtimex --print | grep status
status: 8209
8209's binary representation is like this. This surely have INS bit "100000000[1]0001" (5th LSB). $ ruby -e 'p 8209.to_s(2)'
"10000000010001"
8193 is the value after the clearance of the INS big. $ ruby -e 'p 8193.to_s(2)'
"10000000000001"
Then, let's set it as a current value. Please ensure your ntpd is not running. $ adjtimex --status 8193 SLE9 (kernel 2.6.5-7.325): NOT AFFECTED
SLE10-SP1 (kernel 2.6.16.54-0.2.12): NOT AFFECTED
SLE10-SP2 (kernel 2.6.16.60-0.42.54.1): NOT AFFECTED
SLE10-SP3 (kernel 2.6.16.60-0.83.2): NOT AFFECTED
SLE10-SP4 (kernel 2.6.16.60-0.97.1): NOT AFFECTED
SLE11-GA (kernel 2.6.27.54-0.2.1): VERY UNLIKELY
SLE11-SP1 (kernel 2.6.32.59-0.3.1): VERY UNLIKELY
SLE11-SP2 (kernel 3.0.31-0.9.1): VERY UNLIKELY
Update (06/26/2012): after thorough code review -> SLE9 and SLE10 not affected at all.Running on a 3.2 kernel
Rebooted them all and they're fine.
This fixed it:
date; sudo date `date +"%m%d%H%M%C%y.%S"`; date;I attempted to fix with adjtimex and the script in the linked question, but to no avail, in the end having to restart them all instead. After that, all was good again.
stop ntpd, run ntpdate or sntp, start ntpd
/etc/init.d/ntp stop; sntp -s <ntpserver>; /etc/init.d/ntp start
Unfortunately sntp / ntpdate wrapper is not shipped with squeeze for example. I've used the binary from SuSE 11.4 just fine on squeeze.
apt-get install ntpdate; /etc/init.d/ntp stop; ntpdate pool.ntp.org; /etc/init.d/ntp start
without ntpd restart
My Debian GNU/Linux 6.0 is still standing
Oh well, reading the issue, the machine date is Sat Jun 30 16:11:31 EDT 2012
Stopped ntpd just in case
I think the Xen host takes care of the synchronization and we need not do it in the guest OS. (see http://serverfault.com/questions/100978/do-i-need-to-run-ntp...).
Is this fine or should we run ntpd for better accuracy?
That won't work. The bug is only triggered when an upstream NTP server reports that a leap second was scheduled. Since leap seconds aren't predictable (and aren't even scheduled very far in advance), just setting the time back to the date of a previous leap second won't do anything.
/etc/init.d/ntp stop; date; date `date +"%m%d%H%M%C%y.%S"`; date;
What a bit sucks is that my VPN was affected to (openvpn) causing my computer to do a poweroff. I replaced the poweroff with
ip route add to 192.168.1.0/24 dev lo
hope that saves me when the next leap second occurs.
From the answer: "The reason this is occurring before the leap second is actually scheduled to occur is that ntpd lets the kernel handle the leap second at midnight, but needs to alert the kernel to insert the leap second before midnight. ntpd therefore calls adjtimex sometime during the day of the leap second, at which point this bug is triggered."
More specifically, there is a condition in which the kernel tries to insert a leap second and, in doing so, attempts to acquire the same lock twice causing the spinlock lockup and (effectively) halting the kernel.
> The work-around is to just turn off ntpd. If ntpd already issued the adjtimex(2) call, you may need to disable ntpd and reboot to be 100% safe.
I assumed some cronjob or something similar was to blame.
[EDIT] Ah, covered elsewhere. Fixed by manually setting the date on the box; stopping/restarting mysqld or ntpd doesn't make any difference.