Ntpd won't save you from one particular rogue bit
rachelbythebay.com
rachelbythebay.com
I learned that once tried to implement "universal" datetime library for fun and education. There is wall clock, which leaps for political decisions; universal time coordinated which leaps on schedule to ease astronomical differences; atomic time of few sorts; numerous computer clocks; time zones which can go back and forth for north and south, land and sea; non-gregorian calendars and the fact that there is no 0 AD and no 0 BC; that dates were offset for many days few times in gregorian history; special and general relativity errors; and of course integer overflow issues.
I may have missed few points, but still, time is hard.
"Schedule" is a bit generous there.
Oh and there's TAI for just counting seconds but you're discouraged from using it that way because of some slightly-suspicous reasoning about retrospective calibration.
On Linux, there's now a CLOCK_TAI, but I'm not sure how usable it is. Does it have the correct offset on boot? when ntpd starts? and you have to convert back and forth to interact with the rest of the world, and there's no call to get the time and the conversion factor atomically, which is what you'd probably want if you're using the same timestamp internally and externally. (I think there's some non-privileged call to get the conversion; you could call that, get the time, and then call it again; retry if it's changed. but that's annoying.) And of course this is non-portable.
edit: additionally, if you have a stored, bare TAI value, what can you do with it other than see how many seconds older it is than the current time? nothing on a standard system stores when previous leap seconds happened iirc so you can't convert to UTC (and thus can't convert to civil time) unless you roll your own table (and conversion routines) or always store the offset with every TAI (increasing storage overhead).
"TAI and UTC have a fixed offset relationship, it is true. However, UTC is computed in real time (with several varieties to choose from if you care about the nano-seconds), but TAI is a retrospective timescale that's not computed until after the fact. I get the feeling that the BIPM want TAI to be their baby, free from "production" concerns that UTC has to deal with"
I wonder where that limit comes from, since the RTC in a standard PC stores both the year(00-99) and century(also 00-99) in BCD, so a date in 2153 should be representable.
This reminds me, one interesting thing I've noticed over the years is that early PC's RTCs were pretty accurate, but the most recent ones, as in the past few years, are horrible (as much as +/- several seconds per day), and the ones in smartphones even worse. Maybe because they're assuming NTP, so have cut costs by using less accurate crystals? I've had systems in storage for many years, and the RTC was within a few seconds of the current time when turned on again.
Anecdotally clocks in small mobile devices like phones do seem less accurate though, at least when they don't have access to the network to update themselves. I've not looked into it but assumed temperature variances, which a device like a phone will experience to a larger extent, were part of the problem.
Of course phones syncing with the network relies on their clocks agreeing with the rest of the world - for some months my provider seemed to be a minute or so behind which meant I had to adjust a little when gauging if I was just going to make train I was rushing towards or if I'd likely missed it leaving already...
A GPS is basically just a very, very precise clock and a radio receiver.
Why don't smartphones use the GPS clock?
I'm sure there's a good reason (and some probably do), but it sill surprises me. (Presumably it's battery related -- the GPS clock can't be used unless GPS is fully enabled or something similar.)
That's actually only about 18ms difference in delay. Unless there's something I'm failing to account for here, you should be able to get more than 0.1s precision based off of a single satellite's signal.
Subtract a flat ~75ms to account for the minimum delay (~66ms @ 20,000km) and half of the variable delay, and you shouldn't really be more than about 9ms out.
Oh, I bet my tooth that I can hear my phone playing music slightly faster than usual sometimes. Never experienced that on PC. Maybe it is not connected to RTC and is purely biological condition (never cared enough to test two devices side by side), but if someone experienced that too and has an explanation, it would be great to know.
Which does not prove it's a fault in the device—it could be a fault in my fried brain—but anecdotally, I've never noticed it on any other music playing device, since the time of the original walkman (and in that case the music would usually slow down due to attrition or other mechanical faults, but rarely if ever speed up.)
I'm wondering -- "What might cause music playback to be faster than normal?" I don't think I know enough to tell. Perhaps the clocks on the audio playback device are skewed? Most SoCs use a dedicated DSP which doesn't seem terribly different from PC's. Not sure, but we can say for certain that it's definitely immune to changes in wall clock time (whether in the RTC or its representation in the OS).
Even more fun was that it was a dual-boot machine, and the audio worked perfectly in Linux but shifted up in Windows. I honestly thought I was losing my mind for a while. I'd listen to the same song in Spotify on both OSes and just get this "something is wrong here..." feeling in my gut.
Maybe... but how much money could they possibly be saving? Quartz oscillators like the ones used in cheap digital watches are plenty accurate and completely abundant. I can't imagine they cost very much.
Maybe the issue has to do with miniaturization? If you're trying to make smaller and smaller devices, there could be an acceptable tradeoff between size and precision.
The tuning fork crystal design used in wrist watches has a parabolic temperature coefficient, which means that the clock is only really accurate at room temperature. This isn't a problem for wristwatches, because, presumably your wrist is at approximately room temperature, but it does become a problem for electronics that operate with large temperature swings. (like the inside of a phone or computer, for instance).
https://www.maximintegrated.com/en/app-notes/index.mvp/id/58
A simple tuning procedure could make almost all quartz clocks you own a heck of a lot better. But in many cases this is hard to do because the manufacturer never put the tuning components (a variable capacitor or the like) on the PCB.
Insofar as sensitivity to thermal conditions is an issue, early PCs tended to be big, airy, ventilated boxes and have less bits that would ramp up to high temperature under load but drop down in other operating regimes. So that may be a factor.
Chromium threw a fit because news.ycombinator.com's SSL cert had expired and offered to reset the clock which it could not do, given that I'm not in the habit of running apps as root. My Kerberos tickets all expired so Evolution lost contact and so did quite a few other things. MariaDB dumped core. systemd's timers went berserk so things like my LetsEncrypt cert tried renew itself.
I'm still hunting through some odd looking log files but overall things seem to be back to normal when I reran timeshift and returned to current time.
Scary to think what might happen if some database purge process is running after this bit flip!
A much more common case of when such a check help is "last time run is in the future", that you face every time there is a clock issue. Some scripts... Don't react too well about that.
If all your programs use /dev/rtc and ioctl()s you are probably a little safer because there will be coarse locks around the RTC itself, those will serialize the activity. But IIRC the inb/outb stuff can be done from user space (as superuser) and even if you're only reading you have to write to the address register which could break a write-in-progress by sending its output to the wrong RTC field.
Of course I laughed when I read that, but then I realized I was interpreting that line to mean "it declares shenanigans and stops reporting time so you realize something broke."
Just wanted to clarify - do you mean the above, or "of course software cannot kill hardware!!1" "won't tick anymore"?
Note that it doesn't stop reporting the time in this condition it will just give you back all the garbage fields that you wrote before. So an interesting thing happens when the system tries to transform that into a UTC wall clock basis and it usually ends up with a really wild interpretation of the date (decades/centuries off, similar to the problem described in TFA).
Ah, okay then. Good to know I can't accidentally thousands of dedicated servers :P (ie, via NTP MITM, writing to hw RTC...)
> Note that it doesn't stop reporting the time in this condition it will just give you back all the garbage fields that you wrote before.
I don't know why I didn't remember this last night: Linux uses the RTC as a poor-man's NVRAM that will persist across a reboot. Provides <24 bits of data to work with (yay! ...not). https://wiki.ubuntu.com/DebuggingKernelSuspend, useful info in https://github.com/torvalds/linux/blob/e34bac726d27056081d02...
> So an interesting thing happens when the system tries to transform that into a UTC wall clock basis and it usually ends up with a really wild interpretation of the date (decades/centuries off, similar to the problem described in TFA).
Right.
So if you change that flag without also setting the time, then many of the fields are now invalid. But the tick logic doesn't care. It just sets the invalid fields back to zero and keeps on ticking.
The year 20 17 becomes 1 4 1 1.
The RTC clock is a 1984 part that is still getting embedded, more or less unchanged, into today's PCs. It is maddening.
It shouldn't do that, otherwise it's going to break in 2036.
(The "future time" it's going to clear off will set the time back to 1900 once we are past February 7, 2036)
Edit: I don't know if I didn't explain myself well enough. Because the protocol only has 32 bits of seconds it cannot tell apart September 27, 2017 and November 4, 2153. This means that ntpd absolutely must trust that it is in the correct 68-year span. But according to the author ntpd "clears out the future time" if it is more than a few milliseconds off. This violates the spec and is also inconsistent as it happily keeps up with the future date as long as it doesn't have to step the clock.
Or actually there is this little tidbit from RFC 5905:
> Eras cannot be produced by NTP directly, nor is there need to do so. When necessary, they can be derived from external means, such as the filesystem or dedicated hardware.
So it could be that ntpd looks at some file or the rtc to determine the era rather than assuming the current system time is in the correct era, and it would be allowed by the spec. But it's quite inconsistent if it only does it if the system time is more than a few milliseconds off (presumably beyond the limit of when it corrects the time by time stretching). I'm going to go with it just being a bug in ntpd.
Your message merely reminded me of that and so I put my message here as I saw similarity. Having not read the ntpd spec in question I can't answer you.
https://twitter.com/petecooper/status/911946604759977984
https://pbs.twimg.com/media/DKficrvW4AA1Hxm.jpg:large
Numerous hosts across numerous networks, perhaps two or three an hour.
I've wondered what exactly would be gained by resetting a clock to a different time – this is a useful article.
Are you sure that it is not just NTP servers being served via DHCP once the connection is established and your computer trying to use those provided by your VPN?
I'm not 100% sure about any of the ntpd connection requests, they're not predictable in their appearance. Some sessions are very quiet (zero requests), others I get a bunch of incoming connection requests for smbd, and other odd things. I really should start taking notes rather than just deny the connections.
That could just be different pool addresses coming up.
1. your VPN provider is giving you an actual public IP address (??)
2. people are scanning your computer for NTP vulnerabilities or something (this happens if you have a public IP, regardless of network)
3. NTP is using UDP and so connectionless, and so Little Snitch can't distinguish "ntpd wants to reply to someone who contacted it" from "ntpd wants to connect to someone"
An alternative explanation for 1/2 is that your VPN provider is not isolating you from other VPN users (less surprising than giving you your own public IP) and someone else on the VPN is trying to conduct NTP amplification attacks using you: https://blog.cloudflare.com/understanding-and-mitigating-ntp...
In either case, the solution is basically to make your ntpd not listen for requests from other machines and only handle time from your local computer + initiating requests to time.apple.com or whatever your chosen NTP server is. It shouldn't be trying to reply at all to unexpected packets, even to send a refusal message (again, because UDP is connectionless, it's easy for an attacker on your LAN to send spoofed packets and convince you to send replies to some random computer on the internet, and I guess on this VPN, other customers are your LAN). I'm surprised that macOS's default NTP server isn't configured this way out-of-the-box, though.
This, I think, is most likely.
I'm not sure if this rogue bit can be used to attack TOTP. Can anyone clarify?
However, you could imagine an NTP implementation which hardcodes the approximate starting time to get the right era. You'd only have to recompile every lifetime or so to keep it up to date.
Yep, see "NTP pivot dates" [0]:
> When ntpd(8) receives a unresolved timestamp from an upstream server that timestamp could be based in any era ... To resolve this ambiguity, NTP also uses an internal pivot date ... An ntpd(8) instance’s pivot date will be the date it was compiled and built.
[0]: https://docs.ntpsec.org/latest/rollover.html#ntp_pivots
Was anyone else successful with the PoC code?