2020 Leap Day Bugs
codeofmatt.com
codeofmatt.com
Anyone still using a Zune out there?
I'm definitely going to try it
But now, after reading your link and giving it a bit of thought, what s/he wrote makes perfect sense even without your link.
So... thanks.
Build your time library on (un)signed 64 bit integers representing the number of nanoseconds since the utc epoch. Adjust above sentence to reflect the level of precision and range your use case needs. You are now done for 80% of use cases (perf timing, logging, timeouts, event storage, event ordering within jitter).
If you need to parse/display for humans or have something happen at a particular time in a particular timezone, things get gross. But that's no different than any other situation where you eventually have to interface machine data with humans. Either it's your particular expertise or it's a distraction and you should use someone else's solution.
You poor sweet summer child. Has no one told you about the leap seconds yet? Unix time is not the number of seconds since epoch. It deliberately excludes leap seconds, which happen unpredictably whenever scientists measure the Earth as having spun at a different enough speed for long enough.
Time is fucked on every level:
- Philosophical: What is time? We just don't know.
- Physical: Turns out there is no such thing as simultaneity, and time flows differently at different locations. Time may be discrete at the Planck level, but we don't really know yet.
- Cosmological: The Earth does not rotate at a constant speed, the Earth does not orbit the Sun at a fractional component of its rotation, and the Moon does not orbit at even ratio either.
- Historical: Humans have not used time or calendars consistently.
- Notational: Some time notations are ambiguous (e.g. during daylight savings transitions) and others are skipped.
- Regional: Different regions use subtly different clocks and calendars.
- Political: Different political actors choose to change time whenever they feel like it with little or no warning.
- Religious: Many religions come with their own system for timekeeping, and people don't like when outsiders impose other systems.
haha. Google is way ahead of you and your "leap seconds".
"Since 2008, instead of applying leap seconds to our servers using clock steps, we have "smeared" the extra second across the hours before and after each leap. The leap smear applies to all Google services, including all our APIs."
https://developers.google.com/time/smear
...now write code to convert Google time to any other random type of time.
https://pubs.opengroup.org/onlinepubs/9699919799/xrat/V4_xbd...
and https://en.wikipedia.org/wiki/Unix_time
indicate that leap seconds are not included in Unix time as seconds-since-the-epoch. Leap seconds are included in UTC, and thus the Unix time appears to skip a second relative to UTC.
I assumed that is (part of) why he ended up wanting to build his own. The way UNIX time handles leap seconds was arguably a mistake, GPS (and Galileo) time does it right; have the leap second information in a separate field. So as a time scale I imagine OPs one would be similar to GPS but with different epoch. Of course everyone loves having informally defined ad-hoc timescales around
In fairness, he said to count the number of seconds since the Epoch. That is independent of UTC.
That means using e.g. TAI. Unix time also ignores leap seconds since it counts the number of seconds actually elapsed.
hence it counts the number of seconds that would have elapsed, if they didn't exist. In effect, it's timescale has time that never happened, and time that happened twice.
It may be determined that the computer's clock got too far ahead. For example, it booted and you cared about time, but NTP hadn't yet made corrections. Suddenly the time runs backwards.
Leap seconds may get interesting too, especially if you have to predict ahead or if the OS isn't updated often enough. (there is a 6-month warning) If you want to call leap seconds an issue for humans, then you aren't using UTC at all. You're using TAI. Software interfaces often ignore the distinction between UTC and TAI, and even between UTC and UT1, preferring to pretend these issues don't exist. POSIX is in conflict with international timekeeping, effectively requiring that there are zero leap seconds.
It is my understanding that the way NTP deamons work is that time is never adjusted backward. Instead, the ticks are "slowed down" on the local machine until it is in sync with the NTP time. However, if the difference is too great then I think NTP deamons might refuse to correct the time all together. So then, if my understanding is correct, your machine is "stuck in the future". But it will never make a jump backwards because of NTP.
However, I am not familiar with the intricate details of NTP so do take this with a grain of salt.
Subject to configuration, of course. man ntpd [1] says:
> Sometimes, in particular when ntpd is first started, the error might exceed 128 ms. This may on occasion cause the clock to be set backwards if the local clock time is more than 128 s in the future relative to the server. In some applications, this behavior may be unacceptable. If the -x option is included on the command line, the clock will never be stepped and only slew corrections will be used.
/*
* Maximum delta in seconds which the system clock is gradually adjusted
* (slewed) to approach the network time. Deltas larger that this are set by
* letting the system time jump. The kernel's limit for adjtime is 0.5s.
*/
#define NTP_MAX_ADJUST 0.4Build it yourself and pass -DNTP_MAX_ADJUST=XX
The ntp case is contrived. Either you care and wait until ntp has connected to do your stuff. Or you care and don't let ntp rewind and instead smear. Or you don't care and deal with the consequences.
They are excluded from "unix timestamps". They would still be part of the number of seconds since Unix epoch.
AFAIK, unix time skips a beat or repeats itself to remain in alignment with UTC, and just keeps chugging along.
Unix time does indeed repeat itself so as to remain in alignment with UTC. But unix time is not the number of seconds since the unix epoch.
Leap seconds are not excluded from number of seconds since Unix epoch, interpreted on a physical, TAI, UTC, typical local time scales.
However, they are indeed excluded from "number of seconds since Unix epoch", when interpreted in unix time. In unix time, those seconds simply don't exists (never happened) and the events of those seconds are smooshed sometime. Unix time representation forms a (quirky) timescale.
I'll quote a few excerpts from 'date' man page of linux, openbsd and for gettimeofday:
Convert seconds since the epoch (1970-01-01 UTC) to a date
Print out (in specified format) the date and time represented by seconds from the Epoch.
The time is expressed in seconds and microseconds since midnight (0 hour), January 1, 1970
I think that UNIX time stamps are generally a very good approximation, and if you are comparing long enough time intervals for the error to get over one second, and/or that error to matter, you are doing something wrong anyway.
For exact time interval measurements that you have to get exactly right, don't use UNIX time stamps.
You could store something as '2022-02-02 02:22:22 UTC', but try storing it as seconds since epoch, if you don't know if there will be leap seconds. Not to mention software that may be years or decades old.
https://en.wikipedia.org/wiki/Leap_second#Binary_representat...
It’s frankly amazing how much they changed every year. Different counties, and sometimes towns/cities within counties, would jump back and forth year to year. It would have been awful to manage if computers had been more important at that time.
This also pushes all the madness to the edges and out of the business logic.
That being said, this works well for applications that are not "date intensive" so to speak. If your business logic has to deal specifically with calendar dates, e.g., monthly events, then you have to deal with calendar months and all that this involves, including explicitly dealing with the 29th February.
If you care more than this, you'll be displeased with off the shelf solutions too
10m12s watch. Informative and entertaining.
If that's not true, then there's room to complain, but dates and times fit the bill. In some ways they're worse than other "harder" problems, because people are more likely to think those harder problems are too hard for them. And while the vast majority of companies don't need a proprietary database, it's more likely to be a competitive advantage than your own datetime library.
I think I could eventually write a good datetime library. But I certainly should not, unless I decide that's going to be one of my major efforts to help a language that doesn't already have one.
Instead, what if people tried to hear graciously and "assume good faith"?
The quote is ripped from HN's guidelines.
For instance, I've not placed the burden of "conveying identical information in a manner everyone can understand" on anyone, nor have I assumed bad faith. So it seems like a weird comment to tack onto mine, but I am assuming I just don't correctly understand.
For example, if you live in London then "9 in the morning" means when the world synchronized clocks agree that, locally for you, the time is 9am. But if you say "9 in the morning" to someone in Belize, it means "first thing after you are finished with your morning and ready to start your day," which can mean 1pm in some cases.
Here's a lovely article on these kind of time-keeping differences, around something that you might expect to have a precise meaning:
https://www.businessinsider.com/how-different-cultures-under...
More to the point at hand, however, you suggested that people not speak in hyperbole but instead speak accurately. Although you can request that others adjust their use of language while in your presence to better meet your needs for a certain kind of precision, policing other people's language isn't possible. However, re-interpreting what people say into what they mean is somewhat possible for an astute listener who understands the context.
I'm not walking around demanding people change their communication to accommodate me, I'm suggesting that trying to speak precisely can be a useful exercise.
I am not the person you originally responded to, fwiw, and I agree that trying to speak precisely is a useful exercise, even if I believe it is impossible except in highly formalized languages.
As to explaining more about my thought process: You asked a hypothetical "what if" question which, given the question itself is imprecise, I interpreted as you wishing information was always conveyed to your desired precision/accuracy, and extrapolated that (given this is a public forum) into general communication.
I shared my thoughts, specifically that it seems impossible for a person to always communicate perfectly to an unknown audience, and offered a different hypothetical. The last line explains the punctuation use in my "what if."
Sounds boring but functional.
In Python, if you take a datetime, and call .replace(year=X) on a datetime for Feb29, it'll throw a ValueError.
There's a nice package called "python-dateutil" that includes a "relativedelta" class; adding a month to March 31 results in April 30:
In [7]: datetime.datetime(2020, 3, 31) + dateutil.relativedelta.relativedelta(months=1)
Out[7]: datetime.datetime(2020, 4, 30, 0, 0)
Adding a year to a leap day: In [8]: datetime.datetime(2020, 2, 29) + dateutil.relativedelta.relativedelta(years=1)
Out[8]: datetime.datetime(2021, 2, 28, 0, 0)
The exact duration that relativedelta adds depends on what you add it to. (Hence the name.) But the results tend to match up with human expectations.Also, if you add one month to Apr 30 and two months to Mar 31, you’ll see that month addition is not commutative-y and one should operate on distances from a base date, not in an incremental way.
Edit: So, another leap year would be fine.
> Confirmed! Your mystery Leap Year offer awaits inside. How much will you save?
The second arrived at 9:54 AM, with subject
> Oops! Your code is fixed. (We shoulda looked before we leap-yeared...)
The messages themselves appear to be identical except for some query parameters on some URLs, so I'm guessing that whatever they botched for leap year was on the server when one tried to respond to the offer.
Botching Feb 29 in general date handling code is embarrassing, but there is it least a somewhat plausible excuse that you just forgot about that special case. But botching Feb 29 in code that is meant to only work on Feb 29? Wow.
Experienced people who avoid these issues are mostly experienced with smashing headfirst into one of these tricky areas and approaching them with an attitude of "here be dragons" afterwards. One bitten, twice shy.
I played around with it a little on google maps. It says it pulls route information directly from the King County transit website[1] which lists routes by the day of the week, not by the date, and seems to be displaying the routes for today just fine. My guess is that it's a Google issue, but maybe the King County transit website was broken earlier today and gmaps is just serving the cached routes now.
For anyone curious, this is what it currently looks like to get from Bellevue to Redmond in google maps. https://i.imgur.com/wMTIzBL.png
[1]https://kingcounty.gov/depts/transportation/metro/schedules-...
They did the math for which year is a leap year incorrectly, which wasn't the bug that bit us, they helpfully added the leap day regardless of where in the year it was. The tests they wrote targeted the middle of the year and missed it. I just pulled in date2j and related match from postgres, added more test cases.
Java, for example, famously doesn't even have an unsigned 32-bit integer primitive type. (But it has library functions you can use to treat signed integers as unsigned.) Ultimately not a good design choice, but the fact that it actually wasn't that limiting and relatively few people care or notice tells you that many people have a mindset where they use signed integers unless there's a great reason to do something different.
Aside from just mindset and inertia, if your language doesn't support it well, it can be error-prone to use unsigned integers. In C, you can freely assign from an int to an unsigned int variable, with no warnings. And you can do a printf() with "%d" instead of "%u" by mistake. And I'm fairly sure that converting an out of range unsigned int to an int results in random-ish (implementation-defined) behavior, so if you accidentally declare a function parameter as int instead of unsigned int, thereby accidentally doing an unsigned to signed and back to unsigned conversion, you could corrupt certain values without any compiler warning.
It was a horrific design choice, and people do notice and do hate it.
Not so much because of how it applies to ints, but because the design choice Java made was to not support any unsigned integer types. So the byte type is signed, conveniently offering you a range of values from -128 to 127. Try comparing a byte to the literal 0xA0. Try initializing a byte to 0xA0!
In contrast, C# more sensibly offers the types int / uint, short / ushort, long / ulong, and byte / sbyte.
AFAIK, that doesn't exist? Otherwise, unsigned would also have UB?
I'd even go so far to say that defaulting to signed instead of unsigned was also one of the biggest blunders ever. I would've never defaulted to a type that inherently poses the risk of UB if I have another type that doesn't.
Though it's also possible that precisely that was the reasoning for it.
Some more thoughts on unsigned vs. signed: https://blog.robertelder.org/signed-or-unsigned/
isotime = datetime.strptime(time.get('title'), '%a %d %b %I:%M:%S %p').replace(year=YEAR, tzinfo=TZ).isoformat()
ValueError: day is out of range for month
It's not mission-critical, so I'm just going to wait it out until tomorrow.EDIT: Hacked it out.
isotime = datetime.strptime(f"{time.get('title')} {YEAR} {tzoffset:+03d}00", '%a %d %b %I:%M:%S %p %Y %z').isoformat() Python 2.7.17 (default, Dec 31 2019, 23:59:25)
Type "help", "copyright", "credits" or "license" for more information.
>>> from datetime import datetime
>>> datetime.strptime('Sat 29 Feb 12:00:00 PM', '%a %d %b %I:%M:%S %p')
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
ValueError: day is out of range for monthe.g.:
>>> from dateutil.relativedelta import relativedelta
>>> date = datetime.utcnow().date()
>>> date
datetime.date(2020, 2, 29)
>>> date + relativedelta(years=1)
datetime.date(2021, 2, 28)Thanks.
It also has this incorrect example: https://docs.python.org/3/library/datetime.html#examples-of-...
Also, but probably unrelated, my WiFi stopped working exactly 00:00AM CET but that might be unrelated as it does that all the time :)
the rule I know is divisible by 4 && ( not divisible by 100 || divisible by 400) so 2020 is not even an edge case.
"Leap day bugs, again! People, must we go through this every 4 years?"
So this is a crop of bugs that are either code written in the last 4 years or untested/unnoticed/unreported 4 years ago.
One of the Sprint examples (someone roaming between two cell towers and their date changing back and forth) is especially disturbing.
Seems hard to believe that a mobile app would be allowed to tell the server what time the cheque was deposited!
Why doesn't the syslog protocol (RFC5424) deal with leap seconds (the seconds field goes to 00-59, not 00-60)? Are they using UTC (they would have to ignore LS and have crappier logs) or TAI (doesn't have LS)?
https://mailarchive.ietf.org/arch/msg/syslog/DDLgKsRPITFXYSB...
http://www.madore.org/~david/computers/unix-leap-seconds.htm...
https://tools.ietf.org/html/rfc5424
It's something that does not matter in 99.99% of the cases, and that other 00.01% need specialized hardware and software for dealing with it anyway.
The corollary is that for the most commonly used APIs a Unix second isn't the same thing as an SI second. A Unix second is effectively defined in terms of the civil calendar, not as a fixed quantum of physical time.
Technically this doesn't preclude Unix date-time strings from displaying a 60th second. (And maybe some do.) But it would require unnecessary extra work, introduce inconsistencies (i.e. a generated string for a future date-time that happened to be a leap second would show :59 today but :60 at the moment of the leap second), and invite bugs.
Leap seconds are things for stock market servers, atom clocks, scientific devices and the like.
Throwing shade without evidence doesn't demonstrate professionalism, it demonstrates laziness.
Correlation on high-volume production systems requires precise timestamps all the time.
Who's throwing shade?
Let's see if it happens today since it's now the 1st.
(year % 4 == 0 && year % 100 != 0) || year % 400 == 0
The following test also works: year % 4 == 0 && (year % 100 != 0 || year % 400 == 0)(A & B) | C differs from A & (B | C) in two cases: when A is false, C is true, and B is any value.
But in this case in particular, those differences would create a bug when A is false (the year is not divisible by 4) at the same time that C is true (the year _is_ divisible by 400). Since that can't happen, the bug cannot occur.