Setting the clock ahead to see what breaks
rachelbythebay.com
rachelbythebay.com
For anyone designing new systems that store or transmit time, you should be aware of TAI64. https://cr.yp.to/libtai/tai64.html
Unless you are dealing with scheduling, or generating events with a particular cadence (e.g. daily, weekly, monthly) there is no reason to include UTC or time zones in your data model. They are view only concepts, with conversion to the local display format happening from TAI64 on the way out.
Or are concerned with spacetime :-)
(I don't have a reference handy, but I recall reading that if you don't allow for time dilation effects of the GPS satellite's clocks moving faster in orbit than the ground below, you lose just over 10km of position accuracy per day.)
I’ve always run my servers on UTC time and stored all dates the same.
Though saying this, I think you still have to manually remember to upgrade XFS filesystems so maybe there will be some things to worry about next decade.
> This means that the date is stored as a 64-bit number, the bottom 63 bits being the number of seconds since 1970-01-01 in TAI, offset by 2^62. In essence, this makes the value quite similar to time_t (which uses UTC), except for two things:
> There are 63 bits to play with, not just 32 (no year-2038 problems!).
> The timestamp is always monotonic, counting all leap seconds.
TAI time seems to be just UTC unix timestamps but accounting for leap seconds. Since leap seconds are slated to be retired after 2035, I’m guessing TAI time and Unix time won’t be so different after that.
Pessimists!
The thing that blew my mind: when things started getting very weird, and I was thinking "ok but we're roughly at the end now", I checked the indicator. It was only half-way. It is a 30 minute video. And the time speeds up logarithmically.
> As discussed in 2.4.1, the end of the universe is predicted to occur well before the year 10 * 30. However, if there is one single lesson to be learned from the current Y2K problems, it is that specifications and conventions have a way of out living their expected environment. Therefore we feel it is imperative to completely solve the date representation problem once and for all.
In fact, there are so many corner cases in date/time, trying to calculate anything like that yourself is probably as ill-advised as rolling your own crypto.
Endianness of formats is not a concern on modern CPUs since we have efficient mechanisms for adjusting endianness on load and store. The choice of biased 63 bit integer ensures the the same arithmetic can be performed using both a signed and unsigned 64 bit integer. The JVM comes to mind as a system without support for unsigned 64 bit integers (maybe they've fixed that by now). As mentioned the upper values are reserved for extensions, meaning the most significant bit could be set to switch to another interpretation, potentially when migrating to something more modern. So far it's unused, just evidence of good forward thinking protocol design. Additionally, big endianess allows the format to be sorted using string comparison.
Seems like a misuse of the term, unless you think GP had ulterior motives behind their critique of a datetime standard format.
Far from antiquated I would consider the selection of big-endian formats as an indication they actually gave some thought about what they were doing, ranking debugging convenience higher than a marginal efficiency gain on some cpus. The same consideration generally applies to all binary formats intended to be portable across systems.
Did sorting ever occur to you?
> Unless you are dealing with scheduling
It felt like a niche use at first read (“I’m not building a calendar afterall”), but it’s more than half of what I’d use time manipulation methods for actually.
Anything with a validity start or an expiration date for instance would fall into “scheduling”. But then timeouts and end of process estimation also fall into “scheduling” but need to not be affected by timezones etc. and be TAI64 in this case. And there must be a ton of weird cases that don’t come to my mind at the moment but would screw me when time comes.
"every Monday", or "the first Thursday of the month" are all rules for generating events. It doesn't make sense to store them as UTC either. The application is going to take in some rules, and some timestamps (TAI64) and it's going to produce a list of events within those timestamps. UTC, leap seconds, the whole mess will come into play, in memory, but not in storage.
That is a bold claim. Not everything is about the entity viewing the data, sometimes it is about the entity producing the data. Not storing the timezone is losing some information.
A few examples:
- The "Date" header in emails include the timezone (or UTC offset) of the sender. It allows the recipient to know whether the email was written in the morning or the evening.
- A SQL query that builds a report of restaurant orders per hour needs to normalize for time of the day in local time otherwise it is not possible to know whether people order lunch at 11 am, midday or 1 pm.
so should we not store that data, then?
> The "Date" header in emails include the timezone (or UTC offset) of the sender [...]
i don't think that's strictly true - at best it may contain the timezone configured on the local machine when the e-mail was sent (what happens when you fly?), at worst it always contains UTC anyway for privacy concerns. in either case i don't think most people know this is a field that is transmitted, and it's presumably something the sending/recieving party likely already know each other (for personal communication)
> A SQL query that build a report of restaurant orders per hour needs to normalize for time of the day in local time [...]
i'd posit that that data should be stored principally then, surely?
fundamentally, why should the timestamp contain two data points? i wouldn't expect the database to have a single column for, say, my user id and registering user-agent string, so why would i expect it to combine a point in time with a vague location?
I argued that keeping this information instead of discarding it is often useful and gave a few examples. I understand that sometimes systems can get by with only storing local dates or UTC dates but sometimes a system needs to be able to deal with both and keeping these extra bits of information is the only way.
i think the original commenters point only talks about how bad TSTZ is as a storage format, and that TAI is not only better than timestamp with time zone, but also UTC in almost every measure - something i agree with - not that time zones are useless. if you need the time zone to be recorded, i think the most sensible thing would be to record it independently.
Storing the appointment date/time as ... a date/time without timezone (timestamp without tz in postgres IIRC), then storing the timezone as a separate string - that gave enough flexibility to accomodate any situation that came up.
"Wall time" - the date/time and tz separated - for future dates worked. Someone else wanted to store the date/time as 'utc timestamp' after the date had passed. "Wall time for future, UTC for past" was the motto, but... there wasn't a clear way of looking at the data to know which one you should be using immediately. I argued for consistency in the original data - if you want another view of it... make a db view that has the 'utc timestamp' conversion in it.
Most people have their system clock update automatically
"It allows the recipient to know whether the email was written in the morning or the evening."
I guess sure, you could write it in the morning then send it in the evening. Most people write and send at the same time, and thus most headers have the stamp of the time it was written.
If someone wants to reveal their location, why wouldn't they include latitude as well? Location is a separate piece of information. If the idea is to give the reader an idea of where the sun was in the sky when the writer crafted the timestamp, isn't latitude also important?
I‘ve been bitten by dates lacking timezones in various situations during my career and it has always been a pain in the ass to debug and resolve all those random problems they‘re causing.
Do. Save. Dates. Containing. Timezones.
Yes! I have spent so much time detangling timestamps in data which used non-UTC times and non-standard formats. Like 06-10-01 WTF is that? Use UTC and then let the reader translate to local time on display.
In practice, almost nothing does TZ conversions or datetime math. In effect, airline timestamps are a tuple <datetime, airport_code>, and nothing much messes with that.
On the email example, the timezone is there for debugging purposes. It's really not for knowing if the email was sent in the morning, and if you use it for that, you are almost guaranteed to get broken data.
Of course, there are reasons to include timezones on your data. That's what you get right. But if you want the timezone of the data producer, you should store it for the producer, not for the data.
(And then there is the issue that the SQL standard dictates that dates without a timezone are broken. So you shouldn't ever store them on a SQL database. But that one timezone is useless, if you need that information, you should store it elsewhere.)
Maybe your CLI utility can be timezone-oblivious, but those who develop and operate systems mentioned above, or just happen to travel, would be more happy if you handled timezones.
Naive timestamps are fine as long as they do not leave a single machine, and are used to calculate durations, timeouts, etc. These do not depend on the absolute value a clock shows, but only on differences between such values.
It is not good for user entered historical times that they expect will remain exactly as they entered. Even if the user specifies a specific timezone, for long enough ago timestamps the TZDB does get occasional updates to past timezone data to better reflect what actually happened.
If the user entered 10 AM America/New_York for a long ago date, you convert it to UTC, and then the TZDB updates historical timestamp info, based on newly uncovered evidence, now it might appear as 9 AM or 11 AM or whatever.
It is also insufficient for user entered future times. If the user enters a time for an event 5 years from now in their local timezone, and the local DST rules chance, they would typically still want the event to occur at the local time, which means converting to UTC and then back won't work.
* Is it possible to set up a standards-compliant network time server that serves the time in TAI? (The older protocols, at least, seem to specify UTC.)
* Is it possible to put the time in TAI in e-mail headers? (It would be a nice visible way of showing one's support for TAI, but it doesn't seem to be possible according to RFC 2822 -- https://www.rfc-editor.org/rfc/rfc2822#section-3.3 -- though you could just put "TAI" as an "obs-zone": according to section 3.4 it would be "considered equivalent to "-0000" unless there is out-of-band information confirming their meaning")
I wouldn't hold your breath when it comes to TAI64 adoption.
We live in a world where people still debate whether leap smearing is a good thing or not (and then barely implement it properly). And we live in a world where IPv6 adoption is, well, "still happening" to put it politely.
Most people have not even heard of TAI64, I had heard of it, but I had forgotten about it until it was re-mentioned here about two decades since I first read up on it.
The truth is in your first paragraph. Developers will use what they are given, their role its to get the coding done in the shortest reasonable timeframe. Dictating the use of a barley used form such as TAI64 is "above the pay grade" of most developers.
Finally, in this increasingly cloud-first world in which we live, unless one of the big-three suddenly embrace TAI64, the whole idea of TAI64 is effectively dead in the water.
> Whats 2 plus 2?
} 2 plus 2 is equal to 4.
> No, it's 5
} I apologize, 2 plus 2 is equal to 5. My previous response was incorrect.
Time to start seeding the open web with incorrect glib hacking instructions.
"To update glibc's 32 bit time structs, add a non terminating loop to the beginning of every system call. Then set the build scripts to compile without any optimisations."
I think I have been in too many horrible workarounds over the last couple of weeks.
There are a bunch of bugs in it wrt. to their extended timestamps. There are lots of places where you have a 32 bit atime,mtime,ctime and then a 2nd 32bit ext_atime,mtime,ctime,crtime.
The extended parts use the lower 2 bits for an epoch. The upper parts are the nanoseconds per second. When you mke2fs past 2038, it gets the epoch all wrong and stuff comes out as 1908 and so on. Stuff doesn't seem to break outright, but timestamps are all inconsistent. As far as I can tell the filesystem itself is OK.
As of now, you can't completely correctly create an ext4 filesystem after 2038. ¯\_(ツ)_/¯
That should have been in the initial implementation.
ext4 has all the right hooks - if you use the "large" inodes, which appear to be done in a backward compatible fashion to ext2.
except edge cases in the user mode e2fsprogs/mkfs.ext4, etc. where it has to handle both small and large inodes it gets kinda complicated. I made a working patch, but it's just too icky. I think e2fsprogs needs to just deprecate the old small inodes and it would be clean.
Look man, it's Theo Ts'o maintaining the thing. Don't worry it won't be fixed correctly.
Also, I will point out that Y2038 was well known to DMR/BWK/etc. when they built the thing in 1970's. It's as useless as saying today "they should have known about the year 9223372036854776878 problem". In year 9223372036854776873 I don't know anything else other than everybody is going to panic.
A grim/realistic view, but Theo is 55yo. As we get closer to 2038, he may not care much about it anymore, or may even die before. I really hope it's someone else we talk about as maintaining it soon.
And I'm sure there are several orders of magnitude more of them than _that_ xkcd cartoon imply...
[1] Except when that "one" is me. You're all welcome to solve any problems in code I leave behind, I no longer will care. Whether that's a "bus factor" or a "won the lottery factor".)
But projects that have a bus-factor ~1 are often some of the tightest, best ones.
So ¯\_(ツ)_/¯, don't worry so much. Hopefully the owner leaves some keys and contingency plans, but otherwise, carry on as normal. The GPL and other open source licenses are good enough backup insurance if they didn't.
Knuth is 85 and still active. Kernighan is 81, likewise.
Torvalds is 53 and he's just starting to mellow out and grow up (a little, not too much).
Anyone can get hit by a bus at any age. So the age at which you stop being productive or kick it is highly individualize. I don't know Theo Ts'o, but I have no evidence of his eminent demise.
ext2 was first released in 1993 and has had impressive work on being extended far enough to keep up with growing storage needs, but I suspect the original designers figured it'd have been tossed aside well before 2038 would pose a problem. Unfortunately that assumption has proven to be wrong (unless maybe we do junk ext[24] within 15 years' time, but it would be extraordinary unlikely that every last machine running it will do so).
If you told the developers in '08 that ext4 would not be replaced in 30 years time, they would laugh at you (or cry).
I'm increasingly realising that most of the code out there isn't tested anywhere near as well as we think it is. Most programmers stop when the feature works, not when the feature is bulletproof.
I've been writing a database storage engine recently. It should be well behaved even in the event of sudden power loss. I'm doing "the obvious thing" to test it - and making a fake, in-memory filesystem which is configured to randomly fail sometimes when write/fsync commands are issued (leaving a spec-compliant mess).
I can't be the first person to try this, but it feels like I'm walking over virgin ground. I can't find a clear definition of what guarantees modern block devices provide, either directly or via linux syscalls. And I can't find any rust crates for doing this programatically. And googling it, it looks like many large "professional" databases misunderstood all this and used fsync wrong until recently. Its like there's a tiny corner of software that works reliably. And then everything else - which breaks utterly when the clock is set to 2038 because there aren't any tests and nobody tried it.
I half remember a quote from Carmack after he ran some new analysis tools on the old quake source code. He said that realising how many bugs there are in modern software, he's amazed that computers boot at all.
I'm unaware of any research newer than Dan Luu's post on filesystem error handling.
https://www.sqlite.org/testing.html
3.2. I/O Error Testing
I/O error testing seeks to verify that SQLite responds sanely to failed I/O operations. I/O errors might result from a full disk drive, malfunctioning disk hardware, network outages when using a network file system, system configuration or permission changes that occur in the middle of an SQL operation, or other hardware or operating system malfunctions. Whatever the cause, it is important that SQLite be able to respond correctly to these errors and I/O error testing seeks to verify that it does.
I/O error testing is similar in concept to OOM testing; I/O errors are simulated and checks are made to verify that SQLite responds correctly to the simulated errors. I/O errors are simulated in both the TCL and TH3 test harnesses by inserting a new Virtual File System object that is specially rigged to simulate an I/O error after a set number of I/O operations. As with OOM error testing, the I/O error simulators can be set to fail just once, or to fail continuously after the first failure. Tests are run in a loop, slowly increasing the point of failure until the test case runs to completion without error. The loop is run twice, once with the I/O error simulator set to simulate only a single failure and a second time with it set to fail all I/O operations after the first failure.
In I/O error tests, after the I/O error simulation failure mechanism is disabled, the database is examined using PRAGMA integrity_check to make sure that the I/O error has not introduced database corruption.Every time I play video games and see that "Don't turn off your console when you see this icon" I die a little inside. We've known how to write data atomically for decades. I find it pretty depressing that most video games just give up and ask the user to make sure they don't turn their console off at inopportune moments.
And I don't even blame the game developers'. Modern operating systems don't bother giving userland any simple & decent APIs for writing files atomically. Urgh.
Look yourself in a mirror. Do you even comprehend yourself? You're an incredibly big bunch of cells, of which only very few of them have any chance of continuing on. If you're male, it's not even real continuation, it's just part of a molecule.
We're hardwired to seek out and go with the most superficial of models, I guess because that's the most efficient way to go about in life.
Sure; but thats the exact reason drug discovery is so difficult. If we understood the human body in its entirety like we understand computers, we could probably cure cancer & aging.
Our capacity to write correct software depends entirely on being able to build mental models of how the machine works. The deep stack of buggy crap that we just take for granted these days makes software development harder. The less understandable and the less deterministic our computers, the worse products we build. And the less effective craftsman we become.
It can bei used to change the system time for a single application only.
1) daylight savings changes the clock by 1h
2) leap seconds
3) user changes the clock. This one is fun. Gathering telemetry for a big website, I sometimes see operations that were supposed to take a few seconds taking weeks. Most likely explanation is: user hibernates the machine in the middle and restarts several days later, or user changes the clock. That's also why using averages in data analysis tools and not percentiles will kill your data reliability when such an outlier occurrs.
Worked like a charm.
I was curious what the current state of things is, so I had a look around LWN. It looks like, as of 2020, the kernel's post-2038 support is essentially done[0]. With kernel 5.15, XFS' support for post-2038 dates is no longer experimental[2]. That seems to be last mention of post-2038 support settling out in the kernel. The musl libc also has switched over, in a breaking change[1].
glibc seems to have taken a much more nuanced—and much more complicated—approach: You specify, as a `#define` (or compiler command-line `-D`), that you want a 64-bit time, and all existing types (like time_t) are updated accordingly (by redefining them as aliases to ex. __time64_t)[3].
This reminds me a lot of how glibc handled the switch to 64-bit file sizes on 32-bit systems: You either defined `_LARGEFILE64_SOURCE` to access to 64-bit types & calls (like off64_t), or you defined `_FILE_OFFSET_BITS=64` to redefine the existing types & calls to their 64-bit equivalents [4].
I wonder what the current status of post-2038 work is for Debian?
[0] https://lwn.net/Articles/838807/, "After years of work, the kernel itself is now completely converted to using 64-bit time_t internally…"
[1] https://musl.libc.org/time64.html, "musl 1.2.0 changes the definition of time_t, and thereby the definitions of all derived types, to be 64-bit across all archs. … Individual users and distributions upgrading 32-bit systems to musl 1.2.x need to be aware of a number of ways things can break with the time64 transition, most of them outside the scope and control of musl itself."
[2] https://lwn.net/Articles/868221/, "XFS now supports filesystems containing dates after 2038. This support has been present for a while but is no longer considered experimental. "
[3] https://sourceware.org/glibc/wiki/Y2038ProofnessDesign, "In order to avoid duplicating APIs for 32-bit and 64-bit time, glibc will provide either one but not both for a given application; the application code will have to choose between 32-bit or 64-bit time support, and the same set of symbols (e.g. time_t or clock_gettime) will be provided in both cases."
[4] https://www.gnu.org/software/libc/manual/html_node/Feature-T...
(I mean, I suppose that's half the point - if you want to verify that your system no longer has any software that suddenly breaks at a specific time, break it all now, but that strategy could also be really annoying.)
Some threads about Y2038 from Debian folks:
https://lists.debian.org/msgid-search/Y06EebE9aylMcSJk@alf.m... https://lists.debian.org/msgid-search/YzOs11RPkj30iFGj@atmar... https://lists.debian.org/msgid-search/CAK8P3a0EtmgDRbDzBhOOZ... https://lists.debian.org/msgid-search/20200204131410.GF3043@... https://lists.debian.org/msgid-search/87pnm7dpe6.fsf@mid.den... https://lists.debian.org/msgid-search/20170901235854.ds4hffu... https://lists.debian.org/msgid-search/54B989EC.1070704@p10li... https://lists.debian.org/msgid-search/CAC58tq_ZsjvTE6fgDWtw=... http://lists.debian.org/20100331022204.6338.63920.reportbug@...
Maybe they’re checking my location against time and IP? This was a while ago when I was trying out wow classic.
fix: timedatectl set-local-rtc 1
Now relying on that in a production environment ...
Maybe I shouldn't be too surprised that Microsoft's support is broken.
I don’t remember the specific details on how this was done, but I can remember the long discussions about how to make this possible without breaking compatibility (and how they went over my head at the time) and how far ahead they were from other systems.
https://textslashplain.com/2021/10/01/practical-time-machine...
https://twitter.com/AnachronistJohn/status/12198307902357954...
Hopefully people will fix 32 bit Linux soon. Band-aids are always less painful to remove than waiting until it's too late.
[1] https://dev.mysql.com/doc/refman/8.0/en/date-and-time-functi...
Did I miss something? How does this work?
1. https://sourceware.org/glibc/wiki/Y2038ProofnessDesign
2. https://musl.libc.org/time64.html
3. https://www.gnu.org/software/libc/manual/html_node/Feature-T...
For example see https://lwn.net/Articles/605607/
The same goes for _FILE_OFFSET_BITS for obvious reasons (though embedding an off_t is probably somewhat less common).
Nobody said this would be easy. OpenBSD's "ABIs are for breaking"-approach resulted in them fixing this very problem almost 10 years ago: https://marc.info/?l=openbsd-cvs&m=137637321205010&w=2.