Stop using utcnow and utcfromtimestamp
blog.ganssle.io
blog.ganssle.io
Sure, you can format datetimes in localised formats for end users, but standardising on UTC+0 and RFC-3339 for programming and maintenance purposes just prevents so much confusion and hassle.
Not that I've ever had to spend a day trying to troubleshoot a bug that was alerted by our monitoring system at 2020-11-09T09:11:05Z (as it was defaulting to UTC), but the relevant logs were timestamped at 10:11:05 09/11/2020 (I only wished they'd used the normal German period separators in the date to give me a clue) because of course your German colleagues want their timestamps in German format and defaulting to CET...
When you have colleagues in Europe, Oceania, and distributed from coast to coast in the US, doing everything in UTC means everyone only has to do one mental conversion from their local timezone to UTC, as opposed to trying to remember if they're in LA or New York, or if they're in Mountain Time, if they're in the parts of the US that are Mountain Time but don't do daylight savings.
(I always set CW back to UTC, for consistency)
I also remember that AWS used to display some datetimes in Pacific Standard Time, which is completely useless to somebody outside the US, however I haven't seen that in a while anymore.
Similar to the global language switch I wish the AWS Management Console would allow setting in which timezone to display datetimes globally as well.
It gets even more fun when such events happen during the switch between daylight saving time (like CEST) and standard time (like CET).
Ah, daylight savings.
"2020-11-09T09:11:05Z" uniquely identifies a point in time. "10:11:05 09/11/2020" is some time 11 minutes past the hour on either September 11th or November 9th. It doesn't identify anything.
Edit: Or 41 minutes past the hour.
I don't want my equality operator doing timezone conversions. If one type has no timezone attached and one does, then you probably shouldn't be able to compare them at all. Likewise, for a type that has a timezone attached, if the timezone differs between a and b, then I want a == b to return false, even if it represents the same instant in time.
Equality can be a tricky thing. Named functions are the way to go when two different developers might intuit different behaviors.
Really? So if you want to compare that two times are really the same time for two different timezones, you'd want to convert them to UTC first? What's the use case for this?
Seems like the less-used way to me.
The fact that your intuition of "these two timestamps are equal" is "these two timestamps denotes the same instant" seems problematic when we know that the notions of timestamps and time are not in fact aligned (because, for example, of leap seconds).
But an API should never get into a place where your personal choices should matter and various incompatible interpretations potentially make sense to different people.
No, most people do want to compare times in different TZs without the extra cruft.
I think a better API would be to have a different types for timezone- naive vs. aware datetimes, and a canonical way to convert either of those types into an instant in time (requiring a timezone be attached for the naive type). Then the instants in time could be compared with the equals operator without confusion.
Those who say that don't know what it means. And given how much I have to deal with crap APIs that disregard those principles it's safe to ignore discussions leading elsewhere
I work on the Python API of a timeseries database and it’s very frustrating to have to accept timestamp objects without time zone because you don’t know what to do with them. People aren’t aware of this constraint, which leads to unexpected behavior.
As of now, we just reject any timestamp without an explicit timezone and try to explain to our users why utcnow() is bad. People are generally very surprised that this issue exists at all.
It's going to bite you at some point.
Only time differences do not have a timezone.
When Python continues to use the word "timestamp" for a time value with an unknown time zone, then that is the fault of Python, because the name is completely inappropriate.
I assume that whoever enabled a value "None" to be possible for the timezone did not really expect that anyone will ever use a timestamp with unknown timezone, but that in such cases there is a timezone known to the programmer, which is not stored with the timestamps, and that it is the responsibility of the programmer to restore the timezone whenever necessary.
On the other hand Unix (POSIX) time does not have leap-seconds, while UTC does not, so it is no longer "in" UTC.
What they're saying is correct in the most technical sense, but our industry has landed on timestamps not needing timezones. If they want to work in this industry, they need to accept that.
It's ok to be surprised by something like this, it's less ok to argue about it.
A Unix timestamp is a difference in time, not a time point. It's a number of seconds. It's a duration, not an instant, which is why it doesn't in and of itself need a time zone attached.
That being said, people do use timestamps to represent time points sometimes, and that does requires picking a particular timezone. That's fine too, but it's not the only use or meaning of a timestamp. It doesn't make timestamps without time zones meaningless, and doesn't mean they implicitly refer to the UTC epoch.
Durations and time differences are timezoneless numbers. When you add a duration to a time point (like the Unix epoch in your preferred timezone), you get another time point, in the same timezone.
The number of seconds since what, exactly?
It appears to be the number of seconds since Jan 1, 1970, at midnight... in some timezone.
You want to know how long something takes. You take a timestamp before and one after.
A timestamp since what? Well, by definition, the number of seconds since the unix epoch.
Which unix epoch, exactly? In what timezone?
Unspecified. You could have picked any timezone, and the result would not change, as it is not essential. Even if there is no timezone, not even one that you picked in secret, it still works.
Asking the question for durations is a subtle error. You only need a timezone for a time point. Timestamps are often immediately converted to time points, but that is neither necessary nor inherent.
If I'm comparing system time one year ago to system time now, that's a duration, and the timezones don't matter. But if I'm not comparing, just recording, then the point of comparison is the unix epoch, which is a point time.
Time zones don't matter to durations, only to points in time, okay. The unix epoch is a point in time. Saying it isn't a point in time, or that its timezone doesn't matter doesn't actually make it so.
"It's the very precise number of seconds since an event that took place 52 years ago, and it has a margin of error of 86,400 seconds," sounds crazy to me.
Comparing seconds-since-epoch "now" to the same from a year go: timezone doesn't matter so long as I'm consistent. Pick UTC, pick CDT, pick nothing, just do the same thing both times. This is the vast majority of usages of seconds-since-epoch, I think.
Comparing seconds-since-epoch "now" to epoch: in the abstract, timezone still doesn't matter, except, well, how do use the same timezone as epoch unless you know which timezone epoch was in? You could be off by 82,800 seconds! In practice, assuming that the epoch was Jan 1, 1970, midnight UTC seems to work for those rare cases, leading to the widespread (but technically incorrect) belief that epoch is in UTC.
The big problem with the python API, in my opinion, is that they conflate these two very different concepts in a single type.
> ts = 1571595618.0
This a timestamp, there is no TZ data required.
> x = datetime.utcfromtimestamp(ts)
Now for some reason the developer cares about UTC all of a sudden?
If the developer doesn't care about the TZ, they should use:
datetime.fromtimestamp(ts)
The main issue is that .utcfromtimestamp and .fromtimestamp both return naive datetimes, instead of aware ones with the timezone property set to UTC or the local timezone respectively.
Nobody argues for that. Especially since we’re comparing timestamps (floats), not datetimes (objects). Let’s go line by line :
> ts = 1571595618.0
This is 2019-10-20 18:20:18.000 UTC
> x = datetime.utcfromtimestamp(ts)
I would therefore expect x to be 2019-10-20 18:20:18.000 UTC
> x_ts = x.timestamp()
I would therefore expect x_ts to be 1571595618.0. Which is surprisingly not.
Nowhere the equality operator was involved in the surprise.
>>> from datetime import datetime
>>> ts = 1571595618.0
>>> x = datetime.utcfromtimestamp(ts)
>>> x_ts = x.timestamp()
>>> type(x_ts)
<class 'float'>
>>> type(x)
<class 'datetime.datetime'>I have, however, been using python so long that I’ve gotten used to the way things sort-of-work and developed such a healthy caution around it’s unintuitive built-in time zone handling.
In this case, I believe that it's actually the inverse from what I thought - utcfromtimestamp is doing... nothing; but is different by the fact that nothing is different from everything else. It looks like .fromtimestamp builds a "naive" datetime in local time. So, it assumes the timestamp is UTC and "helpfully" converts it to the local date and time. When you do .timestamp it inverses this process and gives you back the timestamp value in UTC by converting implicitly from your local time zone. Thus, if you do .utcfromtimestamp it doesn't do this conversion, and gives you a naive datetime "in UTC".
This is horrible and I fully support the recommendation never to use it.
Unfortunately, now that the default has been introduced, there seems no way to get to get out of this situation while preserving backwards-compatibility. The lesson here, perhaps, is to not make it easier to use that which should have been deprecated.
Thanks for your clear and to-the-point article!
This is why for anything that uses timezone, I use pendulum: https://pendulum.eustace.io/
It's compatible with the datetime api, but it has sane default, nice tools to convert between timezone, some cool date adjustment stuff, and can humanize time in several languages.
Basically, datetime has the same problem as text vs raw bytes in python 2.7, except it has never been fixed.
If only it returned something that contained the UTC offset in which is was created, then all of this would be much easier to deal with. Conversions to timestamps would just work, and comparisons would just work.
Note that I said "UTC offset", not "timezone". datetime doesn't deal in UTC offsets, it deals in timezones. And timezones are a zillion times more complicated, which is why no one want to deal with them on the critical path, and no one has ever proposed having now() return an object with .tzinfo set. Logical though it may be.
On Java, you can use the forbidden-apis build plugin (https://github.com/policeman-tools/forbidden-apis) to fail the build whenever a timezone or locale or charset is not specified explicitly (it forbids the methods from the Java API which use an implicit timezone/locale/charset). I don't know whether there's something similar for Python; it might be harder because Python is much more dynamic (though it might be possible to use monkeypatching to warn whenever the bad methods are used).
I'm not bashing anyone that's involved as I realize we're all human and prone to mistakes, I just wish we could all do better.
I should mention that the warnings are relatively recent, added in 3.8.
https://docs.python.org/3/library/datetime.html#datetime.dat...
It's surprising how poorly the situation with timezones is understood in the industry.
We have courses about compilers, databases, data structures, algorithms, cryptography. It's surprising we don't have courses about dates and time.
I suspect that the author of Jodatime (which effectively became the java.time API), who was notorious for being almost Linus Torvalds like in his attitudes and approaches, was probably once a very nice and kind individual. Until he started implementing a datetime library.
I mean, he nailed it, but at what cost? java.time.* is my second favourite datetime API, the first being Postgres'.
I'm actually surprised the source code of Jodatime isn't entirely in Zalgotext, technically it'd be valid Java code (I'm pretty sure, anyway).
https://docs.python.org/3/library/datetime.html#datetime.dat...
"If self is naive, it is presumed to represent time in the system timezone."
presumed is kind of a bad word in Python. I would never write an API today that "presumed" something that can be lots of other things and I would not allow any tz conversion on a naive datetime without requiring the existing known timezone be passed.
not following why the "this is the system timezone" must be hardcoded to be an invisible assumption, as opposed to something somewhat explicit. I mean this is literally not far off from a simple namechange of ".astimezone()".
A follow up from April 2022: https://blog.ganssle.io/articles/2022/04/naive-local-datetim...
Inconsistent through various python version, poorly documented and utterly confusing to use, as demonstrated by this article.
It's not hard to come up with an improved library design. But it is extremely hard to improve it while remaining compatible with the billions of lines of Python code out there.
Arrow makes the unfortunate conflation of Date, Time, and Datetime types, which causes all sorts of subtle errors when working with Dates that don't have times attached, and vice versa.
The best Datetime lib I've seen in any language is Chrono in Rust. Maybe there's a way to port it to Python via FFI.
I use the stdlib datetime.datetime with pytz, with a few in-house helper functions, and it's frankly kinda awful across the board but we deal with it.
Pendulum is really nice, and I want to use it as my daily driver, but I unfortunately found some timezone bugs [1] in v2 that make me hesitant to trust it for what we are doing in production (it screws up DST when doing `.utcoffset()`). If you aren't constantly dealing with data timezones from all over the world, it's probably fine, and it's the only one that the least surprising thing at each step.
Pytz has the whole .localize footgun.
The std datetime lib has horrible ergonomics and can't even cope with IANA timezones until the zoneinfo module was introduced in 3.9. Every single day it bugs me that the classes `datetime` and `time` are lowercase, and there is no consensus on datetime.datetime vs from datetime import datetime.
I tried using arrow and while I applaud them for moving away from the horrible datetime class, that makes it pretty much incompatible with any other time library, including things like pydantic and orjson.
I don't have much experience with dateutil.
1 - https://github.com/sdispater/pendulum/issues/655
Seems to have started in 2.x but is fixed in the 3.0alpha. You have to be specifically using utcoffset, which in general is probably not something you should do and use the timezone types to do conversions, but it was enough to give me pause
Mostly I'm not operating at the application / presentation layers.
This does cause friction when taking things at "interface value" (Sherry Turckle): just like everybody (I'm sure) I `grep -E '^2022-10-09'` or `grep -E '^Oct 09'` (looking at YOU, journalctl!). But I can also `perl -ne 'BEGIN { use Time::Local; $TIME = timelocal(0,0,5,7,8,122); } $t = (split /\s+/)[0]; $t > $TIME && $t < ($TIME + 60) && print;'`.
I have a little python script which does pretty much the same thing, with some attempts to recognize and fill in for rando date patterns just to make it a bit more universal.
I would counsel against storing human readable dates in databases and use unix time.
For python the important thing to understand is the difference between tz aware and naive datetime objects. But after you can do depending on your needs.
For example, if you ensure to convert your inputs to UTC and output UTC also, you can easily just work with utcnow by only dealing with UTC internally.
You would have to do that to store datetime in MySQL for example.
Imagine you have a software which deals with CSV files sent by clients. Clients are international corporations. What's the timezone the datetimes in the CSV files belongs to?
# what I used to do:
>>> datetime.utcnow().replace(tzinfo=timezone.utc)
# what I'll be doing from now on:
>>> datetime.now(timezone.utc)And just the naive/aware crap is almost like the unicode conversion previously. Ugh