systemd 100% cpu hang? – Proxmox Support Forum
forum.proxmox.com
forum.proxmox.com
Turns out something in proxmox (maybe a service?) doesn't understand daylight savings time (Dublin).
The only way to know was to google "proxmox systemd 100% cpu" and find above post.
Christ.
Edit: The fix, of course, is `ln -sf /usr/share/zoneinfo/Etc /etc/localtime`
Edit 2: Looks like that just unsets the timezone. It's too late for me to find the real fix, but you basically want to set the timezone to something else (like UTC).
A few decades pass and we now all have highly accurate "clocks" that we carry around in our pockets that receive GPS signals and sync their idea of time to what the satellite thinks. Sounds great, right?
Except, humans don't actually use atomic clocks as the basis for their timekeeping. We look outside and expect the time on the clock to mostly match with the position of the sun in the sky.
Once leap seconds became entangled in them, the whole thing basically a toxic pile of radioactive waste. Timezones are now being updated multiple times a year since the internet means we can update them at will. The whole thing is now a giant mess of toxic radioactive waste that can never be correct for any extended period of time.
We, the software industry, did this to ourselves. We built this, we have to live with it, and it's never going away.
It really isn’t that hard. What makes you pull your hair out is dealing with 3rd party code that handles it incorrectly. Do everything internal in UTC or always attach a timezone to dates, your choice.
When you cross the boundary either ingesting naive dates or displaying them read the system time zone again and convert.
I mean it’s kinda annoying but “naive” times is the time equivalent of storing arrays without their length. I can’t fathom why we even allow it.
Webapps should be timezone aware via the browser or user preferences, internal timezone should be whatever makes sense for logging et al.
[0]: https://gist.github.com/stephanGarland/b7cdd963e0ac53ea42f8e...
[1]: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1019716
I mean, yes, you can do that, but don't pretend that a program not handling unicode is the user's fault for typing unicode (s/unicode/timezones, s/typing/using).
EDIT: I'm just saying that you can't blame the user for using the features that come with the system.
For anyone curious:
> On the flip side, there’s one border where traveling north or south puts you ahead or behind just 15 minutes. That would be the border between India (UTC +5.5) and Nepal (UTC +5.75).
* https://qz.com/357697/time-zone-deviants-part-i-the-stranges...
It’ll smoke out all sorts of fun.
> you could say servers should be airgapped because security is hard
Many effectively do, and I wish more would do, as a default starting point. Maybe not physically air-gapped but firewall all incoming and outgoing then selectively allow packets through as specifically needed.
Translating to timezones: default to UTC and configure hosted software to display local timezones as needed. Only have the system in a different TZ if something won't work properly otherwise (including it can't be configured to display in something other than the system TZ) or is date sensitive and buggy in a way that it gets confused about yesterday/today/tomorrow when local midnight and UTC midnight differ.
Also if considering resources used by people in multiple timezones, you have to pick something and translate for others, you might as well pick UTC as some things assume it.
> Or that you should only type ASCII because Unicode is hard.
Again yes, by default, and always for things like hostnames, and adjust at needed or when you know Unicode is supported end-to-end (or that where it isn't, something just looks wrong and firefighting that is preferable to not sticking to ASCII in the first place). Though it helps from my point of view that the most obvious omission from ASCII in my locale is our currency symbol.
I say the same for spaces in filenames too: while you perhaps shouldn't have to avoid them, in practise you are setting yourself up for potential trouble later if you don't.
> EDIT: I'm just saying that you can't blame the user for using the features that come with the system.
I agree completely.
But, for instance, while I won't victim blame when someone accidentally leaves doors unlocked and has things stolen, I do make an effort to be damn sure my doors and windows are secure and recommend others do too. There is a difference between victim blaming and recommending preventative defaults because you know there is shit code it there.
All of the servers I run are located in Toronto, Canada: why do I care about UTC? What would having all of them being UTC get me?
A few jobs ago I took care of servers that were in: California, Ontario, Quebec, Switzerland, Singapore. UTC made sense there.
If I look at something now (at Mar 25 21:09) the logs would say "Mar 26 01:09" which is not "now" in my mind.
† Though I've been toying with the idea of getting a GMT watch "for fun". However I'll probably go with a chronograph—more useful feature on a day-to-day basis. Not as many choices if I want to get both complications though.
Maybe your logs mention the timezone / UTC offset? Otherwise it would seem possible to get confused.
But I do sometimes have to look at the logs on some random Tuesday, mid-morning, when someone says they can't log into our SSH bastion hosts (but it worked yesterday), I can see there's failed password attempts at 09:59, 10:03, and 10:04 (versus 13:59, 14:03, and 14:04 (i.e., UTC+4), which I would then have to do mental math to bring back to local time), and so it turns out they entered their password wrong a few of times so fail2ban blocked their IP
Wallclock time is a human construct created for human convenience. Having UTC when all my servers are in America/Toronto is not convenient for me.
Having UTC when all my servers are in America/Toronto is not convenient for me.
Until you have to debug something that straddles DST changes.I will perhaps consider doing it at that point, but in ~20 years of being a sysadmin I have never needed to, so… ¯\_(ツ)_/¯
Therefore I optimize for the common case: looking at logs when events do not straddle DST changes.
We could use both dates in logs, but that would eat an unreasonable amount of columns and confuse naive greps. We could also use non-standard time formats, e.g. <localtime8601>[<dst switch warning>] <unixtime>, but standard log handling tools aren’t built for it.
* https://en.wikipedia.org/wiki/Energy_Policy_Act_of_2005#Chan...
From a Unix perspective (Solaris, Linux, IRIX), it was fine. And Ontario has provisionally passed getting rid DST changes (sadly in favour of going to year-round DST), and IMHO it will be fine again if it is ever finalized:
* https://toronto.ctvnews.ca/ontario-passes-legislation-to-mak...
So long as you keep your TZ database up-to-date local time is likely to be not-so-problematic. Once you have that one server in a closet that didn't get updated, all bets are off.
It's only one less thing to worry about if you worry about it in the first place. I do not.
> I'd flip the argument on its head and posit that there's no compelling reason to use anything other than UTC.
There is a compelling, or at least useful, reason for me; from another reply I did:
Using UTC when all of my servers are in one timezones actually decreases clarity, as I live in America/Toronto with my daily existence, and which is what my laptop and (mechanical) wristwatch are set to. If I less(1) the logs I now have to do mental math as to "when" things happen because the hour number it says in the logs does not match my human reality.
If I look at something now (at Mar 25 21:09) the logs would say "Mar 26 01:09" which is not "now" in my mind.
[0]: https://www.ola.org/en/legislative-business/bills/parliament...
Local time is governed by local authorities; UTC, much less so.
I mean, clocks go out of sync, every few (days? Weeks?) I assume the ntp client has adjusted actual local time on your server; not just the time zone your server translates it to when you run “date” we are talking about here…
If it can handle that, it can handle a tz with dls time.
Usually being a few seconds ahead is compensated by running the clock slower for some time. Showing the clock to compensate a whole hour breaks many kinds of assumptions. But jumping the clock actually back an hour breaks even more assumptions that look entirely sane.
Better off without it.
Eliminates all the software bugs due to non-monotonically increasing time.
If you get bought by a company in the EU/UK or vice versa then it ultimately makes it easier to deal with time as you become part of a global company.
As far as practical use of UTC vs local time: I am a developer that happens to work with customers directly. It is difficult enough to get customer to provide timestamps for issues they report in local time, I cannot imagine they would be able to convert to UTC.
Love the local time!
The only question is, is the implementation user focused or developer prioritized
What exactly do you mean by this? As in, what part of your software stack is doing this. IIRC if you set your OS to be in UTC then anything at the OS level will only speak UTC. Databases and your own software will do what they are told to do but systemd etc. will all speak UTC only. Even running date in the shell should give you a UTC timestamp.
I'm not even sure what you mean by the OS speaking a certain timezone. Basic "time()" will return seconds from the epoch UTC regardless of your system timezone. Things that care about timezones will use relevant conversions.
Apart from this issue, I’ve never once see an error due to the tz presentation settings on a server, and I have spent 20 years running many thousands of servers at a time. shrug.
Let them have to adjust in grafana/kibana if it helps them sleep I guess..
> but I can’t imagine it arbitrarily converts anything away from UTC time if /etc/localtime is set to UTC.
It really depends on which part of systemd you mean. There's lots of code in there and quite a few bits deal with timezones explicitly. (Since it sets it for users to begin with)
But for basic display it sure uses TZ and localtime as set by the user - and it can be different than the system default. https://github.com/systemd/systemd/blob/cccc14c5a88f000931c6...
For a system to not boot in a particular time zone is a serious bug, but you should not blame the problem on the time zone, you should blame the problem on whatever is broken in some critical piece of code that probably should not be using local time.
I strongly prefer everything in UTC and changing time zone is an issue for only the vary top layer of the display/user interface/frontend whatever if absolutely needed - generally libraries designed for displaying have much more robust timezone handling than just random daemons and backend software.
There are terminal servers, for example, which are basically workstations. There are servers configured by other people for local time, and we may need to configure our own servers into its network.
There is server software out there that uses the host timezone. Heck, huge vendors do this regularly. This happens in log files, in schedulers, in notification emails, and so on. People want to see their log timestamps in their local time, they want to schedule in local time, and they want their emails to show events in their local time.
Yes, optimally, all times should always stored in UTC or UTC+offset formats, then displayed in local time only "at the end client". Generally, this is actually what happens. Windows, MacOS, UNIX, Linux, and Android all store the system clock as UTC and all system APIs use this.
So in some sense, we are all already using UTC on our servers. The time zone setting doesn't change the system clock to something other than UTC, it just tells user-mode software how the human users expect to see time formatted for display.
The problem is that server software especially is terrible at this. Just... bad. The worst offenders are text-based log files, which almost always get a formatted timestamp with an unspecified time zone. Could be UTC, could be the server's local time, could be Mars time, who knows?
In summary: don't blame system administrators for correctly configuring server settings. Blame the lazy software developers who can't be bothered to use ISO timestamps that specify time zones unambiguously, irrespective of the time zone setting.
Next thing you'll tell me to stop using space characters in file names...
Windows Registry Editor Version 5.00
[HKEY_LOCAL_MACHINE\SYSTEM\CurrentControlSet\Control\TimeZoneInformation] "RealTimeIsUniversal"=qword:00000001
We share this planet with people that don't use UTC. That's just the way it is. If some computers are set to time zones other than UTC, then non-UTC time zones must be correctly handled, even "servers".
This is just patronizing hand-wringing; well done you.
For global systems, a single time convention is the basis of co-ordination. So many people on this thread have seem to have a "pets not cattle" viewpoint, it is quite concerning.
For a terminal server it is not uncommon when different users connect from places with different time zones. It would make sense to have TZ in a user profile but keep UTC as a server system locale.
The correct fix would be timedatectl, but as you can guess, that's broken together with systemd.
> I've noticed systemd (PID=1) is running at 100% CPU utilization. That's not normal. I ran this command:
strace -f -p 1
> The process was reporting continuously attempt to access /etc/localtime file. That's also unusual.> I managed to get the output I was receiving on the console while tracking systemd process.
stat("/etc/localtime", {st_mode=S_IFREG|0644, st_size=3522, ...}) = 0
stat("/etc/localtime", {st_mode=S_IFREG|0644, st_size=3522, ...}) = 0
stat("/etc/localtime", {st_mode=S_IFREG|0644, st_size=3522, ...}) = 0
Maybe if it's endlessly failing to stat the target, any valid target might help.stat returns 0 on successful completion.
My guess it that it effectively unset it, in your case that probably does nothing(system time is utc and dublin is in utc) what was it set to before?
edit: all I can find is localtime(5) and it only says that /etc/localtime is the local time zone file.
I've advocated for over a decade as a systems engineer not to set system time to anything other than UTC. Not all languages or applications compensate for DST very well, and depending on your applications relationship with time it's a non-trivial problem to solve. What is far more trivial to do is to ship logs to a log server and let the UI of the log server translate time for the viewer of the logs. Almost every time I have argued with executives, managers, and SWEs without SE experience and lost something unexplainable and detrimental happens around DST.
Edit:
For the uninitiated I'll try to paint a clearer picture. All system time in Linux is tracked in seconds past Jan 1, 1970. The problem occurs in the translation between UTC and local time, which is handled by a library/package/module in your application. If your application is time sensitive and not looking for a literal time traveling event then things can get real weird, real fast. If there is a bug in that library/package/module things will also get real weird, real fast.
Setting the timezone should have no such effect.
> If your application is time sensitive and not looking for a literal time traveling event...
One of my least-favourite facts is that these statements are misleading. Since Linux uses UTC and not TAI, it no longer tracks an absolute measurement of seconds since the unix epoch as it takes “leap seconds” in to account. As well, because of leap-seconds you can absolutely have time-travelling events.
https://developers.redhat.com/blog/2015/06/01/five-different...
Unix, not linux. I don’t even think they can track TAI, but I definitely have never seen one which does, they all define unix time as 86400*days since epoch + seconds since midnight.
March 2022 https://bugs.launchpad.net/ubuntu/+source/systemd/+bug/19668...
March 2021 (Fedora) https://bugzilla.redhat.com/show_bug.cgi?id=1941335
systemd upstream fix https://github.com/systemd/systemd/pull/19075
When Debian switched to systemd in the default was when I began to seriously use OpenBSD and I've never regretted the decision.
Because the Unix process API is unreliable and unsafe ( http://catern.com/process.html ), managing processes is difficult to impossible in a general way under Unix systems including Linux. Linux gained some features to help with this, but for the longest time, PID 1 was the only process that could manage its child processes in a way that wasn't fraught with race conditions -- meaning systemd or something like it was the only way to have safe, effective process management because it tracked all that state in PID 1.
Blame Unix being broken, not systemd, for the problems systemd solves.
The only thing I dislike about systemd is that it is written in C, but otherwise it is a single purpose solution to the surprisingly complex problem of managing services during boot and running. I fail to see how state is relevant here — if it were written in Haskell it would also have state, almost by definition it has to handle external state. No Monad would help you there, there is just correct and not correct software.
You can't be saying this with a straight face. systemd project subsumes functionality of
- device hotplug
- user session manager
- cron
- logging
- network management
- time zone management
- time sync
- cgroup manager
- deleting /tmp files
- DNS resolver
- UEFI boot manager
and also dozen other things I've forgotten.
The Ubuntu bug is from 2022.
Curious that proxmox is hit in 2023. Proxmox must happen to have a particular systemd timer event setup by default that triggers this.
Backport hacks always come back to bite you in the end when your distro is running unsupported software, and it feels like I see something at least annually about Debian's "ensure everything is old so all bugs are predictable" ends up just causing pain.
* Eire after 1971: https://github.com/eggert/tz/blob/main/europe#L526
* Morocco after 2019: https://github.com/eggert/tz/blob/main/africa#L928
* Namibia between 1994 and 2017: https://github.com/eggert/tz/blob/main/africa#L1167
Looks like the negative DST has already caused problems in the past and applications that can't handle it (ICU & OpenJDK) have to build tzdata as per the rearguard section / ziguard.awk.
Example, California had an energy crisis soon after WWII, and changed when day light savings started/ended. There are time zone rules that track this history exactly, including a note about how clocks lost 6 minutes a day because PG&E changed AC from 60 to 59.5 Hz:
https://github.com/eggert/tz/blob/71faa2a55db2c9f21f4099b58c...
But.
It's 2023. Setting a non-UTC timezone shouldn't break your system (or applications). Bugs happen and this is one of them, but "set your system to UTC" isn't an appropriate response here.
It's perfectly appropriate to provide a workaround before the problem has actually been resolved.
With that the following reproducer is fixed:
TZ=Europe/Dublin faketime 2023-03-26 systemd-analyze calendar --iterations=5 'Sun *-*-\* 01:00:00'
[0]: https://git.proxmox.com/?p=systemd.git;a=commit;h=e7eb2b5864...
(just noticed that I wrote Ireland/Dublin instead of Europe/Dublin in the commit message)[1]: https://pve.proxmox.com/wiki/Package_Repositories#sysadmin_t...
For servers I manage, I absolutely refuse to use an OS with systemd. There have been too many unpredictable cases.
The DNS related issues seen (have they now been resolved?) where it takes over the DNS resolver functions and refuses to honor what you have put in e.g. /etc/resolv.conf ; a web search for "systemd breaks normal dns resolv.conf" will give you a lot of results ; unsure if this is now fully fixed.
Strange things with LXC based VMs moved/migrated between hosts, even if an offline move; I think in 1 case I just ended up creating a new VM from scratch; then rsync'ing everything over from the host node into the VM's directory.
e.g. Most DHCP clients have done that since... well, probably since DHCP was invented. See, for example:
https://wiki.debian.org/resolv.conf
If you want a static resolv.conf, but you're not using a completely static network config and have some kind of dynamic network management daemon, be that `dhclient`, or `NetworkManager`, or `systemd-resolved`, you've always needed to explicitly configure it to leave resolv.conf alone.
Not sure how this is a systemd issue?
In terms of problematic default services, I'd look at GNOME. systemd is just the medium.
Fix was to re-symlink `/etc/localtime` to proper time zone.
Didn't want or cared to dig into this deeper, but interesting to see that there are more issues about this.
For anyone interested, I manage around 20+ Proxmox machines with these Ansible roles:
https://github.com/liv-io/ansible-roles-debian
Example playbook:
https://github.com/liv-io/ansible-playbooks-example/blob/mas...
EDIT: Seems to be back.
Proxmox is based out of Vienna, Austria.
So a Timing glitch.
Because who wants to be bothered with Novice or Intermediate Tickets?
The main issue was always this one, having side functions like dealing with timezones breaking your pid 1. In such a case it is almost impossible to resolve the problem from the system itself you have to mount the drive with another OS or reinstall when possible.
When I shutdown I would like a shutdown sig to be sent to my database and then continue once it gets a successful response. No one wants that to just be killed because that would cause data loss.
When I shutdown I want my filesystem synced so things mid-write don't experience data loss.
The problem is badly written scripts that keeps sending a "one more min" response. But you could override that if you actually cared instead of just whining.
Just because you decided to kill the database while it was still in use, that's your fault.
If it is taking 5 mins for my db to respond with "success" then either fix the db or switch it.
That's not systemd's fault that the database is rubbish.
But as you wrote, If someone really cares, they would fix it.
PS: But what usually gets me about these 90s waits whenever I get them is that the message does not say which thing (unit, etc) is the issue. THAT is something worth criticising.
The concept of setting the time zone is a high level library thing already and all these systems, their standard time functions simply provide a time since some fixed epoch.
Ie see C time(), gettimeofday, clock_gettime, GetSystemTime. They’re all “UTC” based.
If a piece of server software is going out of its way to deal with local time APIs and blows up because of it - well it either legitimately needs to deal in local time, or it’s so poorly designed that simply keeping the system locale set to UTC isn’t somehow a magic fix. That’s just an adhoc assumption.