What does an idle CPU do? (2014)
manybutfinite.com
manybutfinite.com
(Edit: there are still actually some circumstances in which hlt is called rather than mwait, but this is very CPU dependent and is based on vendor recommendations. Also, before mwait showed up, switching to deep CPU power saving states tended to involve performing magic IO port accesses rather than calling hlt, and yes the world was worse back then)
Is there a benefit of mwait vs IPI?
>[...] but sending interrupts from one CPU to another (to break it out of hlt) is kind of expensive. The replacement is an instruction called mwait, which will (like hlt) wake the CPU up on interrupts, but will also wake up the CPU on a write to a specific memory address.
It's probably Windows Store going bonkers with some insane backlog of periodic tasks, because you didn't open your laptop for a few months, that it absolutely at all costs must process for some reason, especially during a power outage when you're running on battery and need to use your laptop for once.
I wonder how much of this will be lost if MS starts enforcing policies via attestation on its services.
As far as I can tell, all the actually widespread successful computers after that have used DRAM too.
It wouldn’t call it “widespread successful”, but he Macintosh Portable also had SRAM (https://en.wikipedia.org/wiki/Macintosh_Portable)
Z80 CPUs had DRAM refresh logic built in. The first two clock cycles it would read the primary opcode byte, and the next two clock cycles it would do a refresh of the next row and then increment the row counter ("R"). That worked well because the minimum instruction took four clock cycles to complete. The timing was a bit surprising though -- it would seem to take four cycles to fetch the first byte of an instruction but only three clock cycles for each subsequent memory transaction. Unfortunately, the R counter was only 7b wide, which was fine for 16Kbit DRAMs, but became a problem for 64Kbit DRAMs which needed to refresh 256 rows, not just 128. Some Z80 based designs would decode the refresh cycles and either mux in their own 8b count value, or they ignored the Z80 and just stall the Z80 every so often to perform a refresh cycle.
https://youtu.be/Q8ph2OVqZeM?t=504
I find it really gives you some context. Memory is refreshed many, many times a frame, roughly once per scan line on a CRT.
Edit: Though it's unclear to me if it's the CPU actually doing the refreshing or if the CPU is just waiting until the refresh is complete.
There are some more optimizations that DRAM controllers can perform. As part of reading a certain address, you first need to open a bank. This is not a free operation, it can take a while (relatively speaking). After the first read in a bank, you can read other addresses inside the same bank for free, without having to open the bank again. This means it makes sense for a DRAM controller to process reads out of order: Instead of processing a command sequence like A1 B1 A2 (A/B are banks here) in order, it saves time to process them out of order, i.e. A1 A2 B1.
Disclaimer: It has been a while since I last looked into DRAM controllers, so I hope I didn't miss anything. But the general ideas should be correct.
Or in some real ancient case (ZX81), generate a video signal for the display (TV) - except in "fast" mode, where the screen would just show noise.
the CPU had to do memory refresh during busy times too!
The Z80 would do it automatically
Well, stone age at least if you run a bit more fresh distro than some CentOS or oldstable Debian.
So are typical modern desktop Linuxes tickless already or not yet?
From Ubuntu 20.04 generic kernel 5.13:
CONFIG_NO_HZ_COMMON=y
# CONFIG_HZ_PERIODIC is not set
CONFIG_NO_HZ_IDLE=y
# CONFIG_NO_HZ_FULL is not set
CONFIG_NO_HZ=y
# CONFIG_HZ_100 is not set
CONFIG_HZ_250=y
# CONFIG_HZ_300 is not set
# CONFIG_HZ_1000 is not set
CONFIG_HZ=250
Without looking up all details now I guess CONFIG_NO_HZ_IDLE=y means tickless when idle?That's of course not a very new kernel. Archlinux has 5.18, but don't have my machine accessible at this moment.
Linux my-host-name-here 2.6.18-128.el5 #1 SMP Wed Dec 17 11:41:38 EST 2008 x86_64 x86_64 x86_64 GNU/Linux
Nearly all the (medical instrument) hardware where I work (a major hospital) is on 2.x, mostly running RHEL 5.
Maybe not all that stuff connects to the internet, but I bet an ever increasing share does. So hospitals might have wose firewalls than others, so I would not surprised about such equipment being successfully attacked.
Of course old versions are unaffected because by newly introduced vulnerabilities. But old ones that have existed "forever" get detected. I guess it's not the kernal alone that is so old, but user space, too.
$ zgrep -E '^CONFIG\w*_HZ' /proc/config.gz
CONFIG_NO_HZ_COMMON=y
CONFIG_NO_HZ_IDLE=y
CONFIG_NO_HZ=y
CONFIG_HZ_300=y
CONFIG_HZ=300
It's been that way for years, I remember seeing that since at least 2015.LWN seems to confirm that feeling:
https://lwn.net/Articles/549580/
> In 3.10, the CONFIG_NO_HZ option has been replaced by a three-way choice:
> CONFIG_NO_HZ_IDLE (the default setting) …
https://en.wikipedia.org/wiki/Linux_kernel_version_history
3.10 30 June 2013
I'm not sure if that polling instead of interrupts is also responsible for their perceived lower input latency relative to other OSs of the time.
There was a post on the oldnewthing blog about this but I can’t immediately find it.
On the other hand, HLT was commonly found in the DOS software of the time, often for quick and approximate timing delays, and I don't remember any specific discussions about that causing problems, so I'm still not sure what the true problem is; my guess is that the combination of protected mode, which increases power consumption over real mode, and the sudden drop in current draw caused by a HLT, was enough to cause marginal power supply circuitry to fall out of regulation, similar to how stress-testing when searching for a stable overclock will sometimes crash the machine when it ends, not when it starts. On the more digital side, interactions with SMM might also be relevant: http://www.rcollins.org/ddj/Mar97/Mar97.html
I wonder how this works. How do you safely check for an instruction that "locks ups unrecoverably" on some hardware? Will Linux hang on boot on such systems? That is still better than hanging randomly later, but I can see why Microsoft didn't want to go for that approach.
EDIT:
The check in Linux literally just called hlt a bunch of times and hoped it didn't crash. There was no recovery in case it failed (note how there's no code path to print anything else than "ok"). Searching for "checking hlt instruction" you can actually find people asking on forums why their boot stoped at that point. The check has been removed in recent kernels, so I had to go back in history a bit to find it.
https://github.com/torvalds/linux/blob/521cb40b0c44418a4fd36...
Yet Windows 95 has a message in its hardware detection dialog which tells you to restart the computer if it stops responding for a long time while it's detecting hardware --- and I believe it writes to the disk its progress, so it'll remember where it was and skip it the next time.
One wonders why neither MS nor Linux did that for HLT. Linux has a no-hlt boot parameter to disable its use.
> If this fails, it means that any user program may lock the CPU hard. Too bad.
[0] - https://asahilinux.org/2021/03/progress-report-january-febru...
Just halt the CPU permanently until the next interrupt.
4ms interrupts are only needed when there are multiple active threads that need preempting.
> The solution here is to have a dynamic tick so that when the CPU is idle, the timer interrupt is either deactivated or reprogrammed to happen at a point where the kernel knows there will be work to do (for example, a process might have a timer expiring in 5 seconds, so we must not sleep past that). This is also called tickless mode.
One of the fun things about building a recreation of the PDP-11/70 is to watch the idle display when RSX-11M+ isn't doing anything.
The following is the source code of the CPU idle loop routine, with blinkenlights [0]. Note the funny quotations, and also note the use of self-modifying code - light direction reversal was implemented by overwriting the ROL/ROR instruction dynamically.
; "A source of innocent merriment!"
; - W.S. Gilbert, "Mikado"
; "Did nothing in particular, and did it very well"
; - W.S. Gilbert, "Iolanthe"
; "To be idle is the ultimate purpose of the busy"
; - Samuel Johnson, "The Idler"
; "I got plenty of nothin', and nothin's plenty fo' me!"
; - George and Ira Gershwin, "Porgy and Bess"
;-
30$:
.IF NE LIGH$T
DEC (PC)+ ;The RT-11 lights routine!
LITECT: .WORD 1
BNE ..NULJ ;Not too often
ADD #<512.>,LITECT ;Reset count, clear carry
40$: ROL 70$ ;Juggle the lights
BNE 50$ ;Not clear yet
COM 70$ ;Turn on lights, set Carry
50$: BCC 60$ ;Nothing fell off, keep moving
ADD #<100>,40$ ;Reverse direction
BIC #<200>,40$ ;ROL/ROR flip
60$: BIT #<LIGHT$>,CONFG2 ;Does CPU have a light register?
BEQ ..NULJ ;No
MOV (PC)+,@(PC)+ ;Put in lights (for 11/45)
70$: .WORD 0, SR
.ENDC ;NE LIGH$T
If the code doesn't detect a light register, it runs the following NOP loop instead. ..NULJ::
.IF EQ RTE$M
;;; WAIT ;Nothin' to do, so don't
NOP ;Nothin' to do, so don't
.IFF ;EQ RTE$M
NOP ;Let the host do the waiting
.ENDC ;EQ RTE$M
NOP ;Second pad instruction
.BR SCNALL ;Drop into ready job scan loop setup
[0] http://www.kpxx.ru/DEC/PDP-11/Software/OS/RT-11/05.07/05.07....I eventually tracked this down to Windows Defender sitting at 40% CPU! Disabling it entirely brought the performance of the laptop back to what it was pre-update.
I find this behaviour extremely disappointing. This is a perfectly serviceable laptop that would have likely been replaced by someone who was less technically inclined had it happened to them and they did not have the wherewithal to track down the offending piece of software.
Your original quote reminded me of a Cormac McCarthy quote which seems to apply to the Windows Defender engineers right now (I kid):
> But when God made man the devil was at his elbow. A creature that can do anything. Make a machine. And a machine to make the machine. And evil that can run itself a thousand years, no need to tend it.
I've found Windows 11 out of the box works just as well as Linux. I never really need to dualboot Windows (I don't play as many games these days), but W11 looks like a lot of hard work has gone into making the scheduler and other tools work well. It's just a shame their UI is absolute garbage.
With games that support a gamepad and accept input while not focused, I've started leaving a browser window focused so that it gets correctly scheduled with essentially no noticable impact on the game.
The transition required a number of compromises and personal workflow re-engineering, but once that hump is overcome it's been smooth sailing. Only ever have to re-visit my dual-boot setup in Windows for the occasional android unlock tool.
What Does an Idle CPU Do? - https://news.ycombinator.com/item?id=8529658 - Oct 2014 (55 comments)
Also every couple of seconds it would 'wash' an ECC memory page. So the machine wouldn't accumulate single-bit errors over weeks of uptime. The hardware didn't do that for you back then.
After putting that stuff in, our incidence of corrupt file systems went way down!
One example I found: https://einsteinathome.org/
I wonder if the same logic applies to CPUs.
Unused clock cycles is wasted CPU.
Case 1: we used this CPU for year with 1% average load. We paid $500 + 365 * 24h * 100W * 1% * $0.1 = $500 + $876 = $1 376.
Case 2: we used this CPU for year with 100% average load. We paid $500 + 365 * 24h * 100W * 100% * $0.1 = $500 + $87 600.
Case 3: like case 2, but we have cheap electricity with $0.01 for kWh. We will pay $500 + $8 760.
In any case CPU cost itself is tiny compared to energy price. So we can consider CPU cost is zero and all computations just require spending electricity.
So when some project asks to use your CPU at idle times, they're asking to use your electricity. And that's about it. You're donating electricity. It would be wrong to think that they're asking to use CPU which is not used anyway.
Now whether you want to donate electricity or not is another matter. But one should clearly understand what he does.
Edit: I forgot to divide by 1000, so those numbers are completely off, sorry.
100W * $0.10/kWh * 1 year = $87.66
You forgot to convert Wh to kWh. Google can evaluate the LHS in one step.
That aside, I agree. The price difference is even larger with 400W GPUs and data center bills which include cooling in the electricity cost. Initial cost matters very little vs efficiency.
Anyway I’ve never used OpenBSD, but you can probably mess with the CPU governor settings to make it idle more aggressively.
Makes more sense to disable unused modules imo.
Or for people who actually use their computers.
It's definitely the case if your CPU is in the cloud and you're being charged by the minute.
If I ever offended Nyan Cat, then I'm very sorry.
On that note, why aren't we calling it "bamboo code"?