I installed thermald on my Lenovo T480 with Debian Bookworm and I get 20% better results in stress-ng. The fans are a bit louder now under high load and off under low load.
Without thermald:
$ stress-ng --matrix 0 -t 3m --metrics-brief
stress-ng: info: [3755113] setting to a 180 second (3 mins, 0.00 secs) run per stressor
stress-ng: info: [3755113] dispatching hogs: 8 matrix
stress-ng: info: [3755113] stressor bogo ops real time usr time sys time bogo ops/s bogo ops/s
stress-ng: info: [3755113] (secs) (secs) (secs) (real time) (usr+sys time)
stress-ng: info: [3755113] matrix 2278812 180.00 1437.43 0.27 12660.06 1585.04
With thermald: $ stress-ng --matrix 0 -t 3m --metrics-brief
stress-ng: info: [3755550] setting to a 180 second (3 mins, 0.00 secs) run per stressor
stress-ng: info: [3755550] dispatching hogs: 8 matrix
stress-ng: info: [3755550] stressor bogo ops real time usr time sys time bogo ops/s bogo ops/s
stress-ng: info: [3755550] (secs) (secs) (secs) (real time) (usr+sys time)
stress-ng: info: [3755550] matrix 2791272 180.00 1404.32 0.57 15507.06 1986.83
I just installed it using apt and did no extra configuration. My system was anyway configured for balanced power mode.
Why is thermald not installed on desktop installations by default?The thermald man page says this:
> In some newer platforms the auto creation of the config file is done by a companion tool "dptfxtract". This tool can be downloaded from "https://github.com/intel/dptfxtract". It is suggested as parts of the install process, run dptfxtract.
The dptfxtract gibthub project (https://github.com/intel/dptfxtract) says Intel discontinued the project.
> Thermald version 2.0 and later has in built parser for thermal tables. So this utility is not required. Make sure that thermald "--adaptive" option is used.
I also use this tool to bypass manufacturer limits in battery mode that are intended to make the system seem like the battery is not undersized for the CPU's power draw. Sometimes I'd rather have more CPU for less time.
I have never used a laptop for a decade, nor have I had the CPU fail. So perhaps faster performance and shorter lifespan is okay.
In all I've spent about 25% of the cost of the laptop over the years in repairs, but the result is that I'm able to keep using a device that's almost 12 years old instead of junking it. Saved quite a bit of money vs a new device, too.
Incidentally, it's interesting in regards to this issue that these old bulky laptops with beefy chips are actually quite good at avoiding thermal throttling. Despite having a traditional loud fan, I rarely even hear it come on.
In the show, trigger is a street sweep who has "kept the same broom" for 10 years, only periodically replacing the brush and occasionally replacing the handle.
The joke is of course, that it's no longer the same broom.
(Response crafted by anti-repair lobby gpt)
Dies of silicons are packaged, socketed, etc instead every part of silicon will get closer in a package and will handle everything.
Faster, cheaper, more energy efficient.
This is my PC right now. Sometimes it does a boot loop, where it can get to BIOS but not GRUB and I'm not sure why, as there are no error messages. Pulling the plug and plugging it back in helps.
I'm not sure whether it's the motherboard, or maybe a disk not wanting to read data or something. Guess I'll have to wait and see, eventually swapping out parts one by one once it fails entirely.
Just replaced my main desktop which was fx8350, but that computer is now on its 2nd life with a friend and shows no signs of stopping.
There was a time when 10 year old computer was absolute rubbish. These days, if you upgrade from spinning drive to ssd, 10 year old well made computers are just fine :)
Basically, in 3 decades of personal computing, I never had a CPU failure other than that one time I crushed atlon xp with the heatsink :O
Of course lots of CPU make it for longer, but 10 years is certainly not a design target.
Or if you put consummer CPUs in embedded equipements that are always (or at least often) under load, Intel will not care much if you have an high failure rate after a few years, even if you buy an high volume. They have dedicated SKUs with better expected durability.
That's it.
I still regularly use this machine even though it not longer is my main portable device. The only thing keeping me from doing so is its 32-bit CPU which now that 32-bit support has mostly gone the way of the dodo. I much prefer its 4:3 form factor 1600x1200 screen, keyboard and construction over the P50 I'm using in its stead.
The computers don't break on Windows, why Linux uses should take this overly conservative approach that limits the performance of their computer?
This check doesn’t work on Linux, so for safety reasons the former limit is enforced.
The whole point is to ignore them because they're horrible and hold back what these CPUs can really do. Fuck the manufacturers playing these stupid marketing games.
Intel warrants their CPUs at TjMax 24/7, they'll automatically throttle when they hit that limit, and disabling all this other throttling crap makes them run that way for full performance.
Mostly I mention it to say that there's not _no_ reason for thermal limits in modern setups...
There's zero reason for all this extra "thermal management" crap beyond corporate greed.
What people want is a box that does computing within spec and without melting, sterilizing their laps/hands, or otherwise introducing malbehaviors they have to work around.
People want stuff that doesn't or is very difficult to break. When it does break, they want it to be a straightforward repair.
but you're now generating more heat than the system is designed to dissipate
I believe that's more accurately rephrased as "now you have more performance than they wanted you to have", because that's how it's being used in practice.
FYI, it is painfully obvious that you're toeing some sort of company line here.
Edit: and getting flagged for pointing out the truth should itself be quite telling. ;-)
"An unthrottled CPU without any OS-level policy will generate enough heat that the case temperature may rise to levels that violate safety regulations" can certainly be interpreted as "now you have more performance than they wanted you to have", yes.
> FYI, it is painfully obvious that you're toeing some sort of company line here.
I work for a company that develops trucks. While thermal dissipation is something that we do need to worry about, I can promise that it is absolutely not at the scale involved here. I've also been pretty consistent in criticising Intel for refusing to document DPTF. Whose company line do you think I'm toeing here?
Despite literally zero evidence of that actually being a problem in practice?
I've also been pretty consistent in criticising Intel for refusing to document DPTF. Whose company line do you think I'm toeing here?
Not Intel, but there's plenty of evidence that DPTF just cripples performance; and why else would you jump in to an article about unlocking manufacturer-crippled products with FUD like "not the correct approach"? It's like posting "may cause processor damage" and advocating for stock speeds on every article about overclocking.
For the record, I've been overclocking (not at extreme or competitive levels) since the early 90s and haven't killed any hardware because of it. The system I'm posting this from has been on a 50% overclock for over a decade and it's still rock-solid stable.
Seems like there are quite a few other reasons why someone would "jump in", including being a relative expert. Maybe you could think of a few others yourself?
Sustained contact with metal at a temperature as low as 45C is sufficient to cause discomfort, and it's a surprisingly small rise above that before you can actually cause burns (https://citeseerx.ist.psu.edu/document?repid=rep1&type=pdf&d...). Skin temperature is a standard input to DPTF policies, and during my testing I was certainly able to get the skin temperature of my test laptop well above that level - DPTF would trigger throttling instead.
> why else would you jump in to an article about unlocking manufacturer-crippled products with FUD like "not the correct approach"?
DPTF defaults to a "safe" configuration, where "safe" is a euphemism for "almost unusably slow". If what you care about is your laptop not being almost unusably slow, it's reasonable to choose an approach that obtains that without introducing any additional thermal concerns. If what you care about is obtaining the absolute maximum performance possible from your hardware with no regard to any safety or longevity concerns, then yes, implementing the DPTF policy is probably not what you want to do. But recommending the latter without providing any information about the tradeoffs is irresponsible.
You choosing to not even look for evidence so you could make an argument in ignorance does not automatically mean it doesn't exist. This is literally the first google result for "laptop case temperature causing burns": https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4292129/
Also look at this fun condition: https://en.wikipedia.org/wiki/Erythema_ab_igne
> Temperatures between 43 and 47 °C can cause this skin condition; modern laptops can generate temperatures in this range. Indeed, laptops with powerful processors can reach temperatures of 50 °C and be associated with burns.
Some more results:
https://escholarship.org/uc/item/4n04r793
https://www.researchgate.net/publication/23164210_Thigh_Burn...
They added thermal throttling because most users only max out their CPU a tiny fraction of the time so the thermal inertia can compensate for inadequate cooling. It’s infrequent enough that most users could even disable it without noticing a difference. The problem is under the wrong conditions Aka inadequate cooling plus high ambient temperatures plus extreme workload = dead CPU.
And no this isn’t theoretical, one summer working without AC I cooked a CPU before they added thermal protection. It actually happened to several people in the same heatwave.
Only fairly recently with Turbo Boost did Intel combine both ideas. It’s actually a fairly clever optimization as single threaded workloads benefit quite a bit from the extra thermal overhead.
Silicon aging is worse at higher temperatures, people are going to kill their 7/10nm chips far quicker if they keep overclocking as they've previously done. But that'll be Big Silicon's fault too, no doubt.
Thermal limits like this are really about managing the manufacturers liabilities, and protecting the expected lifetine of the product.
Like I said, Intel warrants their CPUs to be operational at TjMax continuously, and that's the temperature they'll reach and stay at automatically if not given enough cooling. There's plenty of stories of people with computers whose heatsink has somehow detached or become so severely clogged as to be at that limit all the time, and the CPUs survive just fine.
This article isn't about that; it's about manufacturers artifically limiting performance beyond that to hit a marketing target like battery life or power consumption.
So, instead, you see the latter. This means that, yes, even if using official Intel DPTF drivers, systems will still occasionally thermally throttle. They may do so even if the CPU itself is not at a dangerous temperature, because the rest of the system may still be too hot and the CPU is the easiest tunable knob in terms of heat generation.
So, yeah, you can still obtain better performance by disabling DPTF and overriding power limits. And in the process you'll end up with a system that will become hotter than permitted by international standards, and you'll risk various types of failure ranging from inconvenient (the adhesive sticking things like rubber feet to the bottom of the machine tending to melt) to expensive (the unexpected levels of thermal expansion cycles causing components to fail earlier). It's fine that you're not worried about these things, but it's just straight up wrong to recommend that people override those limits without being clear about potential outcomes.
[1] https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CEL...
(Edit: replaced IEC 60950-1 with 62368-1, which is the up to date version, and linked to EU incorporation of that spec)
Since that isn't the case, however, many people get a laptop that isn't capable of performing the tasks it's advertised for. The excuse of "well, we didn't think you'd really need that much power" doesn't really cut it, when said power was strongly advertised on the box.
I don't think the average laptop user should need to concern themselves with acquiring the necessary knowledge to discern good laptop benchmarks.
They don't. The average laptop user doesn't have sustained workloads that keep the system running at 100% capacity for over an hour, so they don't need and shouldn't want benchmarks reflecting such a use case. They're genuinely better served by benchmarks reflecting workloads that come in shorter bursts.
Trusting Intel to provide accurate info on the actual performances of a chip feels too naive at this point.
It shouldn't be too difficult to correct on Linux. Why is Windows taking the manufacturer's limits into account while Linux basically ignore it?
As for the pros and cons of the `throttled` project, yes, this might not be the "officially desired" approach but I know several people who have been using to great success for years. The reality is, unfortunately, that particularly those of us who use Linux machines for work and order the most recent & powerful Thinkpads or Dells they can get their hands on, often realize later on (when the machine arrives) that the default settings handicap our machine to such a degree that we can barely work. Unfortunately, not everyone is a kernel developer, though, or knows how these things work, so quick fixes are often welcome, even if they limit the hardware's lifetime (after all, we'll buy a new device in a few years anyway).
What exacerbates this problem is that it's all very intransparent: The "official" way to solve these issues (which, as far as I understand you, is installing thermald?) is not really communicated anywhere, nor does thermald come preinstalled on any of the major distributions AFAIK[0]. What's worse, thermald often doesn't even solve the throttling issues without installing further patches. On top of that, BIOS updates by the manufacturers also seem to play a major role as manufacturers like Lenovo introduce different performance modes and things like "lap mode" etc. To be honest, to this day I haven't quite understood how these things interact and whose responsibility it is to fix things.
In my particular case, I have been using a Thinkpad X1 Carbon Gen9 (which Lenovo says "supports" Linux) and at some point, after installing numerous BIOS updates and working my way through hundreds of posts on the Lenovo forums, I just gave up: My machine still regularly throttles down to 800 Mhz per core and 16W under medium load until I hit the secret Fn + H key chore to tell the BIOS to switch back to high performance mode and set the thermal limit back to the maximum.
Do you happen to have a recommendation for me as to where I should start looking (again) for a solution? Does thermald fix these issues these days? (I know that when I last looked into it, it didn't.)
[0]: (EDIT) I take that back, it looks like thermald does come preinstalled on Fedora and Ubuntu these days. At least it's present on my new Ubuntu 22.04.2 installation. Unfortunately, that doesn't really help me since (and now I remember reading this last time I looked into thermald) according to the changelog[1] for v2.3:
> - thermald will not run on Lenovo platforms with lap mode sysfs entry
Great, so I still have nowhere to go from here it seems.
(Essentially we're back again to the situation I described: It's all very intransparent to me.)