The roughest era of AMD CPUs was the FX era. While it was comprable to its mid-range competition it was alos a sure fast way to burn down your house with its power draw.
Ryzen was a huge step forward in CPU design and architecture.
I see this era as Intel's FX era, if they have the right leadership in place they can turn the boat around and innovate.
Ahem. Bulldozer?
>Ryzen was a huge step forward in CPU design and architecture.
First gen Ryzen was kinda mediocre. Second gen(correction: meaning Zen 2 not Ryzen 2000 which was still Zen 1) was where the performance came.
Also let's not ignore how they screwed consumers like me by dropping SW support for Vega in 2023 while still selling laptops with Vega powered APUs on the shelves all the way till present day in 2024, or having a naming scheme that's intentionally confusing to mislead consumers where you don't know if that Ryzen 7000 laptop APU has Zen2, Zen3, Zen3+ or Zen4 CPU cores, if it's 4nm, 5nm, 6nm or 7nm or if it's running RDNA2, RDNA3 or the now obsolete Vega in a modern system.[1] Maddening.
Despite that I'm a returning AMD customer to avoid Intel, but I'm having my own issues now with their iGPU drivers making me regret not going Intel this time around. The grass isn't always greener across the fence, just different issues.
I get it, you're an AMD fan, but let's be objective and not ignore their stinkers and anti-consumer practices which they had plenty of and only played nice for a while to get sympathy because they were the struggling underdog, but didn't hesitate to milk and deceive consumers the moment they got back on top like any other for profit company with a moment of market dominance.
My point being, don't get attached or loyal to any large company, since you're just a dollar sign for all of them. Be an informed consumer and make purchasing decisions on objective current factors, not blind brand loyalty from the distant past.
[1] https://www.pcworld.com/article/1445760/amds-mobile-ryzen-na...
https://www.digitaltrends.com/computing/amd-confusing-naming...
>Ahem. Bulldozer?
Bulldozer is the same as FX.
>AMD FX is a series of high-end AMD microprocessors for personal computers which debuted in 2011, claimed as AMD's first native 8-core desktop processor.[1] The line was introduced with the Bulldozer microarchitecture at launch (codenamed "Zambezi"), and was then succeeded by its derivative Piledriver in 2012 (codenamed "Vishera").
Ha, well that's wrong. This is the first time I find a mistake or more accurately, a contradiction in Wikipedia.
AMD's first FX CPU (the FX-51) came out in 2003 as a premium Athlon 64 that was an expensive power hungry beast, which is the one I assume the GP was talking about. Here, also from Wikipedia:
"The Athlon 64 FX is positioned as a hardware enthusiast product, marketed by AMD especially toward gamers. Unlike the standard Athlon 64, all of the Athlon 64 FX processors have their multipliers completely unlocked."
[1] https://en.wikipedia.org/wiki/File:AMD_Athlon64_FX.jpg
[2] https://commons.wikimedia.org/wiki/File:AMD_FX_CPU_New_logo....
Or the early Athlons that would literally burn down without cooling? https://www.youtube.com/watch?v=YYQSHXNFvUk
Same as Pentium 3 of same era, thermal throttling on socket A was supposed to be implemented by Motherboard vendors using chip integrated thermal diode. Pentium 3 would burn same way if put on a motherboard with non working thermal cutout.
TBirds and spitfire didn't have die sensor, that was first on Palomino/Morgan.
That said I've seen P4s die due to cooler failure so it was still dumb.
"Just like AMD's mobile Athlon4 processors, AthlonMP is based on AMD's new 'Palomino'-core, which will also be used in the upcoming AthlonXP processor. This core comes equipped with a thermal diode that is required for Mobile Athlon4's clock throttling abilities. Unfortunately Palomino is still lacking a proper on-die thermal protection logic. A motherboard that doesn't read the thermal diode is unable to protect the new Athlon processor from a heat death. We used a specific Palomino motherboard, Siemens' D1289 with VIA's KT266 chipset."
Intel suggested Siemens D1289 board for the test, board didnt have thermal protection. Intel suggested (or even delivered) Pentium III motherboard with working thermal protection.
Are you sure? I just looked at Ryzen 5 1600 vs 2600 benchmarks and the difference is around 5%. And I also remember the hype when the first generation was released. I think Ryzen gen 1 was by far the largest step.
https://www.cpubenchmark.net/high_end_cpus.html
Yes, it is deceptive and annoying shenanigans for retail products =3
We are better for Intel and AMD to coexist. But my gamble is on AMD because I've always liked the compatibility of the hardware with variety of technology. You can easily get server grade interfacing on consumer grade parts. For the longest time that wasn't true for Intel. When AMD pulls an Intel I'll be full Intel. There are huge wins for Intel getting new fabs built in the states, because it means a lot for security and development.
I've used both the Ryzen 3 1200 and 7 1700 and all of them seemed fine for their time and price.
Honestly, I had the 1700 in my main PC up until last year, it was still very much okay for most things I might want to actually do, except no ReBAR support pushed me towards a Ryzen 5 4500 (got it for a low price, otherwise slightly better than the 1700 in performance, still good for my needs; runs noticeably hotter though, even without a big OC).
I guess things are quite different for enthusiasts and power users, but their needs probably don't affect what would be considered bad/mediocre/good for the general population.
https://www.techpowerup.com/276125/asus-enables-resizable-ba...
Given that I got an Intel Arc A580 for myself, this was pretty important! Quite bad that it wasn't officially supported if there are no hardware issues and I would have liked to just keep using the 1700 for a few more years, but opted for just buying a new CPU so my old one would be a reasonable backup, path of least resistance in this case.
Would also like to try out the recent Intel CPUs (though surely not the variety that seems to have stability issues), but that's not in the cards for now because most of my PCs and homelab all use AM4, on which I'll stay for the foreseeable future.
We had to use intel cpu/gpu + CUDA gpu simply because of compatibility requirements (heavy media codecs and ML workloads.)
Lets be honest, AMD technically has had a better product for decades if you exclude the power consumption metric. ARM64 v8 is also good, if and only if you don't need advanced gpu features.
The Ryzen chips definitely are respectable in passmarks benchmark value stats rankings. =)
Quiet machines are great, especially when you have to sit next to one for 9 hours a day. =3
And single core performance.
And some other stuff which obviously didn’t matter during the period in question but suddenly became very important when AMD surpassed Intel in that regard…
It looks like some newer motherboards are finally fixing the issue. Note that AM5 is now 2 years old.
Intel has its own issues, Gigabyte told me to pound sand when asking to unlock the bios on my own equipment to disable IME.
There is no greener grass on the fence line... just a different set of issues =3
Since he's taking about iGPU issues, he most likely has a laptop APU, so no RAM to reseat. I'm also having similar issues on my Ryzen 7000 laptop. Kinda regret upgrading from the Ryzen 5000 laptop which AMD obsoleted just 2 years after I bought it, as at least that had no issues. Hopefully new drivers in the future will fix stability but you never know.
What I do know, is that this will most likely be my last AMD machine if Intel shows improvement to match AMD, since their Linux driver support is just top notch.
I'd try a slower cheap set of lower-bandwidth/higher-latency ram sticks to see if it stops glitching up. If you are using low latency sticks (iGPU means this is usually recommended), than dropping the performance a bit may stabilize your specific equipment.
Of course, I'm not that smart... so YMMV... =3
sudo apt-get install cpu-x
sudo cpu-x
We may still be able to use this information to compare with other users glitches to see if there is some underlying similarity.
Unfortunately, if it is a thermal stress/warping on the PCB cracking open RAM BGA balls on chips or shifting traces... One won't really be able to completely identify the intermittent issue.
We were actually looking at buying a similar economy model earlier this year (ended up with a few classic Lenovo models instead)... so please be verbose with the make/model to help future searchers =3
Please dump the problematic cpu/ram chip model numbers to help other users. These chip manufacturer numbers is not really personally identifiable information, as they are shared between hundreds of thousands of products.
The classic cpu-z for Windows users is here if you don't run *nix:
https://www.cpuid.com/softwares/cpu-z.html
Best regards, =3
That snarkyness is uncalled for. I repasted the laptop, ran benchmarks and checked the temperature sensors plus used my FLIR. It's no thermal issues. It's just AMD iGPU driver buggyness.
Processors Information
-------------------------------------------------------------------------
Socket 1 ID = 0
Number of cores 8 (max 8)
Number of threads 16 (max 16)
Secondary bus # 0
Number of CCDs 1
Manufacturer AuthenticAMD
Name AMD Ryzen 7 7840HS
Codename Phoenix
Specification AMD Ryzen 7 7840HS with Radeon 780M Graphics
Package Socket FP7
CPUID F.4.1
Extended CPUID 19.74
Core Stepping PHX-A1
Technology 4 nm
TDP Limit 54.0 Watts
Tjmax 90.0 °C
Core Speed 2761.5 MHz
Multiplier x Bus Speed 27.71 x 99.6 MHz
Base frequency (cores) 99.6 MHz
Base frequency (mem.) 99.6 MHz
Instructions sets MMX (+), SSE, SSE2, SSE3, SSSE3, SSE4.1, SSE4.2, SSE4A, x86-64, AES, AVX, AVX2, AVX512 (DQ, BW, VL, CD, IFMA, VBMI, VBMI2, VNNI, BITALG, VPOPCNTDQ, BF16), FMA3, SHA
Microcode Revision 0xA704104
L1 Data cache 8 x 32 KB (8-way, 64-byte line)
L1 Instruction cache 8 x 32 KB (8-way, 64-byte line)
L2 cache 8 x 1024 KB (8-way, 64-byte line)
L3 cache 16 MB (16-way, 64-byte line)
Preferred cores 2 (#1, #3)
Max CPUID level 0000000Dh
Max CPUID ext. level 80000026h
FID/VID Control yes
# of P-States 3
P-State FID 0x898 - VID 0xBF (38.00x - 1.194 V)
P-State FID 0x858 - VID 0xAB (22.00x - 1.069 V)
P-State FID 0xA50 - VID 0x97 (16.00x - 0.944 V)
PStateReg 0x80000000-0x49AFC898
PStateReg 0x80000000-0x45AAC858
PStateReg 0x80000000-0x4425CA50
PStateReg 0x00000000-0x00000000
PStateReg 0x00000000-0x00000000
PStateReg 0x00000000-0x00000000
PStateReg 0x00000000-0x00000000
PStateReg 0x00000000-0x00000000
Package Type 0x4
Model 00
String 1 0x0
String 2 0x0
Page 0x0
Power Unit 0.0
SMU Version 76.73.00
TDP/TJMAX 0x36005A
TCTL Offset 0x0
PMTV 004C0008
Package Power Tracking (PPT) 54.0 W (current)
Package Power Limit #1 (long) 35.0 W
Package Power Limit #2 (short) 25.0 W
DMI Physical Memory Array
location Motherboard
usage System Memory
correction None
max capacity 64 GB
max# of devices 4
DMI Memory Device
designation DIMM 0
format Row of chips
type LPDDR5
total width 32 bits
data width 32 bits
size 8 GB
speed 6400 MHz
manufacturer Micron Technology
part number MT62F2G32D4DS-026 WT
serial number 00000000
voltage 0.500000
manufacturer id 0x2C80
product id 0x0
Display adapter 0 (primary)
ID 0x2180003
Name AMD RadeonT 780M
Board Manufacturer Lenovo
Codename Phoenix
Cores 768
ROP Units 16
Technology 4 nm
Memory size 1024 MB
Current Link Width x16
Current Link Speed 16.0 GT/s
PCI device bus 99 (0x63), device 0 (0x0), function 0 (0x0)
Vendor ID 0x1002 (0x17AA)
Model ID 0x15BF (0x3819)
Revision ID 0xC7
Root device bus 0 (0x0), device 8 (0x8), function 1 (0x1)
Performance Level Current
Core clock 800.0 MHz
Shader clock 400.0 MHz
Memory clock 800.0 MHz
Driver version 32.0.11021.1011
WDDM Model 3.1Note BGA chips were never initially intended to be larger than 20mm wide, and can still put enormous shear forces on the contact bonds as the solder solidifies post re-flow (and the bimetallic cantilever PCBs form start to pop-back.) A certain percentage of products will thus fail when they warm up as the PCB will locally heat/warp the area again, and foobar a few random connections in the process. A low-heat paint-stripper heat-blower might be able to replicate the crash to eliminate this theory, and you might be able to RMA the board/chips if you are still under warrantee.
Could indeed also just be software as some suspect, but that is a lot harder to find in kernel drivers.
It is hard to read peoples emotions online, but do assume if someone is reaching out to help they probably think you are worth respecting too.
Thanks for posting data other users may find useful, and have a wonderful day. =3
Might be interesting if the bug is codec dependent, as YT can stress some browsers configs (software codecs etc.)
Does it do this with the windows 11 driver set as well? =)
Increasing the VRAM size (UMA size) to 4 GB fixed the frequent driver timeouts for me.
Reverting to older driver (driver cleaner -> driver v23.11.1) fixed the memory leak. This memory leak is weird since PoolMon doesn't show anything unusual. Nothing shows as using too much memory anywhere, except committed memory size grows to over 100GB after few days of uptime and RamMap shows a large amount of unused-active memory.
Should get interesting soon =)
https://developer.nvidia.com/blog/nvidia-transitions-fully-t...
Yet to personally try it out, but this should eventually enable better integration with the library ecosystems. =3
sudo apt-get install cpu-x
sudo cpu-x
I think comparing your specifications may help other users narrow down if a manufacturing or software defect is present.
Thanks in advance =3
I also increased UMA to 4 GB, it reduced the crash frequency, but it still happens.
The discrete NVIDIA GPU I use at the same time is fine.
If there is enough data here, we may be able to see a common key detail emerge. i.e. if the anecdotal problem(s) remain overtly random, than a solution from the community or OEM may prove impossible.
Thanks in advance, =3
Stability has been fine since. The bug report has since been closed but I haven't tested in a while to see if disabling PSR is still needed or if the issue has actually been fixed.
I haven't seen significant stability issues on Windows, although I don't use it much on the AMD device.
Your tip may help some folks in the future. =)
With PSR in the mix, is the system really hanging or is it just failing to update the screen somehow? I.e. can you tell the difference with logs or a remote connection or configure and use an unprompted shutdown via the power button?
I can't remember the details of it. It effectively hung in the sense that I couldn't get the system into a usable state again locally without rebooting. I'm not sure if the system responded to the power button or not, or whether there was useful log output.
I didn't bother trying with a remote connection since the hang was frequent enough that it wouldn't have been of any use as a workaround anyway. I'd guess switching to another virtual console probably didn't work because I'd probably remember it if it did.
I can try re-enabling PSR and see if the problem is still there if you're interested.
So, I don't really know if the system was fully hanging, or if the display was just unable to update any more, but it was likely exactly the same that happened to other people with Parade TCONs in that bug discussion.
I also did Prime95 CPU stress testing a few times without issues.
All issues seem to be related to either BIOS or drivers.
What is your current ram chip model, maker, and configuration on your machine?
sudo apt-get install cpu-x
sudo cpu-x
Cheers, =3
Running RAM at default speeds (4800MHz) or using XMP profile 5600MHz C36 doesn't affect these issues (they are no more or less frequent).
EDIT: XMP profile, not EXPO.
CPU Ryzen 9 7950X. Family: F (ext.: 19), Model: 1 (ext.: 61), Stepping: 2, Revision: RPL-B2.
iGPU: Raphael, revision: C1.
MB: ASUS TUF Gaming X670E-PLUS WiFi. Rev 1.xx. Southbridge rev.: 51.
Swapped board out assuming it was that. Same problem. Turned out to be the CPU which was a pain in the ass getting a warranty replacement for.
Ended up buying a new open box Intel 12400 Lenovo lump off eBay and using that.
Which board was it?
But at that point I was using the Lenovo box. So I just sold all the crap on eBay for the next victim.
For my Ryzen 5000 series build (a while ago now) I went with an ASRock board for ECC support, and also ECC ram.
It's been mostly flawless, though as I'm undervolting the ram it does let me know about an ECC corrected error once every 6-9 months or so. ;)
And there's one thing you can NEVER trust and that is objectivity from gamers when looking at failure and reliability statistics. It's one huge cargo cult.
Notably my kids both have Ryzen 5600G + MSI B550 boards with no problems.
My motherboard is an ASUS ROG Strix variety with 4x32GB ECC RAM and the Ryzen 9 5950X works just fine.
Oddly, this Intel build somewhat restored my faith in humans to build hardware and software as the thing seems to work quite well.
Couple that with the underwhelming software support for AI/ML on their own hardware for about a year after CPU and GPU launch, and I wish I'd just stuck to AMD.
I don't think either are perfect, but it's the devil you know, and I've grown to trust that even when AMD cocks something up, they'll listen to customers, coordinate engineering efforts with OEMs, and handle it. Intel are either too high and mighty or don't empower their engineers to treat partners like partners without layers of management getting involved to be able to do something similar.
What support did AMD have?
No particular straw broke the camel's back, they just haven't managed to justify their price premium in a very long time.
Didn’t AMD also have issues with different cache types/sizes on dual CCD chips? Meaning that it basically didn’t make any sense to buy anything more expensive than the 7800x3d if you only care about gaming.
This fiasco has me convinced that I will not be building an Intel system again, and I haven't even (yet) had any problems with either of the Z790 i7-14700K systems I put together in March.
I go with AMD because they make the best desktop CPUs right now. When Intel gets their shit together (hopefully) and the pendulum swings back I'll go with them.
These are all corporations at the end of the day. They get better and worse overtime. They certainly never stay the same forever and deserve any kind of loyalty to their brand.
The leading theory at this point is not really voltage or current related but actually a defect in the coatings that is allowing oxidation in the vias. Manufacturing defect.
https://www.youtube.com/watch?v=gTeubeCIwRw
It affects even 35W SKUs like 13700t, so it’s really not the snide/trite “too much power” thing. Like bro zen boosts to 6 GHz now too and nobody says a word. And believe it or not, if you look at the power consumption, both of them are probably fairly comparable in core power - both brands are consistently boosting to around 25-30W single-thread under load. AMD's highest voltages will occur during these same single-core boost loads, these are the ones of concern at this point - if it is just voltage that is killing these 35W chips, well, AMD is playing in the exact same voltage/current domains.
Furthermore, if it was power it wouldn’t be a problem that is limited to 10-25% of the silicon, it would be all of them.
There was a specific problem with partners implementing eTVB wrong, and that was rectified. The remaining problem is actually pretty complex and potentially there are multiple overlapping issues.
It just has become this lightning rod for people who are generally dissatisfied with Intel, and people are lumping their random "it doesn't keep up with X3D efficiency!" complaints into one big bucket. But like, Intel actually isn't all that far off the non-x3d skus in terms of efficiency, especially in non-AVX workloads. "140W instead of 115W for gaming" is pretty typical of the results, and that's not "burn my processor out" level bad. 13900K has always been silly, but 13700K is like... fine?
https://tpucdn.com/review/intel-core-i7-13700k/images/power-...
https://tpucdn.com/review/intel-core-i7-13700k/images/effici...
https://tpucdn.com/review/intel-core-i7-13700k/images/power-...
https://old.reddit.com/r/hardware/comments/yehe1s/intel_rapt...
(granted this may be launch BIOS, and it sounds like part of the problem is that partners have been tinkering over time and it's gotten worse and worse... I'm dubious these numbers are the same numbers as you'd get today, but in fact they are pretty consistent across a whole bunch of reviewers, ctrl-f "CPU consumption" and the gaming and non-AVX power numbers are in broadly unconcerning ranges, 57-170W is broadly speaking fine.)
Again, even if there is a power/current issue, at the end of the day it's going to have a specific cause and diagnosis attached to it - like AMD's issues were VSOC being too high. Saying "too much power" is like writing "died of a broken heart" on a death certificate, maybe that's true but that's not really a medical diagnosis. Some voltage on some rail is damaging some thing, or some thermal limit is being exceeded unintentionally, and that is causing electromigration, or something.
You might as well just come out and say it: intel's hubris displeased the gods, they tempted fate and this is their divine punishment. That's really what people are trying to say. Right? Don't dress it up in un-deserved technical window-dressing.
https://www.youtube.com/watch?v=gTeubeCIwRw
No need to make shit up, things are bad enough already.