Microsoft's latest Windows 11 24H2 update breaks SSDs/HDDs, may corrupt data
neowin.net
neowin.net
But, if the update hasn't been completed and the hammering continues, then I think chances are it might happen again.
My guess is that the firmware of the ssd needs to be updated first to account for the HMD bug?
Oh well...
(I’m still wondering if the successor, SN7100, is affected, as it’s not mentioned in that thread either way. The upmarket DRAMful SN850X seems not to be affected.)
Or not? The bug here looks different from that one, which the same outlet reported on[1] earlier.
[1] https://www.neowin.net/news/wd-ssds-still-block-windows-11-2...
This could just as well as a bug, be the Windows kernel now being better optimized to take advantage of fast storage media. Perhaps the software cache, scheduling, and what not became a bottleneck, and they fixed that. Unless it's part of the NVMe standard or what have you that the OS IO scheduler should not attempt writing "too fast", it's hard to blame it in good faith.
https://support-en.wd.com/app/answers/detailweb/a_id/20968/~...
My point is various drives are bad and causes issues not only with Windows, be it with some or all firmwares from the manufacturer. Among them the only specific drive mentioned in the article. If WD advertises this mode then it ought to work.
If you push an update that bricks your customers machines you have failed and you have no QA clearly.
The best they can do to save their ass is say "actually you know what these SSD and HDD are all incompatible with W11" and start a new hardware certification program.
"enterprise" doesn't imply reliability, eg. https://www.engadget.com/2020-03-25-hpe-ssd-bricked-firmware...
The article is light in the details of which are those enterprise grade SSDs
This feels like an extrapolation by the article author based on the fact Phison microcontrollers are also used in enterprise models.
They were talking about Microsoft. QA ? Testing ? That's why they have users. /s
and i see the same error under windows update on my backup pc now:
2025-08 Cumulative Update for Windows 11 Version 24H2 for x64-based Systems (KB5063878) (26100.4946) Failed to install on 17/08/2025 - 0x80073712
articles says August but I had this update stuck on failed for 2 months. bricked last pc now waiting for it to brick my current one.
how long does it take MS to fix something..
Flakey WiFi at the start? Just flat out broken now.
Weird stuff like nearly impossible to find/configure network connections? Worse over time.
It’s like whatever criteria the MS PM’s are using is prioritizing ‘engagement’ (aka how frustrated and angry someone gets at the OS) over anything else.
I'm finding it super concerning that it seems like Microsoft isn't even pushing driver updates properly nowadays. Everything just sits as an 'optional update' hidden away buried under a couple more menus, and the latest drivers are jumbled in the list, and you can't tell which driver is for which device.
And that's how I discovered my WiFi was breaking out of the box because Microsoft had newer drivers they didn't push down...
I wouldn't be surprised if the related dev/QA teams were part of recent layoffs/offshoring.
Sure, we can always contract a divinacy specialist to guess which hardware may fail next, but I thought the (paid) Windows testers had that job.
Windows is therefore going to be used by a wide variety of systems, including those with low grade hardware parts. So it either has to be part of windows quality control to test these low end systems or it should make it very clear about the system requirements.
Either way, the problem was patched, which makes it a problem on windows side.
If you have an old car, you shouldn't press the gas pedal. /s
Linux still has tons of driver issues and is a massive time vampire with its "type magic spells into a terminal for hours to get something done" GUI.
Mac is way more expensive and doesn't run a ton of apps available on Windows.
cough Linux, BSDs, Solaris clones cough
If this was a widespread problem I'd agree it should've been caught during testing and mitigated, but it doesn't sound like that's the case ("We looked around and could not find other reports resembling such situations"). As far as I'm aware it's just one Twitter thread reporting this, not all/many low end systems. Only so much Windows can do if it's just a faulty batch of SSDs, or even an incorrectly installed part in the user's setup.
> Either way, the problem was patched, which makes it a problem on windows side.
Has it been? I can't see Microsoft even verifying that the issue exists yet.
Sure, Microsoft can issue a SW workaround to mitigate this issue as they did for Intel Skylake CPU issues, but that doesn't change the fact that the core issue is faulty designed HW.
Why don't SSD manufacturers test their shit before shoveling it out the door to consumers?
Most of them do, not that it stops people from buying Temu PCs anyways.
Microsoft should be more careful, but also such blatantly garbage hardware never should’ve shipped. So I’d say hold MS responsible, but also push back against manufactured e-waste like that.
What does "more careful" even mean? For instance, if your crappy SSD has some bug that causes data corruption if there are more than 16 writes in flight at a time, and microsoft pushes an update that causes it to do writes more aggressively, should it really be microsoft's fault that your crappy SSD broke? Should microsoft keep throttle windows IO at 2000s HDD throguhput on the off chance that some IDE drive bricks itself when it encounters SSD level workloads?
see sibling comment for something similar happening on zfs. linux is also filled with workarounds for crappy SSDs, so "SSDs breaking due to software" isn't exactly exclusively a windows problem. It just gets more attention because of the bigger user base.
>This suggests that this kind of strain on the hardware is unnecessary given the task, which then paints a picture of negligent sloppiness in engineering and QA culture at MS.
"straining hardware" shouldn't be a thing for PC components. If it can't handle a given workload, it should throttle/queue the requests, not brick itself. I can grant some leeway for long term wear (eg. heat damage or electromigration), but that's clearly not what's happening here. Can you imagine a network card that bricks itself if you send too many packets, and you have to baby it with how many packets you send in any given amount of time?
I don't think Windows has the most hardware compatibility issues.
It seems increasingly these problems (including the NVIDIA one) are only cropping up on 24H2 machines which is why I've decided to stick with 23H2 until it's out of its support window. Unfortunately that's coming up pretty soon.
[1]https://www.neowin.net/news/nvidia-fixes-windows-11-24h2-dri...
SSD makers should make better products.
Microsoft should do enough testing so it doesn't push updates that trigger data-losing hardware flaws.
Damn, i hate to defend Mickey$oft here but, if your hardware breaks under work loads, better buy a better one. Windows is not able to use all cores on a CPU properly, so if you have a workload problem, it is your hardware.
Linux is beautifully optimised for performance. Linux is more likely to write the same data more quickly than Windows, when an application has a large amount of data to write. So if the problem is the SSD fails when a large amount of data is written quickly but within specs, then it's likely to fail on Linux too.
Unless it's a bug in the Windows driver. But it sounds from descriptions that it isn't a bug in Windows, it's a bug in the SSD that went unnoticed because not many people wrote large amounts of data quickly on systems with those SSDs.
So it sounds like there may be a model-specific blacklist required in the driver, to detect particular SSDs and reduce the speed they are written to, because they fail when run at the speed they advertise to the OS.
Or, alternatively, it sounds like those models may require a firmware upgrade from the SSD vendor.
If either of those are required, a similar workaround will be required in the Linux driver too, to avoid the same problem as soon as someone runs a similar application on Linux.
Unfortunately, even with a speed-limiting blacklist in the driver, whether in Window or Linux, those SSD models probably still corrupt data from time to time, because speed alone is unlikely to be the underlying cause of corruption. A vendor firmware update, or vendor confirmation that a specific change in the driver avoids the SSD bug, are what's required.
The linux install base is orders of magnitude smaller than windows, so there's less of a chance that it gets picked up by tech media. Moreover if it's really an issue with the underlying hardware itself, a more technical user base would be able to trace the issue to the actual hardware, rather than going straight to social media with "ZOMG windows update bricked my PC!"
>Why did it started happening after a windows update then?
Some sort of code change that changes the behavior of the drive, but is technically within spec.
>And why did it need a patch?
Because putting in a patch doesn't imply you're at fault. The linux kernel is filled with workarounds for crappy hardware as well. The existence of those workarounds don't imply linux was somehow at fault for those bugs.
So what you are saying is that users went "ZOMG windows update bricked my PC!" And that prompted the article?
> Because putting in a patch doesn't imply you're at fault
If you designed a system for certain range of hardware, your latest update broke that design and you had to go back and put a patch, I'd say it's pretty much Microsoft at fault here, even though your comment is right
Most tech "journalism" is basically regurgitating stuff from reddit/twitter/other tech sites, so yes.
>If you designed a system for certain range of hardware, your latest update broke that design and you had to go back and put a patch, I'd say it's pretty much Microsoft at fault here, even though your comment is right
So there's no accounting for which party broke the spec? Whoever touched it last is at fault?
Although that bug report is for ZFS on the same SSD as the Windows article (WD SN770), the faults described in the Linux bug report look like would also happen with ext4 or btrfs.
So I searched for reported problems with SN770 and ext4, and found enough results to convince me.
That particular SSD fails with Linux and ext4 too.
Blaming the end user for the update corrupting an existing drive that has had NO ISSUES for over 4 years is about the dumbest and most condescending thing I've seen. Thanks for being absolutely no help at all, everyone but wildpeaks. The very first reply is someone who apparently went over to the FB post and spammed every comment with "CLickbait (sic) doesn't apply to anyone blah blah".
Considering the giant SHOVE that Microsoft has subjected us to in the last few years to upgrade to the latest and greatest OS, you would THINK that SOMEONE would remember that a lot of those older machines that are now running Win11 should be considered.
Also, anyone replying to a Windows problem with "well, you should be using Linux, too bad so sad" needs to check their attitude at the door. What an entitled life you must lead.
Going back to wrestling with trying to access the now external HDD. Do I sound bitter? You're d**ed straight.
which part of the article leads you to believe this?
> The report speculates that this could be due to a malfunction in the drive cache subsystem. Symptoms are said to recur predictably after a system reboot, which temporarily restores drive visibility but does not address the underlying fault. Affected users are said to be consistently experiencing failure under similar workload patterns within minutes.
> Further analysis has suggested that SSDs built on Phison NAND controllers especially DRAM-less models exhibit failures at lower write volumes. Reports suggest that select enterprise-grade HDDs also display comparable symptoms under intensive writes.
> The issue definitely bears high similarity to the WD SN770 host memory buffer (HMB) flaw, and in this case, too, restricting or disabling HMB yields no improvement. A suspected memory leak in Windows’ OS-buffered cache region could be the problem.
Speculation almost entirely revolves around the Windows disk cache subsystem.
That being said, I completely disagree with most of the replies you are getting. If Intel/AMD released a CPU that exhibited faults at high instruction throughput (making good use of all execution ports, etc.) but within the advertised power limits, I would in no way blame whatever software exhibited the problem, I would blame the chip.
Not an argument for staying on an unsupported OS, but an argument for staying on a stable release as long as possible if you prefer stability over features
I think most software developer gets to understand this once they're working with production environments sooner rather than later. Most of the times stuff breaks is when someone changed something and failed to consider something else, not a lot of times things break by themselves. Sometimes, things break by themselves because someone in the past forgot to account for something that will happen far in the future though.