Retrieving 1TB of data from a faulty drive with the help of woodworking tools
blog.jgc.org
blog.jgc.org
Boss walked by. “What the heck are you doing?”
“I’m milking the data out of this drive!”
Shims were plastic cards, subway tickets, and occasionally toothpicks as an adjustable clamp while I was trying to figure out where the failure was.
My theory (based on the sound the drive had been making, which prompted me to try imposing different kinds of sideways acceleration on the laptop until I could hear the drive spinning ”freely”) is that it wasn’t a problem with the disks/heads but a messed up bearing in the spindle motor.
I would like to call out to the younger readers that this is not a typo. Think about how minuscule the capacity of this hard drive is, and also that it's entirely possible it was a full height 5.25" drive, which is 3.25" tall!
Here is a fun tour of old drive technology that goes into just how amazingly simple some of it is: https://www.youtube.com/watch?v=8LbFKV_pPAE
It also meant lots of highly unreliable floppies !
Seriously though: the amount of miniaturization in storage is something I'll never get used to. From punch cards and papertape via cassette to floppies, harddrives (in various incarnations and densities) and now to solid state so compact that you could store all of the data that I've created in my whole life in something about the size of your smallest fingernail. Incredible.
The MFM drives were… fun… in that if you connected one of the two data cables upside down, you’d crash the armature and destroy the hard drive.
Someone gave me a 30GB drive the other day. I still haven't figured a use for it?
Actually there is a typo there, it's Conner. But yeah, 40MB was even a luxurious capacity, my first HDD was 20 (later reformatted to 30 with an RLL controller)
-g=c800:5
We had three levels of hard drive data recovery that we did in order. Usually didn't have to get past #2.
1. Put it into another computer
2. Cold spray the heck out of it and see if it works for an hour
3. Remove the circuit board of the drive and replace it with the circuit board from a matching known working drive
If it didn't work after that and they 100% needed the data we sent it out to Ontrack. They'd put it in their clean room, remove the platters and read the data directly.
I didn't know anything about cold solder joints at that time, but I did discover that my flaky motherboard would start working when I flexed it a certain way. So I wedged a bamboo chopstick between the motherboard and the case to keep it in a little bit of flexion. Lasted the rest of the year.
Was so relieved as newer socket designs came out.
Eventually I read a random post on Reddit[0] how a guy tried putting it his SP3 in the freezer (in a sealed freezer bag) and it eventually came back to life.
I skeptically tried it. I thought freezing was more likely to destroy something else on it, but what did I have to lose? It was already dead.
I put it in a sealed freezer bag with as much air removed as possible, and then put it in the freezer for a couple of hours. I took it out, plugged power in, and I was able to turn it on!
The first time I did it, it only worked for a few days. I tried freezing it again, and that worked. It still works to this day, many years later. My theory is that it seemed like a power problem, and that freezing the battery put the chemistry through some kind of cycle that repaired it.
I did keep the Surface wrapped in layers of towel to warm it back up slowly. Mostly to help prevent any moisture from building up somewhere before it got to temperature.
[0] https://www.reddit.com/r/Surface/comments/5tficj/how_i_reviv...
Do you think that replacing the battery would have fixed it?
Instead of this i would suggest a fan. Even small ones create much more airflow than needed when please next to a device.
I have two Compute Modules (CM4s) and found a passive heat sink that is finned aluminum that covers the entire module. There is no fan to fail. I can overclock the CM4 to 2 GHz and load it for stress testing and it remains well within limits. One of these is equipped with an NVME SSD and I was surprised to see it reach excessive temperatures under heavy loads despite the PCIe x1 connection on the Pi. I got a "hunk of copper" NVME cooler, cut it to fit a 2230 NVME and it keeps the NVME SSD within reasonable limits.
Actually, almost nothing in a wood shop is stainless because it sharpens poorly[0] compared to O1, A2, or any of the other tool steels you'll find in a wood shop.
[0] Yes, I realize that stainless and carbon steel are two whole classes of alloys, all with different characteristics. Based on my experiences with kitchen knives, I'll stand by my claim that at least the common ones sharpen poorly compared to the common carbon steel alloys.
https://microchip.my.site.com/s/article/SAM-Cortex-M3-M4-M7-...
I took pieces of wood, and a clamp and applied pressure there and it worked.
The same pressure is applied if you use washers on the heatsink clamp - I did that and it's been running fine for months.
The solder joints on the GPU get brittle under the stress of heat and compressing it restores the connection. Pressure or re-flowing is really just a band-aid, the real fix is to re-ball the GPU but that gets expensive.
Never actually had to escalate to reflowing the BGA mount underneath the GPU but recall tutorials of how to do that in a consumer oven.. thank goodness I never tried that one at home.
Tech is already pretty reliable in my experience (At least the cheap stuff that isn't as high power and doesn't thermal cycle as much I guess), but getting rid of solder as a failure mode and making chips swappable would be so cool.
If they could somehow make production grade Z axis tape we could just tape the parts on with a 3d printed frame for position, and anyone could do component level repair without much skill.
And all the chips from dead devices could be reusable, if there was a way to automatically sort them all.
I found that if I manipulated the axle of the read-write-head arm where it came out through the bottom of the drive, it would "unstick" the head from the surface of the disk, and the thing would boot! I imagine there was some kind of lubricant in there that would congeal when the machine was off for a certain amount of time.
So I left the drive slot cover off, and "fingering" the drive would get it to start reliably for a number of years after that.
Oh so the fridge trick now also works for NVMe drives? I saved files out of an old HDD (about 20 years ago) by putting the HDD in the fridge for about half an hour too: got the files out of the drive then trashed the drive.
Eventually I realized it was overheating.
Cue getting ice from the hotel dispenser and using a fan and a metal tray to keep the laptop alive until my work was done.
But more importantly, this is a reminder TO HAVE BACKUPS. Many of my friends love their Synology systems. I am thrifty so I use urbackup and it has saved my butt a few times.
A buddy just gifted me a decomm'd Synology RS2416+ w/ expansion unit full of drives that could, just, back up what I have... and I am definitely not going to do that.
> TO HAVE BACKUPS
Agree. I admire the diagnostic skills and ingenuity to recover the data, but was thinking I would just restore from backup in the same situation. But I'm weird. I have a "home lab" file server with a true server H/W that my desktop and laptop back up to as well as a remote server that the local one backs up to.
e.g. Samsung https://download.semiconductor.samsung.com/resources/data-sh...
But nowadays a drive controller is far more complex, e.g. it might implement transparent hardware AES encryption, in which case swapping the board loses the key. And I've no doubt there are many modern manufacturing-related tricks for yield that go into them as well, any of which might make a different set of platters than what was shipped unreadable.
Note that many controllers store parts of their firmware/config on the disk platter, so without the disk platter the board may not show up via sata.
The drives are 3.5" 2TB WD Green purchased around 2012, and had been in use for about 1 year when they both died.
If that's not a typo, then it seems like those drives have been powered off for about 10 years.
I think powered off hard drives are commonly said to retain data for um... maybe 3 years (from rough memory). So, your drives have probably lost their magnetism (and thus the data). :(
The only issue that may happen on an unpowered for several years hard disk is so called bearing seizing, the (fluid) bearing of the motor/platter may become stuck, but it is relatively rare, though some particular make/models are more prone to this, and though (usually) fixable, in some cases it can be made to rotate freely again, but you need the services of a specialized service, as the disk needs to be opened, in some other cases the bearing can be replaced, and some specialized tools are needed:
That reminds me of a work colleague a few years ago. He got an ancient drive working again by tapping it on the side with a screwdriver while powering it on, to get it "unstuck".
HPe had two rounds of enterprise SSDs that failed because the uptime counter overflowed, but I never saw content about fixing those after the fact. And I think I had seen a different SSD uptime based failure a year or so before.
IMHO, it's best to avoid same batch storage, and if that's not possible, stagger the online time to try to give enough time to notice a failure, obtain replacement storage, install replacement stotage, and migrate data. Backups are important too, but it's nicer to have a path towards mostly online recovery. And some mostly replacable data is hard to justify backups for (do I need three copies of format shifted media? probably not, if my online storage fails, I can re-rip)
I don't recall hearing about this for Western Digital drives, but there's some xbox360 stuff that I thought involved the ttl serial on WD drives... It's certainly worth exploring. WD green drives do also have a very short default timeout to park the drive, and as a result can experience a large number of parking cycles in some applications, and the parking ramp can wear out; I don't think this is really recoverable, the heads are likely to get damanged and debris may damage the platters.
This happened only on some disk drives because it was initially triggered by a defective testing equipment only on some production lines, see "Root cause" here:
https://msfn.org/board/topic/128807-the-solution-for-seagate...
A good half of "failed" drives can be made to work for at least a few hours longer with that method.
Unlike OP, If you see signs of life, don't mess with clamps or reflowing. Just leave it in the freezer/oven while you take data off it, with longish power/sata cables to a machine just outside the freezer/oven door.
I recommend GNU ddrescue for getting data off - when you only have a few hours of service life left till it is dead-dead, it maximizes data recovery in a given time. There are various ways to generate a mapfile to skip recovering free blocks, which are worth using if you suspect the drive is mostly empty.
That said, if you ever have a drive with absolutely critical, must-have data, don't bother with any of this and just ship it to professionals. You'll pay dearly, but they'll get your data out.
If you've got the controller board your drive wants and still nothing, then its time for professional help or considering the data lost, imo.
I don't know how complex HDDs have got, but I recall giggling at someone installing linux on an HDD controller board several years ago. So I bet its much worse now.
There is specialized hardware (and software) to be able to extract the data and save on another board's memory, but the poorman's way is that of transferring the actual chip from the old board to the new (identical) one.
This, commonly referred to as "ROM swap" is not particularly difficult[1] as the chip is usually a rather simple 8 pin one, if you are not into this kind of things a hardware repair shop (like a phone repair one) will normally make this work for you.
However newish hard disks may have not this separate chip, it has to be seen which model yours is.
Here is a site with some more info:
https://hddpcb.eu/gb/content/how-to-swap-hdd-pcb
[1] meaning that it can be done DIY if you are familiar enough with soldering/desoldering components
Ah, so you don't really need that data...
Because if you really need the data then you go to people who makes a living by recovering data.
But if you are okay to lose the if unsuccessful then it's okay to try, but you should know/tell that beforehand.
Reading through the other comments - you have a very low chance to succeed, because if you want to swap controller boards then you need to move adaptive data too, as other had said.
But I'm curios what exactly happened, WD Green from 2012 are not the worst drives out there. How exactly they failed, what happens now when you power them on, with SATA connected, without? Did you try external USB2SATA converters/boxes?
I'm pretty decent with soldering iron and hot-air gun. Migrating the SMD flash/EEPROM chip shouldn't be too hard.
>"But I'm curios what exactly happened, WD Green from 2012 are not the worst drives out there. How exactly they failed, what happens now when you power them on, with SATA connected, without? Did you try external USB2SATA converters/boxes?"
The discs spin up but the controller no longer communicates with the host. The computer doesn't see that the drives are attached to the SATA bus. There were no signs of problems coming, they just suddenly stopped working from one power-up to the next. I tried with different motherboards and a couple of SATA-USB bridges, all same result.
Well, good luck then.
> The discs spin up but the controller no longer communicates with the host
Now this is strange, if the controller would be dead then there would be no spin-up. If you hear the heads working than the controller is definitely not dead.
You tried to search forums dedicated to data recovery with your exact P/N?
G-clamp makes total sense, I'm just curious!!
(Wiktionary does give the clamp meaning too, not with a UK qualifier though. https://en.wiktionary.org/wiki/cramp)
In the 90's they had this policy that you can send them a broken HD and they would replace it with a new one. Since they wouldn't make the old version anymore, the new drive would have more capaciy.
So at first you could just send them your hard drive and you would get an upgraded back.
I was probably not the only one who heard about it, and so they started testing the incoming drives to see if they were really malfunctioning.
I once tried to break such a drive. Formatting the hard drive while bashing it on the table. The thing kept going, no problem.
Sorry school!
I discovered at some point that a few swift kicks to the front of the case would get it going again for a couple of days. This lasted for a good year or two before no amount of kicking would revive it.
So, I did the right thing and stacked a bunch of electrical tape between the chassis and the SSD, and it has been working ever since. It's OK if the SSD dies though; it's running NixOS, so getting it back up and running with a new SSD would be a very short ordeal.
I assume this occurred after the manufacturer's warranty expired.
The fix was to disassemble the back of the laptop, mask around the GPU with tin foil, and hit it with a hair dryer on max setting for about 20 mins.
The fix would last like 3 months. Learnt it off some guy on YouTube who was such a hero.
Just scrolled through the Ontrack website:
Data recovery myths: why you should avoid the temptation of a DIY repair
Suggestions we’ve seen online that definitely _will not_ help you recover your data include:
- Putting your hard drive in the freezer overnight
:')I offered a very cheap alternative, and took his spare wifi antenna cables and cut the off. Then I soldered them DIRECTLY to the motherboard and routed them to hang out the side.
It was like jumper cables, you tap the and the laptop would turn on
Compared with my samsung ssds which have (knock on wood) never died in a decade.
I know all brands fail but is there something of lesser quality with the firecuda line?
There are some products where some start dying after 1 year, and nearly all of them are dead after 2 years.
Other times it's a combination of a bad design and some specific use case. For example "these fail after 3 years if you turn them on and off every day, but don't fail if you leave them running 24x7".
Basically one will suddenly decide it hates life and will then only show up as having only 1gb of space and the firmware version set as “ERRORMOD” (“error mode”). This is not a time counter rollover issue as far as I can tell (though samsung has plenty of those too [0]) as there wasn’t anything in my smart logging that would indicate a time value they approached and died. Samsung’s firmware is just super buggy and can get caught in a bad state. You can find business purchasers complaining about these issues as well. [1]
Someone once decompiled the firmware of the EVO 840 before they started encrypting it. Have a look at the “Bugs” section to have a laugh: http://www2.futureware.at/~philipp/ssd/TheMissingManual.pdf
So when you see a samsung enterprise/oem drive on ebay with 99% of it’s life left, what you are really buying is a drive with bugs but no way to obtain fixes, since samsung will give you the business version of go fuck yourself by telling you to “contact your vendor”.
Part of the problem is samsung cultivates a complete shitshow where vendors will “customize” firmware, so an identical drive from lenovo will need different firmware than one from HP. The other part of the problem is Samsung’s consumer drive branch is basically completely separate from their oem/enterprise branch despite what is mostly the same hardware and firmware. So while the consumer branch has historically had to eventually face the market consequences of gross negligence, the enterprise branch is shielded by misdirection to the vendors.
Fortunately mine where quite cheap so not much loss, but still, they lost any illusion of competence over other companies in my eyes.
Note there is a collection of samsung firmware here, though it is hardly complete: https://github.com/lolyinseo/samsung-nvme-firmware
P.S. Also note there is a poorly documented but common form of ssd “failure” that can happen. If you have a drive suddenly not show up on next boot (especially after power loss), using the power cycling technique can often recover a drive: https://dfarq.homeip.net/fix-dead-ssd/
[0]: https://www.tomshardware.com/news/samsung-990-pro-health-dro...
[1]: https://forums.servethehome.com/index.php?threads/pm9a3-firm...
Many many drives, even from big companies like HP, came with the damn Sandforce controller that tends to brick itself if your computer ever goes to sleep. I know organizations that ended up replacing entire swaths of drives because they were dying so often. It was enough of a problem that some people went to extreme (legally dubious) lengths to try to recover the drives. I mean just look at the procedure:
https://computerlounge.it/how-to-unbrick-sandforce-ssd/
Of course Sandforce was completely unhelpful in trying to fix the problem.
1) A "faulty connection" and why it fails at a certain temp? Is it a partially broken circuit that doesn't work anymore if resistance is too high?
2) Why airblowing fixes it. Because it melts the crack together? Any why it doesnt do any other harm to the SSD?
It may have started with a hairline defect that got worsened by thermal cycles.
If it is a bad solder joint, reflowing bridges and fixes the joint. It won't harm the chips as long as the temperature is correct, since it needed to be soldered in a reflow oven in the first place. However (I believe) there is some risk of excess heat corrupting some data, and/or worsening the defect, so if you can back up first using the jury-rig, that's certainly preferable.
It WILL do other harms to the SSD. Semiconductor parts have limited reflow counts allowed to meet specifications such as failure rates, longevity, maybe power consumption too. It will be a factor at scale or over time. But it's free extra half life for semi broken parts too.
It's not a matter of if your storage media will fail, just when.
Always keep backups.
Yep, exactly -18C cool.
Was the drive getting too hot? Maybe under the gpu or in the bottom slot one of those piggyback boards?
That's why. HDDs aren't failure free but at least most of the shenanigans are solved already. Except SMR.
Years ago, the drives had safety systems to prevent this.
It started when I threw a very expensive (at the time) 1.5TB HDD. It had various things going on with it but I recall it wouldn't spin up and once I was able to get it to spin up, it wouldn't function for more than a few minutes before it would error out and the drive would become unrecognized by the PC.
I remember I made it spin up again entirely by accident. Figuring the drive was toast, I'd removed it and while fumbling with the cables it slipped out of my hand and landed on the hardwood floor. I hadn't noticed when I was holding the drive, but it felt like something was "stuck"[0] and after picking it up off the floor, it felt like the components were moving again. On a whim I plugged it back in and it worked.
The second problem, I knew, was thermal. The drive got incredibly hot very quickly. Being that it was a 1.5TB drive at a time when "that was big", I'd need the drive to be functional for hours to complete copying the data.
I had a dorm fridge in my office and debated putting it in one of my external cases and running a cable from there but knew that the moisture wouldn't play well. Because I kept very little in this fridge, it was filled with a few cans of soda, coffee creamer and about ten of those blue gel bags that people freeze/put in coolers in lieu of ice. I did this mainly because the fridge was really loud and keeping it stocked with anything made it run a lot less.
The upshot was that the size of these "bags" was just a little larger than a 3.5" HDD, if left out they maintained their temperature for hours[1]. I grabbed two, placed the drive on top, put two more on top of the drive and managed to image the whole thing.
It was so simple, I popped an ad on Craigslist and $200 "I'll get your data back or you don't pay." Out of 30-40 drives, there were less than 5 that were beyond help. About half were software issues, many of which were simple to resolve. Very few required a tool like photorec, but it did the job when needed[2]. The rest were handled with an extremely stable power supply, four gel bags and patience. Except for one of my own drives, I never popped screws and generally didn't resort to "applying physics and gravity" in hopes that if I couldn't recover anything, I'd at least give them a drive that was "no more damaged than when I received it" (they still signed a document covering me for any liability).
The whole experience made me realize how critical keeping "components other than the CPU" cool is. Every drive I've experienced hardware failures on[3] has had something heat-related coinciding. I remember one time I couldn't figure out which of the 8 drives was indicating failure; I saw one of the three fans in the drive cage wasn't spinning. I shut the server down, pulled out the fan and its 3-drive set, figured has to be the middle drive, swapped it and it started rebuilding the array on reboot. I guessed right. :)
[0] If you take a drive, lay it flat on a desk and spin it, you'll hear things move ... this one didn't.
[1] Provided it's just sitting on a wooden desk.
[2] There was one customer that I'll never forget and he was a reason I stopped doing this entirely. I told him I could recover photos/videos/specific file types from his drive but it will recovery everything it sees including files that may have been intentionally deleted. He said "Oh, no! Don't do that. I'll pick up my drive." I still wonder, to this day, what kinds of horrors I would have been in for (I rarely did more than a spot check on the data, anyway).
[3] I still have a pretty massive custom built storage server with an older LSI MegaRAID controller and a mess of SAS HDDs. For what it's used for, the speed is more than adequate the individual drive costs are substantially less (even taking into account SAS over SATA) and they outlive my SSDs.
They got the bright idea to make the tabs for the sleds out of plastic to make them easier to slide in and out, or save a couple pennies, or both. Problem was that metal drives in metal sleds on metal rails going into metal bays do a pretty reliable job of keeping the motherboard and the drives on a common earth ground. Plastic rails meant a floating ground, which electronics especially do not like. A little static electricity or inductance and things get ugly.
If memory serves they added some sort of little grounding cable they would send you, making it harder not easier to get the drives out. So dumb.
Ah, a man after my own heart <3
With the modern widespread trend in tech of treating the owner/user as a security threat, it's easy to feel like the hacker spirit is dead or dying, and then posts like this rekindle my hope for humanity
I'd be ok with this in a machine with regular backups, where you're looking at worst case the data from a day or two is lost, plus some downtime while getting a replacement and restoring a backup, and best case another trip through the hot air station. Seems pretty ok to me in that use. Would probably be fine for a use case with easily replaceable data too: if you've got a fast network connection, it could be your steam drive.
Good to know it's still useful.