Intel’s Atom C2000 chips are bricking products, and it’s not just Cisco hit
theregister.co.uk
theregister.co.uk
"The problem in the chipset was traced back to a transistor in the 3Gbps PLL clocking tree. The aforementioned transistor has a very thin gate oxide, which allows you to turn it on with a very low voltage. Unfortunately in this case Intel biased the transistor with too high of a voltage, resulting in higher than expected leakage current. Depending on the physical characteristics of the transistor the leakage current here can increase over time which can ultimately result in this failure on the 3Gbps ports."
http://www.anandtech.com/show/4143/the-source-of-intels-coug...
Since I needed that machine functioning, I never replaced it (the mobo had some extra SATA ports handled by a different controller, so they kept working and I switched to using them). I suspect a lot of people are in the same boat. I'll never do business with that store again.
I have 2x C2758 Supermicro boxes running core routing services and a Synology RS2416+ for storage - all on the affected CPU list - guess I better double check my backups are working and allocate some funds for replacement kit in case things go belly up!
http://www.cisco.com/c/en/us/support/web/clock-signal.html#~...
No replacements unless its under warranty
It seems that this isn't entirely the case.
From TFA: "if your product was under warranty as of November 16, 2016, you are still eligible to replace your products", and "Cisco is offering to provide replacement products…even if they have not failed".Without more information, I wouldn't like to speculate on the cutoff date of November 16, 2016, but so far it doesn't sound too unreasonable. We don't really know the scope of the failure, nor relevant batch numbers, etc, so I'd give them the benefit of the doubt, until clearer details emerge.
100%, this wording is to make sure that grey market buyers aren't covered under the replacement. Basically anyone who bought from Cisco will be able to get a replacement.
Cisco is like Apple in that regard; You pay premium but the service is nowhere near premium.
Also, I'm reasonably sure many of these were sold already permanently affixed to the motherboard, so the fix may be worse than swapping out just a CPU.
From what I can tell, all of the Avoton chips were sold in a BGA package that required them to be soldered to the motherboard. There isn't a socketed version of Avoton.
C2308, C2338, C2350, C2358, C2508, C2518, C2530, C2538, C2550, C2558, C2718, C2730, C2738, C2750, and C2758.
How to tell what CPU your Synology NAS has:
https://www.synology.com/en-us/knowledgebase/DSM/tutorial/Ge...
My 3-month old 1815+ is on the list...
See the last couple of posts in the thread as well.
Sounds like you can get an RMA, but it's a slow process.
Instead, I'll shut down the NAS and buy a replacement from another company (QNAP probably) and transfer the data. The other options feel too risky.
(Synology 1515 specs: https://www.synology.com/en-us/products/DS1515+#spec)
Are Cisco products worth the money ?
Personally i haven't had that much experience with their stuff, but i remember seeing a brand new router running hot with two fans blowing in it (other routers at the time were 2x smaller without fans). I understand that Cisco should be the de facto networking standard, but is it really worth the name ?
That said, they are definitely more expensive than their peers. But then again, an Escallade is more expensive than a Tahoe.
And some technical specifics behind the problem: https://slashdot.org/comments.pl?sid=10214953&cid=53819967
Can't post to The Register, since they don't have ACs.
Anyway, the issue is damage to the LPC (low-pin-count) bus clock line. This is a secondary bus where you hang old ISA-style devices, like the system FLASH. If the FLASH is the only thing in there, it will mostly render the system unbootable (so, stuff that never gets power-cycled would just keep going). But LPC can generate interrupts, and one often hangs other crap to that bus, such as i2c controllers for hot-swap bays, motherboard management controllers, and other sensors. In that case, you can expect severe runtime misbehavior.
The issue is caused by "continuous degradation due to use", so repairing it is easy, if costly: replace the motherboard with a new one under warranty (and even if out of warranty period wherever this kind of "stealth" manufacturing defect is not subject to warranty time period limitations, such as in Brazil). It will "reset" the counter. This is your zero-day solution to the issue.
Depending on time-to-market for the new stepping (hardware revision) B1/C0 of the Atom C2000, you might need an interim solution, which is the "platform-level change", i.e. redesigned board with extra components that work around Intel's hardware design error. As soon as you have these, you start using these to replace any boards returned due to the defect, or start a "recall" to preemptively replace boards.
Depending on the total cost of the board plus other components, you keep the old boards you replaced around, and when revision B1/C0 of the Atom C2000 is out, you BGA-replace them in a factory (about US$ 25 per board in large volumes, if that much), maybe replace any liquid electrolytic capacitors and other crap that ages badly, and use the boards either as new or as refurbished, depending on your corporate/regulatory ethics. This kind of repair almost always really resets the boards MTBF. If Intel supplies the replacement Atoms at no charge, the cost of repair might well be far less than the cost of the production run for boards you'd want to keep around for warranty services, anyway.
Mind you, at 1.5 years per failure, it will be rare the legislation/contract that forces more than one replacement... so, let's hope they don't replace a faulty board with a brand-new virgin but-still-timebombed board. You'd have trouble to replace it a second time if it fails after the warranty period.
NMI ... going to debugger
Stopped at acpicpu_idle+0x22d: nop
ddb{0}>
I googled it and found one similar report on the OpenBSD misc mailing list from September 2016 [1]. Interestingly, the person who reported the bug was running the same Supermicro board as I was. The report didn't get anywhere other than a vague suggestion that it might be heat related. These boxes run very cool and I didn't think that was likely. I thought it might be a RAM issue and that it was probably just a coincidence that the other person had the same hardware as I, but now I'm inclined to think that both of us have experienced the issue described in TFA.Seems like I'll be looking for new firewall hardware.
[1] https://www.mail-archive.com/misc@openbsd.org/msg149348.html
"Other vendors using Atom C2000 chips include Aaeon, HP, Infortrend, Lanner, NEC, Newisys, Netgate, Quanta, Supermicro, and ZNYX Networks. The chipset is aimed at networking devices, storage systems, and microserver workloads."
I'm guessing that may represent what OP meant (?)
https://www.amazon.com/ASRock-Rack-Mini-Motherboards-C2550D4...
https://www.google.com/search?q=ASRock+C2550D4I+failure+rate...
http://www.intel.com/content/dam/www/public/us/en/documents/...
and search for AVR54
I'm not claiming that the chip isn't failing; I'm disappointed that the title makes a claim that the article doesn't deliver on.
Quite a few units have completely died without explanation. From the descriptions given by the users it does sound like dead cpu.
The BIOS is a piece of shit. It's buggy, the legacy-BIOS support is unstable, the Win7-EFI and Win8-EFI modes are not good either. I patched a Win7 DVD with Win8 files, so that I could install Win7. Now Win7 runs great and stable - but only after I installed various Intel drivers that fixed the hardware flaws.
I am seriously looking forward to the upcoming new AMD CPU - Intel dod barely anything the last five years, a 2011 highend CPU is almost as fast as Intel 2017 flagship, and costed a lot less back then, had less DRM or other shit that is broken. Intel needs a proper competitor, so a comeback of AMD on the one side, and Apple notebooks with ARM CPU are very welcome to stop Intel from siting on their quasi monopoly chair.
No it doesn't. Did you read the errata? It completely stops. There's no weirdness. It's just dead.