Pi 5 overclocking: Silicon Lottery
jeffgeerling.com
jeffgeerling.com
When I started to help friends overclock theirs, I quickly realized the "silicon lottery" variance. Some would only run reliably at 78 or 76 MHz. I bought a bunch of fixed frequency clock generators (that were drop-in replacements for the original on the motherboard) in 2MHz increments due to the variance.
This was back before CPU's had heat sinks or fans, so we quickly figured out that adding those gave better margins. We even made some 10-LED bar temperature display that had a thermocouple glued to the CPU case and indicated 10 degree C increments (green=0-60c, yellow=70-80c, red=90-100c).
The reason that this story is interesting is that in most cases you could just yank C9 entirely and with nothing more than a resistor between the clock pins, you'd get a roughly 300% performance increase. I guess the parasitic capacitance was enough to still oscillate a bit although mostly it would have been random. Looking back, this was basically a CPU being clocked with 50mhz noise and still running happily! Amazing!
I held my hand over the solar panel at various lengths until the screen cut out and while doing this I just kept hitting random keys and while lifting my hand up.
One day when I did this I must have hit the one in a million chance. It started rapidly counting up by itself!
I think I only got it to happen once more. I suspect the fluctuating voltage and it trying to do calculations while I was pressing keys was just enough to get some gates latched into the wrong state, somehow.
I have a recollection of a guy who got a 486 DX-33 up to 133 MHz by putting the entire computer in mineral oil and floating chunks of dry ice in it. Watch out for asphyxiation.
What effect would running an overclock like this permanently have on longevity? Is it even worth thinking about "longevity" of chips?
Obviously stability suffers but for example... how much? Author was able to get a Geekbench 6 benchmark to pass. If they tried 100 times, would it be expected that a non-zero amount would fail?
If voltage is raised we enter the realm of electromigration, though not sure how relevant it is for such a minuscule OC.
As for stability, yes. If the voltage is not sufficient there will be stability issues which will require further raising it, thus raising temps and requiring more power which you can't be sure if you could deliver. And then of course electromigration.
This was mostly happening during the transition to lead-free solder. Today, component should fail earlier than BGA.
But in general, the way you tend to end up with wearout is through electron migration ("electron wind") which is damage to interconnects from electrons slamming into metal atoms over and over and slowly ripping apart the wires. Modeling electron migration correctly is really hard, but (from memory) a general relationship is that voltage linearly increases failure rate and temperature exponentially increases it. The constants for these models are determined empirically and, of course, are NDA'd.
In general, I wouldn't worry about it. The MTTF for semiconductors is already very high even under awful temperatures 100+ C, and so as long as you cool it properly you'll be fine.
> would it be expected that a non-zero amount would fail?
The failures you describe here are going to be due to setup time violations. These issues shouldn't be transient (assuming identical temperature and voltage) since the performance characteristics of an individual device don't really change over time. Of course, the issues can seem transient as the failures may not actually always cause noticeable corruption (maybe you generate a wrong FPU result but that specific path is only exercised under rare uarch conditions).
So, it's a non-answer, but: yes and no. Maybe your chip is perfectly okay and never has any violations at the parameters you selected. Maybe it does. Neither you (nor in fact the manufacturer, though they do have a better chance since they know the process/design) can ever really be sure--all that's left is empirical burn in testing and hoping for the best :)
There’s a whole bunch of other failure modes that aren’t captured by the 10C rule. It’s more for estimating chip failure due to things like electromigration. You can observe this if you run a desktop CPU overclocked for many years. I had a 2600k that I had to keep bumping the OC down on, and jt eventually bit the dust after a decade.
We had cards in the worst of the worst environments and they ran fine for years on end.
Reboot the system, device just disappears, never to be never seen again. It generally starts after ~6 year mark.
Sometimes device starts to corrupt things silently, but not always. However they too disappear after some time.
Oh, sometimes GPUs do that, too.
Thus expected lifetime quickly becomes long enough that it's effectively not an issue for CPUs and GPUs if you provide sufficient cooling.
Both yes and no.
I still have an old AMD Athlon XP system, which works at 2200MHz (200x11), which is completely out of spec for that generation of AMD systems (2200MHz parts had 166MHz bus), and it still performs as on day one since it's not overclocked and cooled well.
On the other hand, we change parts which fry because they feel like it even they are not even close to their thermal limits, because they're kept in well cooled data center.
Sometimes, things go bzzt even without extreme heat. It's really interesting. Something is working at full throttle with no problems, you update a couple of things, reboot, and the device is gone for good.
Same for the RAID card. The processor has a couple of failure modes (no cache or no card), both directly related to RAID processor itself. Again same for the Ethernet cards we fry. They lose their MAC addresses, all pointing to in silica problems.
There’s a trade off between single threaded performance and power, right? I’d expect the increase in power cost to be between the performance increase squared or cubed. If you expect a one-to-one trade it is never worth it to increase frequency, haha.
The universe will give you throughput at a fair rate, but it is very stingy about latency, in general.
Silicon lottery, where the chip was on the wafer (edges tend to be less reliable), manufacturing batches, component batches, heat, cooling, power supplies, etc... the list goes on and on...
It is understated how much all of this is a huge impactful thing on performance and stability.
Anyone know which SBCs use this chip?
The rk3588 is a nice chip, but support just isn't there yet if you want to do anything with the GPU. The "Panthor" GPU driver, which is the FOSS driver which supports its GPU, was just merged in to Linux and mesa this month[1] (yay!) which means you're probably gonna have to build your own kernel if you want it.
The old mali proprietary driver is borderline unusable on anything remotely modern, only really working on Linux 5.10 and special X11 builds with legacy features re-enabled.
It's crazy that the rk3588 has been on the market for many years at this point and is just now starting to be usable on Linux, but it's exciting that things are taking shape.
[1] https://www.collabora.com/news-and-blog/news-and-events/rele...
This should be the central lesson learned from the Raspberry Pi by open-source projects.
There will be faster, there will be smaller, there will be cheaper. But if the user can go on the web and find the _exact_ thing they're looking to do spelled out, they'll buy that product, every time.
Are you in the US, Canada, or EU? Outside of those places, there may still be some delays in getting stock to meet demand.
I have the original and updated Pi Zero W-- unbelievable bargains at ~$15 but if I needed any horsepower I think I'd rather have an x86_64 so I can run whatever.
Adafruit FT232H Breakout - General Purpose USB to GPIO, SPI, I2C - USB C & Stemma QT: https://www.adafruit.com/product/2264
Adafruit MCP2221A Breakout - General Purpose USB to GPIO ADC I2C - Stemma QT / Qwiic: https://www.adafruit.com/product/4471
The downside is that code becomes slightly more complicated, the upside is that it makes easier to replace the Mini PC with a totally different one, or even emulate it in a VM then connecting the GPIO board using USB pass through.
Rethinking GPIO can also open more possibilities such as putting them beyond a network (say Arduino+ENC28j60 and similar ones for Ethernet, ESP32 for wireless, etc); of course having them outside the main CPU will imply some speed and latency issues, although I'm sure they would remain unnoticed for many non critical use cases.
My question would rather be why the RPi when there are better SBCs out there, but that is my personal taste and the Armbian and DietPi communities being more than enough for me.
Anyways, I ended up making my own and it works great. No fuss, no wait, no complications.
As for bandwidth, well, it's one lane of PCIe gen 2. This won't win any races but can be useful to access exotic hardware not available in usb or if you don't care about bandwidth. (e.g. HBA with many drives for mass storage without speed requirement).
What SmartNIC are you using? Most SmartNICs that I'm aware of suck a decent amount of power, many more require significant external airflow. Are you using the Mikrotik active cooled one? https://mikrotik.com/product/ccr2004_1g_2xs_pcie
The 12V comes from an external power supply through a barrel jack, from which I also derive the 3.3V rail. The Pi provides no power whatsoever. I should publish the design files somewhere.
As for the NIC, it's a Netronome Agilio-CX which is fully programmable using eBPF and such.
But so far there's no straight PCIe expansion board available to purchase yet. It is a slimmer market than 'NVMe on top' or other more standard use cases, but it's one I think could expand as people do weird things with the Pi 5.
Looks way more compact and well thought out in terms of mounting so definitely seems like the better option once it's released.