Intel updates mysterious ‘software-defined silicon’ code in the Linux kernel
theregister.com
theregister.com
"You're a consumer, you don't need error correcting memory! Only the upper class, I mean... your superiors... err.. sorry, enterprise customers need that feature. It's reserved for them, it is not for the likes of you."
Meanwhile AMD basically sells silicon by unit area. It's like buying gelato at the ice cream shop. You can ask for one scoop, two scoops, or three. You get to decide how hungry you are, that's it. There isn't a flavour with broken glass in it[1] served only to the working class stiffs, and with the glass-free gelato reserved only for the gentry.
[1] This is pretty much what a CPU without ECC is. It randomly crashes and corrupts your data. Look. Not every bite of ice cream has glass in it! There's very little glass in the ice cream. It's super rare, and Intel Pty Ltd tells you that this is an acceptable amount of risk for you to take, because you aren't as important as other people.
AMD does the same thing too. Ryzen APUs don't have (even unofficial) ECC support, except if you pony up for a Ryzen Pro one.
(and so do many other SoC providers...)
I've been expecting them to just simplify to a core count/memory channel/socket matrix for a while because they are so close. There really isn't any reason these days to segment by frequency beyond maybe a couple golden parts where all the cores run at the peak frequency.
For Ryzen CPUs (ones without an iGPU), they don't fuse it off, and as such it can unofficially work.
For Ryzen APUs (parts with an iGPU), they fuse it off, and as such it will never work there.
Just about every company in the world that sells things (including AMD) engages in market segmentation.
A classic - https://www.joelonsoftware.com/2004/12/15/camels-and-rubber-...
It has nothing to do with the "class of person" -- that whole notion seems pejorative in this context -- but maximizing how much you can extract in the aggregate.
In the case of Intel, for years (decades?) they were their own primary competitor. They wanted to be sure that for a given customer they could extract the maximum amount possible, and they did this by gating some features (ECC, AVX, etc) to try to avoid business customers deciding to get by on "lesser" processors. And the simple truth is that the overwhelming bulk of consumer products will never, ever have an ECC relevant error event, which is how businesses managed to get by without it.
My PC doesn't have ECC memory and it isn't crashing and corrupting my data. I think you are vastly over exaggerating.
Generally by the time you see actual program crashes or notice data corruption your system is really good and screwed. That is part of why people in the know are so afraid of RAM corruption. It can be persisted and exist silently for a very long time before someone tries reading some file/transaction that was corrupted during a write years back, or the database suddenly starts crashing after it updates some index as part of a GC pass/whatever, while the actual bug/HW failure will never be reproduced.
Problem is if you start having unreliable hardware it could be any component of your system. ECC memory helps track down a class of errors that could be in the ram chip, dimm, dimm socket, motherboard, CPU socket, or CPU.
My desktop cost another $100 or so because I got a E3-1230 xeon instead of the similar clocked i7. I bought it in 2015 and it's been fast, and reliable. Sure I might manage 6 month uptimes anyways, but I also might track down a dimm problem in hours instead of weeks.
I use ECC RAM in systems with error counters.
Error correction is an extremely rare event, if it happens at all. This idea that non-ECC computers are crashing all the time due to memory errors isn’t true. I've also had unstable systems with ECC RAM and no error correction events, so it's not a magic bullet.
You can run something like memtest86 for days on end. You shouldn’t see any errors at all, unless your memory is bad.
The reality is that the average consumer doesn’t really want to pay the extra amount for ECC RAM. As you said, ECC has been available on AMD consumer platforms for a long time and there’s barely any uptake outside of people building servers and workstations with consumer-grade AMD CPUs.
The AMD (Consumer) ECC story isn’t perfect, either. It’s not officially supported so there’s a lot of debate around whether or not it’s actually working on certain consumer motherboards. It’s definitely not equivalent to the official ECC support of their higher-end parts.
Intel, on the other hand, actually did roll out a lot of i3 processors with full, official ECC support. These are (or were) a favorite for low-cost servers for this reason. It worked well and the support was great.
The parent comment saying that they never see ECC errors in the wild is missing a few things:
- Server memory tends to be clocked lower than consumer memory, so errors are less frequent to begin with.
- The errors are not evenly distributed. Some memory sticks have a high error rate, others are virtually zero. There's batch-to-batch variations.
- I've done my own tests on hundreds of servers. We run burn-in tests for about 24-48 hours. About 95% have zero errors of any kind, but 5% have a high enough rate that putting them into production would be a mistake. ECC allows us to catch those bit errors instead of silently accepting them and allowing data corruption to creep in.
- I've personally had 3 different personal computers experience high memory error rates, to the point of multiple BSODs per day and data corruption. They all started off "good" and slowly turned "bad". The only reason I knew to look for memory corruption as the root cause is because of my extensive industry experience. A grandma using the same PC would have just blamed Windows for being unstable.
- Vendors like Microsoft simply ignore all crash error reports sent back by telemetry with only 1 or 2 samples, because those are virtually guaranteed to be caused by memory corruption, not programmer error. I've done similar memory dump collection and found that easily 30% of all crashes were unique in this way, suggesting that ECC memory could improve PC stability significantly.
- Suggesting that ECC memory is not needed because "good" memory doesn't need it and only "bad" memory is a problem is missing the point. All memory is bad, it's just that the bit error rates are different!
Eventually I got a corruption in a large Git repository with photos. First I suspected a disk error but reruns of "git fsck" reported different bad objects, run to run. How odd, so I ran memtest86 and it reported a bad 32 MB memory region at offset 4 GB. Never saw any kernel issues or other instability. Booting up the computer does not use up 4 GB of memory (Linux) but starting a bloated web browser does.
The computer is stable after mapping away the bad memory region by using GRUB_BADRAM. That was an new takeaway for me, a machine with ECC memory can do this bad blocks mapping automatically. It's not just beneficial in correcting single bit errors.
I would love to have ECC memory in my machine. I used it for my previous build in 2012 but I think the situation has gotten worse since then. Bigger price difference and ECC memory is not even available at the same clocks as non-ECC memory.
It's true that ECC will catch errors caused by cosmic rays, but in practice most memory errors are just from faulty memory. Large studies in Google datacenters showed that the errors were heavily concentrated in a small number of DIMMs: http://www.cs.toronto.edu/~bianca/papers/sigmetrics09.pdf
The catch is that you don't really know if you have one of these faulty memory sticks. I always run memtest86 overnight to check for obviously faulty memory, but some of these errors could take months to manifest.
I had a laptop with a RAM cell that failed. It would show up in memtest86 in a matter of minutes, yet surprisingly I didn't notice the issue for a very long time in day to day usage. I always wonder if there were random bit flips in anything I worked with during that period, but I'll never know.
With ECC, it's just one less thing to worry about.
There are some bugs that get raised on linux once every few years that are so obscure they are suspected to be random memory errors.
Sidenote I seem to have experienced Cunningham's Law for the first time semi-consciously, as I was aware my claim is most probably wrong. Still, I wonder if data corruption is a more common scenario rather than crashing, because ever since I use Linux I don't recall experiencing many crashes I couldn't attribute to a driver issue or known instability in the software I used, especially on compatible hardware. Windows days were another story... But maybe the memory modules were worse back then and perhaps having less memory made it more likely to crash?
Without ECC any of the above will crash an app, or even the entire machine. With ECC you'll get a log message, and if it's a single bit error it will be automatically repaired. If it's not fixable and it's in userspace that application is killed (and the kernel logs the error), if it's in kernel space the kernel will panic.
So generally it makes your machine more reliable, and easier to debug. Generally repeated errors are a fault of some kind and you can easily tell which dimm it is. Without ECC you end up troubleshooting all causes of crashes, and even if you know it's memory, you can't be sure which dimm it is.
So I think it's well worth the minimal premium to make your machine more reliable, and if it's unreliable it's much easier to track and fix the problem.
Amusingly without ECC heavy CPU use with parallel gcc compiles causes a particular error. Not sure if it's in the FAQ, but it's well know that the particular error means you have a CPU/RAM problem, common with memory errors, CPUs that are too hot, or overclocked CPUs.
"Studies by IBM in the 1990s suggest that computers typically experience about one cosmic-ray-induced error per 256 megabytes of RAM per month"
https://en.wikipedia.org/wiki/Cosmic_ray#Effect_on_electroni...
But these weren't just cosmic-ray-induced errors. The DIMMs with errors were far more likely to have more errors in the future, pointing to faulty memory cells.
"About a third of all machines in the fleet experience at least one memory error per year"
"The median number of errors per year for those machines that experience at least one error ranges from 25 to 611."
The extra cost for ECC is basically noise in the cost of a desktop system. I buy unbuffered ECC DIMMs for my AMD systems and they're basically maybe 5-10% more than a non-ECC DIMM. The SSD premium is more than that. Registered DIMMs have a higher premium mostly because they need to have additional ICs on the DIMM to handle the higher memory capacities that servers tend to support.
I really wish CPU vendors would just flat out require ECC on every system shipped. Knowing that main memory is failing would save so much time speaking as someone who has encountered DIMM failures in otherwise stable systems. Days of debugging what turns out to be a hardware issue is no fun.
Where? Are you talking DDR4-3200? These are 40-50% higher than non-ECC UDIMMs. It usually gets worse as capacity increases. Additionally, there are only one or two models of 1x32GB UDIMM in production. These sticks are quite rare. You don't just go to Newegg or Amazon unless you want to pay out the arse for them. And in addition to that, ECC has worse timings which are important for Ryzen.
> Registered DIMMs have a higher premium
You have this backwards. RDIMM are common and cheaper because that's what servers generally use. UDIMM ECC cost more because, well, they barely exist as a thing.
I have a feeling the cost is not just a money-wise 5-10%.
The battle scars are as follows: - The 2MB expansion card on my Amiga 500 turned out to have a couple of bad bits, and because I was a green techie back then, I just attributed it to general software instability. Years later a RAM test showed there was a problem. Sadly there were no memory test utilities shipped with Amigas back then. - My first 486 system had a bad bank in the SRAM on the motherboard. Took years to figure out as it only really hit when heavy DMA occurred from a VLB SCSI card while the CPU was heavily loaded. Swapped the SRAM out and it was stable for a few more years of service. - One of my K6 systems had a DIMM go bad. At least it failed badly enough that the system couldn't boot. - I've had an Intel Xeon in a colocated server start throwing ECC errors in the L3 cache. Replaced and RMAed the CPU before it caused an outage. - One system started throwing ECC errors because the power supply was marginal.
Now, please, show me how the extra cost of ECC is worthless to you when things like this happen in the Real World. Is your time spent debugging hardware failures really not worth the cost of ECC?
> DDR4-3200 is an expensive boutique product.
You have to be fucking kidding me. On a desktop Ryzen?? Most desktop PC builders are gamers. That's just a fact. They drive the market. That's why every motherboard looks like a gamery stealth bomber. Outside of Asrock Rack and one or two Asus Pro boards, there is simply nothing out there that matches what the server world is seeing with SuperMicro.
I have 1x32GB DDR4-3200 UDIMM ECC on my Ryzen. But I have no delusions about the cost. It was expensive as fuck compared to non-ECC RAM.
edit: On a second look, Mushkin will actually sell you 3200 CL14 ECC memory... 14-18-18-38 that is. At almost 10 quid on the gigabyte. Meanwhile you can get real CL14 non-ECC memory at around half that.
BankGroupSwap: Disabled BankGroupSwapAlt: Enabled Memory Clock: 1800 MHz GDM: Enabled CR: 1T Tcl: 18 Tras: 39 Trcdrd: 22 Trcdwr: 22 Trc: 83 Trp: 22 Trrds: 5 Trrdl: 9 Trtp: 14 Tfaw: 38 Tcwl: 18 Twtrs: 5 Twtrl: 14 Twr: 26 Trdrddd: 4 Trdrdsd: 5 Trdrdsc: 1 Trdrdscl: 5 Twrwrdd: 6 Twrwrsd: 7 Twrwrsc: 1 Twrwrscl: 5 Twrrd: 3 Trdwr: 9 Tcke: 0 Trfc: 630 Trfc2: 468 Trfc4: 288
Yes, they are somehow really hard to get in the US for decent pricing, but in the EU they are quite affordable (according to geizhals.eu ).
I do question though how OC'ing non-ECC is a better idea than OC'ing ECC. Also, just for the record, I based my timings off of an XMP profile for a non-ECC version of the same die/capacity configuration (32G 2Rx8 w/ Micron Rev.E 16Gbit).
I also haven't seen 32G sticks that do better than CL22 at 3200 with just 1.2V.
How so? For example I can see when the higher summer temperatures becomes an issue and reduce my timings, then tighten them again as it gets colder. No more random crashes where I wonder if it's my RAM or something else.
Currently doing 3333MT/s CL14 on 4x16GB of 2666MT/s ECC. I use CL16 in the summer. Since the memory sticks barely are affected by the heat, and that ECC errors occur on the same slot of the motherboard regardless of which stick is in there, I assume its the memory controller on the CPU in combination with four sticks of memory that's the limiting factor here.
I mostly just did that test to check if I'd have to give up the 3600 speed if I'd upgrade from 64GB 128GB on this 5950X; I'm well aware that one shouldn't be running it single-channel, especially not if that would mean quad-rank operation.
The likelihood that a RAM error is corrupting data silently is slim, given the same corruption could affect the RAM storing the OS kernel or your application stack.
If you have bad non-ECC RAM, you will notice issues that warrant further investigation (i.e. memtest). But that assumes you're savvy enough to understand what failing memory looks like.
Being kernel developers, we were majorly concerned. There were lots of changes to filesystem, VM and other parts of the kernel, and we couldn't rule out that we had a bug where code was using a stray pointer or some other problem. The stress tests had successfully passed last release...
Ultimately we tracked it down to network data being corrupted. The shiny new ethernet switch was rarely flipping a bit in packets, then happily fixing up the packet's CRC and IP checksum so that the computers on either end of the network link had no idea that the data was mangled in the network fabric. Oh the hours of brain wracking pain that caused!
shudder
Furthermore, your comment reads like someone that has never spent weeks debugging issues that only happen under extreme load stress tests. A single marginal bit is not so simple to track down when memory allocation patterns are non-deterministic. I have no desire to go through that again if it can be avoided by spending an extra $20 per DIMM. My time is worth more than that. Hardware is not perfect, and anyone who trusts hardware blindly should really take a look at what's going on behind the curtain.
Fwiw, the nature of DRAM errors is different now compared to the studies published decades ago when cosmic rays would only disturb a single cell. These days minute physical manufacturing defects and electrical disturbances are the dominant failure modes. Educate yourself: read the papers published about row hammer. Hammering a row with reads can flip bits in adjacent rows, and even DRAMs that have supposed mitigation features can suffer from disturbances a row or two further away. It has already been proven that the DRAM vendors' hardware cannot be trusted! What more proof do you need that this is a real world requirement in 2021?
I'm going to claim it wasn't entirely my fault. ARCnet did automatic resends for you so if fixed what the CRC detected, and silently what it didn't detect through. I had no idea the errors were happening.
ECC is nice, but as you say I don't need fancy hardware to correct a malfunctioning RAM chip. I can do that perfectly well with a hammer. What I need is the hardware to tell when and where to deploy the hammer. Things were actually better 30 years ago - back then PC RAM had parity. It's cheap, fast and effective enough. I have no idea what went wrong.
your systems must all be running at sea level.
For the more common story of IBM enabling present but disabled memory, they couldn't magically teleport a new memory board into your mainframe. But they could ship your mainframe with extra, verified but disabled memory.
Customer UX for memory upgrades then changes from "Wait for it to ship, etc" to "We'll have it done tonight."
I used to run b2c teams which dealt with big banks. One bank's IT department used to bill internally for data center costs using RAM as a proxy for energy usage. A big spark cluster was purchased for the project I worked on and the minimum RAM for the spec of machine that was "standard" was 128MB, so to "save costs" they had the data center guy take out half the memory in every machine before racking it up. :-(
You bought a 60 kWh car. If it performs as advertised, you not entitled for more, no matter what hardware is inside. And if you find a way to exploit the full capacity of the battery, you can, it's yours, but don't count on Tesla to support it.
I have no problem with market segmentation. It allows more people to have access to a product while protecting the margins of the manufacturer. Because, without margins, there is no product.
And in fact, I am happy when I am traveling and the student next to me paid half of what I paid, I could afford the higher price ticket in exchange for more convenience, so no problem for me, but if the student had to pay the same price as I did, he wouldn't have been able to travel at all. I was like him years ago, and our children will probably be too.
It’s the same thing with spinning hard drives and SSDs. The reported capacity is almost always lower than the actual capacity. That allows dead blocks to be “replaced” without the user even knowing (“remapped sectors”). Could you imagine the PR disaster if your SSD lost visible capacity as time went on?
You do have a point about the weight. The cost of shipping an SSD with 0% overprovisioning and one with 25% is negligible; the ICs weigh just a few grams. However, a car that has to haul an extra 15 kWh of unusable capacity does use extra energy. I’d be curious how much is needed to haul those extra cells.
I understand in case of Tesla you don't like the idea of hauling those extra unused cells around. But you are buying a car model with a specific mass, acceleration and range. It makes no difference whether how exactly the producer fit into those parameters.
The batteries are backed by a warranty, which in turn is part of the cost of the car. It makes perfect sense to derate the batteries in order to lower warranty costs. Of course, we can debate whether the consumer ever saw the benefit of those lower warranty costs, but the core idea is sound enough.
As for "legal title," you have "legal title" to whatever the sales contract promised. If access to full battery capacity at the expense of service life wasn't in the contract, then you're not entitled to it.
That's definitely not how physical property works.
I mean, there are implicit guarantees that don't have to appear in the sales contract, such as the requirement that the car will meet the applicable pollution-control regulations and won't actively try to kill me, but that's obvious enough. Did you have something else in mind, particularly vis-a-vis the batteries?
You own the car. You own the battery. No sales contract changes that fact of ownership. They delivered a 75kwh battery into your possession. Since it's your battery, you should be able to use your battery however you want without having to pay Tesla. People are rightfully upset by a PR move that proves that the manufacturer is treating a car that was sold and delivered as if it still belongs to the manufacturer.
https://nyctomachia.wordpress.com/2018/09/02/riglol-unlock-e...
https://www.makermatrix.com/blog/hacking-the-siglent-1104x-e...
..etc.
I mean... this seems like a very good point, but from the rest of the article, and the previous one, it sounds like the patch was approved?
Why was this allowed into the kernel? looks like drivers/platform/x86/intel/* mostly. Does intel just "own" that section of code? who accepts these patches?
I'm not particularly familiar with this stuff so would appreciate someone filling me in on this.
Intel actually said, in a short snippet where The Register refers to them as "Chipzilla", "If we plan to implement these updates in future products we will provide a deeper explanation of how they are implemented at that time.". One interpretation could be they are saying they haven't figured out if they are going to implement this in anything so they just added it for fun... I'd call that the "makes a good article" interpretation. The other interpretation is they plan to use the feature but have not yet finished planning which products will use it to unlock what and will talk about it in more detailed once planned, I'd call that the "corporatespeak decoded version". If you were a literal interpretation bot I would still say the corporatespeak version fits better than the article filler version hence why the extrapolations with that version don't match up to what is happening.
> Why was this allowed into the kernel? looks like drivers/platform/x86/intel/* mostly. Does intel just "own" that section of code? who accepts these patches?
I didn't look if it was actually accepted yet or just a proposed patchset but generally Intel will be the maintainers for Intel code (but not always) and they are one of the larger contributors to the kernel overall. You can see more details about maintainers and reviewers (as well as maintainers for sections of code) here https://www.kernel.org/doc/linux/MAINTAINERS#:~:text=One%20j....
It seems that a lot of the changes in the patchset are in the intel pmt driver, which is maintained by the author of the patch, so that makes sense why the maintainer included it.
And you can search the mailing list[2] by commit hash/author name/etc and look at relevant discussion trees to see lively patch discussion and criticism/final pull requests/etc.
I don't believe the changes in question have been merged.
[1] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
It could run in slow-mode or fast-mode. Always-on fast-mode cost more per month than slow-mode. (Yes, these machines were leased, not sold.) However, you could get "fast-mode" on an as-needed/hourly-basis, for a "small" fee of course.
ACPI allows random x86 board vendors to literally have hundreds of products themselves that not only boot linux, but a pile of other OS's because they aren't wasting their time writing drivers for every single board and voltage regulator.
So, call it shitty, but understand that its the reason you can boot a single linux image on everything from an atomicPi, to HP superdome flex, and the thousands of desktops/etc being customized by people in their bedrooms.
Its also largely the reason why linux works on 20 year old PC's that no one is actually testing on anymore.
The alternative is a monoculture where the HW is basically provided by a single vendor that doesn't change much and provides a half dozen supported configurations. I would point to the zseries here because that is effectively how it works, but it also has a huge hw/firmware abstract machine that allows IBM to change the underlying HW details without having to rewrite all this low value garbage.
There's literally tens of thousands of devices for the PC platform and Linux has no problem supporting them if the hardware can be documented. Why can't the vendor document that instead of hiding this stuff behind opaque binary blobs that can have security vulnerabilities? Whether this is ARM or x86 doesn't matter.
> ACPI OTOH allows random x86 board vendors to literally have hundreds of products themselves that not only boot linux, but a pile of other OS's because they aren't wasting their time writing drivers for every single board and voltage regulator.
I'm parsing this as "ACPI OTOH allows random x86 board vendors to develop shitty hardware, cut corners on documenting it, and not care as long as Microsoft OS's boot on it."
> they aren't wasting their time writing drivers for every single board and voltage regulator.
Document the damn registers and hardware interfaces and then someone out there will write it for them.
> Its also largely the reason why linux works on 20 year old PC's that no one is actually testing on anymore.
Well Linux still supports hardware that predates ACPI, like floppy drives, so this is making a connection where there is none.
> The alternative is a monoculture where the HW is basically provided by a single vendor that doesn't change much and provides a half dozen supported configurations.
Wrong. Before ACPI, for example, there were many chipset vendors--Opti, VIA, etc. ACPI didn't kill these off but there wasn't a monoculture before ACPI.
> I would point to the zseries here
So ... do we want an IBM-mainframe like monoculture where you have to depend on IBM when the hardware changes?
> allows IBM to change the underlying HW details
The job of abstracting the hardware details is the operating system and operating system drivers. Literally that's 50% of the reason why you have an OS in the first place is to abstract I/O into things like open(), read(), write() or other interfaces that are portable because of underlying drivers. If the hardware is so different that existing I/O calls can't handle it, you need to develop new ones - that's what happened with Berkely sockets - NICs are not block or character devices.
If the drivers are open source then they are bug-correctable and useable even well after the hardware manufacturer goes away. Embedding that in closed-source platform-firmware makes you dependent on that platform manufacturer. A strong contributor to the Wintel monopoly.
Wrong. Before ACPI, for example, there were many chipset vendors--Opti, VIA, etc. ACPI didn't kill these off but there wasn't a monoculture before ACPI.
That is where your wrong, x86 PC's had BIOS and APM which also provided minimal platform abstractions. But there was a monoculture, you either provided PC/AT HW and BIOS compatibility or your x86 didn't work. There were HW "standards" for everything, be that CGA/EGA/VGA, or IDE controllers. Yes, you might make your own video card, but its absolutely supported those standards. Similarly with chipsets, there were closed source early boot firmware, but by the time the MBR was being loaded it looked like a 1980's PC/AT.As far as there being enough driver developers to further fill the kernel with all these drivers belies a fundamental misunderstanding of how complex even those arm machines are behind the scenes, even the ones that work likely have huge stacks of binary code running places that aren't visible to linux/DT. You need only look at some of the open or reverse engineered arm boards to see that. the RK3399 specs were published, what 4 years ago at this point, and there are still rk3399 fixes landing. Should it take 5+ years from the release of a piece of hardware before linux can boot and work reliably on it?
edit: See https://en.wikipedia.org/wiki/Option_ROM for where all that binary code used to hide in the 1990's when you plugged in random "vga" boards and storage controllers.
This was awesome when your platform consisted of an 8-bit CPU, a serial port or two, a printer port or two, a disk drive or two, and a text-based video display.
ROM and BIOS is a poor place for drivers if
- your system has any notion of plug-and-play at all
- your system has expansion slots and arbitrary people you don't control might develop hardware for it.
and these two things above are desirable if you want a free-ish computing platform not monopolized by one company.
Hardware interfaces are not the same as the firmware gunk ACPI foists upon you. Hardware interfaces are simply a way for the CPU to talk to a device outside of the system, but it doesn't result in the CPU running unknown code behind your back. A CPU that's cordoned off behind a peripheral interface running closed-source code is fine - where that's not fine is on the same CPU that's running my OS kernel.
> Should it take 5+ years from the release of a piece of hardware before linux can boot and work reliably on it?
I mean if the hardware manufacturer won't document their devices it's something they bring upon themselves. Embracing the open source community here would have substantial benefits unless something like market segmentation is taking place.
Option ROMs? Yeah, those are BIOS extensions - your OS isn't dealing with those ROMs once the BIOS has booted the OS unless it's a CP/M era operating system - like actual DOS. Except for the modesetting - but that's just as much bullshit as ACPI. Document the damn registers that do the modesetting change so we don't have to thunk back into 16-bit mode just to change the screen resolution.
So, all that said, your mental model of how a modern machine works, seems like its stuck with the idea that linux is the center of the machine and can access all the hardware, and that the HW looks like a 1980's PC with "registers" that actually modify HW states. Which is provably false, and will continue to be that way as long as people want inexpensive and/or high performance machines. The mainframe guys go on about channel processors, but that concept (using a small cpu/etc to manage a piece of hw or communication) is fundamental to a very large percentage of HW produced in the past couple decades. There is code buried in pretty much every single USB device to manage the bus, as is true for nearly every storage device where those microcontrollers manage everything from queue scheduling, signal processing, etc on spinning media, to the flash translation layers and error correction on SSDs. Then, there is all the power mgmt code running to control internal bus power/frequency, as well as cache power mgmt, etc, etc, etc. Even things you probably think are simple register model HW devices (say XHCI or NVMe) have microcontrollers buried in them actually driving the bus and maintaining connections. Overwhelmingly what you think of as HW registers are actually mailbox interfaces to micro controllers running proprietary firmware.
So, as I mentioned pretty much the entire HW docs for the RK3399 were released years ago, and that SoC and the dozens of boards you can find it on, still in general won't work out of the box with a random upstream linux kernel on any random board you can find. And when it "works" the power mgmt tends to be terrible. That is because it actually takes engineering time to make these things work well, someone has to write the drivers, device tree's, etc and go through the pain of getting them merged to mainline. And its very obvious that while sometimes there are people willing to spend their holidays and weekends making it work, those people are few and far between. And that's just for a few pieces of HW, if there were hundreds of manufactures making variations on say the pinebook pro, pretty much none of them would work outside of the hacked up debian/whatever that the manufacture shipped on the device.
And then, say linux actually works well on it, what happens if you want to run netbsd?
That's because the BIOS is not an operating system--or at least used to not be. And it goes back to the CP/M architecture that's even older than the 1980's PC. Option ROMs are there for the BIOS and single-tasking "operating systems" that use it CP/M style. Modern operating systems don't need the layer there.
I have a Guruplug (ARM platform) that boots Linux and what's in flash is U-Boot - a bootloader that loads Linux, the initrd, then it gets out of the way. The only reason why the PC platform can't work like this is because of ACPI.
> Overwhelmingly what you think of as HW registers are actually mailbox interfaces to micro controllers running proprietary firmware.
Oh I know. SATA is a communications interface (as was IDE, ATAPI, SCSI). USB is a communications interface. NVMe is a elaborate tagged/queued communications interface. Et cetera.
I'm not sure why you are conflating CPU-facing hardware interface (registers) with anything that is physically peripheral side stuff except for this which I will address:
I know a CPU I don't control and don't know the code it's running is on the other side of those links. That's fine--because it can't directly access my OS's RAM unless the OS allows DMA - and it's probably going to cause a device-level issue instead of a machine-level issue if there's a bug in that firmware.
Not seeing the value add--other than saving some poor, poor Microsoft-aligned platform firmware developer's bit of time--of having to jump to a closed-source firmware routine to use those communications interfaces instead of letting the OS directly talk to them.
> So, as I mentioned pretty much the entire HW docs for the RK3399 were released years ago, ...
So what makes the PC platform avoid this mess is not the presence or absence of closed-source firmware running on the main CPUs, but simply the hardware platform itself being standardized. This is because people copied it from IBM. The BIOS was copied because DOS needed it, not because the hardware needed it other than a CPU requires ROM at it's initial boot address. There were well known addresses for each device, such as DMA channel 0/1, IDE channel 0/1, FDC 0/1, serial port 0/1/2/3, paralell port 0/1/2/3. I really want to know why we couldn't have a hardware standard for laptop power control interface instead of APM. The industry managed to settle on PCI in reaction to IBM's attempt to grab back the platform with MCA. PCI doesn't requre firmware, and it's registers to scan the bus are well known and standardized in hardware. So why did everyone say it was OK to hide the power-controlling hardware behind APM? DOS of course needed something like that but you had more than a couple commercial operating systems on the platform besides that (Xenix I think, OS/2 likely still kicking a bit, NT).
Early 90's when APM started taking hold was also the time when Intel started to not publish certain things in its CPU manuals.
Device trees should be easily obtainable from device datasheets.
Overall the real problem with the ARM boards is an economy that values time to market and treat-your-first-X-customers-as-beta-testers over quality. Separate issue that we really should be unwilling to accept as an absolute requirement for unauditable code running on the same CPU my operating system is running on.
I have a guruplug too, or rather an openrd, because all the guruplugs died. But that is 1980's PC level of simple HW, and its not SMP, doesn't support virtualization, or any power mgmt to speak of. The list of things it can't do is longer than the list of things you can do with it compared to a modern piece of HW.
Modern PC's with ACPI aren't "standardized" outside of a few interfaces being used by the OS, that is my point. All those regulator drivers, clock controllers, pin muxing, I2C'ing, SPI'ing, to manage the platform is whatever the vendor put there in PC land (and arm land for that matter) is still there, only the OS doesn't have to worry about those details because it simply asks to power something on, and it happens, and some other processor takes over picking perf profiles and idling links/etc when needed. If there was a standard PMIC, wired up in a standard way it might make sense to attempt to use it, but there isn't. Instead its a bucket of parts wired up in every way imaginable.
ACPI, doesn't dictate what is on the other side, it could be a SMM trap, or it could be a mgmt processor, or a BMC, or the function can be completely written in AML. That is the point, its just an OS API surface, it could just as well be a standardized pile of HW registers but that would remove the ability to run in cases where the vendor wants to save a penny and run everything on the main core as you suggest doing with DT. A large part of why you provide these interfaces is because having the main core waking up to fittle with some 100Khz SPI bus to flash a led is dumb, wastes power, and eats perf doing work that would be better handled with a small mgmt core. Your basic argument seems to be that you don't want the OS wasting cycles in the firmware, but your perfectly ok with the OS wasting even more cycles all the time doing these functions sub optimally in a kernel that doesn't understand the intricacies of power management on the most pwoer hungry core in the system. You seem to think ACPI is always just a SMM trap, and that isn't the case. The reason I conflate is its because in the case where the remote is a BMC/etc your effectively just poking a mailbox to get some other piece of hardware to do the work and those mailbox interfaces are no more standard than PMICs, so now you have thousands of them needed to talk to battery mgmt controllers, power/standby buttons, you name it. Instead of standardizing all that garbage, a software API was created. Its no different from openGL or any other standard except that it uses a bytecoded function call interface (ala openfirmware's forth). Do you hate openGL too because your game isn't fittling fake HW registers?
And now that I point out what happens when a vendor just tosses those register maps you want over the wall you change the topic to how they are just creating beta level products while refusing to acknowledge that someone has to do the work. From the perspective of a vendor it costs a tiny fraction to have some closed source firmware that works across dozens of OS's and allows them to redesign their hardware vs hiring experts in a half dozen OSs to write piles of custom power mgmt code accessing dozens of drivers talking on dozens of SPI/I2C/mailbox/etc interfaces for a single machine. Its actually a little bit crazy that people in the arm space are trying this while simultaneously complaining about the HW vendors creating piles of patches that they can't get upstream so they build their own custom linux forks. Its the natural result of making the same claims your making. A vendor can either spend years fighting with kernel maintainers, or they can fork linux and ship it to their customers with a pile of patches that only work in linux. The middle ground is where a number of them have been going, its the rpi's proprietary mailbox interface that talks to the videocore to set the processor frequency. Multiply that mailbox by a few dozen SoC providers and you have the future of DT on arm and risc-v. In a decade or two they will end up standardizing the mailboxs and reinventing ACPI and openfirmware in order to support cross platform mailbox interfaces.
Intel Software Defined Silicon: additional CPU features after license activation
TPM and SecureBoot came directly from large customers requesting exactly that.
HN (and more broadly, PC users) aren't all of the market for CPUs.
I do not really believe this is their intent thought: people are way too much continuing to act like mean 5% better perf on arbitrary benchmarks has any significance (even regardless of power consumption). So Intel will likely continue to ship broken but allegedly fast processors, and will continue to kill the perfs to remove the bugs once they sold enough.
Plus, in a competitive market, it would be an equilibrium hard to achieve if competitors are not making pay for their own deferred enablement.
You seem to be terribly confused. End customers demand features, they do not demand limitations.
AWS and GCP don't want to spend money acquiring or powering chips with dead areas of silicon.
If there's unused silicon in a chip, but it costs the same or cheaper and has the features they want, I think you'd be hard pressed to find a customer who's going to complain.
At the end of the day, I don't understand the umbrage at companies segmenting by software instead of hardware. As long as it's obvious what you're getting when you buy, and what you buy lives up to what you were promised, why should I care what it actually is?
If the company chooses to look the other way and allow relatively easy unlocks without promising stability, well, that's nice. But nothing I'm owed just because it was physically on the chip.
Now on the other hand, if a consumer company switches from a product model to a lease/HaaS model... that's an entirely different can of worms, and they can go to hell.
It's a relatively niche interest, but it exists.
Hmmm. Intel has built whole ISAs that no one asked for and didn’t want.
Some good ideas work. Some good ideas don't. Doesn't mean they were bad ideas, only that unknown at the time factors or future developments made them bad ideas.
Not sure how the rest of your comment relates to my point though. Customers may want better performance. Doesn’t mean they want your new architecture. They certainly didn’t want IA-64. That’s just a statement of fact.
And as for the link between architecture and performance, I'll turn the common quip: "You can have performance or architecture changes, pick two."
The debate between runtime parallelism vs compiler parallelism, and which would result in greater performance on real world workloads, was an open question at the time.
As it turned out, the market preferred answer was "Screw it, we'll push superscalar and add more cores." But that's non-obvious in foresight. See: the famous P68/NetBurst/Pentium 4 vs P6+/Pentium M struggles.
I was really (semi humorously) trying to push back against the parent comment saying that everything these firms have done has arisen directly from customer demand. Obviously firms must think that there is demand for new products or features but sometimes they just misunderstand the market.
To give another example was there really any demand from customers for x86 cores in smartphones or was it Intel just trying to establish a market presence?
To answer your question iAPX432.
But in the early-80s, Symbolics was making a lot of noise with Lisp machines, the late-80s AI winter hadn't yet set in, and there were certainly worse bets than CISCy "we need to integrate up the stack" ideas.
x86 on mobile is a hard one. I see why Intel did it: it's what they had. And I'm sure Microsoft was whispering some demand in their ear behind the scenes.
But it really only made strategic sense pre-App Store volume (so say pre-2010?). Once the mass of code existed on ARM and built for iOS, that genie wasn't going back in the bottle.
As far as I've heard it told, that was more of a financial business decision though. Intel wasn't willing to cut its margins, because it couldn't see trading volume for margin as mobile exploded.
If they'd been able to offer the market cheap, sufficiently performant, power efficient chips for mobile, I think present would look a lot different.
The i860. High peak performance on paper, realizable in approximately zero real-world code.
Secure boot was mostly just a toy until Google implemented it in Chromebooks and the large customers you speak of pointed and said "I want that!"
It's the opposite of being capable - it's silicon that refuses to play to it's full capacity lest you buy an additional license.
It's like we are going backward in progress.
https://www.pcgamesn.com/nvidia/CMP-170HX-price-double-RTX-3...
It made them a lot of money, but they (rightly so) look at it as margin that could evaporate tomorrow. And in the mean time pisses their core users off (through volume availability issues) and makes an already difficult production -> retail supply chain even more bursty.
So how do you deal with that? Well, a not-terrible way is artificially segmenting your market. Crypto folks, over here in this pool with huge waves. Everyone else, back to business as usual.
Miners would probably find a way to use that new entrant's GPUs for mining too.
I think the real threat is that mining could chill the gaming industry as a whole. Either by pricing tons of people out of gaming so they find new hobbies and maybe never come back, or by influencing the direction of game development away from advanced graphics. If most gamers can't afford top-end GPUs anymore, then the perceived value of [Latest Game] having the most photorealistic graphics yet could be flipped into the negative by a change in gamer culture.
It's not a coincidence the software lockout landed right as Nvidia released an unlocked mining card without a display output ( https://www.nvidia.com/en-us/cmp/ ). And since it doesn't have a display output, all those used mining cards become ewaste instead of something that cuts into nvidia's revenue.
No, they are trying to piss off the miners so that their actual customers (gamers) have a chance of getting GPUs again.
It's like the first Bitcoin craze many years ago when the first CUDA miners came out - you could not find any GPU anywhere because people bought them for mining. Gamers were pissed as the miners outbid them everywhere. Once ASICs came out, the complete GPU market crashed - suddenly there was a massive supply of second-hand GPUs from miners... the gamers snatched these up, and NVIDIA suddenly had the problem that their new GPU sales crashed.
The only difference is that this time it isn't Bitcoin anymore but a whole bunch of random shitcoins.
Somehow nobody shed a tear for the poor enterprise customers until miners came along. But that's what it is, miners are enterprise customers, they use them to make a profit.
Similarly, almost all R5 consumer processors are made by artificially disabling fully functional cores on R7 chiplets. AMD already had nearly 80% of the chips coming off the line with 8 functional cores at Zen2 launch, and numbers have only gotten better since there. Yeah, maybe a few fail clock bins or other things, but those bins are chosen very loosely such that they're not throwing away any significant amount of chips anyway. The overwhelming majority of R5s are perfectly good chips being gimped and sold as lower chips, because the availability and demand are exactly opposite of each other - they have the highest availability of 8-core chiplets and the highest demand for 4- and 6-core chiplets. How do you square the two? Not by lowering prices of your high-margin products - you keep those high to extract the maximum price from price-insensitive consumers, and you artificially gimp a bunch of them for the rest of the market to protect your profits in the higher segment.
https://en.wikipedia.org/wiki/Price_discrimination
The thing to remember is that home users are almost always the beneficiary of market segmentation, because we are the most price sensitive. The alternative to Core-vs-Xeon-E segmentation isn't that everyone gets Xeons at i5 pricing, it's that the i5 now costs almost as much as a Xeon. The R&D has to be paid back one way or another, or else R&D will be significantly reduced (most likely meaning much longer product generations and slower moves to new nodes, etc). Price discrimination means enterprise customers pay more than their share and home customers pay less than their share, so if market segmentation goes away and those equalize then home customers will have to bear more of their fair share of the product cost.
Again, who wants to race to pay more for their gaming graphics card so that Boeing and cryptominers can pay less? It would be great if everything were free, or priced at its true cost, but in the real world R&D needs to be paid for somehow.
Oh, and the other thing is Enterprise customers actually love this, it's called Capacity on Demand and I'm sure the reason Intel is doing this is because their customers keep asking for it. Sun, Oracle, Fujitsu, IBM - all the big enteprise vendors will sell you a machine with a bunch of cpu and memory that you're not allowed to use or touch! But if your workload scales up, instead of having to shut down your mainframe to upgrade it, you can just pay money and turn the cores on! And sure, it would be nice if all software were just microservices that you could re-instantiate onto another machine, but some stuff (like databases) can't be trivially scaled and partitioning (read: stale data reads) can't be tolerated. The only surefire answer to CAP theorem is One Big Machine such that you never have to deal with splits in the first plac, and that's what some situations need, and they love this.
Someone cracked it and enabled them with custom driver. (And of course nvidia patched it out in the latest vbios to make the bypass impossible).
There is the story of this on HN few month a go.
This was never popular with the Free Software crowd, who believe in the principle that the owner of the hardware has the right to do whatever they like with it. They would say you have the right to flip the secret switch yourself.
This is the same idea, but updated for the 21st century. My laptop came with hardware support for virtual machines - In a few years time, perhaps features like hardware-accelerated virtual machine support will be separately licensed premium features, instead of being stock features.
Some companies already do this - for example, some Teslas have autopilot and long range batteries, but they're disabled in software. And if you buy a Tesla with those features second-hand, Tesla can disable them until you pay them.
Everything I've seen says this is not the case. As I understand it, the story in the news about "something like this" was misleading. The car went through a dealer and was sold certified pre-owned; and it was sold as not having those features. It turns out they forgot the disable those features. I you buy one from a third party that already paid for the feature, it stays with the car.
At least, that's the understanding I was left with.
Was the car sold to its original owner with those features enabled? If so, it's effectively the same as the original story reported, meaning that Tesla expected to be paid twice for the feature.
- Car was sold to original owner, who enabled the features
- Car was sold back to dealer/Tesla (presumably a trade-in?)
- Car was sold to second owner, with information that it did not include the features
- Turns out the car had the features enabled (they mistakenly forgot to turn them off)
- Features remotely disabled
There's nothing inherently wrong with Tesla charging both the original and second buyer for the features, as those features make the car worth more. Conceptually, its the same as paying more for a car with a better entertainment system, then the person who buys it next (from the dealer after you trade it in) pays more for it than they would for one without that system.
> - Car was sold back to dealer/Tesla (presumably a trade-in?)
a used-car dealer bought the car from Tesla and was given (legally required) documentation about the car, which indicated the car had the features. Then Tesla removed the features remotely while the car was owned by the dealer, without informing them (because they decided during an audit that it shouldn't have had them). The dealer then showed the documentation to the customer (as required) and sold the car, both sides in the understanding it had the features as documented (because why wouldn't it, if Tesla documented it as such). The final customer then found the features disabled (or being disabled on an update/sync).
So what prevented you from finding and flipping the secret switch yourself? Or is it that merely being a business you couldn't afford to take this risk and lose IBM support?
You don't want to get letters from Oracle, IBM, Intel and Co. because you infringed some of their EULAs. They may be invalid, but it will be a very costly decade long court case. And the other part will pump thousands of times more money into it than you can afford, because when they loose they will loose a big income stream in the future.
You mean, completely anonymously, and without any association of a machine to an owner?
I install software through `adb install software.apk`, or through `make`, or similar, of packages stored on drive...
I am not sure this can be the model, for this new "unlocking hardware features" scheme. Many will just say, "No "thank" you" - and I am not sure about all the practical consequences. I am even more concerned that this facilitates preparing the way for taking for granted some practice like "register after purchase to be enabled for using the product", which I am informed some already attempted.
It’d need measuring the hash of the loaded firmware itself to hardware managed PCRs that affect the hardware crypto engines. Something that’s done on the Secure Enclave for Apple A13 onwards AFAIK.
(also… managing DRM systems, they’ll have to be put on a coprocessor somewhere.)
It's really stupid that other vendors can't just do the same.
I assume you meant to be a help by assuring people their computers aren't backdoored, but I don't think you have the certainty to claim ME is sound and secure. It just reminds me of how well-meaning people insisted the USG wasn't interested in capturing everyone's communications, back before it was accepted fact, pre-Snowden.
https://www.youtube.com/watch?v=jmTwlEh8L7g This is one of many amazing videos this guy did before they hired him.
Perhaps funny, but inevitable due to the lack of knowledge. Why don't you do us a solid and get us some documentation for our hardware using a cell camera?
...and hardware manufacturers make money selling hardware, so why keep the details of how to use it secret? In the days of SDRAM and DDR(1) a lot of companies, including Intel, were far more open about documentation than today.
I've analysed the init code from BIOSes and it is indeed not that complicated.
There are also a lot of businesses that act as patent trolls , e.g. RAMBUS and stifle memory innovation by suing everyone left and right and enforcing their patents. Part of licensing agreement is to keep stuff secret and proprietary. Otherwise if you give something of value away for free, you can’t make money off of it anymore , can you?
This smells a lot like laying the groundwork for locking processor features behind a license key. Get ready for a monthly CPU license plan...
MPEG 2 has been patent free since 2018 in US and 2020 worldwide. Along with AAC-LC. They are charging for license key for the usage of that specific decoder block only. i.e The IP of the actual HW decoder, not the codec itself.
MPEG 2 and VC1 is so low in complexity you are software decode it with minimal CPU resource on RPi 4.
The patents were still active, and the technology was still useful, when the Pi Foundation started selling the license keys in 2012. Nowadays... yeah, they're irrelevant.
On the other hand if I have no guarantee the license is valid forever and I own it, or that I can use it in any conditions not just with internet connectivity, or it's tied to the rest of the machine, or only until I want to sell my hardware, then it does become fundamentally worse.
This might be a new way to enable more flexibility for cloud providers so they can customize each offering without having different sets of hardware, just a license management component. I also see this as something OEMs may be interested in, tiering otherwise identical products could be just a FW switch away, greatly simplifying logistics and manufacturing. OEMs could customize the CPUs they receive while Intel would shrink the number of SKUs. Being Intel though means none of the savings will ever reach the consumer.
If they can sell another 16 cores at no incremental cost post sale, that's attractive. As long as it doesn't prevent that customer from buying a whole new 48 core chip from them, anyway.
The only reason why we tolerate such clear signs of market failure is because in the larger picture, Intel et al. need to keep their margins high and consistent to keep Moore's Law alive(ish). Maintaining such a fast pace of progress is probably better for customers in the long run, even if it means getting ripped off in the short term. But without such a benefit to make the tradeoff worthwhile, customers and economic policymakers would be right to be incensed when Intel uses wasteful tactics like disabling defect-free silicon in order to insulate their price sheet from the true supply and demand dynamics of their market.
Let's not forget our history here. I'm not a big fan of Intel but any techie will remember we successfully "tolerated" such practices decades ago when the Celeron 300A ran at double frequency, the Radeon X800 series got 50-100% more pixel shaders, or the Phenom X3s got 33% more cores for free. If anything customers were rejoicing.
The market was always artificially segmented. The only reason we don't know today whether disabled blocks are good or not is that manufacturers have gotten better at keeping them disabled.
You'd be hard pressed to tell me the actual real life difference for you between a $300 4-core CPU with 2 cores fused off, and a $300 4-core CPU with 2 cores licensed off.
> Intel uses wasteful tactics
Intel fashion dictates the user will see little benefit but differentiating SKUs in software is certainly less wasteful from the manufacturer's and OEM's perspective. The logistics alone might be worth it.
Software manufacturers ship full installers with features disabled in licensing. Streaming services store all content but serve you a lower quality one based on your particular "license". Car manufacturers sell essentially the same engine with different power outputs.
As I said, there are no bad products, just bad prices.
The fact that artificially crippling chips is not new does not mean it's a good thing. Having accepted a tradeoff doesn't erase its downsides.
Why?
If their yield improves, should lower core count chips become scarce and then approach the cost per core of higher end chips? That's not exactly in the interest of a consumer anyway.
Also, they absolutely already do that today - all chip manufacturers make viable cores inaccessible so they can have the right part mix that is selling the way they want.
If your demand curve perfectly matches your supply/yield curve, great - but that's rare over time.
As yields improve, consumers should expect to see changes other than dwindling stocks of low core count parts. High core count parts should be getting cheaper, or low core count parts start showing up with higher speed bins or lower voltage bins.
Chips have been known to have perfectly functioning blocks disabled, or frequency limited not for improved yields but simply because the demand for a lower end SKUs was so high. Unlocking them was a matter of penciling some fuses, flashing a firmware, or simply changing a setting in the BIOS.
If anything this kind of licensing opens the door to "hardware piracy" where under the correct conditions you may be able to hack older, or unsupported hardware to gain the full functionality.
You have to buy some stupid dongle to enable "Intel® Virtual RAID on CPU" feature, and it's like $100 or something silly.
HN cracks me up.