Whether it is the unfortunate materialization of Spectre-style bugs or the deliberately insecure-by-design ME, Intel's inability to support its products is dismaying.
Whether it is the unfortunate materialization of Spectre-style bugs or the deliberately insecure-by-design ME, Intel's inability to support its products is dismaying.
We have some 7 year old Dell servers that are still chugging along, performing their duties as well as ever.
The main motivation for upgrade is software support (usually for the OS, driven by Microsoft), or failure rates for the older systems.
And that's for the desktop side. For servers they tend to be taken out of commission when the service they provide is migrated to a whole new platform and the legacy one goes to a better place along with the metal. Servers are more reliable than your regular desktops, they get better support, and most companies don't upgrade running systems.
P.S. I don't know of any enterprise environment where ME is used for managing servers. It's usually ILO, ILOM, DRAC, IMM, RMM. I think ME is mostly for desktops or SOHO, Intel has RMM for servers.
Edit: a quick search yields a common core or division at Intel behind the ME and SPS, so that makes a bit of sense. There has been at least 1 exploit in the wild for that SPS, it also lists TXT and ME, so I guess it's a shared (MINIX?) kernel that had the bug.
Or when the service contract expires or is too expensive to extend. You can't run a server of any importance without a service contract; it can be the difference between all the server's users and services being down for hours or a more than a week, and between IT management keeping their job for hours or longer.
It's hard to imagine not being able to find a replacement in more than a week, especially if one skipped service contract and just bought a spare or two with a fraction of the savings.
It's why moving to cloud infrastructure can be so much cheaper for these shops.
Of course, if it is a trivial size, like a single server at a small business, that's a different story.
> makes no sense for a commodity hardware installation beyond a trivial size
It's not the commodity hardware - x86 servers have been mostly commodity hardware for decades - it's virtualization (or other rapid recovery and migration tech) that makes it work.
My assertion is that service contracts for hardware have zero value for commodity hardware installations (beyond trivial size). To whit, they are, invariably, a scam.
Server hardware fails vanishingly rarely, with specific, notable exceptions of certain components. Those exceptions have predictable [1] failure rates that are therefore straightforward to budget and/or engineer around. Managers don't know to insist on this, so the scam persists.
This was true even before the advent/popularity of virtualization (and therefore "cloud" techniques).
It's also true regardless of whether or not one keeps spares on hand. With commodity hardware, vendors always have plenty of spares.. one just hasn't purchased them in advance, so there's an increased latency (not entirely unlike the spares available under a service contract).
> It's not the commodity hardware - x86 servers have been mostly commodity hardware for decades - it's virtualization (or other rapid recovery and migration tech) that makes it work.
That's where I disagree, at least partly, if I understand correctly the kind of tech you're referring to.
Those software solutions are generally just about convenience, or, ideally, reducing recovery time in the case of a non-redundant architecture. Even then, they might turn hours into minutes, not weeks into minutes.
More importantly, it's not that software that makes this work. It's the commodity hardware (and, arguably, commodity software/firmware hiding within) that makes it work. If a disk fails, any same or larger [1] can be used as a replacement. Same for RAM or even full servers. It doesn't have to be the same brand because that's the point of it being a commodity. That fact completely eliminates any problem of being down for a week due to hardware failure. In a major tech hub city, getting something delivered same day might not even require paying a premium.
Before virtualization or any similar abstraction layer, one could just move disks to a new server (best case) or restore from a backup (worst case).
[1] Other than "black swan" events like the flooding in Thailand that wrecked predictability (and reliability overall) for a generation of HDDs. Even then, if one made assumptions based on warranty length changes, those would have been close enough.
[2] Well, OK, one does have to be careful to meet minimum performance, especially for SSDs, but even a failure there won't kill basic functionality
https://aws.amazon.com/blogs/aws/ec2-instance-history/
I didn’t have one, but it seemed like a worthwhile thing to have and so I spent a few minutes putting the following list together (these are all announcement dates):
August 2006 – m1.small.October 2007 – m1.large, m1.xlarge.
May 2008 – c1.medium, c1.xlarge.
October 2009 – m2.2xlarge, m2.4xlarge.
February 2010 – m2.xlarge.
July 2010 – cc1.4xlarge.
September 2010 – t1.micro.
November 2010 – cg1.4xlarge.
November 2011 – cc2.8xlarge.
March 2012 – m1.medium.
July 2012 – hi1.4xlarge.
October 2012 – m3.xlarge, m3.2xlarge.
December 2012 – hs1.8xlarge.
January 2013 – cr1.8xlarge.
November 2013 – c3.large, c3.xlarge, c3.2xlarge, c3.4xlarge, c3.8xlarge.
November 2013 – g2.2xlarge.
December 2013 – i2.xlarge, i2.2xlarge, i2.4xlarge, i2.8xlarge.
January 2014 – m3.medium, m3.large.
April 2014 – r3.large, r3.xlarge, r3.2xlarge, r3.4xlarge, r3.8xlarge.
July 2014 – t2.micro, t2.small, t2.medium.
January 2015 – c4.large, c4.xlarge, c4.2xlarge, c4.4xlarge, c4.8xlarge.
March 2015 – d2.xlarge, d2.2xlarge, d2.4xlarge, d2.8xlarge.
April 2015 – g2.8xlarge.
June 2015 – t2.large.
June 2015 – m4.large, m4.xlarge, m4.2xlarge, m4.4xlarge, m4.10xlarge.
December 2015 – t2.nano.
May 2016 – x1.32xlarge.
September 2016 – m4.16xlarge.
September 2016 – p2.xlarge, p2.8xlarge, p2.16xlarge. October 2016 – x1.16xlarge.
November 2016 – f1.2xlarge, f1.16xlarge.
November 2016 – r4.large, r4.xlarge, r4.2xlarge, r4.4xlarge, r4.8xlarge, r4.16xlarge.
November 2016 – t2.xlarge, t2.2xlarge.
November 2016 – i3.large, i3.xlarge, i3.2xlarge, i3.4xlarge, i3.8xlarge, i3.16xlarge.
November 2016 – c5.large, c5.xlarge, c5.2xlarge, c5.4xlarge, c5.8xlarge, c5.16xlarge.
July 2017 – g3.4xlarge, g3.8xlarge, g3.16xlarge. September 2017 – x1e.32xlarge.
October 2017 – p3.2xlarge, p3.8xlarge, p3.16xlarge.
November 2017 – x1e.xlarge, x1e.2xlarge, x1e.4xlarge, x1e.8xlarge, x1e.16xlarge.
November 2017 – m5.large, m5.xlarge, m5.2xlarge, m5.4xlarge, m5.12xlarge, m5.24xlarge.
November 2017 – h1.2xlarge, h1.4xlarge, h1.8xlarge, h1.16xlarge.
November 2017 – i3.metal.
June 2018 – m5d.large, m5d.xlarge, m5d.2xlarge, m5d.4xlarge, m5d.12xlarge, m5d.24xlarge.
July 2018 – z1d.large, z1d.xlarge, z1d.2xlarge, z1d.3xlarge, z1d.6xlarge, z1d.12xlarge, z1d.metal, r5.large, r5.xlarge, r5.2xlarge, r5.4xlarge, r5.12xlarge, r5.metal, r5.24xlarge, r5d.large, r5d.xlarge, r5d.2xlarge, r5d.4xlarge, r5d.12xlarge, r5d.24xlarge, r5d.metal.
Not only can the price:performance spread vary between generations, but this can change over time, particularly because the model availability within a generation broadens over time.
Add to this the dimension of low power versions of certain processor models (whose selection is therefore strictly a cost/longevity decision, presumably invisible to someone like a cloud end user), one can't safely generalize.
The other problem is that the CPU isn't even, necessarily, the majority of the purchase cost of a server.
Specifically as it relates to the resources required to provide the same compute power for any given pool of instance capacity.
I know for a first-hand fact that AWS will not reveal its compute capacity, but in conversations with them, when we were actively monitoring the availability and continual use of all spot prices across every region, we had the data, according to them, to infer what their capacities were.
We had the data at the time, but not the interest/need, to be able to surmise ballpark compute capacity across all instance types in the spot pools - which is to say, what AWS' spare (i.e. non-fully reserved/dedicated) capacity was, and what the demand was for it. Additionally - we could have also bee publishing AWS' hourly revenues, dammit - I wish I would have thought of that - as people would have been interested in that data...
But we still don't, so we still can't infer anything at all from the announcement schedule.
Moreover, assuming AWS is growing fast enough, new instance type announcements could be entirely decoupled from deprecations. Even with both sets of data, it may not tell us anything about overall replacement rate, without also knowing the overall growth rate.
Even AWS's current mainstream "M4" offering is using a CPU from 4 years ago.
But with x86 being basically commodity and virtualization being used everywhere this is less of an issue in most cases. It means your VMs can mostly run on whatever metal you throw under them and that x86 metal can just be propped up to keep working for many, many years.
And TBH replacing servers after 2 years because you're worried about uptime feels like a horrible overreaction and self harming at the same time. Servers that are meant to provide 99.999% uptime (so 5 nines or above, or maximum 6min downtime per year) are built to run for far longer than 2 years. And they must be supported by the manufacturer of the system and the manufacturer of every sub-component for longer than that.
So unless I'm missing something I really can't see the benefit of such upgrade cycles.
I agree. Such a tactic, in the face of modern high-availability systems design (including the notion of servers being "cattle not pets"), seems actively harmful, other than, perhaps, introducing a something akin to Netflix's "chaos monkey" into the system.
I could understand pre-emptive replacement of non-hot-swappable components, but only if they're known to degrade over time [1] and only if the server is already otherwise out of service. Even then, I'd consider 3 years the minimum.
> Servers that are meant to provide 99.999% uptime (so 5 nines or above, or maximum 6min downtime per year) are built to run for far longer than 2 years.
This kind of design sounds like it's from a different "world" (mainframes) or era (proprietary, even if x86-based, Unix hardware of the 90s), not commodity x86 servers.
[1] so, maybe RAM, which has increasing CEs with age/usage, though UEs appear to be skewed toward early RAM life and therefore the bad apples are eliminated early. non-pluggable PSUs. fans. very old internal-only HDDs.
But even going some orders of magnitude lower about, 85% of x86 servers achieve at least 99.95% availability according to statistics [1].
[1] https://www.statista.com/statistics/515441/worldwide-survey-...
> Share of servers with four or more hours of unplanned downtime due to hardware flaws worldwide in 2017/2018, by hardware platform
Systems like IBM's System Z had 0 servers with more than 4 hours unplanned downtime over this period. At the other end Cisco and Fujitsu x86 servers had ~16% of servers experiencing 4 or more hours of unplanned downtime over the period.
We can infer that the vast majority of the x86 machines will be Intel considering AMD's market share in the server space.
P.S. The article is about Intel's ME flaws and patches so probably our whole discussion branched off in the wrong direction :).
Maybe that's why I was confused? I thought we were talking about commodity, x86 servers all along, particularly the GGP about "where uptime really matters" and other comments by the same person about warranty lengths (which seemed irrelevant, other than being indicative of commodity hardware).
Of course, that commenter's very terse remarks, combining "safety", "uptime", and "insurance" (and, later, billions in damages) into an environment where commodity (or not?) x86 servers are proactively replaced started me off in a state of confusion, like there was some very major assumption about the system design that I (and everyone else here, too, it seems) was missing.
I'm not sure where safety or insurance comes into this discussion at all.
People upgrade sooner because they stand to gain more than the upgrade costs not because its not safe to operate.
Your cpu is probably fit for 10-20 years of service. Every other part of your computer will fall down around it first.
But the XP example you gave kind of undermines the point. Nobody should be using XP. Not even on air-gapped networks, totally cut off, with a special support contract from MS, etc. If anyone is still using it they clearly have nothing but disregard for any kind of rules (like all the XP ATMS still out there).
I think you're severely over-estimating the criticality of stuff that's still running XP. It's mostly used for antiquated industrial hardware that gets used 5x a year, maybe (old CNC mills and whatnot), often times with no network. If it gets crypotwalled via flash-drive then someone will reinstall it and carry on with life.
So we're talking about 5 million machines with an OS designed in the late '90s (20 years ago) and that stopped receiving any meaningful updates 4 years ago.
Most of them are actually ATMs in developing countries like India and they are definitely not air-gapped [1]. The POSReady XP with the bare modicum of support until April 2019 made companies take it as a green light to keep using XP in embedded systems.
Other systems running XP: many of the NHS systems hit by WannaCry last year, many of the systems in UK Police stations, most electronic voting machines and gas stations in the US, most digital signage in train stations, airports, hospitals, or cinemas, parking garage payment machines, so many POSes, even passport security in some airports!
Does this put the magnitude of the problem in perspective? It's a threat from so many perspectives. Your safety, your data, your money, you name it.
[1] https://www.rbi.org.in/scripts/NotificationUser.aspx?Id=1131...
I'm hesitant to switch to AMD since Intel internal graphics play nicely with Linux. However that kind of doesn't matter if my machine isn't mine.
AMD has their own kind of management system available on some machines (DASH, using "smart" NICs like Broadcom), but the PSP isn't even involved when DASH is available and in use, as far as I know.
However, on my Intel machine with AMT, there's a network port opened by the ME itself (TCP/16992). It can use the same IP as the main OS, or a different IP entirely if desired. It uses the same ethernet port as the main OS though, splitting the packets that are directed to one of the ME ports and selectively allowing the rest to continue to the main OS (there's a low-level ME firewall[1]).
On that port, there is a full remote desktop with mouse and keyboard, the ability to remotely connect small drives and/or ISO files, a remote serial console, a low level firewall configuration utility, and power/reboot controls. Even on machines where AMT is not even supposed to be available, there have been PoC demonstrated using one of the ME flaws to turn it back on[2].
[1] https://i.imgur.com/Wphopk1.jpg
(slide 68) [2] https://www.blackhat.com/docs/eu-17/materials/eu-17-Goryachy...
That said the amdgpu driver is still relatively new and major changes and improvements are still ongoing but overall it's definitely production stable and for dedicated gpus performance is generally very good.
Not trying to sound like a fanboy it's just nice to see someone else push as hard for open source mainlined graphics as Intel has for all these years.
Case in point, the ATI Radeon HD 6310 that came on EEE PC models, the video hardware acceleration no longer works as it used to be.
Sure, I can probably hack the older driver into newer Ubuntu releases or track down someone that has already done it, but that is exactly what I don't want to spend my time doing outside work.
[1] https://www.ifixit.com/Guide/Asus+Eee+PC+1008ha+Graphics+Car...
The graphics card was working perfectly fine before they decided to reboot driver support.
Now with the legacy driver I have to force enable acceleration and even then I sometimes get the feeling it isn't really working, given how the fan behaves when watching movies on the go.
If you want zero-copy video playback for optimizing battery life use mpv with --hwdec=vaapi. Or vdpau or whatever API is supported with that driver. You can also try switching -vo to vdpau/vaapi from OpenGL.
> I know, and this kind of attitude regarding drivers is what as graphics oriented person, eventually pushed me back into the Windows/OS X world.
On Windows you get legacy drivers for older architectures as well. Also the case with NVIDIA. It’s legacy hardware, after all… AMD’s new driver is completely open source, but it targets GCN, which is a completely different type of hardware.
The issue is that AMD's legacy driver is not the same code as the driver it replaced.
https://www.omgubuntu.co.uk/2016/03/ubuntu-drops-amd-catalys...
On Windows and OS X, the legacy drivers keep working.
The fact that they are open source is of course a good thing whats not is that they were at one time less than half the performance.
The new drivers from AMD gpus are both open source AND performant.
Basically the proper strategy 2003-2017 was to buy nvidia and install the binary drivers.
At present you can go with either so long as you aren't buying hardware too old to be supported by the new amd drivers. I'm still using nvidia on all my hardware but maybe I will give amd a try again next time around.
It sucks that its complicated but its not as complicated as it seems.
Eventually one gets fed up and wants the laptop just to work.
Asus used to sell their netbooks with Ubuntu pre-installed on the German Amazon store.
It already had a phase where I couldn't use WiFi for a couple of months as Ubuntu decided to replace the binary driver for an open source implementation partially working, and then fix issues as they came.
Now it is the same story regarding ATI drivers.
As long as you avoid hostile vendors like Realtek and NVIDIA everything should just work.
TL;DR Ryzen 2700X is great value for money, would recommend it any day :)
System runs cooler, quieter and faster. Gentoo absolutely flies on it as well