Best CPUs for Workstations: 2017
anandtech.com
anandtech.com
In truth, there are plenty of good reasons, both for data integrity and security, to have ECC. However, there are real costs to ECC too - 10-30% immediate performance overhead (depending on your application) and much more than that for memory-bound workloads (due to lower frequencies and subtiming issues), not to mention hardware cost efficiency. Instead of oscillating between "ECC for everything" and "ECC is useless" people should rationally compare the costs and benefits for their specific workload and choose accordingly.
In terms of speed it's ~2–3 percent slower worst case, and memory is only one component of overall performance lowering the impact further.
More importantly, bit errors are not evenly distributed with some users seeing vastly more errors.
I am one of the principle engineers working on a game that was able to support over 1.2 million players concurrently, so I feel that I have some merit behind what I'm about to say.
ECC Memory does not have such an extreme overhead. Our gameserver is memory bandwidth bound. To put that into perspective a change from DDR3 to DDR4 saw a jump in performance of 34%~. Using non-ecc memory (which is obviously very available to the development team) only gave a penalty of roughly 3% in the worst case. (both DDR4 systems).
Why are we memory bandwidth bound? because we allocate 250GB of "game world" and phases and make teeny tiny modifications to them when you run around or pick up loot and shit. It's as taxing as you can be on memory bandwidth since you can't batch set or load anything.
On a further note, while the implementation of ECC itself is very fast for modern RAM, the same is not true of VRAM on GPUs. At least as implemented in NVidia Tesla GPUs, enabling ECC has a huge performance penalty by itself. [1]
[1] http://geoceano.fr/wp-content/uploads/2015/01/cpu-gpu-lbm-in...
When you start looking at multi CPU computers with 1+TB of ram the story changes.
PS: Some of this stuff is also really hard to test. Without ECC it's hard to data in terms of stability for petabytes of ram over years. And no extrapolations don't work very well.
Now, non ECC memory might have better specs / a higher overclock rate etc, but doing either of those increases the rate of errors. And sure you might not notice the problem directly, but that does not mean things are ok.
The value of ECC is often more to detect errors than to correct them. Remember, if you have 10 errors an hour then you unlikely to see them in a 20 second memory test.
In practice out of 10,000 machines you might have 1/2 your errors on 10 of them. Critically, if you have ECC memory and you both detect and replace those 10 machines the apparent rate of failures then drops.
PS: This is a little dated (09) and only from a single set. But, still an interesting introduction: https://www.cs.toronto.edu/~bianca/papers/sigmetrics09.pdf "Around 20% of DIMMs in Platform A and B are affected by correctable errors per year, compared to less than 4% of DIMMs in Platform C and D. ... About a third of all machines in the fleet experience at least one memory error per year (see column CE Incid. %) and the average number of correctable errors per year is over 22,000. "
EDIT: the least out of date reference I could find is a blog post by Jeff Atwood, where he concludes that based on the available literature ECC-correctable errors are quite rare. [1]
"The median number of errors per year for those machines that experience at least one error ranges from 25 to 611." "We find that for all platforms, 20% of the machines with errors make up more than 90% of all observed errors for that platform."
A minority of that 20% again have even more issues.
The worse case is literally 10,000x more than 1 bit error per year on a single machine. Without using ECC you really can't tell other than some vague not sure what's wrong just replace it.
Hmm, like what? I've run many simulations too and had the opposite experience. It's hard to implement checks because your algorithm defines how the simulation should permute over time.
While a bit error could certainly crash the simulation, long running simulations can be resumed from checkpoints with minimal loss of time and the probability of a crash is still very low for most real workloads.
Not really. Your algorithm determines how the simulation plays out, so you can't really notice small errors until much later. By then it's hard to roll back to an earlier state unless you're saving 64GB snapshots.
[*] There are exceptions in specific circumstances most of which would produce obviously wrong results or crash the simulation outright. However, I'm not claiming that single bitflips are never an issue for simulations, only that it happens much less often than people sometimes think.
Also, it's not just physics simulations that are at risk. Voxel renderings are especially susceptible to one-bit errors.
Anyway, I haven't argued against using ECC memory in any of these comments. I've argued that you should know your needs and choose accordingly, and that ECC memory isn't free and therefore cannot be a universally optimal default for everyone.
For much of my physics work I've found ECC memory not to be worth the cost at all. For other things, like most database servers (and perhaps voxel rendering, I don't have enough experience there to comment) ECC memory is clearly worth it. I just think people should invest some thought into the question before evangelizing for or against ECC in the general case.
It’s some kind of snobbery to advocate that good technology should be held back until non-enterprise plebs demonstrate that they are deserving of it.
Just give everyone the good technology as much as is possible. Who’s it hurting if some subjectively underutilise it?
Everyone who buys the product and doesn't require the technology, because they'll be spending more money on something they don't need or use.
It's a political decision by Intel to make ECC an Xeon-only product which has resulted in the vast majority of the price difference. The silicon die area required for it really doesn't cost that much more - often, the area is present on the consumer dies but just disabled so they can charge a lot more to the enterprise customers who actually require it. And political decisions like that make a lot of HN users unhappy.
As Joel Spolsky said in https://www.joelonsoftware.com/2004/12/15/camels-and-rubber-...:
> Working my way backwards, this business about segmenting? It pisses the heck off of people. People want to feel they’re paying a fair price.... And God help you if an A-list blogger finds out that your premium printer is identical to the cheap printer, with the speed inhibitor turned off.
Everyone knows that ECC parts are nearly identical to the cheap parts, with the ECC module turned off. But there's no boycott, because there are (were?) no other options, like there might be for a printer company. ECC segmentation is just something everyone knows and resents.
Computer reliability should be a rising tide. Up-and-to-the-right is the goal.
Scientists, Engineers and Economists shouldn't use a Phone or MacBook? Your handle is ironic.
The lack of ECC is a method to artificially segment markets and nothing more.
Not ECC, but "single bit errors" can be fatal in other domains [0]
[0] https://gizmodo.com/382026/a-cellphones-missing-dot-kills-tw...
"During the logging period there were a total of 52,317 bitsquat requests from 12,949 unique IP addresses. When not counting 3 events that caused extraordinary amounts of traffic, an average of 59 unique IPs per day made HTTP requests to my 32 bitsquat domains."
http://dinaburg.org/bitsquatting.html
Overheating can be a cause of bit errors, and heat is a concern in mobile devices and laptops.
That's a misconception, to quote Matthew Ahrens, cofounder of ZFS: "There's nothing special about ZFS that requires/encourages the use of ECC RAM more so than any other filesystem."
See also [0].
[0] http://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-yo...
https://research.cs.wisc.edu/wind/Publications/zfs-corruptio...
ZFS is a lot more active... checksumming and caching mean that important information is spending time in RAM. It's not good if that information gets corrupted.
ZFS doesn't cache in any special way, or more than your typical filesystem.
Lack of ECC just lessens the benefits of ZFS, it doesn't exacerbate any problems.
It's not that ZFS without ECC is completely unsafe... it's just not as safe as it could be. ECC RAM becomes the next biggest concern once you've addressed on-disk bitrot.
I was hinting that you still get a bit of extra protection with ZFS over other filesystems, both on Non-ECC RAM. There is a chance with a stuck/corrupt bit that checksum will fail and you will get a read error. I interpreted OP as saying ECC is a must for ZFS, but I don't think you are more prone to corruption that any other FS?
I'm basing my understanding on this: http://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-yo...
These checksums are fixed exactly once, in the zio pipeline, as a dirtied buffer is being passed through it during a write operation. The checksum then stays on disk in that form in block B for as long as block A is reachable.
In the read case, checksum validation is done at the vdev_{raidz,mirror,...}.c level and repaired there if possible e.g. from another element of the mirror; recovery is also possible in the case of copies=N (N > 1). Generally, the checksums of read data and metadata are not used once a buffer is successfully in the ARC. (However, compare arc_buf_{freeze,thaw}(), which one can enable, and which will happily cause panics if ARC buffers are corrupted whether through physical problems (e.g. in memory, in a bus, in the CPU) or software error in the kernel.)
Block B may already be in cache for some other reason, in which case the in-memory checksum is used (this is the "yes" part of "yes-and-no"); otherwise, block B will have to be fetched before block A, because block B is where A's checksum lives. Block B may be evicted from ARC long before block A is evicted; any read() or similar call that needs data from A will not result in B being read back into memory.
A buffer kept in ARC and used as the source of subsequently-dirtied data will have its checksum generated anew, just like the fresh write above. Obviously if the data it is dirtied with is bad, it will go to disk and its parents will have a checksum reflecting the bad data. However, the source of the badness need not be within the transient dirtied buffer, nor in a long-resident ARC buffer; userland memory can be corrupted too.
If one has bad memory than bad data may end up on disk, checksummed in such a way that the badness will not be detected by the ZFS subsystem. But this is really no different from any other RAID-like data-validating system or a traditional one-disk filesystem which doesn't checksum. If bits are flipped in userland data that is ultimately subjected to a write(2) call (etc., think mmap()), no filesystem can reasonably expect to do anything other than write out the data handed to it via the system call.
Like any other filesystem, ZFS has its own metadata for tracking allocations, object metadata, and in generating the branch-to-root Merkle tree. The wrong memory corruption (and this has happened in software during ZFS development) can result in data loss or even a pool too damaged to be imported (go to backups). ZFS is not especially more exposed to that than any other filesystem.
Since the arrival of compressed ARC and crypto, only a small fraction of ARC buffers are kept uncompressed or unencrypted in memory at any time (this is controlled by the dbuf cache tunables). Bitflips in any encrypted-in-memory ARC buffer will be caught (noisily) if the buffer is subsequently read or modified, since a decryption will be done at that time and would fail under a single modified bit. Many possible corruptions in compressed-in-memory data will also be caught depending on the type of compression used when the buffer was first written out to disk. In neither case is the ZFS checksum involved in the catching of such corruptions.
Finally, this still won't protect against corruptions in userland, corruptions which hit the various data structure that point to buffers in ARC, or corruptions in the text segments in the zfs subsystem or elsewhere in the kernel. However, ZFS is not realisically more fragile to such corruptions than any other filesystem.
What ZFS really needs is native encryption to be generally available.
Does anyone have an idea of the cost increase for using ECC RAM in a mobile device for instance? Are there also power usage concerns to worry about?
AFAIK there are basically zero consumer devices using ECC RAM.
I want to say I had ECC in my old AMD K6-350, but that was a lot of years ago. Either way, I'm pretty sure they have supported it for a long time and in many of their CPUs, even the CPUs aimed at consumers.
Here's a comment I made last night, if you want to dig in and find some current offerings:
Same reason ECC is not supported on consumer processors from Intel. Here it would be almost trivial to add, but they don't.
Simple business strategy.
So, they're okay with it on consumer parts, as long as you're not looking to do anything resembling professional work I guess.
Technically. A fraction of a percent increase in die size for the memory transceivers, basically nothing for making and verifying checksums.
It's all business strategy.
manufacturing cost is non-linear. Obviously intrinsically word gunk.
And in a lucky coincidence it turns out that those who want ECC are willing to pay the moon for it, which of course gatekeepers like Intel exploit.
Basically there's no evidence that ECC memory is signficantly slower or more expensive than non-ECC memory. Obviously it's slightly more complex than non-ECC memory, but it's not a particularly high-tech addition. For all its supposed complexity/slowness, Samsung is using it in its caches on their Exynos SoCs.
Everything can be perfectly explained in terms of market segmentation.
However, in many applications we find that using (forward) error correction almost always increases data density (for storage) or bandwidth (for transmission), simply because a FEC stream does not require a nearly-perfect channel any more. This is the way hard disks, SSDs, WiFi, LTE, DSL, ..., satellite communications, ...[, ...][, ...] are able to cram incredible amounts of data into very noisy channels. Thus, ECC significantly lowers cost in many dimensions (be it frequency spectra, storage prices, not having to re-cable entire countries...).
(And if you don't use the extra noise margin to increase density/bandwidth, then you can use it to increase reliability, like we usually do with ECC memory)
Thinking about it for a few minutes, the memory bus will most likely be the only bus in your computer that has no error correction/detection. USB, SATA, PCIe, all of them require it. The main memory will also most likely be the only storage that doesn't use it (apart from firmware flash chips and the like, but these often use a checksum at least).
Do you mean that ECC has those benefits, or that other applications of error correcting codes has them?
Personally, I've always found this to be very annoying. The only people using non-ECC memory should be Xtreme Gamer type people who don't care about their data and only care about squeezing that last 2% out of their system to land at the top of the penis size chart. Non-ECC should be a feature like liquid N2 cooling.
The moral of this story is that ECC is only ever advertised explicitly, when it is special to have, i.e. because there are segments in the market where it is not used.
(A somewhat curious case are small microcontrollers. Often these control relatively important aspects of machinery, for example, drives [say a bit-flip in the drive controller turns STOP into RUN, which can get rather ugly on e.g. a lathe]; these usually don't have ECC. However, they are manufactured in very coarse processes and generally don't use DRAM.)
The real question though for a large scale consumer product manufacturer is how much margin they could convince the RAM makers to give up on ECC. Once it became standard, it probably would not cost much more than non ECC. Maybe 15% more? (1/8th extra memory, plus some extra for other expenses.)
It is nice to see "EDAC amd64: DRAM ECC enabled." in my dmesg output. :) That, and the NVMe SSD is screaming fast.
I'm debating whether to try my hand at overclocking my R7-1700, but even leaving it at the stock clocking it's pretty speedy, and it runs cool with the stock cooler.
See e.g. this paper by Walker and Betz from the San Diego Supercomputing Centre (paper from XSEDE '13) on the effect of ECC on GPUs for MD simulations: https://dl.acm.org/citation.cfm?id=2484774
Even if it's just in a float, it looks like a single bitflip could change FLOAT_MAX to NaN. That could hurt.
You can also use ptrace to corrupt the memory of a process more subtly.
I am in the market for an upgrade to my home PC. I use it for Compiling and Gaming.
I am finding the sheer variety of CPU's and their weird naming conventions utterly confusing. Then combine that with the almost-as-confusing choice of motherboards and the bit-better-but-not-much choice of RAM and I am completely lost.
Any tips out there for finding a path through this maze, how do I upgrade my PC without requiring an advanced degree in Intel/AMD marketing speak?
EDIT: Added use of PC.
The site is tilted for gaming, but GPUs are mostly interchangeable - so if you want more compute power, you can drop to a lower GPU and spend the money on the CPU instead.
Compare their "Best CPUs for Workstations" and "Best CPUs for Gaming" articles, and choose a platform. Probably going to be Ryzen 7 vs. X299, and you'll probably get at least one more CPU upgrade from either of these motherboards.
Get an NVMe Samsung 960 Evo or Pro because fast storage is super important and these are currently blowing everything else out of the water. Get enough RAM for now, and leave a slot or two free for when the prices for DDR4 drop.
Pick your graphics card from the spectrum, or honestly, just keep using whatever you have right now because cryptocurrency is currently severely distorting the market and limiting availability. I don't see a "Best graphics card" summary article right now, but I think this chart is somewhat current:
AMD Price NVIDIA
RX Vega 64 $500 GeForce GTX 1080
$450 GeForce GTX 1070
RX Vega 56 $400
RX 580 (8GB) $300 GeForce GTX 1060 (6GB)
RX 580 (4GB) $200
RX 570 $170/$180 GeForce GTX 1060 3GB
$130 GeForce GTX 1050 Ti
RX 460 $100/$105 GeForce GTX 1050
The most important thing to understand is that getting the exact ideal processor doesn't really matter. The difference between competing processors will only be discernible with a stopwatch and carefully controlled tests.For choosing everything else many websites put together system guides laying out a set of compatible components they think are good at the same part of the market. Those tend to be very gaming focused but I think that's a good starting point. Intel just released a new top end chip so they're all out of date but the recommendations here[4] are where I'd advise you to look if you want a new system now. The prices have come down a bit since the guide was written, though.
Or there's the Build a PC subreddit wiki as an information resource. [5]
[1]https://www.phoronix.com/scan.php?page=article&item=intel-co... [2]https://www.anandtech.com/show/11859/the-anandtech-coffee-la... [3]http://techreport.com/review/32642/intel-core-i7-8700k-cpu-r... [4]http://techreport.com/review/32474/the-tech-report-system-gu... [5]https://www.reddit.com/r/buildapc/wiki/index
The best part is I never hear it, even when gaming the fans seem to barely kick in.
Things that do not parallelize: the final linking step (especially with mingw), compilation if all your code is one giant CPP file (prefer small files).
Assuming you have an SSD, your project is configured for parallel build, and your code is organized into many small CPP files, more cores == more better. I look for the highest core count CPUs with the lowest power. Last year I built a dual CPU Xeon workstation / home server using 1.8 GHz 14-core chips (mostly from second-hand used hardware to save money). I feel that that low clock speed is actually an advantage, because even with 2 sockets and 8 sticks of 8 Gb RAM, it idles at only 75W (at the outlet, as measured with a Kill-a-Watt). With 56 total threads, it builds my C++ project in 2.0 seconds flat, something which takes almost 2 minutes single-threaded. (I think the fact that there are 60 seconds in a minute, that 56 ~= 60, and that the build time went from 2 minutes to 2 seconds is not coincidental.)
Unfortunately a hardware problem with one of the used components seems to be causing the machine to lock up under Linux, and I never had the time to troubleshoot it, and moved on to a different codebase which doesn't require as much compilation so haven't yet had motivation to track down the problem.
To compensate for a lower IPC, AMD will give you more cores (sometimes a lot) for the same price.
It all depends of what you need your processing power for.
0b0111111111101111111111111111111111111111111111111111111111111111 == 1.7976931348623157e+308
0b0011111111101111111111111111111111111111111111111111111111111111 == 0.9999999999999999
Do a lot of people really run >512GB RAM in their workstations rather than running a "thin" workstation and running simulations on servers or EC2?
I just built a new workstation a couple months ago: i7 7700 (not the K), 64GB RAM, one of those closed loop CPU water coolers... As a workstation, I wanted it quiet, and performance has been great. Long running jobs I run on our dev/stg cluster (4 machines, 512GB total RAM, 48 total cores.
I can think of some reasons to have twice, four times, or eight times as much though. You can almost never have enough block cache, after all. Local testing databases come to mind. Also, tracing/profiling Elixir programs takes up enormous amounts of memory. My colleagues at my last company did not have enough RAM to generate meaningful profiles of our system, 64GiB was just barely enough to get a signal out of it. Rather than spending a month working on the profiler, it was nice to be able to get some data out of it as-is.
Ryzen is for Programmers | https://news.ycombinator.com/item?id=14243350 (May 2017, 272 comments)
the problem is you can't easily buy a 7401, there is no video online verifying the claimed cinebench score is a pretty good example on how hard to get access to one.
Discussion here: https://news.ycombinator.com/item?id=15408850
Money may not be an object, but choosing the first option seems a little stupid. Even rich people don't throw away money like that, because if they did, they probably wouldn't have gotten rich in the first place.
So the total performance (to first order) is almost twice as good for the 28 core part. Sure you need to use all cores for that comparison to hold, but usually software either scales up to 8 cores or less or it scales up to N cores.
A lot of "workstation" use is long stretches of interactive use on a few cores followed by periods of offline use (NN training, Graphical rendering, Code compilation/testing etc). The boost clocks on these higher end cpu's should really be able to go even higher, even if it means the base clocks would have to go down somewhat. I'd rather have a 28core CPU where it runs either 20 cores entirely OFF and the rest at 4.5Ghz, than one where it runs at 2.5-3.8Ghz.
Is this just a thing of dreams?
I don't have any precise requests, but I would like to be able to pretty much hook up an oscilliscope or logic analyzer anywhere and in principle be able to understand what's going on.
I'm looking to get into hardware by way of tinkering and my mental model is of open software. Currently, any time I have a burning question about one of my tools, I can just open a man page as well as start digging into the source. I think it would be cool to do the spiritual equivalent with hardware.
For example, I am currently digging into the bootup process--everything after CPU POST until a fully ready userspace. However, the details that between pressing the power button on the power supply and CPU POST are completely opaque to me, and some are completely opaque even in principle on my current system. I'd like to have the ability to fully grok the electrical underpinnings of what's going on or at least as close as possible.
I was thinking of dual-CPU Z840; are there better options?
http://www.velocitymicro.com/desktop-workstation-pc.php
Might be worth a look, unless you can make do with a second-hand dual xeon workstation from eBay (possibly adding in an nvidia 1080 or two).
1. For a lot of workloads, it's more convenient to run in on a desktop. Sure, with the right tooling, you can bridge your cloud and desktop pretty seamlessly, but a lot of people can't do that or it's not possible with their tooling. Sure, it doesn't scale, but it often doesn't need to. 2. A Threadripper CPU may cost $1000, but but look at what that buys you in the cloud. The machine pays for itself within a year if used extensively. 3. Sometimes it's not easy to just upload code and data (!) to the cloud. Sure, there's all sorts of security layers and certifications and yes, clouds are very often more secure than desktops, but still, some people can't afford to let their data leave their machines.
Two of these were a thing at my former employer (but we still used the cloud for other use cases). So no, I wouldn't say these workstations are obsolete.
Finally, cloud computing is still really expensive compared to dedicated hardware. A halfway decent EC2 'workstation' like the p2.xlarge is over $5000 a year (plus storage) if you pay for the whole year up front and that 'only' gets you 4 cores and 60 GB of RAM (and a pretty good GPU). $5000 will buy you a really nice workstation with much better specs, and you get to keep it at the end of the year
$0.4/hour * 24 hours * 365 days = $3,504 per year. You can build the same machine for a fraction of that cost and you'll get more than a year's use out of it. Reserved instances would bring the price to rent down a bit further, but I still don't think you can beat the economics of owning a workstation if you're going to be using it most of the time.
- Don't care about latency.
- Don't need bare metal performance.
- Don't mind ongoing costs.
- Either don't care, or don't need to care about their data sitting outside their organisation.
- Have internet connections that can handle massive data transfers for months on end.
Then there's still latency, and the need for some kind of local client for the remote session.
It's also a move away from "micro-computing" and back to the "mainframe-client" model that is so popular with Cloud-and-Phone.
They mention the Ryzen 5 is better bang for the buck, but dismiss it for having "low overall performance." I guess I'm wrong in thinking that, for 98% of people, a $250 CPU would still be wasteful.
Most people will either buy a laptop or desktop (if their smartphone is not sufficient) and not spend nearly as much money, they only need to run a browser and a couple of other small programs.
The workloads are fairly standardised and non-diverse. Not necessarily small.
1950x is slow for multithreaded applications, it has a low cinebench score of ~3,000, you pay some serious $ for a fancy motherboard full of LED lights, then you don't have officially verified ECC support.
As a comparison, you get the same mutithreaded speed, much cheaper system cost if you just buy second hand dual 2696v2 processors with real ECC support. if you need real processing power in a single box and have a tight budget, there is a flood of E5-2696v4 processors on the market at $950-1,100 each. for a much nicer budget, you can always go dual/quad Intel 8180.
For 1950x, details explanations were provided in my post -
1. too sloow, cinebench score is 3k, you can get it from 4-5 years old ancient processors already declared as EOL. 2. no official ECC support
I can list more actually - single socket system with 8 DIMMs only. As explained, it is a good platform for kids to game/overclock.
if it is about PCIE lanes, you should at least be buying dual socket systems, that gives you more PCIE lanes than 1950x.
if it is about memory, well, 1950x is limited to 8 DIMMs.
Says who? The Intel marketing department?
what is your first AMD processor? mine was AM486DX4-120
as comparison, I can order a pair of 8180 processors, its mb and RAM now and expect the package to arrive in 12 hours.