Sure, Intel's clocks are slightly higher than before, and their parts use less electricity, but they've been standing still in virtually all other areas.
So you do all the lithography, and vapor deposits on a wafer. That wafer has ~100 physical processors on it (100 just to make rounding easier). You split them (into individual chips), and you start testing.
Say on ~10 Hyper Threading, all the cache cells, and all the magic virtualization stuff works. These become some pro-sumer Xeon type deal.
On another ~10 Hyper Thread, and all cache cells works. This is your i9's
On another ~30 no Hyper Threading, and only some cache cells works. This is your i7's
The rest there is no Hyper Threading, only some cache cells work, and wow only 4 physical cores work. This is your i5's and i3's (kind of).
The idea is yeah, whole parts of a CPU are defective, or unperforming. So they just get disabled and "binned" as another lower tier CPU of the same micro-architecture. All of these get solid at >100-5000x markup to offset the $50bil+ in R&D Intel spends each year. Yes their margins, are... amazing.
Also, the reality is that the consumer market is a massive beneficiary of this whole scheme. The server market is effectively subsidizing consumer processors to the tune of billions, if consumers had to pay full freight on their dies prices would be several times higher than they are.
if consumers had to pay full freight on their dies prices
would be several times higher than they are.
This is doubtful, Intel's margins are amazing as they're fully integrated vertical monopoly. They make the wafers, own the fabs, cut their own masks, etc. Very literally sand comes in one end, and chips come out the other.The fundamental processes of producing an equal die space SoC as a Xeon on the same (node) is likely roughly equal cost (or I imagine the fabs as a service would go out of business). So saying Intel -needs- the consumer market to subsidize their server line is a total lie.
Intel puts an extreme markup on their server class processors, and a milder mark up on the consumer segment
Looks like they could just be trying to brute force the loss of speculative execution simply by augmenting the number of cores.
AMD managed to innovate at a price point where most people can afford. But its a bit disingenuous to ignore the advances that Skylake-X has brought forward.
--------
Skylake-client is very similar to Haswell, but still has larger ROPs, decoder, and branch predictor. It doesn't lead to much better clocks, but IPC is still up over the last few cores.
I do like AVX512 in principle, but in practice the runtime cost of using it is so high that I've offloaded anything highly parallel to other hardware (i.e. GPGPU) -- there's always work enough to keep the CPU busy anyway.
Have you found AVX512 great in practice? I'd love to hear about it!
But what I can say, is that AVX and AVX2 are rather limiting. And I want those new instructions that are found only in AVX512.
Not even to run 512-bit computations mind you. I want to run 128-bit computations in AVX512. You can use XMM registers in AVX512 ya know!!
The main benefits of AVX512 are the mask registers, scatter-instructions (AVX / AVX2 have gather, but no scatter), and an extension all the way to 32 XMM / YMM / ZMM registers
True, using ZMM registers causes severe downclocking issues. But XMM (128-bit) and YMM (256-bit) registers are still sufficient for CPU-based tasks.
-----------------
Practical, pragmatic use of XMM registers include:
* Cuckoo Hashing (See: https://github.com/stanford-futuredata/index-baselines/blob/...) -- Have 8-bins per hash value for a total of 16-bins. Use SIMD to perform 8x comparisons at once.
* Huge number of database applications: http://www.cs.columbia.edu/~orestis/sigmod15.pdf
Bloom Filters are the most obvious "SIMD" data-structure, with applications to databases and many other tasks. Sorting networks are best implemented in SIMD.
* See this discussion: https://news.ycombinator.com/item?id=16171806
In effect, you have Base64 encoding / decoding that is faster than a PCIe x16 slot. You can encode / decode Base64 using AVX registers faster than you can even pipe data to the GPU.
---------
True, the GPU is the ultimate SIMD processor. But in many cases, the SIMD task is too fast to go to the GPU. In particular, the Base64 encoder was encoding 20GB/s, while PCIe can only transport 15.6GB/s!! The CPU is done before the data even gets to the GPU.
Ditto with latency-specific code, like Cuckoo Hashing. SIMD speeds up the overall hash-table, but you're only parallel by 8x. There's no point to actually offload a Cuckoo Hash to the GPU.
In effect, you look for "small" parallelism of size 8 to 16 or so, and that's where AVX / AVX2 / AVX512 really shines. Its too expensive to move to the GPU, but you still get a HUGE speedup when you process it on the CPU.
That seems like a good case for using integrated graphics that has access to the CPU's memory controller and so can access data directly without copying it over PCIe.
Though if you're running that kind of volume you may have PCIe as the bottleneck in any case to access the data on whatever storage/network device.
The Xeon Phi (KNL) used a mesh topology and the KNL was released before Skylake-X.
But I'm a developer on Linux who uses a couple of languages that compile to machine code. Which means that I don't care if AMD has something twice as fast, I'm going to buy Intel for one reason alone: their performance counters support running rr.
That's at least a factor of two improvement in my productivity right there.
Last I heard, Ryzen's perfcounters weren't deterministic enough for rr, so it's dead to me.
Core i9-9900K 8/16 Cores 3.6-.5.0 GHz - $488
Core i7-9700K 8/8 Cores 3.6-.4.9 GHz - $374
Ryzen 7 2700X 8/16 cores 4.3 GHz - $320
Easy pick for me, Intel isn't even a consideration.
I am not fanboying AMD here either. The facts for me are clear, AMD is going to continue to eat Intel for breakfast for the next 2 years at least. Intel isn't going to be competitive until they release a modular designed chip, which although they haven't announced they ware working on, I am 100% certain they are. Why am I certain, first they now have Jim Keller, who designed the Ryzen chip. Second, they acquired NetSpeed, which has IP around modular CPU design (likely a play to ensure they don't get sued by AMD when they release a modular chip). I know a lot of people are looking at Intel right now, thinking it is a bargain prices and a good time to buy. I think they have a long way more to fall. When the chips that Jim Keller is working on are about to hit the market, that is when it will be time to buy Intel. Until then I have no interest in anything Intel has to offer.
The tech is called "infinity fabric", you can look it up there is lots out there about it. It basically allows AMD to make several smaller dies and have them function together as one CPU. Here is an excellent video that explains it all (and also goes into why this is such a huge advantage that allows AMD to have significantly better yields with their wafers)
The Intel is probably faster... but enough to justify literally double the money? I seriously doubt it.
Only in specific IPC heavy workflows.
In any case, bit flips are much more common than were suspected: https://arstechnica.com/information-technology/2009/10/dram-...
I believe strongly that ECC should be standard, because you can't safely assume that your users are doing worthless work. Apple got this right on (non-Mini) desktops a long time ago. Not yet on laptops, unfortunately.
EDIT: At -3 so far, does anyone want to explain the downvotes? I saw the google slides first hand, and there are comments from 2009 in that article saying the same thing.
Like I said though, just a guess.
Your comment provided no substantiation of your claim, merely hand-waving, while casting aspersions on someone else's work.
Most people don't buy desktops with 16 or 32 or 64-thread processors designed to maximize throughput.
Those who do tend to want to max out how much RAM they can shove in their box.
Bitsquatting: DNS Hijacking without exploitation
When bit-errors occur they can change memory content. Computer memory content has semantic meaning. Sometimes, that meaning will be a domain name. And applications utilizing that memory will use the wrong domain name.
Also, are cosmic rays really the main source of single bit flips as apposed to just bad ram maybe?
Some of the time the instructions don't match up, indicating corruption _somewhere_.
For the specific case of crashes in JIT-generated code, the contents of registers and the instructions can be related in various ways (e.g. if you have a jmp instruction the register better contain your code location). And if you know where your code locations might be (because you're a JIT, and are generating the code and aligning it in memory yourself) and the register with the code location looks like the sort of address you would end up with but with one extra low bit set, say...
I am having trouble right now finding the bug report where some of the JIT engineers were analyzing crashes in jitcode, but about 1/3 of those were due to bitflips if I recall correctly. What that means in terms of absolute numbers (or numbers per user-hour, which would be even more useful), I don't know.
Note, by the way, single-bit flips can be a consequence of a bad memory chip, not just of cosmic rays.
If you're only doing primarily single threaded things (i.e. gaming) then the intel chips will give you better performance.
Lets say you can split up your threads to AI, Physics, Rendering, Networking, etc. etc. But lets say Physics dominates: then your game is still single-threaded bound. You only get faster if you make your physics faster.
You can split your game up into work-queues, thread pools, and such. Except not everyone is up-to-date with the latest techniques yet. Furthermore, thread-pools aren't always cache friendly and may hamper your performance. (If Core1 works on something, then Core4 completes the work, you have to transfer all the data out of your L1 and L2 cache to continue working on it on Core4).
So its not exactly an easy thing to program.
But many games are still single-thread bound (despite being multithreaded). As such, Gamers should still prefer single-thread performance.
Gamer's who prefer cutting edge performance should prefer multi-thread performance.
Optimizing for MT-heavy games is the gaming equivalent of premature optimization - everybody can run Doom at a million FPS, but Fallout 4 or PUBG shit all over every system and you'll be begging for every frame you can get.
You can "prefer" whatever you want but developers don't care. This site of all places knows that time-to-market is what really matters. If you want to play those titles, you have to deal with it. Either optimize for the shitty titles, accept that you're going to be losing a fairly significant amount of frames (the 2700X is behind as much as 30% in some titles), or don't play those titles.
I mean that if the software is implemented in the correct manner, "more can be done" by utilizing multiple threads.
Games make use of up to 6 threads, but single-threaded performance still determines overall framerate to a large extent. Games are not exempt from Amdahl's law: as long as you've got enough threads to offload work, it comes down to the single-threaded portion of the workload. The faster you can run the main game loop the faster the game runs.
That's because once you are down the critical path, the only way to go faster is by improving single core performance.
There is a reason reason why Core i5 are so popular with gamers. They have just enough cores not to limit a game engine, excellent single-thread performance, and they are affordable.
But what application are you needing more than 10Gbps externally?
Laptop docking. Running a 4K monitor over a USB Type-C cable leaves you with only the USB 2.0 lanes available for data. (And 5K is impossible.)
[EDIT - removed AVX-512 claim that was wrong]
Overall perf is likely a lot better with the i9 for most workloads.
It's true that Coffee Lake has a performance edge over Ryzen when running AVX/AVX2 code though.
Ryzen doesn't have 256-bit wide units, so 256-bit instructions take 2x as long.
The benefit of that is that Ryzen doesn't throttle the clock speed of the whole chip when executing AVX2 instructions :P
Intel's widening of units has brought a giant downclocking problem: https://blog.cloudflare.com/on-the-dangers-of-intels-frequen...
AMD is offering a 16-core / 32-thread 4.4GHz monster for $899, and a 32-core / 64-thread for $1799.
If you wanted the best under $5000 (total cost of a computer), it seems like AMD Threadripper is the best. If you wanted the best below $1000, it seems like AMD Ryzen is the best.
Only if you are single-thread bound (5GHz clocks!!), AVX2 or AVX512 bound, or PEXT / PDEP bound (Stockfish 9) should you consider Intel. But otherwise, AMD is offering more performance at all price points up to the EPYC 7601 (~$3000 CPU: 32-core/64 threads / 8-memory controllers / 64 PCIe lanes direct to CPU, support for dual-socket).
--------
With that being said, I'm definitely interested in Intel's Xeon Silver platform. If Intel pushed dual-socket out better, then they would be cheaper AND faster. IMO, its a bit weird that Intel isn't taking advantage of their dual-socket solutions to counter AMD Threadripper (I mean... Threadripper really is just a dual-socket or quad-socket NUMA chip combined into a single socket).
As it is, Xeon Silver is hard to find and seems to be ticking up in price unfortunately. But their nominal prices are actually quite good, although their clocks are kinda low. But Xeon Silver really seems like Intel's price/performance champ (even if its still a bit more expensive than Threadripper or EPYC).
What do you think it's more profitable to sell: one CPU with many PCIe lanes that you can attach 8 GPUs to (8 NVIDIA GPUs for 1 Intel CPUs), or more CPUs with fewer PCIe lanes (8 NVIDIA GPUs for 2 Intel CPUs)
Workloads limited by a single thread.
Overclockers.
That's about all I see.
This is a critical security bug that was discovered over a year ago. Intel just ignoring the problem and casually releasing yet another (underwhelming) upgrade that doesn't address it all...
The fact anyone accepts that just goes to show how low our expectations have dipped when it comes to Intel.
See for example https://en.wikipedia.org/wiki/Spectre_(security_vulnerabilit... :
> Spectre proper was discovered independently by Jann Horn from Google's Project Zero and Paul Kocher in collaboration with Daniel Genkin, Mike Hamburg, Moritz Lipp and Yuval Yarom. Microsoft Vulnerability Research extended it to browsers' JavaScript JIT engines.[4][20] It was made public in conjunction with another vulnerability, Meltdown, on January 3, 2018, after the affected hardware vendors had already been made aware of the issue on June 1, 2017.
There were also stories they were aware of the problem even before, and basically ignored it. If you check out that history section, there were multiple public presentations about the feasibility of an attack for years before the practical exploit was discovered. The only way Intel didn't at least suspect it is if it was very sloppy, or didn't care at all.
That said, the next gen Ryzen is rumored to be 4.5 GHz chip on turbo. So, definitely the steam is picking up.
If you look at the cost difference, it is pretty dead on placed with ~$130 higher than Ryzen 8 core. People trash around AMD/Intel like its their home town sports team. I find that everywhere including on HN.
Let's wait for the benchmarks, specific load comparisons to see if the price difference makes sense. The 8700k is ridiculously fast and it beats the hell out of Ryzen's 1600X.
The i7 has the same amount of cache as its predecessor and the i9 has more.
9-series has hardware fixes for Meltdown variant 3 as well as the L1 terminal fault fix.
I will only buy Intel if I'm forced to by other people (think: employer, laptop). By my own choice, I will buy Ryzen/Threadripper whenever I can. And I will do so gladly, knowing I give my money to the David who's fighting Goliath and giving me a much better bang for the buck at the same time. Win/win.
[1] https://www.phoronix.com/scan.php?page=article&item=amd-athl...