AMD lands Google, Twitter as customers with newest server chip
reuters.com
reuters.com
AMD 10K employees and a $1 billion annual R&D budget.
Intel 100K employees (including 10K just in software) and $20 billion R&D.
Nvidia 11K employees and $2 billion R&D budget.
AMD CPUs now surpass Intel and their GPUs are competitive with Nvidia except at the high end.
The Rome is a nice chip nonetheless, I hope they can rehire a big enough Linux team to get over these humps it isn't that much work on the CPU/platform side.
On the GPU side though, at least with the modern APIs (Vulkan, D3D12 and Metal), AMD isn’t too far behind Nvidia. I actually prefer RenderDoc to Nvidias Nsight because I can capture multiple frames instead of pausing the application to look at one frame at a time. That being said though, the OpenGL tooling for AMD is abysmal and so is their OpenGL driver. All things being equal, just swapping an Nvidia card for an AMD card gives around 40-50% speed up when running OpenGL purely by getting rid of the driver overhead. Although with Vulkan AMD and Nvidia are actually on par, so at least that’s solving itself for us.
Good excuse to build another machine!
Have you tried the AMD mesa drivers? The open source driver performance doesn't look terrible, and in my own experience, it tends to work far better for OpenGL than AMD's proprietary GPU drivers.
https://www.phoronix.com/scan.php?page=article&item=rx-5700-...
I have been working with performance tools on Linux for over 10 years and occasionally debugged graphics issues.
I should note that the OpenGL stack on macOS isn't better than the Windows one either. Probably not surprising anybody, given the lackluster support that Apple has for it. We are seeing similar improvements with Metal as we do for Vulkan.
Windows OpenGL driver is so bad that running it through ANGLE with DX backend is faster.
I built a Threadripper 2 machine recently for home, but I don't know that it would have similar utility in a work environment where I could use something like Incredibuild to farm out compiles. Is that not an option for you in your work?
Mobile is also not Intel.
The only case to left are the Wintel PCs - and even that just makes Intel look bad since AMD CPUs run the games just fine. For the actual gaming experience it basically amounts to no difference at all.
I'll still need to keep some Intels around for VTune and optimizing deeply for AMD is going to be a chore... But given how the market is moving, I don't think I can avoid making AMD the primary development focus.
Ryzen < 2(!) has a much narrower AVX unit which could easily cause a 50% slowdown.
Branchy code like that of compilers tends to run somewhat slower on Ryzen, though closer to 15% than 50% (comparing AMD and Intel CPUs with otherwise same performance in mixed benchmarks). The latter really depends. Core and Ryzen have different branch predictors, both so complicated that they can't be fully described in documentation anymore. Ryzen 2's first level cache configuration also moved closer to Core's, presumably for software that has been fine-tuned for Core.
Specifically, in both a physics engine solver and some machine learning backwards passes I was working on, there are areas where a little false sharing is likely. It's relatively uncommon and a negligible concern on the smaller Intel chips I tested on, but Zen cores will sometimes have to synchronize with another CCX's caches and it seems to come with a penalty vastly larger than naively expected latency or IF bandwidth.
A fully loaded 2950x ended up getting beat by a 1700x in the solver despite having virtually the same architecture, twice the memory bandwidth, and an observed clocks advantage. Cutting the used thread count down to the same as the 1700x helped a little, but it looked like Windows was scheduling the threads on as many CCXs as possible and it ended up still being slower.
On the upside, the 2950x blasts through friendlier workloads like collision detection and inference with perfectly reasonable scaling.
I'm hoping that the dramatic redesign in Zen 2's memory architecture (unified IO die and whatnot) will help things a little. If not, I'll probably have to rework some stuff.
I use ECC on a (consumer) Ryzen chip/board and edac-util seems to give me the same information that it does on Intel - what's missing?
One the CPU/platform side of things, my biggest annoyance is how far behind k10temp is (Zen2 support not in mainline until 5.4?), and how bad sensors support is in general on the boards (requiring reverse engineered non-mainline modules for my Zen/Zen2 workstations).
While I agree that on the GPGPU-front Nvidia is still ahead, I'm very happy these days on my workstations with the state of AMDGPU in the mainline kernels and much prefer AMD GPUs for my workstations now vs Nvidia cards, which has led me to pay a bit of attention to ROCm - it looks like they are making very steady progress and TF and PyTorch support seems pretty good at this point (also, stuff like MIVisionX/OpenVX, CenterNet, BERT support seem to all be working relatively painlessly [3]), although it'd be useful if anyone has a resource that does continual benchmarking comparing on-prem/cloud perf of the various platforms, it'd be nice to get good $/perf and W/perf numbers over time. (I assume that anything running with Nvidia's tensor cores still completely blows away what AMD has to offer atm).
[1] https://developer.amd.com/amd-uprof/
[2] https://github.com/RadeonOpenCompute/ROCm/commits/master/REA...
> I use ECC on a (consumer) Ryzen chip/board and edac-util seems to give me the same information that it does on Intel - what's missing?
Event-based sampling on Intel is accurate to an instruction-level (while event-based sampling on AMD is less accurate. You're forced to use the more complicated IBS metrics if you want instruction-level accuracy of events).
Intel also has branch-history data stored. Super useful for some developer tools, but I forget which tools those were...
-----------
I think AMD uProf is certainly usable. And the price is good (free). But Intel vTune is just light-years ahead.
AMD vs CUDA on the other hand is... closer than I think most people realize. CUDA has a bunch of libraries (Thrust, TensorFlow support, etc. etc.) which helps. But if you're doing high-performance coding, you'll likely have to write your own specialized data-structures. At least, that's the approach I'm doing with some GPU hobby code I'm writing.
TensorFlow (due to Tensorcores) and BLAS are solidly NVidia advantages. But general purpose libraries (ex: Thrust) is more of a convenience.
AMD's main disadvantage is documentation. But the tools are actually quite usable. AMD documents the lowest level well (the ISA), but their HIP / HCC / etc. etc. documents are lacking and difficult for beginners to follow.
AMD should work on updating their beginner guides (their OpenCL guides) to their ROCm framework. Even if its ROCm OpenCL 2.0 stuff, its important to get beginners to use their platform. Or at least, update their beginner guides to reference GPUs that have come out within the past 5 years...
Sadly also the third-party hardware ecosystem is (at the moment) pretty bad.
For example you can choose among almost 300 different motherboards with an Intel 1151v2 Socket [1]. They come in all possible sorts and combinations of form factor, chipsets, ports, etc. In comparison there are less than 100 motherboards with an AMD AM4 Socket [2], most of which in the large ATX form factor and with basic I/O ports.
Let's hope that all these third-party companies felt the change of the tide and are hard at work on new AMD-based products.
[1] https://geizhals.de/?cat=mbp4_1151v2 [2] https://geizhals.de/?cat=mbam4
[0] - https://www.macrotrends.net/stocks/charts/AMD/amd/revenue
As an aside: why won't everyone migrate to Vulkan compute shaders? I hate these "special" compute stacks >_< Clearly I'm not alone in thinking this: Tencent's ncnn uses Vulkan as the only GPU option, some Googler is working on clspv (OpenCL to Vulkan SPIR-V compiler)..
1. No libraries. cudnn, cublas, cufft, are all the fastest available (except maybe magma sometimes), plus no one writing an actual application wants to reinvent a fast gemm. Also cutlass, cub, thrust, ...
2. No c++. The "standard" seems to be glsl, and a "prototype" opencl c -> spir-v compiler doesn't give me much confidence in that approach.
3. No one wants to use vulkan apis directly, and there are approximately 5 billion different "vulkan compute" wrapper utility libraries. i.e. No consistent platform.
Vulkan compute and OpenCL are not entirely compatible (both are backed by SPIR-V (but different flavors of it.)) Khronos has chosen to maintain OpenCL (and OpenCL-next) separate from Vulkan.
TSMC 48K employees, and $13 - $15B R&D.
And even that is excluding all the ecosystem and tooling companies around TSMC. Compared to Intel which does it all by themselves. Not to mention Intel does way more than just CPU, GPU, also Memory, Network, Mobile, 5G, WiFi, FPGA, Storage Controller etc.
https://www.theverge.com/2019/7/25/8909671/apple-intel-5g-sm...
I guess Intel current handcuffs on customers is them owning Thunderbolt.
Intel did announce in 2017 they planned to release thunderbolt royalty free in the future; however, they still have not released it as of yet. And given their new competition from AMD - I doubt they plan to now anytime soon
They gave it away to usb form but maybe they won't for anything AMD wants to do?
The annoyances are such that I'm waiting for a 3900X to be available in-store so I can build my full workstation finally after sitting on a NUC + eGPU setup for over a year.
You'll be going for 3900X + X570 Taichi? Or another board? Ideally I'd also want ECC RAM.
The Asrock X570 board mentioned definitely has support, as do the other high-end X570s I looked at like the Aorus Master, Asus C8H, Pro WS, etc.
Based on X570 BIOS updates so far (and their overkill VRMs this gen), I think the Gigabyte/Aorus boards would be my pick atm (I went w/ a C8H and I'm somewhat unhappy w/ the CPU/memory voltage wonkiness when tweak, but it does give me 30 IOMMU groups, so I'm able to do GPU passthrough via VFIO to a Windows VM w/o any issues).
Note: I haven't dealt with sound yet, which is another potential sticking point if you're doing gaming (most people seem to use a separate sound device, although I've seen some people use network sound when running into issues; HDMI sound I assume would work fine if you're doing GPU pass through) since I was just focused on getting my VR HMD working, which has it's own sound output already.
IOMMU groups are defined by AGESA and the X570 boards in general seem to have much finer grained groupings vs earlier AMD chips/boards (although people seem to have done ACS workarounds). My recommendation is to find something that someone has gotten working already and just follow along (the VFIO subgroup at level1techs and r/VFIO on reddit seem to be the best resources).
Caveats:
- Anti-cheat systems used in multiplayer games generally don't like running in VMs. - Consumer Nvidia cards don't get reset on VM reboot, you need to reboot the machine. - IIRC Nvidia cards need a KVM hack to get the driver to work.
Nvidia cards do require a config workaround for VM detection but the workaround is relatively straightforward (and works w/o further mucking).
Yes, I was rather hoping for some surprise, with 2 Socket, 128 Core Max Mac Pro. I guess Apple has to keep Intel happy until they get their hands on Intel's modem unit.
The honest truth is that AMD is unreliable and their performance specs are only higher currently because they haven't fixed serious security vulnerabilities that would result in reduced multi-processing efficiency.
I hope that's the biggest outcome from this. My inside sources tell me that Intel has their briefs in a bind over AMDs latest technology, and are certainly kerfuffling over it, but I'm not sure how quickly they can respond. It seems like the disparity in this cycle has grown wider than in the past. But that is just my subjective take - I'm not really conversant on the manufacturing part. I'm just historically reviewing the power consumption and benchmarks.
True, it has more cores, but it seems Xeon's have to clock down when executing AVX-512 ops to stay within their power budget.
It's also doing this at almost half the cost and consuming less power ... that's nothing short of astounding.
Xeon's do have a benefit for workloads that can stay in L3 cache as their mesh means latency is stable, whereas Epyc is 16MB per CCX.
Overall memory latency is also lower, which is the tradeoff with the IO die, but it will be interesting to see what real world numbers for DB's etc. say about that.
* With heavy avx instructions churning, AVX should still be a net positive performance-wise even when clocking down.
* It's when you're running AVX mixed with non-vector code that you can see the performance effects of clocking down, since you can drop way down from Turbo on non-vector instructions. Here are the different ranges [1]
[1] - Pages 14-21 https://www.intel.com/content/dam/www/public/us/en/documents...
So there's both AVX-512 vs. clockspeed & AVX-512 vs. silicon budget going on. For many usecases having extra cores is a good trade-off.
The problem with the AVX2/AVX512 clocks with Intel is not the fact that they must clock lower to use them (for pure AVX code, running wider but at the lower clock speed is still worth it!), it's that they need to clock lower pre-emptively for any such instructions, and must remain at this lower clock for a while. This means that code that executes a few AVX instructions every now and then mixed in with a lot of integer code needs to run the whole program at a lower clocks.
In contrast, AMD runs their chip power supply from a huge mimcap built in the chip, meaning that they have margin so they can start executing exceptionally power-hungry instructions, and only clock down reactively if it's actually needed. And also do the clocking down and up with a much finer granularity, clocking up immediately after the need to clock down passes.
AMD won't be in the position to just rest on their laurels for years to come. So the fact that Intel is scrambling to come up with a competitive response shouldn't be a concern as far as competition goes in the short-medium term.
I'd say we have at least a few CPU generations coming from AMD without worrying that they become complacent, even if Intel still comes up short.
[1] http://www.daemonology.net/papers/htt.pdf by none other than our resident cpercival!
I've only followed casually in the last decade+ but I was under the impression that the ATI merger was meant to support a long term bet on better chips by AMD and we're starting to see the fruit of that labor.
But the chiplet architecture, supported by fast Infinity Fabric interconnect is what puts AMD ahead. Intel will have to go same route to stay competitive. And now they are two steps behind.
The smaller dies and modular setup even on the same die gives them so much flexibility in their binning and package integration and must keep defect rates much lower than Intel’s monolithic high core count chips.
Exactly this. Funny Intel is no longer mocking this as dies "glued together".
* Integrated graphics in their CPUs is important for lower-end systems and APU/SoCs. Intel had an integrated GPU at the time, AMD didn't.
* It allowed them to provide the CPU+GPU for all Xbox and Playstation consoles since the acquisition
* It might have allowed them better deals with chip factories because their volume increased
Also, most computers come with integrated graphics today. It's the no brainer choice for most uses.
So it would appear AMD was ahead of the game.
I guess, at the time, the thinking was that GPUs were a core threat to the CPU business itself, since a lot of the buzz was around using the GPU to do CPU activities.
To be honest for a while I thought that was where everything but servers would be going. You can tune your OS and applications around the slower GDDR5 and it seemed a no brainer: instead of 8GB of system memory topping out and then 4GB of VRAM doing jack shit you could have 12GB that could be dynamically assigned to whatever you wanted!
Alas, that dream was never meant to be.
Edit: I posted about this a while a back and apparently Arch Linux has a hacky way to use your VRAM as system RAM. Good stuff.
One thing missing from discussions as the results of Zen 2 EPYC is ARM on Server is now pretty much dead. I don't think any HyperScaler really want ARM on server per se, they simply want better pricing. And having AMD competing with Intel will provide just that.
I don't see ARM Server being a viable alternative in the next 5 years.
Ampere has a chance too, if they get big gains from 7nm and/or massively improve their microarchitecture while keeping the current prices, they'll be unbeatable in price/performance.
> I don't think any HyperScaler really want ARM on server per se
There definitely seems to be a concern with the AMD/Intel duopoly, and Amazon clearly just wants absolute control, since they've already deployed in-house chips that mix ARM cores with <s>Bezos Backdoors</s> fancy cloudy networking and verified boot stuff.
Seems AMD is back
Was the I/A cycle intended just to keep Intel and AMD on their toes, knowing that Google's data centers and software seamlessly supported both platforms? Was there any technical benefit, other than ensuring Google's software was portable?
IBM quoted them as willing to switch to POWER if they could save 10% in energy costs
In most cases, I suspect developers will see improved wall-clock times with substantively worse FLOPS/watt. Good for developers, bad for data-centers.
This junk justification has no longer been relevant for years. Most developers don't care because (1) they rely on core applications that are already multi-threaded (web servers, SQL engines, transcoding, etc), or (2) in today's age of containers, VMs, etc, it doesn't matter to them. Now we scale by adding more containers and VMs per physical machine. Bottom line, data centers always need more cores/threads per machine.
Unsure what the percentage of VM's that use no time sharing or oversubscription is though.
When you have variable length requests, you will find cores will not always be balanced, it is simply a statistical reality. And in those cases, the kernel will have to migrate your process to a different core, and if you have 256 cores, that core might be really far away.
Edit: my source is this German article: https://www.heise.de/newsticker/meldung/AMD-Server-CPUs-Epyc...
Everything goes through the central crossbar on the I/O die, where Zen1 had memory attached directly to each CPU chiplet which would relay as necessary. On Zen1 if you accessed direct attached memory you wouldn't pay the latency penalty from relaying the data. In Zen2 all data is relayed via the I/O die with the associated delay that entails.
The speed of light is constant, and some cores will always be a little closer to various resources.
The main improvement is the max number of hops is log(n) instead N/2.
But yes, writing correct and performant highly parallel code is difficult & error prone, often prohibitively so.
I guess you get paid, and can think of it like a hobby project for your own technical chops. But still.
"Upper management want us to be able to offload burst capacity to AWS, MS, Google or other public provider, do what you can to make it work but I reckon in-house can beat them on pricing"
Six months later -
"Congrats, good work! We showed them, we're getting a new data centre!"
In addition, this kind of work does indirectly help keep Intel competitors viable, which helps keep Intel in check for everyone. Stuff like that is pretty exciting in its own way.
Then I have my professional work, where the timeline is the primary focus, and correctness can only be pursued where it moves the timeline forward.
These projects you’re talking about are an interesting mixture. There is still a timeline that must be hit, because you need to do your demos, and you need to be ready to shift to a production timeline if negotiations go south. But since there are no customers the business model isn’t changing. And you don’t need to do any polish. So you can stay focused on the raw architectural problems.
It’s like my hobby projects in that you can focus on readiness over completeness, but there still is some timeline pressure.
Interesting to think about. Strange for me, but I guess it’s every day for others!
The second is important. Customers need to adopt AMD EPYC. To our readers, it is important when you get a quote to at minimum quote an AMD EPYC alternative on every order. More important, follow through and buy ones where Intel is not competitive. If AMD EPYC 7002, with a massive core count, memory bandwidth, PCIe generation and lane count, power consumption, and pricing advantage cannot take significant share, we are basically done. If AMD does not gain enormous share with this much of a lead, and easy compatibility, Intel officially has a monopoly on the market and companies like Ampere and Marvell should shut down their Arm projects. If AMD does not gain significant share, there is no merit to having a wholistically better product than Intel.
As for bettering cost-performance, the full review gives plenty of context that the new Epyc 2's soundly beat out the current Intel Xeon lineup (often by 2X or more), but I think AMD is also doing what they need to do get marketshare (while still raising their ASPs):
When it comes to the top-bin SKUs, the value proposition is simple, just get a higher-end SKU and consolidate more servers to save money. AMD is extracting value for the higher-core count SKUs. For AMD a chip with 64-cores, 256MB L3 cache, 128x PCIe Gen4 lanes at just under $7000 compares favorably when its nearest Intel Xeon competitors are two Intel Xeon Platinum 8280M SKUs (M for the higher-memory capacity) that run just over $13,000 each. AMD at around $7000 is essentially saying Intel needs to start their discounting at 73% to get competitive, and that is not taking into account using fewer servers.
On the AMD EPYC 7702P side, AMD is calling Intel that if it wants to be performance competitive, it needs to discount two Platinum 8280M’s by 83% plus the incremental cost of a dual-socket server versus a single-socket server. This is a big deal.
* A number of errata (not just 298) delayed full production, sapped performance, or negatively impacted idle power. Take a look at doc 41322, DR-BA step for many samples.
* It was late and didn't achieve performance targets; it missed clock rate targets and 2 MiB L3 was insufficient.
* Intel delivered a very compelling server part (Nehalem) during the lifecycle of family 10h.
Are there performance benchmarks that are designed to measure server application performance?
Google probably has a whole team internally to benchmark their own applications on different hardware.
Intel really messed up, and has no one to blame but themselves.
But anytime there is renewed competition I wonder if the new competitors are better security wise, or just haven't been tested much yet security wise.
I hope AMD is doing better, but I'm not sure how to tell just yet. Things like the speculative execution problems seem to be a general issues inherent to speculative execution, so if AMD is doing it and they become a bigger target I would expect new issues to arise.
The combined issues of these two aspects of the mistake in the theory of speculative execution as well as the implementation is what makes it so bad for Intel vs any one else on the market.
I remember seeing the Linux kernel devs discussing some massive 10+% performance hits back around Meltdown/Spectre patch time, and I’m now wondering what the final impact has been.
https://web.archive.org/web/20180106011413/https://www.epicg...
[0]: https://www.phoronix.com/scan.php?page=article&item=swapgs-s...
If you're doing something that's syscall heavy, you're going to see a big negative difference. If it's something that's CPU heavy without making many syscalls, you're not going to see a lot of difference.
Google cannot have a breach between the services running gmail and the services running adwords for example, even if those are running in the same server on an internal cloud that has strict permissions being enforced by software.
This is especially even more relevant in any kind of datacenter application, even if the company is the sole tenant, because they are working with Client Data - which is data that does not belong to Google, but to their customers.
All of their processes run together on the same machines so you wouldn't want to risk one compromised process access data on possibly any other process.
That proved pretty successful for them. Now they have done it again by commoditizing "high core count" processors.
Each time they do this I wonder if Intel will ever learn that you can't "get away" with selling something for a lot of money that can be made more cheaply forever. Processors are not a veblen good.
They've done much more than that. Intel's current server CPU lineup is tightly siloed into different segments to limit every feature that some customers would pay more for into it's own line priced to match. That's why they currently have Xeon Scalable {Bronze, Silver, Gold, Platinum}, Xeon {D, W, E} lines, with 402 different Xeon CPUs actively being sold.
In contrast, AMD has two EPYC lines, P and non-P, only differing in that P is for 1-socket servers. The models in these lines differ in that they have more/less cores and different clocks, all those features that Intel gates and segments by, are found in every AMD CPU.
Huh, TIL
https://en.wikipedia.org/wiki/Veblen_good
"Veblen goods are types of luxury goods for which the quantity demanded increases as the price increases"
I think the keyword is "forever". They know they can get away with that for a long time though. Because they historically have.
I hope AMD turns their attention to machine learning tasks soon not just against Intel but NVIDIA also. The new Titan RTX GPUs with their extra memory and Nvlink allow for some really awesome tricks to speed up training dramatically but they nerfed it by only selling without a blower style fan making it useless for multi-GPU setups. So the only option is to get Titan RTX rebranded as a Quadro RTX 6000 with a blower style fan for $2,000 markup. $2000 for a fan.
The only way to stop things like this will be competition in the space.
The last time I looked at ROCm (which was something like a year ago) it was "supported" as in "there's reports on the internet that someone got it to work" but when I tried, I couldn't get it to work and it really wasn't worth the effort. If I'd dig out a machine with an AMD GPU right now, can I get it working (like, train MNIST or some other helloworld'y system) within an hour, and is there documentation available on how exactly that should be done?
If you're asking why AMD doesn't make a compiler for CUDA source code that targets their own GPUs--that's basically what ROCm currently does. They're pushing their CUDA alternative, called "HIP", which is essentially just CUDA code with a find-and-replace of "cuda" with "hip". (And other similarly minor changes.) They have an open source "hipify" program that does this automatically (https://github.com/ROCm-Developer-Tools/HIP/tree/master/hipi...).
So, basically, AMD GPUs are already sort of CUDA compatible: just run your CUDA code through hipify, then compile it using the HIP compiler, and run it on a ROCm-supported system (which, for now is the most spotty of all of these steps IMO).
I think there have been a few projects that try to translate CUDA stuff into OpenCL or other AMD-compatible compute platforms.
for $2000, surely you can install an aftermarket cooler or watercooling loop?
[1]: https://forum.level1techs.com/t/cooling-8x-titan-rtx-in-a-se...
I mean people seem to get non-commercial software but why is non-commercial hardware weird?
What weirds me out is that you could have purchased the card, eventually decide to use it in a way that goes against their EULA, and they could just decide to take away your access to the software required to use the hardware.
I guess it’s more or less the same feeling as certain people have about Windows licensing or buying games on certain platforms. You’re buying a license to use whatever it is rather than actually buying the thing, and this kind of feels like an extension of that to hardware.
Now look, Nvidia wants to make more money, let's not pretend that there's really any other primary motivation here. And segmenting the people who make money with their cards and the people who use them for entertainment is a pretty solid way to do that. However, the secondary reason for this is that on the consumer side people complained that stores were constantly sold out of new cards from large businesses buying them all up. Some stores implemented rationing schemes but the 'final solution' it seems is to just stick a line in the license that says you can't stick these cards in your DC.
A kind of clause which might very well be void in tons of jurisdictions, by the way.
Edit: source https://www.phoronix.com/scan.php?page=article&item=amd-epyc...
Jokes aside, I know H.264 was at this point in the past, I just wonder how long it's going to be before we see AV1 hardware encoders that produce good quality video (and hell, hardware decoders at that as well).
AV1 is designed to give you the ability to throw more CPU cycles at the encode side to achieve higher quality/byte while maintaining reasonable decoding performance. You don't have to use it that way, you can encode faster but give up quality/byte. But without AV1 we would not even have that choice, the previous generations reach diminishing returns at some point.
And hardware decoders are not that great. They use less power and can achieve realtime encoding, but if you want the maximum quality/byte (at the expense of encoding time) then software encoders generally reign supreme due to years of iterative improvements that you don't see in hardware. This is important for streaming services which spend those cycles once and then streams to millions. The asymmetry makes it worth it.
When H.264 was introduced, you needed a state of the art CPU just to play it back... You can imagine how slow encoding it was.
Higher resolutions were another matter. 720p needed at least 2GHz Athlon/2.8GHz P4 for smooth H.264 720p 5Mbit playback. Either top of the line 2003 CPUs, or 2006-2007 budget ones.
H.264 1080p 30Mbit bluray released in 2006 could be decoded purely in software on top of the line 2006-2007 CPUs (dual core P4, A64 X2).
It was basically a perfect storm that is unthinkable few years ago. ( To me it is still very much unreal even with today's announcement. ) Intel 10nm cant be fixed in time ( In fact for 24 months, they just keep lying both publicly and in investor's meeting ) , that is Intel's Fab problem. And Icelake couldn't arrive on time because of 10nm, their Design could not adopt to 14nm or other node, it was stuck with 10nm, compare to Apple and AMD which has adopted the train development method where their Design were less fixated on a Node Schedule.
And AMD managed to execute in perfection. Naples set the tone to the industry, Rome ( Zen 2 ) were a huge leap in performance, the Chiplet design gave AMD the advantage in cost ( Smaller Die, Higher Yield, Mass manufacture in volume ), so while it had a lower price than Intel, they are not hurting their margin to fight this battle. Very important for the long term survival of AMD. Along with TSMC 7nm were running in perfect harmony. Not to mention TSMC were willing to fight and get 7nm capacity for AMD. Along with risking more CapeX and building more Fab.
And to add a fifth thing to all these, Intel had major securities problem just months before AMD's Zen 2 launch.
As if the whole thing were scripted to play the perfect Counter Attack by AMD and TSMC. But No, it is the Hard work and dedication of AMD and TSMC, the will to fight and deliver against all odds, compare this to Intel in the past 4 - 5 years.
So if you loathe Intel, now AMD is not only an alternative, but also possibly the best option on Server.
And if you love Intel, you should buy AMD to teach them a lesson for milking and sitting on their butt not Innovating.
Damn thing is faster than my 2700X (which I'll be upgrading to a 3900X when I get back from holiday).
AMD is straight killing it at the moment.
Intel's development process is such that a lot of the components of the chips were expected to be introduced on 10nm. With 10nm in such production trouble, these components mostly weren't backported to 14nm design. As a result, Intel's chip design has mostly stagnated as a result of its 10nm fabrication issues.
If GF had kept up with the competition that wouldn't necessarily be a problem but for the past 10 years GF has struggled to deliver new nodes on time making the WSA a millstone around AMD's neck. With 14nm GlobalFoundaries ended up licensing Samsung's 14nm process and with 7nm they gave up altogether. While the details are confidential the latest amendment to the WSA presumably gives AMD much more freedom to manufacture their newest products elsewhere as GF literally can't.
You pretty much got it. It all basically comes down to the yield you get on waffer. The larger the die, the less you can fit on a waffer, the more chance there is a defect in the die. Using a few smaller dies is a smart way to get high yields out of your waffers.
We've had engineering samples of Rome for quite a while. However, there are very few available boards with PCIe4 right now. The one we've tried (under NDA) has a busted BMC that won't accept network settings, which has really hampered Rome testing. We've actually done most of our Rome testing using 1st generation boards, with slower RAM and PCIe Gen3.
This is true for the vast majority, but there are niche cases were there are differences even though both are both x86_64(like Intel's FlexMigration and AMD-V Extended Migration).
The context of recent news is that Intel has been dominant for many years, prices of processors have been high and growing, and performance improvements minimal. AMD is now offering cheaper and faster processors, which characterises a resurgence of competition in the desktop and server compute market. Ideally Intel will improve their offering in response, otherwise facing loss of the market to AMD, and further continue the competition. No impartial consumer should want either company to prevail, but AMD's present lead signals an end to Intel's dominant position.
Developers benefit directly from cheaper and more plentiful computing, but also indirectly as more applications and approaches become practical for their target platform. As devices become cheaper, the potential user base for an application becomes larger.
Win-Win for consumers either way. And for the last what almost 10 years Intel has been ahead in many use cases - and their new offerings are rather stale IMO.
1. It pushes up the price which helps incentivize AMD executives and key employees who have an equity component to their remuneration
2. It makes it easier/cheaper for AMD to raise funding via equity or equity linked markets if they ever need to
At scale, either or both of those could help them though sales might be better
Buy some stock instead, don't just throw your money away. How would you even 'donate' to them?