AMD to Acquire Xilinx
amd.com
amd.com
My contrarian speculation is that this is a move driven by Xilinx vs. Nvidia given Nvidia’s purchase of Arm and Xilinx’ push into AI/ML. Xilinx is threatened by Nvidia’s move given their dependence on Arm processors in their SOC chips and their ongoing fight in the AI/ML (including autonomous vehicles) product space. My speculation is that this gives Xilinx an alternative high performance AMD64 (and possibly lower performance & lower power x86) "hard cores" to displace the Arm cores.
Interesting times.
Xilinx already has a RISC soft core in their MicroBlaze architecture so they don't have a pressing need for a low power, reasonable performance RISC soft core. Ref: https://en.wikipedia.org/wiki/MicroBlaze
AMD has high performance CPUs being fabbed by TSMC (same foundry as Xilinx), so (theoretically) AMD CPUs can be grafted onto the Xilinx FPGA as a hard core.
With AMD and the MicroBlaze, they have the high performance and low power processor spectrum covered with no need for 3rd party licensing costs.
Note however that Xilinx has a dsp slice (UltraScale) which is a prefabbed adder / multiplier. This would be PyTorch in the analogy.
FPGA LUTs cannot compete against ASICs, so modern FPGAs have thousands of dedicated multipliers to compete.
The LUTs compete against software, while the dedicated 'Ultrascale DSP48 Slices' competes against GPUs or Tensors.
--------
It's not easy, and it's not cheap. But those UltraScale DSP48 units are competitive vs GPUs.
It's still my opinion that GPUs win in most cases, due to being more software based and easier to understand. It is also cheaper to make a GPU. But I can see the argument for Xilinx FPGAs if the problem is just right...
In contrast: the NVidia A100 has 19 32-bit TFlops. Higher than the Xilinx chip, but the Xilinx chip is still within an order of magnitude, and has the benefits of the LUTs still.
-----
It should be noted that Xilinx Versal "AI engine" is a VLIW SIMD-architecture: https://www.xilinx.com/support/documentation/white_papers/wp..., effectively an ASIC-GPU hardwired into the FPGA.
In my experience FPGA>GPU for inference, if you have people who can implement good FPGA designs. And inference is more common than training. Much of this is due to explicit memory management and more memory on FPGA.
ASICs (in this case: a fully dedicated GPU) obviously wins in the situation it is designed for. The A100, and other GPU designs, probably will have higher FLOPs than any FPGA made on the 7nm node.
But not a "lot" more FLOPs, and the additional flexibility of an FPGA could really help in some problems. It really depends on what you're trying to do.
------
At best, 7nm top-of-the-line GPU is ~2x more FLOPs than 7nm top-of-the-line FPGA under today's environment. In reality, it all comes down to how the software was written (and FPGAs could absolutely win in the right situation)
The original comment by brandmeyer said "ASIC", not "GPU".
Take the same RTL. Synthesize it for ASIC and for FPGA. Observe a 20x difference after normalizing for power, area, and clock speed.
I think it also works better to think of the DSP resources as a big systolic array with lots of local connectivity and memory and only sparse remote connectivity. The SIMD model doesn't really apply.
That's an interesting assertion. FPGAs are better at moving data from one place to another than any other general-purpose device I can think of.
How many GPUs have Xilinx GTx-class transceivers, for instance? A GPU with JESD204B/C connectivity would be an extremely interesting piece of hardware.
But its vastly easier to get started with CUDA or OpenCL than it is to get started with a big FPGA.
But its certainly more advanced than a DSP slice (which was only somewhat more complicated than a multiply-and-add circuit).
-------
I guess you can think of it as a tiny 32kB SRAM + CPU though. But its still missing a bunch of parts that most people would consider "part of a CPU". But even a GPU provides synchronization functions for its cores to communicate / synchronize together with.
Are they really competitive from a price / performance perspective? Based on my limited understanding, nvidia GPUs, for example, are several times cheaper for similar performance?
One of the few architectures to beat x86 in price/performance was ARM, because ARM aimed at even smaller and cheaper devices than even x86's ambitions. Ultimately, ARM "out-x86'd" the original x86 business strategy.
-------------
GPUs managed to commoditize themselves thanks to the video game market. Most computers have a GPU in them today, if only for video games (iPhone, Snapdragon, normal PCs, and yes, game consoles). That's an opportunity for GPU-coders, as well as supercomputers who want a "2nd architecture" more suited for a 2nd set of compute problems.
-----
FPGAs will probably never win in price / performance (unless some "commodity purpose" is discovered. I find that highly unlikely). Where FPGAs win is absolute performance, or performance/watt, in some hypothetical tasks that CPUs or GPUs don't do very well. (Ex: BTC Mining, or Deep Learning Systolic Arrays, or... whatever is invented later)
Computers are so cheap, that even a $10,000 FPGA may save more electricity than an equivalent GPU, over the 3 year lifespan of their usage. Electricity costs of data-centers are pretty huge.
The ultimate winner is of course, ASICs, a dedicated circuit for whatever you're trying to do. (Ex: Deep Blue's chess ASIC. Or Alexa's ASIC to interpret voice commands). But FPGAs serve as a stepping stone between CPUs and ASICs.
------
If you have a problem that's already served by a commodity processor, then absolutely use a standard computer! FPGAs are for people who have non-standard problems: weird data-movement, or compute so DENSE that all those cache-layers in the CPU (or GPU) just gets in the way.
Is this flow filled with divide-and-conquer algorithms with very low work per step? Yes. Is that particularly ill-suited to FPGA logic? Yes. Is it unfair to the FPGA? Not in my opinion.
I stand by my claim: If you normalize a general circuit's speed in units of time instead of cycles, then you'll find that ASICs come out much much farther ahead.
[0] https://www.xilinx.com/support/documentation/ip_documentatio...
I'm still debating on getting the Icicle development kit.
A smaller AMD core could be supplied as a hard core on the Xilinx part but would that really be worth it?
Hard cores are not part of the configurable logic matrix but are separate resources on the FPGA. That means they can't be tuned to the use case in the same way as a soft core. The trade-off is that they typically are better optimized with regards to clock frequency and power consumption since the components are made to be a CPU and not generic configurable logic. One example of an FPGA with a hard core CPU would be the Xilinx Zynq devices.
Anyway in a Xilinx chip the bits of data you put in the memory cells determine what logic function gets executed. That works because in general any particular stored memory -- in any computer, anywhere -- is (conceptually) just a logic function, and conversely all logic is implementable with the stuff we conventionally call memory.
But we typically don't do that, because "real" logic made of fixed-function transistors is much faster than logic built with changeable memory cells. However, there's a market for fully-changeable logic--even if it's slower--and that's what Xilinx chips are.
Every CPU is just a bunch of registers and logic. If you hand me a few million discrete NAND gates, I can use them to build an X86, a RISC-V, and ARM, or whatever. It will be the size of a house and it will be very slow, but it will run the binary code for that processor. With a Xilinx chip, you have a few million NAND gates (or NOR gates or inverters or whatever you like) at your disposal and they're all on one chip and you can wire them up however you want with nothing but software. Bingo: You can build an X86 out of pure logic, and it's all on one chip rather than being the size of a house. That's a soft core.
The nice thing about soft cores is that you can build whatever CPU functions you want and leave off the functions you don't need. If you want to change the design, you just download a bunch of new bits to the Xilinx memory cells. Thus you can change an ARM into an X86 in an instant, without changing any hardware.
Soft cores are very flexible, but they're also slow, because implementing logic with static RAM cells is slower than doing it with dedicated transistors.
That's where hard cores come in: A hard core is a dedicated area of silicon on the Xilinx chip carved out to only implement an ARM chip or a PowerPC or other CPU with fixed-function transistors. So it's fast. The downside is you can't change its functionality on-the-fly. If you decide you'd rather have a PowerPC than an ARM chip you have to change the whole chip.
In both types of cores, you still have a bunch of memory cells left over that you can program to do whatever kind of logic you like.
https://hackaday.com/2019/06/22/fpga-soft-cpu-is-superscalar...
1. That's no longer the case, so sucks for Altera / Intel
2. AMD doesn't have a fab, so any advantages are necessarily on the design / architecture / integration side.
And I believe AMD are good with using calculators.
I'm not sure this is worth $35B, but if Lisa Su thinks so, it probably is. She's proven herself to be one of the most capable CEOs in tech.
I'm trying to imagine an x86-based Ultascale+ style processor.
Hopefully AMD can help fix the mess known as vivado and the petalinux tools. lol.
AWS based ARM processor looks to be widely deployed in the cloud.Nvidia, the leader in GPU compute in buying ARM.Intel, which has suffered deeply because of their 10nm fab problems are going to work with TSMC.And AMD's P/E ratio is at 159, higher than Amazon's!
So Maybe AMD is looking to convert some inflated stock with a predictable business.
And it's better to invest in a predictable business that may have possible synergies with yours. Otherwise it looks bad to the stock market.
And Xilinx is probably the biggest company AMD can buy.
it's a relatively safe bet now that intel has more or less conceded leadership through 2023 but it's not zero risk. The market generally doesn't have an appreciation of that, P/E was still nuts even before the release of Zen2 when AMD's success was far less clear (Zen1/Zen+ were far less appealing products and scaled far less well into server class). It's a lot of amateurs (see: r/AMD_Stock on reddit) buying it because they like the company rather than trading on the fundamentals.
Right now the stock market is just nuts in general though, there's so much money from the Fed's injections sloshing around and looking for any productive asset, and tech companies look like a good bet when everyone is stuck at home, building home offices, consuming tech hardware and electronic media. Housing is getting even more weird as well.
If Intel gets their house in order in a couple years, AMD won't have much time to gain market and raise prices. I've rooted for AMD since the K6 days but I think there's a risk that they'll always be #2(or less).
This can be a relatively fast operation. Seconds or less depending on complexity.
I've always been pretty skeptical of their approach though, in order to be usable they'd need excellent tooling to support the feature, and if there's one thing that existing FPGA software isn't it's "excellent".
Getting FPGAs to perform well is often an art more than a science ("hey guys, let's try a different seed to see if we get better timings") so the idea that non-hardware people would start to routinely generate FPGA bitstreams for their projects is so implausible that it's almost comical to me.
Maybe one day we'll have a GCC/LLVM for FPGAs and it'll be a different story.
In theory, you could even page out code, but I guess the speed of that will be slow. Also, paging in probably would be challenging because the logical units aren’t uniform (if only because not all of them will be connected to external wires)
An enormous crossbar could solve that, but I would think that would be way too costly, if practically possible at all.
Repurposing FPGAs to different tasks means loading a new bitstream into the device every time. So it is much more efficient to grant exclusive access to each user of the device for long stretches od time. The proper pattern for that is more like a job queue.
Where FPGAs win are new architectures, like Systolic engines. Entirely different computer designs from the ground up.
Basically you analyze the code for candidates, select a candidate, upload your custom hardware design, run your operation on the hardware, and repeat.
The difficult part is that uploading your hardware to FPGA is in the order of tenths of seconds, which is ages when compared to the nano and micro seconds your CPU works. So your specific operation must be worthwhile to upload.
A bit of FPGA on your CPU makes it more flexible, for example your could set a profile such as 'crypto' or 'video' to add some specific hardware acceleration to you general purpose CPU.
Imagine your CPU being able to switch your embedded GPU into another CPU core.
Let's say the current zen 2 had an FPGA onboard. AMD could sell you an upgraded design with AV1 support for a few dollars. Most people aren't going to buy a new CPU on the basis of a video decoder, but they'll buy an upgrade to the chip that auto "installs" itself. That's a sale AMD otherwise wouldn't have made.
But maybe I'm being overly optimistic. (Probably because—disclosure—I'm long AMD. Been long for years.)
If I remember correctly, there was something similar back in the early HyperTransport days...
I'm just looking at this from a logical chain of "who needs FPGAs in their computers?" => "cases with loooots of specific data crunching" => "want a controlling/driving CPU for the complicated parts, but then just concentrate as much FPGA in as possible." => Multi-socket with 1 CPU & rest FPGAs.
(There currently is no commodity Quad-socket SP3 mainboard, not sure if this is a design limitation or just no one made one yet? I'd still say the approach works great with only 2 sockets.)
They had to redesign ReefShark and cancel dividends. It was a huge setback.
It was not Nokia SoC just plain Stratix10. They moved to own SoC after that glorious project.
https://semiaccurate.com/2018/07/02/intel-custom-foundrys-10...
It'd be interesting to see how AMD will execute and integrate this acquisition, considering they are less of a madhouse company than Intel.
Those toolchain disasters are not quite as hilarious when you have to use them daily....
I'm really surprised that Lattice hasn't tried to go around Xilinx and Altera by doing exactly that. You would think that an open bitstream format and a couple million dollars thrown at academic researchers (Lattice makes about $200 million per quarter in gross profit) would produce some real progress, but I digress ...
SystemVerilog, on the other hand, was specifically created because Verilog and SystemC got loose to the end users and the EDA companies were not going to make that mistake again. So, yeah, SystemVerilog is pretty bad.
These "tools" have no target so no incentive to improve. To use them you have to basically push their results back into a Cadence/Synopsys/Mentor toolchain anyway, so you might as well stick to the supported toolchain.
> The reality is SystemVerilog is a huge language, and already an open standard yet no open source project supports it fully
Most commercial systems don't support it fully. And its not clear that SystemVerilog is that superior to VHDL. And, for quite a while, SystemVerilog wasn't open and had some fairly obnoxious patents surrounding it. I don't know when/if that has changed as I have been out of semiconductors for about 20 years now.
Icarus Verilog has been slowly supporting features from SystemVerilog but doesn't have a lot of manpower.
In general, the consolidation of the semiconductor industry and EDA has hurt open-source EDA improvements. There's not very much money coming from companies to fund EDA research. EDA startups can't really get venture funding since VC's all want to fund the next pile of viral social trashware. And anyone with good software skills left the semiconductor industry eons ago because the pay differential is ridiculous.
> The reason is if it has not happened for the first (and arguably easiest) step in the chain i.e. System Verilog, they why would it happen for the others?
Ehhhhh, I don't think I buy this at all. There are dozens of alt-HDLs out there, many of which are quite powerful, designed by solo users. People had working, simple-but-practical PnR for real devices in a ~7k C++ LOC codebase written by an individual (arachne-pnr) and many individuals have independently reverse engineered small-ish scale device families for packing utilities. nextpnr was written by a very small group (solo?) in a year or something. I don't think you could fit an equivalent parser for SV2017 in ~7k LOC, much less elaboration, type checking, a netlist database, to all go along with it. SystemVerilog might actually be the most difficult part of the whole equation because it simply has so much surface area. PnR tools are limited by their target: only targeting small iCE40 devices? Your PNR algorithms don't need to be cutting edge. Targeting SV2017? Your job is hard no matter what device you synthesize for. And I can't think of even a single commercial tool I know from any vendor that supports all of it, up-to-date with SV2017.
All that said, I use SystemVerilog as my "normal" RTL when using commercial tools for stitching together IP, wiring up top modules, etc.
The free part is valuable not in that it's cheap, but in that it saves you from having to deal with licensing.
DevOps pioneers hailed from the likes of Google, Amazon and Facebook, who are not exactly short on cash, but you simply couldn't do what they did if you had been nickeled and dimed at every VM and container.
Bitstreams are closed. There's little to no point in doing an open source compiler if the target is not just proprietary, but deliberately opaque.
Overall your comment strikes me as what a proprietary compiler advocate would say in the 90s. "GCC? Lol"
Since then, Microsoft had to include Linux in Windows just because they absolutely needed Docker. DevOps was invented based on free/open source, it just couldn't be done proprietary style by a company as large as Microsoft.
The limitation here is writing the SystemVerilog parser and compiler.
As for the incentive I'm fairly pessimistic. There is definitely no money to be made for a start-up in this space, it is way too conservative. Maybe the hobbyist intellectual challenge of working on some hard problems like constraint solving or formal property proving? There is a massive task of writing a SystemVerilog parser before you get there though and the SAT solving and property proving problems are present elsewhere with lowers barriers to entry.
Having said that, there has been some promising F/OSS work on the small Lattice devices. It allows for a decent, modern workflow, and it's possible because the devices are approachable, but also because Lattice hasn't been hostile. Why they haven't been more supportive is a mystery to me however.
SystemVerilog is a good examle of an organically grown language with no 'benevolent dictator'. A few pet peeves:
* Why is the simulation delta cycle split into 17 regions? Exactly when does the Pre-Re-NBA region happen and what assignments take place there?
* Why can't a function return a dynamic/associative array or a queue? This is clearly possible, since the array find functions return a queue, but it's not possible to define a user function with this return type.
* It has way too much cruft. E.g. what problem does the forkjoin keyword solve? Who thought that was necessary and why? Not a fork-join block, the forkjoin keyword.
* Why can't you have a modport inside a modport? This would be great for e.g. register interfaces, but modports are not composable.
* What is the difference between a const variable and a localparam and why does the language need both constructs?
* Is a covergroup a class or what? It behaves very much like it is, it has a constructor, some class local information and at least one class local function (the sample() function), but you can't extend it.
* Why are begin-end used for scope delimitation everywhere except in constraints where curly brackets are used? I know it was a Cadence donation, but why wasn't the syntax changed before it was merged? Backwards compatibility can only justify so much...
//rant off
edit: formatting
As for SV - a lot of your gripes are Verilog issues, and SV has tried to fix some of them. I agree the blocking / nonblocking is a mess but most folks just learn the rules to avoid issues, but delta cycles can be a pain. The syntax limitations/quirks you point out are intersting, though not enough to say the language is terrible, it's extremely powerful with very good composability of types, constrained random is very powerful, the coverage is extensive, assertions again are very powerful. In a way its line a few seperate languages bolted together so sure there is some duplication, but it works surprisingly well in the whole.
You only start paying for FPGA tools when you need the really big FPGAs.
And, I'll go out on a limb, but, at this point, I think Arduino causes more harm to beginning embedded developers than good. Yeah, the ecosystem is wonderful if you aren't a developer.
However, Arduino is now weird compared to mainstream embedded development. Most things have converged to 32-bit instead of 8-bit. Arm Cortex-M is now mainstream so your architectural understanding is useless. 5V causes a lot of grief given that everybody else in the world is at 3V/3.3V.
A developer basically has to unlearn a bunch of things to move up from an Arduino. I still recommend Arduino to non-developers or somebody just trying to throw together a project, but I no longer recommend them to someone actually trying to learn embedded development.
FPGA are being used in many type of applications where real-time is necessary and non-recurrent engineering (NRE) cost need to be minimized, for example here [1].
One classic example is that if you poke under the hood of any signal generator like AWGs, you will probably find an FPGA inside. As you probably aware since you in hardware business, AWGs are probably one of most common equipment in any electronic and electrical labs or companies.
[1]https://www.electronicdesign.com/technologies/fpgas/article/...
1) You underestimate how critical prototyping has become, again likely since you say it's been a couple decades. Time to market has become more important, and verification has become harder as CPUs have gotten even more complex. FPGAs enable cosimulation and emulation, leading to faster iteration of both design and verification efforts and thus better TTM.
FPGAs are so important in the hardware development process that I would even say you're not a serious hardware company if you don't have any FPGA frameworks to design silicon.
2) As others have mentioned, FPGAs are also critical for low-latency workloads that require constant tweaks-- high frequency trading (ugh...) comes to mind. The need for "constant tweaks" could also be satisfied with just "normal" software, but that has higher latency as opposed to an FPGA, and FPGAs can get some crazy performance if you're willing to pay the price (south of 7 figures).
Overall sure, usage of FPGAs might be niche compared to, idk, Javascript; but it's commonplace/practically essential in hardware.
Configurable as in one SKU is in several products, but not necessarily reconfigurable by the end user.
FPGAs are already incredibly popular. They're just mostly in things you are unlikely to personally own or know about. You're going to find at minimum one, but probably more FPGAs in things like big routers and other telecom equipment, e.g. cell towers, firewalls, load balancers, enterprise wifi controllers, video conferencing hardware, test equipment like oscilloscopes, sensor buoys, scientific instruments, MRI machines, LIDARs, high end radio equipment, or even just glue logic tying together other components, like in the iphone.
You are precisely correct. FPGAs are useful when your volume doesn't reach volumes where an ASIC would get amortized.
Networking companies (Cisco, Juniper, etc.) are classically big consumers of FPGAs.
Tektronix seems to make quite a bit of money and there is at least one FPGA in practically every test instrument they make. This holds true for practically all test instrument manufacturers.
I know a LOT of industrial automation and testing companies that generally have FPGAs in their systems. Both for latency and for legacy support (Yeah, GPIB still exists ...).
Yes, they aren't "Arm in a cell phone" type volumes, but that doesn't mean they aren't quite profitable if you can aggregate them.
The big benefit of having FPGA closely attached to CPU is that you can access the memory and internal buses quickly. Transferring stuff over PCIe hurts a lot. So you could make an argument for jobs using small work units requiring fast turnaround; CUDA kernels take milliseconds to launch.
I worked with some of the early Xeon+FPGA parts and there just wasn't that much we could do with them. There wasn't enough fabric to build anything meaningful and we had an abundance of CPU cores, so the best we could do was specialized I/O accelerators.
The problem is that if a task is common then someone is just going to make an ASIC to do it. And if its uncommon then the terrible FPGA software ecosystem and low prevalence of general purpose FPGAs in the wild mean that people will just do it on a CPU or GPU.
This is true, but keep in mind that that sort of algorithm runs insanely well on any CPU or GPU because they, too, do not want to touch main memory. You would be blown away by how much work a CPU can do if you can keep the working set within L1 cache.
Re. ASICs, it's a continuum:
- "flexible, low performance, cheap in small quantities" (CPUs)
- "reasonably flexible, better performance, cheap-ish in small quantities" (GPUs)
- "inflexible, best performance, expensive in small quantities" (ASICs)
FPGAs fit somewhere between GPUs and ASICs -- poor flexibility, maybe great performance, moderate small-quantity price.
If your problem is too big for GPUs, as you say, sometimes it's easiest to jump straight to an ASIC. But it's such a narrow window in the HPC landscape. The vast majority of customers, even with large problems, are just buying a lot of GPUs. They're using off-the-shelf frameworks even though a custom CUDA kernel would give them 10x performance and 10% cost. The cost to go to an FPGA is too great and the performance gain simply isn't there.
https://www.nextplatform.com/2020/01/31/when-will-fpgas-outw...
https://www.nextplatform.com/2018/03/19/fpga-maker-xilinx-sa...
If AMD wants to get serious in the datacenter/AI/ML space they need a xilinx-like approach to developing tooling. Cuda, nvenc, cudnn etc craps all over amd's offerings in the same space where they are even available.
AMD is prepping to take over datacenter terf and this puts them in a good place to bring a bigger offering.
With PIM your CPU resources grow with the size of your memory. All you have to do is partition your data and then just write regular C code with the only difference being that it is executed by a processor inside your RAM.
Having more cores is basically the same thing as having more DSP slices. Since those cores are directly embedded inside memory they have high data locality which is basically the only other benefit FPGAs have over CPUs (assuming same number of DSP and cores). Obviously it's easier to program than either GPUs or FPGAs.
FPGAs are not an assembly line at all; the assembly line analogy applies much more closely to a processor's pipeline.
FPGAs are just a massive set of very simple logic units which can be interconnected in many different ways. FPGAs are best used in situations where you want to perform a series of simple operations on a massive incoming dataset, in parallel, especially in real-time situations. Performing domain transforms on data coming in from sensor arrays is one very good application for FPGAs.
So CPLDs will have some kind of NVRAM wear-out concern, and this is almost always specified as a number of maximum erase & program cycles.
GP is also correct that DSP/SRAM blocks are critical to performance. FPGAs are not very efficient at raw compute if you have to synthesize everything out of LEs.
The performance benefit of FPGAs, which PIMs also share (in theory, there aren't any PIMs ready for real-world deployment AFAIK) is that they can leverage much larger memory bandwidths than general purpose CPUs can. An FPGA might run at a lower clock rate (low 100s of MHz), but be able to operate on several kb per clock cycle. This can work really well when paired with off-chip logic to convert high rate serial interfaces to lower clock rate parallel interfaces, then back after the FPGA is done processing.
There is also a lot of work going on in the space of time-division multiplexing FPGAs effectively. The two main approaches are overlay architectures and partial reconfiguration. The former implements another high-level fabric on top of the FPGA which will be less general-purpose, but can be reconfigured faster. The latter is a feature vendors have added to some high-end chips where specific regions of the FPGA can be reconfigured without affecting other regions.
One common, and very good application for FPGAs is for use in Active Electronically Scanned Array radar, sonar, or camera image processing. You can perform parallel filtering and transforms with various frequency and phase settings, which would be impossible for a similarly-sized processor to do.
FPGAs have the potential to revolutionize sensor arrays, by making them much more useful and affordable.
FPGAs are what you really want when you need to deal with high resolution data that is coming in at very high data rates. Often even a very fast general-purpose processor with hand-tuned assembly simply won't have even the theoretical memory throughput to process your data without "dropping frames". They also have the benefit of deterministic performance, which with modern caching/branch prediction systems you can't guarantee (AFAIK, my computer architecture knowledge isn't that cutting edge).
They can also work really well if you have some computation you want to do that is so far off the beaten path for general-purpose processors (or so memory bound) that FPGAs can take the cake.
There is also some work in sprinkling even more hardlogic into the FPGA dies, like processors or accelerator cores for various applications. FPGAs are great for implementing the glue logic to move data between those.
Also agree that additional hard logic or peripherals will be a game-changer for FPGAs, though they would make each design more domain-specific. Alternatively, we may see a shift in how the interconnects are done, which allows for flexible use of these 'modules'. It's also possible that we'll see continual increases in LE counts which make more specialized hardware unnecessary. I don't know which way things will go.
The surface area for security vulnerabilities is already impossibly high. Do we really want to add "firmware running on a DIMM exfiltrating key material" to that list?
The same holds for any other vendor.
From what I understand, open sourcing the bitstream format in its entirety will only do so much but it would certainly help. It's not just building GCC for FPGAs
My point is Xilinx have already proven ARM CPU+FPGA on one die and I think AMD CPU+FPGA is very likely to be a success.
Between this, ARM adoption, Apple Silicon and similar offerings (which kind of skipped ARM+FPGA for ARM+ASIC), RISC-V, it's like 1992 again with exciting architectures. Only this time software abstraction is much better so there is not a huge pressure to converge on only 1-2 architectures.
Edit: technically, the arm part is also ASIC, but you get what I mean
What in the world FPGAs have to do in a datacentre?
Think ML, networking and other such uses...
In our future datacentre we want to say how many cores, connected to how much ram, how much GPU resource, some NVME etc. etc. and there's going to be a whole lot of very specialised switching and tunnelling going on. This needs to be as close to the cores/cache as possible, a good order of magnitude faster than we run our present networking stuff, and probably an area where there will be a significant pace of development ie a software defined solution would be nice.
So, a software defined north bridge, in essence. And an FPGA is pretty much the only thing we have right now that could do the job.
From what little I’ve seen in this space, FPGAs have not made large inroads in the ML space or datacenters in general. This seems partly due to their inefficient nature compared to ASICS and moreover their software.
Unless AMD is planning something really ambitious (e.g., true software-based hardware reconfiguration that doesn’t require HDL knowledge) and are confident they’ve figured it out, I’m not sure what they hope to achieve here.
I don't know that they have actually made large inroads into those spaces, but Xilinx is indeed pushing hard for that. For years now.
This has been a holy grail for at least two decades. Its very much like asking for a programming language that can be used by non-programmers.
To asnwer my question maybe the market is so volatile that they cannot do strategic planning like that?
Often it is an expertise thing, especially when buying smaller companies. See https://en.wikipedia.org/wiki/Acqui-hiring
With larger purchases like this one that can still be part of the equation, though there is also the matter of lead times needed to bring a significant team and related infrastructure needed for the project(s) online and up to speed.
Also if a company is seen as ripe for buying, it can sometimes be done in part to stop a competitor getting a chance at the above advantages.
I suspect a mix of all three is at play here.
The more risk option is the just acquire a company that's done it all for you.
Of course that does leave the merging part. But on paper it looks fast and easy.
In my opinion, I would also like for AMD to invest in ML tooling while they have the cash.
I hope one day Pytorch, XLA, Glow would have native AMDGPU support, and I will be able to buy a couple Radeon 6000 series cards, undervolt them and make a good ML box.
I think AMD gpus on TSMC 7nm, then maybe even 5nm, will have the best performance or watt. Even though they might be 10% or 20% slower than the alternative. For me performance per watt and dollar is more important.
Anyway, it's sad that they couldn't make a 5 to 10 people (I might be too optimistic) engineering team that would make their product relevant in this market.
Lattice is probably the next biggest. There's also Microchip (< Microsemi < Actel), Quicklogic, and Gowin.
Nobody really came close to competing with Altera / Xilinx at the high end, though.
There's a couple of upstarts in China like Gowin and Anlogic, but they haven't made much of an impact in the larger market yet.
"... Across the different market segments where it operates, Xilinx brought in revenues of $767 million last quarter."
https://siliconangle.com/2020/10/27/official-amd-snap-xilinx...
I think the more interesting news is what they are going to do pro-actively with these mergers rather than just sitting on it.
I really hope their respective CEOs will take a page from the open source Linux/Android and GCC/LLVM revolutions. I'd say the chip makers companies are the ones that benefit most (largest beneficiary) from the these open source movement not the end users. To understand this situation we need to understand the economic rules of complementary goods or commodity [1].
In the case of chip makers if the price of designing/researching/maintaining OS like Linux/Android and the compilers infrastructure is minimized (i.e. close to zero) they can basically sell the hardware of their processors at a premium price with handsome profits. If on another hand, the OSes and the compilers are expensive, their profit will be inversely proportional to the complementary elements' (e.g. OSes & compilers) prices.
Unfortunately as of now, the design tools or CAD software for hardware design and programming, and also parallel processing design tools are prohibitively expensive, disjointed and cumbersome (hence expensive manpower), and if you're in the industry you know that it's not an exaggeration.
Having said that, I think it's the best for Intel/AMD and the chip design industry to fund and promote robust free and open source software development tools for their ASIC design including CPU/GPU/TPU/FPGA combo design.
IMHO, ETH Zurich's LLHD [2] and Chris Lattner's LLVM effort on MLIR [3] are moving in the right direction for pushing the envelope and consolidation of these tools (i.e. one design tool to rule them all). If any Intel or AMD guys are reading this you guys need to knock your CEO/CTO's doors and convinced them to make these complementary commodity (design and programming tools) as good and as cheap as possible or better free.
[1]https://www.jstor.org/stable/2352194?seq=1
[2]https://iis.ee.ethz.ch/research/research-groups/Digital%20Ci...
[3]https://llvm.org/devmtg/2019-04/slides/Keynote-ShpeismanLatt...
It's really obvious when you think about it. If you sell nails, you want to make sure that everyone has or can afford a hammer, and hammer manufacturers like to make sure that there is a large supply of compatible nails.
As much as I would like to see it, I am not sure the equation is that simple in the case of CAD software. Sure, that would make it easier to use FPGAs, but it would also make it easier to create competing products, as a stretch.
I still think it's worth it, and wish bitcode format was documented, at the very least.
I remember the time when Elon Musk said to an analyst that he's asking boring questions to fill in his spreadsheet, and I'm feeling the same thing while listening to the earnings call.
I've become a huge AMD fan, both because of their hardware, and because of their commitment to open source. But while the battles they have one against Intel on the x86 side are impressive, it seems that CUDA is leaving them far behind.
What's funny is that the same strategy (leaving out specialized instructions from consumer level hardware) that worked extremely well for CPUs won't work for GPUs in my opinion.
If you look at ray tracing hardware (I have it on my RTX 2070 Max-Q card in my laptop), it sucks right now, but it's improving very fast as machine learning algorithms improve.
I just found this:
https://www.tomshardware.com/news/amd-big_navi-rdna2-all-we-...
One thing that I forgot is that AMD can just focus on inferencing hardware (INT16 operations), and leave out tensor cores...so actually you are right, I'll just stay with NVIDIA GPUs.
It seems like Lisa Su thinks that there's a separate ,,gamer market'' and ,,accelerator market''.
Jensen Huang understands that the same person can like to play games and train machine learning models on the same machine.
I'd love to switch to AMD CPU to have a portable laptop with low resource usage, as I spend most of my time travelling, but as GPUs in the cloud are overpriced (thanks to Jensen with separated pricing for servers), and internet in hotels are unpredictable, I don't want to train models in the cloud.
Anyways, Lisa said that she reads all comments about AMD, so I hope she'll listen :)
Judging from watching the Financial News and reading analyst's comment for years. My feeling was that their Job was not to push for hard question or an honest answer. Their job is to push whatever interest they had with the company. So a spin for better long term prospect and downplay risk.
I was happy with the Enterprise results ( +116% YoY ) until I read this
>Revenue was higher year-over-year and quarter-over-quarter due to higher semi-custom product sales and increased EPYC processor sales.
Semi-Custom is definitely PS5 and Xbox.
Basically I still dont see EPYC making enough inroad in the Server Market. And this is worrying, while the Stocks, Reviews, Hype are all going to AMD. No Results so far have shown Intel is hurt or AMD is making big gains in market shares and revenue shares.
The only good part I guess is Ryzen Mobile contribution to Computing and Graphics segment.
AMD EPYC on AWS is 10.42% cheaper than Intel per hour across the board for m5 instances (9.6% cheaper for t3 instances). 7nm EPYC saves more than 10% power vs Intel 14nm and per-chip savings from buying AMD are way more than 10% vs similar Intel offerings. Why can Amazon spike prices that much? Because people will pay and still consider it a deal.
AMD's main issue still seems to be available due to do much competition for 7nm and tsmc being reluctant to build new fabs.
The more fabs they have the slower their node progression will be and the cost of each new node these days seems to be almost exponential.
TSMC dont have capacity problem. They are very much willing to built Fabs if their customers are committed to it. But no company other than Apple has ever done that. AMD could have place a large order over the course of the year. Basically a bet that they will sell that many chips. And TSMC will adjust or built accordingly. That problem is no company is willing to place that bet. What if they dont sell and stuck with big pile of chips?
This is simple Supply Chain Management, it is the same principle in every other industry.
Semi-Custom is indeed PS5 and Xbox. This increase was expected of course. It's lower margins than other segments though.
Epyc adoption is indeed slower than I would've liked. But so far it has matched or beat short and long term projections from AMD.
Enterprise is weird. AMD has better price and performance? Let's buy Intel. Meltdown lowers performance? Let's compensate by buying more Intel. Intel is supply constrained? Let's complain, and still buy Intel.
My guess is Intel is still pressuring OEMs to favor Intel. Cloud providers are increasing adoption though, and there have been a few nice HPC wins.
And lets not forget that AMD is selling every chip they can, TSMC production is fully booked.
If you look at the incentives a bit, management gets to decide which analysts can ask questions, so analysts need to stay on management's good side.
Analysts know the company's figures inside and out (often have a 10-tab spreadsheet with an extensive operational model of the company), and are asking questions to tweak key model assumptions.
So analysts ask pointed questions in shared jargon with management. You don't ask 'are you seeing a big sales drop because the crypto bubble blew up?', you say 'can you provide some color on when you excess inventory will clear from the channel?'
Analysts get their answers. Management avoids bad headlines written by casual listeners.
There's an additional layer, which is the analysts know the industry and company very well, so general bad things are already background knowledge (there's no reason to ask about them). If you want to know what the analyst already knows, then pay for their report—they're not reporters fishing for a sound bite.
Are there any public resources dissecting an earnings call? Doesn't have to be recent or AMD.
Just as an example, I remember reading a lot about AirBnB here at HN when it wasn't even known in Eastern Europe. I suggested him to use it to rent out his luxury apartment that he just bought there, and he was the first one to rent out a luxury apartment in the country (also somebody from AirBnB-s management flew there personally)...he made lots of money from rental fees of course, but also AirBnB got incredibly successful. There are lots of other examples of course, this was just the least controversial that I can write here :)
Take everything you read with a big grain of salt though. You get what you pay for, especially in finance—'free' content usually comes with an agenda.
AMD's performance is fantastic, which is why I was fine getting it for a home gaming desktop, but I'm not sure I'd be willing to pull the trigger on AMD on an enterprise server buildout. Intel still simply (actually or otherwise) feels more reliable.