AMD Is in Advanced Talks to Buy Xilinx
wsj.com
wsj.com
But things has changed since 2016 - 2018. I used to see the move to FPGA as going offensive attack in the server space, now both AMD, and Intel are doing what ever they can to fence off ARM.
You should sit on some TSLA for now also.
I never heard much more about that, but I can imagine it might've been popular with certain specialized workloads, in an era before GPU compute took off, or perhaps workloads that don't fit GPUs well.
I wonder if those specialized applications might be a small but important market.
Really cool to hear this used to be a thing with HT (though long before chiplets), gives me some hope we'll see something like that again.
what makes you believe that is has anything to do with "values"? ;)
AMD really needs to focus on software if they want to catch Nvidia though.
Hope for this, the Nvidia monopoly is hurting ML, I wished to have something else but am forced to buy Nvidias offerings b/c hardware and tool support.
This could just be a strategic acquisition for internal purposes: I bet a lot of design and architecture validation is done on big FPGAs. Maybe they wanted some custom ones?
There was a talk at 32c3 (2015) by David Kaplan from AMD about developing and testing/verifying real world x86 CPU designs:
https://media.ccc.de/v/32c3-7171-when_hardware_must_just_wor...
Around minute ~11/12 he briefly mentions (among other methods) emulation on massive FPGA workstations costing something in the order of $1M.
It is.
And the board and software used for this are so damn expensive they probably account for half of xilinx income despite the low quantities
[0]: https://www.hpcwire.com/2011/07/13/jp_morgan_buys_into_fpga_...
Edit: Actually I'm thinking of Nvidia, with their ML supported by TensorFlow. AMD is lacking here.
From an operation perspective, this combines two non-overlapping TSMC customers that can potentially negotiate better wafer prices together. I assume that AMD has significantly higher wafer counts than Xilinx so this would primarily benefit the Xilinx business.
From a technology perspective, this acquisition may succeed where the Intel/Altera acquisition fizzled due to AMD's chiplet approach. Swapping a processor core for an FPGA chiplet might be an "easy" win in some markets.
Intel has been using EIMB quite effectively to allow for very interesting fpga configurability, letting purchasers pick from a variety of "front end" transcievers to match their app[1]. This doesn't combine multiple big processors, but is still a very interesting tech!
Intel's also been working to standardize & accelerate mulit-chip connectivity, by donating their Advanced Interface Bus to the open group the CHIPS Alliance[2]. This saw a 2.0 draft emerge this summer.
Agreed that Intel has been fairly ineffective at driving interesting new fpga adoption. There has been interesting work but it's hard to see what adoption has looked like. There is the Xeon 6138P for example, which uses a UPI link and dual 8x PCIe links of a 20 core xeon chip to talk to an on-package Arria 10GX[3]. I'm not a market researcher, but my guess is that adoption hasn't been stellar. I haven't seen any follow ups. I do hope there have been some interesting & good uses though! I feel like AMD will face similar challenges to adoption. Frankly, tech like Xe and Radeon seems more appealing to the world we live in; big gpus that can do bfloat16 &c. One of the big advantages of fpgas has been great connectivity, powerful transcievers & other i/o capabilities, which, if you put them in an existing platform socket, is probably going to squander some of those capabilities, I feel like?
[1] https://www.anandtech.com/show/14211/intels-interconnected-f...
[2] https://www.anandtech.com/show/15434/intel-joins-chips-allia...
[3] https://www.anandtech.com/show/12773/intel-shows-xeon-scalab...
Xilinx had an opportunity to be in the position NVidia is in today and it was not obvious in 2005 who was going to win the high performance computing (HPC) market because the inherent advantages that FPGAs had and still have today for high speed I/O, RAM throughput, and hard timing requirements. NVidia's CUDA, on the other hand, produces generally painlessly portable results, with easy improvement on new devices. They have essentially won in HPC application development.
Fundamentally the FPGA companies need to adopt an open-source, cloud-first ethos towards their stacks and especially be focused on making agile low-level application development. Making a counter control some LEDs should be a 10 second process. Changing the clock speed should be a 1 seconds process. Adding a button to reset the counter should be a 1 second process. None of this is true. Good luck getting this to work within one day as a newbie on a vanilla machine. Good luck even installing Vivado and compiling any bitstream in one day (thank God they have AWS images -- good luck getting that going within 2 hours). Good luck coordinating with source control and generally merging work with a team.
There was some hope when Intel bought Altera that the compilers engineering expertise might help improve virtual CPU stacks, or at least some investment in the usability of their software tools might have been thought of as a competitive advantage, but all the FPGA vendors have seriously dropped the ball on improving the eco-system with usable software. FPGA software is so bad that you have to justify a 6 month development cycle for something that should probably take a few days if things were even remotely sane. It's hard to enumerate all the examples of broken-ness but small changes end up taking a long time to re-build, things like inverting your reset in a module and then waiting two hours for trivial changes to resynthesize and place-and-route. (edit, two hours later the trivial reset polarity inversion failed timing, because it also required me to add synchronization stages to send the inverted reset signal across clock domains, shouldn't be doing this at night).
This is probably good news only for Lattice Semiconductor.
I feel this pain!!!!
Absolutely love the flexibility of Xilinx's System-On-Module architectures but their Software toolchain is arguably the most time-wasting frustrating piece of software I have ever had to work with.
We need an open-source LLVM-like toolchain equivalent for FPGA's. I hope AMD has enough muscle to push for this.
I would also like to point out that there are projects out there pushing this boundary, David Shah has made amazing contributions towards this, with projects like SymbiFlow and yosys.
Being involved with FPGAs for 20+ years now, and Xilinx for the most part of it, I have witnessed the progressive decay of their software quality. The legacy Xilinx tools (Foundation) were far from perfect but at least useable -- and hackable to a degree if you needed to implement workarounds. Those were many times even suggested by Xilinx own employees that were easy to approach and responsive back in the days.
Altera Quartus was always better in terms of project organization and user interface, including their hardware tools. Third party software such as Synplify (before Synopsys bought their maker Synplicity) were light-years ahead, especially on the synthesis introspection and support for design and RTL troubleshooting.
People tend to justify the software quality and vendor lock-in because these are niche markets. I think that is a lame statement. You could tell both Xilinx and Altera, being "fabless hardware companies", treated their software departments as a necessary burden and not a critical part of their business model, which by de facto it was. The software started going downhill, progressively accumulating technical debt while it kept getting even more cryptic and expensive. I think they offshored a significant part of its maintenance in the last decade. Each release fixed bugs but astonishingly introduced very stupid new ones that a minimally decent software verification process should have caught.
Vivado is a mess. You can optimize your workflow digging down deeply with TCL scripts (yikes, its 2020!). However who has the time for learning their badly designed API and fighting with the quirkiness of their tools? I am amazed some of us use it in regulated environments where our tools need to be properly validated. I have had designs that take more than one day to synthesize, map and route. Having being bitten by some bugs and my own errors dealing with the cumbersome IDE, I always painfully go through all the logs of each step, and also many times recreate the project and rebuild just to check if the output is reproducible.
You can always tell when someone is a seasoned FPGA veteran by the quality of their rants about the tools.
I recently encountered a bug in the Vivado IP integrator packaging whatever crap where if you left the "Vendor name" field at its default value '(none)', the packaging aborted with an error because that field contained invalid characters: the parentheses. That's a bug that any kind of testing should have caught (two bugs, actually).
Also: Why does Vivado have to treat any little IP core I create like it's some full-fledged library that exists independently of my main project? I just want to attach a 50-line HDL module to an AXI interconnect in the block diagram – I don't want to create a new project for this, ffs.
How would you envision an "FPGA on board the CPU" differing from this? Would it still be a separate thing the way GPU is today, or would you see it more as like a new set of CPU instructions which enable little blobs of HDL to be somehow passed in and invoked, kind of like how a shader works today on GPUs?
Beyond that, there's been research in the past that envisions using a FPGA fabric as an execution unit within the CPU itself: https://www.microsoft.com/en-us/research/project/emips/ Intel hasn't pursued such a thing but that doesn't mean AMD couldn't.
Have you ever thought of building your own processor or maybe just defining your own machine instructions? With eMIPS now you can.
Microcode strikes me as the natural way to define arbitrary instructions.
There's SoC devices that have a soft processor on it that runs a linux kernel and can communicate with the FPGA.
My impression was that you tended to define pretty static blocks of functionality (you know, hash calculators for your bitcoin mining or whatever), and then just communicate to those from the OS using interrupts and a shared memory interface like any other peripheral.
Basically, I want it to implement what are essentially custom CPU instructions.
Rather than implementing your full algorithm on a big FPGA over a PCIe link, you should be able to implement just the few custom specialized instructions that regular CPU is missing, while still utilizing the regular CPU resources.
Hide the registers or set a flag if the FPGA is not configured for it.
Moreover, put another configuration flag to FPGA part so that it either shows up a device or works as a set of registers.
It's easier said than done but, it's not impossible.
The other thing to note is that this only covers using the FPGA for compute. While that's the obvious use-case for an "FPGA in CPU" concept, in the real world, IO is a pretty important part of many non-cryptocurrency FPGA applications, whether you're just counting edges on some input like an encoder, providing a highly time-sensitive output signal, such as for an SDR, or implementing some custom serial protocol.
The thing is, I don't think AMD is going to have much of an advantage. Back when AMD64 was becoming a thing, a long time ago, there was a lot of excitement around HyperTransport (HTX), around an inter-processor bus that people could make interesting accelerators out of & connect fpgas with. It never... really went anywhere. It's still supposedly quite prevalent as the basis for AMD's modern Infinity Fabric, but AMD seems to have no interest, no energy to make Infinity Fabric or HyperTransport interesting to the world, anything beyond an implementation detail.
If AMD does want to start integrating FPGAs, they either need to go back 17 years & start pushing Infinity Fabric / HyperTransport, or they need some other plan to make the fpga useful. Perhaps CXL or CCIX or GenZ or just like they do now gobs of pcie is enough, is how an fpga can mesh with core complexes & peripherals, but wow AMD would be in a much better place to start integrating FPGAs if HTX has kept up steam.
[1] https://www.anandtech.com/show/12773/intel-shows-xeon-scalab...
https://www.nextplatform.com/2018/05/24/a-peek-inside-that-i...
From what I remember, the problem is that there is a fairly large amount of work to get useful performance out of a hybrid like that. It's easier to just, say, learn CUDA and use a GPU than it is to have HDL folks work with SW. I have no doubt a system like that would be useful, but it would probably be expensive and limited to specialized, low-volume applications(like big FPGAs are now).
The modern ARM cores, starting with Cortex-A55 (2017), use the corrected ISA ARMv8.2-A.
On the ancient cores used by Xilinx, when writing multithreaded applications, it is annoying to lack atomic instructions.
Yes, sure atomic instructions can be simulated with only the base ARMv8.0-A ISA, but it is not possible to guarantee worst case timings for the simulated atomics (because of retries).
https://www.wsj.com/articles/amd-is-in-advanced-talks-to-buy...
The second thing was bundling FPGAs into the same package as a CPU. There were two problems with this, firstly it's a load of dark silicon when you aren't using the FPGA part and you have to make a load of tricky decisions like "Am I going to design my thermals to let me run all the Xeon cores at max speed whilst running my FPGA". The other problem being that it's really bloody hard figuring out the programming model, where the FPGA wants a deterministic data flow and you've got these CPU nutters throwing memory at you out of order, stalling and screwing you up with really bizarre cache behaviour. Since the CPU guys designed the interface it leaves the FPGA developer an almost impossible task to build efficient processing pipelines.
Finally we've got the moonshot - the idea that you could come up with a high level design language to programme FPGAs like software. Intel very heavily invested in this when they bought Altera but I'm still not seeing any forward progress on this, and let's be clear, Intel poured huge resources into making that happen. The last I heard they had over 100 engineers on that project and the sum total of their achievement is a handful of "partners" who wrote OpenCL/HLS/whatever and then worked with the engineers to go through a grueling process re-writing their code over and over and over until it looked like the RTL they already wanted. At one point Intel were going to re-write their entire FPGA video IP suite in HLS, I don't know how they went, but the acquisition of Omnitek probably wasn'ta good sign. It's been 4 years since the acquisition, and they were working on it long before then. The project is still basically just a load of marketing guff on their website that no one can actually use. I would say with the innevitable restructuring that Intel will have to do due to their various other fuck ups, this project is on thin ice.
The problem is that AMD is so likely to fall into the trap of points 2 and 3, and we're going to lose the final big independent FPGA company so that AMD can kill themselves trying to compete with Nvidia. And the real danger is that whilst all that's happening, actual innovation will disappear in the FPGA space. Xilinx's ACAPs are actually interesting (one way of solving the programming model), but Intel's last piece of innovation in the FPGA space died with the failure of hyperflex.
FPGAs are such a useful tool for a variety of engineering (and product) problems that fall in-between general computing and ASICs, unfortunately these problems are usually so disparate in nature that you can't target them as a single market. So the datacenter / HLS game will keep going until either a viable solution for the programming model is found or until everyone gives up on it once and for all and decides that FPGAs should go back to their niche (which is fine by me)
But yes, a software stack would help.
At least AMD is fabless so there should be no such issue with Xilinx.
Both Altera and Xilinx sell premium devices in terms of cost. If you make a high volume product based on an FPGA and care about profit margin, you are better off with Lattice (assuming their devices are performant enough for your application). It would be nice if any of these acquisitions made the FPGA prices more competitive, but I doubt it.
GHz limits force companies to move more and more into parallel computing and FPGA might seem the right move, but it is only partially a right move because for computations we don't need to simulate gates, we need arithmetic operations thus rather something like Field Programming ALU Array (FPALUA)
> You clearly saw a few logic gates
Well, perhaps 5 years of studying digital electronics engineering is worthless here... ;-) Saw a few logic gates when was 15, now I'm 40, and saw them already more than few ;-)
Don't get me wrong, I'm not underestimating issues related with entire way from design to manufactured product. But FPGA in not any "magic" different in terms of how you design it and how you manufacture it from other digital ICs.
In case of AMD, they have well established internal process for pushing from design to production of new ICs. Design of LUT or internal routing logic is really trivial, especially when given a perspective of designing something like RYZEN. I'm sure AMD could do it itself without significantly noticeable financial effort to the company.
Given above - software (and possibly software patents) is the only real advantage that XILINX might bring to AMD with that deal (unless AMD has plans to develop on new markets - which I don't really think so).
> where is the RISC-V of the FPGA world?
sounds like a challenge ;-)
I would love an RFSOC type peripheral though. Sort of a Spectrum Processing Unit or SPU. If you extrapolate the cu:Signal work to its logical extreme of a purpose built auxiliary processing unit the possibilities are pretty amazing.
Altera is doing about the same before Intel's acquisition. Or may be losing to Xilinx a little ( within margin of error )
But one thing for certain is that it hasn't grow. ( At least on paper )
So it depends how you define well. Considering the performance of Intel's past acquisition ( Look at Infineon ) and the total failure of Intel Custom Foundry I think it is doing quite well.
AMD seems to be doing well.
Intel is indeed faltering, yes. Especially with process technology. Given what they've got to work with in the process department, they've been building neat good stuff. Xe just launched & in a mobile form factor is quite competitive, no longer leaving AMD APUs to dominate. These same mobile parts have far higher integration than anything AMD can offer, with dual USB4/TB4 ports good for 2 x 40Gbps. Intel has a range of neat things few others have. Intel has a pretty competent wifi that integrates nicely. Their new Tremot Atom cores are a great value, really nice light-weight cores, & scale up to 24 cores on expensive but interesting cellular radio platforms. Xeon-D is a wonderful connectivity platform that there is not much competition for. Alas Intel cancelled OmniPath, which was a brilliant thing to integrate & give away on server CPUs almost for free (~$100 on a >> $1000 chip). And alas the process tech really hampers big server cores, but intel has been getting more competitive.
At this point, Linux need competition. Nvidia thus-far has never played well, Linux4Tegra and CUDA are walled gardens & driver support is otherwise awful, still no good way to run Wayland. AMD is doing great & Linux support is very high. But we kind of need Intel to keep AMD from growing lethargic & un-competitive, from exploiting their new market position as, well, better.
But, if AMD really wants to get into a new market, it could try going into mobile. The 4000-series CPUs are great laptops CPUs which could go further down the scale, and over a couple generations maybe in phones too. Unlikely or stupid idea (RIP Broxton)? Still worth a shot, I feel.
Not my area of expertise, but I get the same feeling regarding their GPUs. Everything is built using CUDA, so AMD is out for ML/AI, etc.
How many f8, f16, f32, f64 units you want now? Do you need some special instructions like add+multiply, or some bit permutation? You can add some.
There's no need to reconfigure it very fast, or per VM; not having to reboot the hypervisor OS would be enough. But it e.g. would allow to change the type of instances a particular server can offer, adjust to demand, and so overprovision less.