AMD Introduces World’s Largest FPGA-Based Adaptive SoC for Emulation
ir.amd.com
ir.amd.com
77 mm x 77 mm package (a bit larger than my old Pentium Pro).. I'm amazed they have any yield for a chip this large at 7 nm.
18 M LEs. 6864 DSPs. 6865 balls.
How long does place and route take?
$57K for the smaller VP1802..
https://www.digikey.com/en/products/detail/xilinx-inc/XCVP18...
I'd honestly think that if you just want to throw money at it, those absurdly-overclocked gamer rigs that can hit sustained +5GHz boosting on just a few cores would be better if you want the fastest P&R times? Cloud servers will probably have more RAM and memory bandwidth but you can shove 128GB of RAM in the 13900K, memory requirements are relatively easy to meet IMO. Synthesis can definitely scale with some more cores in Vivado (OOC synth) but P&R never seemed to scale beyond what you can just get in a desktop. YMMV I guess.
Is it possible to upload just a small bugfix to bitstream?
If you want to upload a small bugfix without it taking 10+ hours, you need to floorplan your design and use "partial reconfiguration" techniques in order to change only part of the bitstream at a time. This requires a more advanced design flow and a bit more work, but is necessary for things like multi-tenancy or lowering turnaround times.
If you want to avoid 12 hours builds you have to use floorplanning, incremental synthesis/implementation, and dynamic function eXchange (swappable components). I've been able to incrementally change designs to the scale of ~50K LUTs (look up tables, 50K LUTs is a big change) in about 15 minutes of synthesis and P&R with the rest of the design locked.
(The chip here is about 400x the size of a single 50K LUT change so I imagine you could have a 100 hour builds if not using floor-planning and partitioning)
"Emulation" refers to emulation of chip designs. When a chip has billions of transistors that are all connected to each other to perform higher level functions, you need to test out designs in order to make sure those higher level functions actually do what you expect them to. Most modern chip designs (at least the digital portions of them) are designed using what's called a Hardware Description Language or "HDL" such as System Verilog or VHDL. It's somewhat similar to traditional programming, except threads are nonexistent, so it's essentially an unholy bastardization of a markup language and a programming language. But I digress... These functions are then synthesized into their low level building blocks such as buses, gates, and registers, and are placed and laid out on a chip.
Going back to the part about billions of transistors; as you might imagine, testing of such designs must be rigorous and thorough. Only problem is that it's really, really computationally intensive. There are 2 ways of doing it: Simulation (CPU Based) and Emulation (FPGA Based) Simulation is slower, but yields more data like waveforms and whatnot (though you can get waveforms with emulation, but there are more limits.) Also, simulation takes less time to compile, since there's no place and route step.
I should also point out that there are 4 key simulation levels: Simulation Model (usually compiled C code) RTL Simulation (Functional level simulation) Gate level Simulation (simulating the individual gates) and physical simulation (where the transistors and trace routing are all accounted for.) Each level being roughly an order of magnitude more computationally intensive than the previous. Most time and engineering effort is spent on RTL level simulation, which is what can, and is often emulated on an FPGA, and why this product is so relevant to hardware designers.
If you're a software engineer interested in the hardware side of things, there's a great article for explaining the difference between the programming methodologies: https://nandland.com/lesson-7-what-every-software-programmer...
To add some extra context, usually when a hardware designer is testing something on an FPGA they'll be testing just one specific IP core (a reusable unit of logic, cell, or integrated circuit layout design that is the intellectual property of the company). Let say, an Ethernet MAC for example. You don't necessarily need a chip like the one in the article to test that. You basically validate what all the inputs and outputs of said IP are going to be; you know that some signals are going to the PCS/PHY and some to a processing core/memory, and provided everything is done correctly then you should be able to slot the IP into your final chip design and everything works. That's not always the case though, which is why some companies end up have many stepping's (revisions of the semiconductor photomasks used to pattern an integrated cirucit). Many companies, my own included, make a business out of designing and selling these IPs to other companies without ever really knowing what the end customer wants to do with it. We just tell them what signals to put in, and what to expect out.
With some of the larger FPGAs, you can connect many multiple IPs together and emulate the whole IC, but it depends on scale. You can easily emulate a whole Intel 4004 chip, but you're not fitting a whole Apple M1 design on a single device. With the chip in the article though, we're a step closer to that. Generally you might have 16+ FPGAs on a single PCB to test that scale of functionality. Even that's not a "true" test of a whole IC though, because of the delay in chip-to-chip communication. You might send your design to the foundry and find you've failed on timing. An Agilex 7 M-Series, one of the larger FPGAs in the Intel portfolio, has 3.9M LEs. This chip has what looks to be 18.5M. It's not a 1:1 comparison because they have different architectures, but you get a rough idea of the jump in scale here. This Versal Premium chip will be a massive boon to the likes of Apple, Tesla, NVIDIA, etcetera, because they can potentially emulate the entire chip function on a single die, meaning less steppings to get production quality silicon.
Not really. In each generation of devices, Xilinx had one focused on emulation for while now: huge number of LUTs, huge number of general purpose IOs, relatively few DSP elements, a moderate amount of block RAM. Before this monster, it was the Virtex Ultrascale+ VU19P, and before that, the Virtex Ultrascale XCVU440, and before that, the Virtex 7 XC7V2000T.
This isn't "FPGA in your CPU" sort of thing but rather just a generational improvement on existing products.
https://www.xilinx.com/products/silicon-devices/acap/versal-...
Seems there's still a relative dearth of 3rd party articles, but probably will see at least some more fairly quickly (w/ lifting of embargo etc.).
It is essentially an FPGA packed together with some ARM cores and interconnects into a single chip.
(after looking at corescore probably not, as it'd be way too expensive. that being said, the DE-10 Nano is getting old, are there any potential successors coming)?
But, in better news, it looks like Intel is finally getting off their asses and actually expanding their newest Agilex series to meet their customers across the spectrum, and finally offer new alternatives to their Cyclone series, which is used in the Nano.
Sometime later this year they are planning to announce the "Agilex 3" line, and that -- or the already announced "Agilex 5" -- will likely be what you will see in the successor to the DE10-Nano. It will still take work to target this chip and get it into a board people can readily buy.
https://www.digikey.com/en/products/detail/amd/SK-KV260-G/13...
The price is very good for Ultrascale+.
I think something could made cheaper using an Efinix Titanium chip, someone should do it. Their own dev. board is too expensive:
https://www.efinixinc.com/products-devkits-titaniumti180m484...
I am keen on the MiSTex project though. The plan is to come up with something that is more flexible than MiSTer to support different targets. So for example, you might be able to target a simpler core to a simpler device and vice versa.
… OP didn’t mention they were optimizing for price.
Some might say it makes sense to optimize for availability, recent updates to hardware software stack, etc.
"Palladium Z2 emulation based on a new custom emulation processor offers fastest, most predictable compiles and most comprehensive pre-silicon hardware debug capabilities
Protium X2 prototyping based on latest Xilinx UltraScale+ VU19P FPGAs offers highest performance and fastest bring-up times for pre-silicon software validation of billion-gate designs"
https://www.cadence.com/en_US/home/company/newsroom/press-re...
A few predictions/concerns:
* I can't find a price, so I'm predicting that it will be over $1,000 which will prevent it from going mainstream (and also be ~10 times more expensive than it should be).
* There will be poison pill(s). Maybe flash memory that can only be written 1000 times, preventing its use for evolutionary hardware and genetic algorithms. Maybe the place-and-route software won't be good enough to prevent short circuits, so certain configurations will burn up. Maybe some aspect of the software will be proprietary and/or encumbered by patents, preventing hobbyists from thinking outside the box and "getting real work done".
* Any dedicated hardware like memories, ALUs, etc may be misaligned for various use-cases. I just want an array of RISC-V, Arm, DEC Alpha, PowerPC 601/603, something like that, starting with 2 and topping 100 or 1000 cores eventually. So where I'll need memories near CPUs, something in the FPGA will lack the interconnect to allow that.
I hope I'm wrong about any or all of these. Price I can live with, as long as economies of scale or competition kick in and eventually deliver something under $1,000. The rest of it.. eh, I'm not holding my breath. I've been underwhelmed by all previous FPGAs, but maybe they didn't count. Maybe this is something new that finally manifests the original vision of what FPGAs could be.
Having thousands of dumb cores is pointless unless you want to work with sparse data. An out of order core isn't bottlenecked by compute, it is mostly bottlenecked by memory access latency and also bandwidth if you do end up using vector instructions. This means most of your core will be memory and your thousand core chip will turn back into a dozen core chip. If you need a dumb accelerator, then GPUs already exist and you don't need a custom chip.
So the only remaining usecase is sparse data. The expectation is that you are going to get cache misses all the time anyway, so the benefit of a large cache is negated by the fact that the same data is rarely accessed again. The problem with this idea is that sparse workloads are pretty rare. The only usecase that could possibly benefit from a custom chip is sublinear machine learning (SLIDE) which basically does nothing but predict which neurons are activating and ignore everything else.
Oh also I am already assuming you want to tape out your chip and that the FPGA is just a stepping stone. If all you do is insist on running softcore processors with no special architecture (e.g. the Reduceron) on an FPGA then the whole exercise is meaningless.
AMD CPUs see major performance increases every generation. The 7700XT is more than 100% faster than the 2700X and that is at the same core count. Higher core counts beyond 8 cores have become much more accessible.
You can check my math, but single-threaded performance has only increased by a factor of about 3 since 2000. So a modern 8 core machine would be about 24 times faster than say a MIPS or PowerPC 601 at the same clock speed, when CPUs had 4 pipeline stages and didn't suffer the kinds of cache miss penalties we see today, so didn't need excessively complex branch prediction. A 16 core machine would be 48 times faster. But if you look at transistor count, CPUs in 2000 had 1-10 million transistors, while today they have 1-10 billion and GPUs have 10-100 billion. So CPUs should have 1,000-10,000 times the performance, not 24 or 48. There's simply no way for traditional CPUs to scale to the level I'm talking about, which I realized in the late 90s while getting my computer engineering degree from UIUC.
https://en.wikipedia.org/wiki/Transistor_count
Having thousands of dumb cores is pointless unless you want to work with sparse data. An out of order core isn't bottlenecked by compute, it is mostly bottlenecked by memory access latency and also bandwidth if you do end up using vector instructions.
I agree, so the chip I'm envisioning would have an array of local memories (one in or beside/above each CPU) which negates this argument. Then the problem becomes one of orchestrating those memories to appear as one coherent address space. I want to use a content-addressable memory with copy-on-write, so that the memory works like a hash tree (BitTorrent). A cyclic redundancy check (CRC) or similar could be used for the hashing, with a fallback to lower clash hash like SHA when a block clashes. This would all be encapsulated in an abstraction below the level of the code, following the same principles as cache coherence. We'd mainly use auto-parallelizing higher-order methods along the lines of map-reduce and scatter-gather arrays within a runtime similar to Go/Erlang, Octave/MATLAB or Haskell/Julia (or vanilla C or Lisp even) to make it appear that we are programming a single synchronous-blocking thread of execution. I've had this approach fully-formed in my head since about 2010, with the original idea coming from when I was programming games back in the 90s and Apple crippled its memory busses for cost reasons and to prevent competition between its entry-level machines and its flagship lines, so I saw that the memory bus is the only real bottleneck in computers today.
You do touch on one major limitation though: there would be a 10-100x overhead to build my design on FPGA with multiplexers and LUTs, then another 10-100x for the hash memory, and at least 50% of the die lost to local memories. So I might need 4 orders of magnitude more transistors to achieve existing performance. Then the hashing could add 10-100x the latency, putting us at 6 orders of magnitude worse performance in the worst case. I'd plan to get around that by de-prioritizing latency and optimizing the best case (sort of the non-branch prediction miss or cache miss case) and then throw hardware at it. It's much simpler to just increase the die size by 10-100x on a side than develop the next architecture or VLSI design rules. So at scale, this new chip simply wouldn't face the linear speed increase limitations that all processors today face. It might be as simple as just plugging in another chip on a PCI bus to get another 1000+ cores under the same hash memory, just like how we program the web with hash id caching at the edges with stuff like CloudFlare.
So the only remaining usecase is sparse data. The expectation is that you are going to get cache misses all the time anyway, so the benefit of a large cache is negated by the fact that the same data is rarely accessed again.
Exactly, which is why each core would have little or no cache. If some number of cores want to crunch away on sparse data, while others run the OS, there is no conflict. Note that this strategy runs counter to just about all processors since the DEC Alpha, which I recall had 3/4 of its die area devoted to cache, which it succumbed to because DEC was never able to figure out multicore, and their demise coincided with the loss of R&D funding after the Dot Bomb around 2001. Note that we also lost cluster computing (remember the Beowulf cluster jokes) because GPUs were "good enough" for the primarily graphics-driven workloads of desktop publishing and gaming. Never mind that GPUs fail for most other real-world business logic use-cases.
Oh also I am already assuming you want to tape out your chip and that the FPGA is just a stepping stone. If all you do is insist on running softcore processors with no special architecture (e.g. the Reduceron) on an FPGA then the whole exercise is meaningless.
Thank you, I hadn't heard of the Reduceron!
https://www.cs.york.ac.uk/fp/reduceron/
https://github.com/tommythorn/Reduceron
Without getting too far out in the weeds, I appreciate what they're trying to do, but wide busses won't work for that. This strikes me as perhaps an extension of SIMD and VLIW, which are what I'm trying to get away from so that we can get back to ordinary desktop computing.
--
Which brings me to my main point. You're concerned with your use-cases around machine learning, it sounds like. And that's great, I'm not knocking that, and GPUs or AI cores for TensorFlow are fine for that.
But what I need is horsepower for unpredictable workloads. I want to explore genetic algorithms, and simulated annealing and k-means clustering and a host of algorithms beyond that which don't work well within a shader paradigm. I need to interact with system calls, and disks, and the network. I need to run multiple types of workloads simultaneously. And I can't be dropping down to intrinsics from a high-level language like Python to do that. Although the rest of the world seems to enjoy that mental tax for reasons which I may never understand.
Other than internet-distributed stuff like SETI@home, Folding@home and BOINC, or EC2 on AWS, there is simply no machine available today that can do what I need. So you're talking past me with efficiency and optimization concerns, while I'm unhoused with no meal ticket.
Now, a handful of people over the last 2 decades have grokked what I'm getting at, and are trying to achieve desktop computing on GPU:
https://en.wikipedia.org/wiki/Transputer
https://en.wikipedia.org/wiki/Multiple_instruction,_multiple...
https://www.microsoft.com/en-us/research/video/mimd-on-gpu/
Due to an almost complete lack of industry support, these projects are destined to fail.
Which is why I've all but given up on the status quo ever changing. It's like I can see an entire alternate reality where kids could build stuff like C3P0 like Anakin did, with off-the-shelf multiprocessing hardware. But because we're focussed on SIMD and going to great lengths to achieve even the slightest parallelism, we can't even build neurons in hardware. Which is really where I'm going with this. A way to emulate 100 billion neurons, where each one has the computing power of perhaps a MOS 6502 or Zilog Z80. Then let a genetic algorithm evolve that core into something that can actually think emergently with its neighbors.
Vs the million monkeys approach we have today, where teams of scientists work frantically for decades spending billions of dollars to build stuff like large language models (LLMs). For all of the excitement around that, I just find myself tired and dejected reminiscing about an alternate history that never came to be.
Anyway, sorry for the overshare, I'll just go back to my day job building CRUD apps on the web, running as fast as I can in place to make rent like Sam on Quantum Leap, never able to exit the Matrix. Maybe Gen Z will pull off what Gen X was blocked from doing at every turn. But I digress.
Try 50x to 100x that at quantity=1. This is not a hobbyist device: this FPGA is the largest you can get on, and it's tailored to the needs of companies that spend 7-9 figures on the verification and validation of their ASIC projects.
> Maybe this is something new that finally manifests the original vision of what FPGAs could be.
Even small to mid-sized modern FPGAs have far surpassed the "original vision of what FPGAs could be".
I should probably start with an open source setup, but last I heard, pretty much all FPGAs had a proprietary layer like GPU drivers on Linux, which was such a huge turnoff that I never bothered proceeding further. If that's changed, I'm all ears!
It sounds like you're frustrated with hardware manufacturers for a slight against supporting your use case. But I think that on closer inspection, you're really frustrated with economics, physics, and the nature of the universe.
It's part of the nature of semiconductor manufacturing that manufacturing a custom ASIC (or massively-general-purpose IC?) especially a large and complex one, is going to be very expensive. There are huge non-recurring costs that can get amortized over each unit, plus economies of scale that kick in, when your masks can be used for production runs of thousands or millions of chips. It's not 10x more expensive than it should be so they can make 90% profit when you compare it per unit of silicon area against the millions of Zen3 chiplets that AMD has sold, it's 10x more expensive because it has fixed costs and they're only going to sell a few thousand of them.
As a result of that, you want to spend as little as possible to make a product that's usable by as many people as possible. If a thousand companies can use this FPGA with flash/EEPROM/fuses that have relatively low write endurance, but it's more expensive and difficult (or eliminates thousands of use cases!) to use SRAM with huge write endurance for settings that will typically be written once and, on dev units, a few dozen to a few hundred times, that's not an intentional poison pill sabotaging your career but just reasonable economics and good sense. Plus, if you're trying to recover NRE, an obvious avenue for recovery of some of those funds that you spend writing the software is to make the software proprietary so you can sell Vivado licenses to each of those customers.
On a 2-dimensional planar die, you can add a few layers, but you can't change the mathematics of graph theory to get 3D or 4D or infinite connectivity. Interconnect is just plain expensive. It's always going to be an engineering tradeoff between power dissipation and speed, so they make a best guess as to where the most applications will need interconnect, and build that. The fundamental software gates that are used to build them are always going to be slower than dedicated hardware gates, because it takes additional transistors to run a signal through them. That's why memories, ALUs, CPU cores, and peripherals on an FPGA SOC are hard-coded into the chip. They're hardcoded in ways that allow you to flexibly make use of them with custom logic, but they're burned into the die because that's better performance for their users.
You can't physically call into being a few cubic nanometers of doped silicon to form a new gate when you downlonew bitstream to an FPGA - and putting P or N doped silicon in just the right spot is critical to designing a chip, so how in the world do FPGAs actually work? They cheat, instead of building actual logic gates they just move charges to build the same truth-table out of generic SRAM-based look-up table, which means working at a layer of abstraction several steps ad a up the stack. The idealized, oversimplified, conceptual model of an FPGA as identical to a mask-programmable integrated circuit, except field-programmable, doesn't exist.
I tried to address some of your concersn about scaling in my response to imtringued.