They had to redesign ReefShark and cancel dividends. It was a huge setback.
https://semiaccurate.com/2018/07/02/intel-custom-foundrys-10...
It was not Nokia SoC just plain Stratix10. They moved to own SoC after that glorious project.
Those toolchain disasters are not quite as hilarious when you have to use them daily....
I'm really surprised that Lattice hasn't tried to go around Xilinx and Altera by doing exactly that. You would think that an open bitstream format and a couple million dollars thrown at academic researchers (Lattice makes about $200 million per quarter in gross profit) would produce some real progress, but I digress ...
SystemVerilog, on the other hand, was specifically created because Verilog and SystemC got loose to the end users and the EDA companies were not going to make that mistake again. So, yeah, SystemVerilog is pretty bad.
These "tools" have no target so no incentive to improve. To use them you have to basically push their results back into a Cadence/Synopsys/Mentor toolchain anyway, so you might as well stick to the supported toolchain.
> The reality is SystemVerilog is a huge language, and already an open standard yet no open source project supports it fully
Most commercial systems don't support it fully. And its not clear that SystemVerilog is that superior to VHDL. And, for quite a while, SystemVerilog wasn't open and had some fairly obnoxious patents surrounding it. I don't know when/if that has changed as I have been out of semiconductors for about 20 years now.
Icarus Verilog has been slowly supporting features from SystemVerilog but doesn't have a lot of manpower.
In general, the consolidation of the semiconductor industry and EDA has hurt open-source EDA improvements. There's not very much money coming from companies to fund EDA research. EDA startups can't really get venture funding since VC's all want to fund the next pile of viral social trashware. And anyone with good software skills left the semiconductor industry eons ago because the pay differential is ridiculous.
> The reason is if it has not happened for the first (and arguably easiest) step in the chain i.e. System Verilog, they why would it happen for the others?
Ehhhhh, I don't think I buy this at all. There are dozens of alt-HDLs out there, many of which are quite powerful, designed by solo users. People had working, simple-but-practical PnR for real devices in a ~7k C++ LOC codebase written by an individual (arachne-pnr) and many individuals have independently reverse engineered small-ish scale device families for packing utilities. nextpnr was written by a very small group (solo?) in a year or something. I don't think you could fit an equivalent parser for SV2017 in ~7k LOC, much less elaboration, type checking, a netlist database, to all go along with it. SystemVerilog might actually be the most difficult part of the whole equation because it simply has so much surface area. PnR tools are limited by their target: only targeting small iCE40 devices? Your PNR algorithms don't need to be cutting edge. Targeting SV2017? Your job is hard no matter what device you synthesize for. And I can't think of even a single commercial tool I know from any vendor that supports all of it, up-to-date with SV2017.
All that said, I use SystemVerilog as my "normal" RTL when using commercial tools for stitching together IP, wiring up top modules, etc.
The free part is valuable not in that it's cheap, but in that it saves you from having to deal with licensing.
DevOps pioneers hailed from the likes of Google, Amazon and Facebook, who are not exactly short on cash, but you simply couldn't do what they did if you had been nickeled and dimed at every VM and container.
Bitstreams are closed. There's little to no point in doing an open source compiler if the target is not just proprietary, but deliberately opaque.
Overall your comment strikes me as what a proprietary compiler advocate would say in the 90s. "GCC? Lol"
Since then, Microsoft had to include Linux in Windows just because they absolutely needed Docker. DevOps was invented based on free/open source, it just couldn't be done proprietary style by a company as large as Microsoft.
The limitation here is writing the SystemVerilog parser and compiler.
As for the incentive I'm fairly pessimistic. There is definitely no money to be made for a start-up in this space, it is way too conservative. Maybe the hobbyist intellectual challenge of working on some hard problems like constraint solving or formal property proving? There is a massive task of writing a SystemVerilog parser before you get there though and the SAT solving and property proving problems are present elsewhere with lowers barriers to entry.
Having said that, there has been some promising F/OSS work on the small Lattice devices. It allows for a decent, modern workflow, and it's possible because the devices are approachable, but also because Lattice hasn't been hostile. Why they haven't been more supportive is a mystery to me however.
SystemVerilog is a good examle of an organically grown language with no 'benevolent dictator'. A few pet peeves:
* Why is the simulation delta cycle split into 17 regions? Exactly when does the Pre-Re-NBA region happen and what assignments take place there?
* Why can't a function return a dynamic/associative array or a queue? This is clearly possible, since the array find functions return a queue, but it's not possible to define a user function with this return type.
* It has way too much cruft. E.g. what problem does the forkjoin keyword solve? Who thought that was necessary and why? Not a fork-join block, the forkjoin keyword.
* Why can't you have a modport inside a modport? This would be great for e.g. register interfaces, but modports are not composable.
* What is the difference between a const variable and a localparam and why does the language need both constructs?
* Is a covergroup a class or what? It behaves very much like it is, it has a constructor, some class local information and at least one class local function (the sample() function), but you can't extend it.
* Why are begin-end used for scope delimitation everywhere except in constraints where curly brackets are used? I know it was a Cadence donation, but why wasn't the syntax changed before it was merged? Backwards compatibility can only justify so much...
//rant off
edit: formatting
As for SV - a lot of your gripes are Verilog issues, and SV has tried to fix some of them. I agree the blocking / nonblocking is a mess but most folks just learn the rules to avoid issues, but delta cycles can be a pain. The syntax limitations/quirks you point out are intersting, though not enough to say the language is terrible, it's extremely powerful with very good composability of types, constrained random is very powerful, the coverage is extensive, assertions again are very powerful. In a way its line a few seperate languages bolted together so sure there is some duplication, but it works surprisingly well in the whole.
You only start paying for FPGA tools when you need the really big FPGAs.
And, I'll go out on a limb, but, at this point, I think Arduino causes more harm to beginning embedded developers than good. Yeah, the ecosystem is wonderful if you aren't a developer.
However, Arduino is now weird compared to mainstream embedded development. Most things have converged to 32-bit instead of 8-bit. Arm Cortex-M is now mainstream so your architectural understanding is useless. 5V causes a lot of grief given that everybody else in the world is at 3V/3.3V.
A developer basically has to unlearn a bunch of things to move up from an Arduino. I still recommend Arduino to non-developers or somebody just trying to throw together a project, but I no longer recommend them to someone actually trying to learn embedded development.
But maybe I'm being overly optimistic. (Probably because—disclosure—I'm long AMD. Been long for years.)
If I remember correctly, there was something similar back in the early HyperTransport days...
I'm just looking at this from a logical chain of "who needs FPGAs in their computers?" => "cases with loooots of specific data crunching" => "want a controlling/driving CPU for the complicated parts, but then just concentrate as much FPGA in as possible." => Multi-socket with 1 CPU & rest FPGAs.
(There currently is no commodity Quad-socket SP3 mainboard, not sure if this is a design limitation or just no one made one yet? I'd still say the approach works great with only 2 sockets.)
1) You underestimate how critical prototyping has become, again likely since you say it's been a couple decades. Time to market has become more important, and verification has become harder as CPUs have gotten even more complex. FPGAs enable cosimulation and emulation, leading to faster iteration of both design and verification efforts and thus better TTM.
FPGAs are so important in the hardware development process that I would even say you're not a serious hardware company if you don't have any FPGA frameworks to design silicon.
2) As others have mentioned, FPGAs are also critical for low-latency workloads that require constant tweaks-- high frequency trading (ugh...) comes to mind. The need for "constant tweaks" could also be satisfied with just "normal" software, but that has higher latency as opposed to an FPGA, and FPGAs can get some crazy performance if you're willing to pay the price (south of 7 figures).
Overall sure, usage of FPGAs might be niche compared to, idk, Javascript; but it's commonplace/practically essential in hardware.
FPGAs are already incredibly popular. They're just mostly in things you are unlikely to personally own or know about. You're going to find at minimum one, but probably more FPGAs in things like big routers and other telecom equipment, e.g. cell towers, firewalls, load balancers, enterprise wifi controllers, video conferencing hardware, test equipment like oscilloscopes, sensor buoys, scientific instruments, MRI machines, LIDARs, high end radio equipment, or even just glue logic tying together other components, like in the iphone.
FPGA are being used in many type of applications where real-time is necessary and non-recurrent engineering (NRE) cost need to be minimized, for example here [1].
One classic example is that if you poke under the hood of any signal generator like AWGs, you will probably find an FPGA inside. As you probably aware since you in hardware business, AWGs are probably one of most common equipment in any electronic and electrical labs or companies.
[1]https://www.electronicdesign.com/technologies/fpgas/article/...
You are precisely correct. FPGAs are useful when your volume doesn't reach volumes where an ASIC would get amortized.
Networking companies (Cisco, Juniper, etc.) are classically big consumers of FPGAs.
Tektronix seems to make quite a bit of money and there is at least one FPGA in practically every test instrument they make. This holds true for practically all test instrument manufacturers.
I know a LOT of industrial automation and testing companies that generally have FPGAs in their systems. Both for latency and for legacy support (Yeah, GPIB still exists ...).
Yes, they aren't "Arm in a cell phone" type volumes, but that doesn't mean they aren't quite profitable if you can aggregate them.
Configurable as in one SKU is in several products, but not necessarily reconfigurable by the end user.
The big benefit of having FPGA closely attached to CPU is that you can access the memory and internal buses quickly. Transferring stuff over PCIe hurts a lot. So you could make an argument for jobs using small work units requiring fast turnaround; CUDA kernels take milliseconds to launch.
I worked with some of the early Xeon+FPGA parts and there just wasn't that much we could do with them. There wasn't enough fabric to build anything meaningful and we had an abundance of CPU cores, so the best we could do was specialized I/O accelerators.
The problem is that if a task is common then someone is just going to make an ASIC to do it. And if its uncommon then the terrible FPGA software ecosystem and low prevalence of general purpose FPGAs in the wild mean that people will just do it on a CPU or GPU.
This is true, but keep in mind that that sort of algorithm runs insanely well on any CPU or GPU because they, too, do not want to touch main memory. You would be blown away by how much work a CPU can do if you can keep the working set within L1 cache.
Re. ASICs, it's a continuum:
- "flexible, low performance, cheap in small quantities" (CPUs)
- "reasonably flexible, better performance, cheap-ish in small quantities" (GPUs)
- "inflexible, best performance, expensive in small quantities" (ASICs)
FPGAs fit somewhere between GPUs and ASICs -- poor flexibility, maybe great performance, moderate small-quantity price.
If your problem is too big for GPUs, as you say, sometimes it's easiest to jump straight to an ASIC. But it's such a narrow window in the HPC landscape. The vast majority of customers, even with large problems, are just buying a lot of GPUs. They're using off-the-shelf frameworks even though a custom CUDA kernel would give them 10x performance and 10% cost. The cost to go to an FPGA is too great and the performance gain simply isn't there.
Basically you analyze the code for candidates, select a candidate, upload your custom hardware design, run your operation on the hardware, and repeat.
The difficult part is that uploading your hardware to FPGA is in the order of tenths of seconds, which is ages when compared to the nano and micro seconds your CPU works. So your specific operation must be worthwhile to upload.
A bit of FPGA on your CPU makes it more flexible, for example your could set a profile such as 'crypto' or 'video' to add some specific hardware acceleration to you general purpose CPU.
Imagine your CPU being able to switch your embedded GPU into another CPU core.
Let's say the current zen 2 had an FPGA onboard. AMD could sell you an upgraded design with AV1 support for a few dollars. Most people aren't going to buy a new CPU on the basis of a video decoder, but they'll buy an upgrade to the chip that auto "installs" itself. That's a sale AMD otherwise wouldn't have made.
In theory, you could even page out code, but I guess the speed of that will be slow. Also, paging in probably would be challenging because the logical units aren’t uniform (if only because not all of them will be connected to external wires)
An enormous crossbar could solve that, but I would think that would be way too costly, if practically possible at all.
Repurposing FPGAs to different tasks means loading a new bitstream into the device every time. So it is much more efficient to grant exclusive access to each user of the device for long stretches od time. The proper pattern for that is more like a job queue.
Where FPGAs win are new architectures, like Systolic engines. Entirely different computer designs from the ground up.
I've always been pretty skeptical of their approach though, in order to be usable they'd need excellent tooling to support the feature, and if there's one thing that existing FPGA software isn't it's "excellent".
Getting FPGAs to perform well is often an art more than a science ("hey guys, let's try a different seed to see if we get better timings") so the idea that non-hardware people would start to routinely generate FPGA bitstreams for their projects is so implausible that it's almost comical to me.
Maybe one day we'll have a GCC/LLVM for FPGAs and it'll be a different story.
This can be a relatively fast operation. Seconds or less depending on complexity.
It'd be interesting to see how AMD will execute and integrate this acquisition, considering they are less of a madhouse company than Intel.
https://www.nextplatform.com/2020/01/31/when-will-fpgas-outw...
https://www.nextplatform.com/2018/03/19/fpga-maker-xilinx-sa...