Intel to set its FPGA unit free to pursue its own path
nextplatform.com
nextplatform.com
They focused solely on the high end, but it turns out nobody really wants FPGA fabric on a CPU. You can already do acceleration over a PCI express link, and that's what you more often do with embedded applications where the CPU is acting more like a dispatch controller than doing the real work.
Intel also have completely ignored the low end of the market. The only true lowend part they have is the Cyclone 10LP, which is literally the exact same part as the cyclone 3/4 from 2008. Just slightly die shrunk. No hard IP support like ddr3 controllers, no MIPI, nothing that people are getting from the competition now.
Intel did realize this, which is why the new AgileX family includes some "low-mid range" parts, but they will be still much more expensive. Low-end to Intel means "under $1k unit cost" which ignores a huge part of the market.
They have better tools, documentation, and support than Gowin, who is a recent Chinese FPGA upstart using stolen Lattice IP and hires. But they will lose to Gowin by default in the commodity space unless they do something.
The entire story of Altera inside Intel can be summarized as:
Intel fabs make amazing promises about process performance and availability. Altera builds their product stack on that. In the end, the fabs fail to deliver either performance, or sufficient amount of manufacturing capability. Now Altera has to pick which products they want to ship. They obviously can the low end. Even the high end that ships is horribly late, because of manufacturing issues.
There would have been massive demand for the combined Intel+Altera products. Many large customers built their future based on the marketing promises Intel made, and when they couldn't deliver, those customers had to redevelop everything on something else. As an example, look up Nokia Reefshark.
Not true... Xilinx sells a lot of their Zynq, MPSoC, and RFSoC chips. Though I guess those are more CPU on an FPGA.
If they'd been made a performant (as in, non-x86) low-end CPU, then they wouldn't have lost mobile and be competing with ARM now.
Intel of the past few decades learned the wrong lesson from their history. You can climb the value chain to riches, but once you're at the top you still need to seed the earth for the next generation of profits. And that means cheap entry points and volume.
https://www.theregister.com/2006/06/27/intel_sells_xscale/
https://www.networkworld.com/article/3574013/marvell-exits-t...
Edit: By the way there is some niche research of the FPGA in core (which don't work if the FPGA is over PCIe). This paper describes a way to rearrange memory using a programmable FPGA-based memory controller to optimize database table accesses: https://arxiv.org/pdf/2109.14349.pdf
I want multiple FPGAs in the CPU. Replace the PMU with an FPGA, replace the MMU with an FPGA, and by replace I mean augment and extend. Of course an FPGA can't do all the work of fixed blocks.
The ultimate problem is politics inside organizations and the lack of any sort of concrete creative thinking.
That Intel isn't incorporating Altera into IFS and allowing IFS customers to include FPGA fabric into their designs is predictable but both astonishing.
x86 and Altera as library IP only available on IFS would give them another 10 years of runway.
Intel had the same market segmentation and lack of imagination in how they applied Optane. Heck, Optane should have had FPGA fabric in each dimm. It appears that each sub-sku has their own fiefdom and their own P&L that ignores ecosystem effects.
Intel needs to be cleaned out.
Naive of me to ask, but how are they allowed to sell anywhere outside of China?
If this is true, then the bitstream for Gowin and Lattice parts would be similar and both are being actively reverse engineered to be supported by OSS tools.
It would be nice to see more evidence of this.
Not only are the patent expired, but there are three projects that generate FPGA fabrics programmatically. FPGAs are extremely easy to construct as they are a repeating lattice.
I do not understand this part though:
> There was talk of hybrid CPU-FPGA packages, which never seem to get > commercialized because no system architect likes static ratios of compute – > unless they are determining the ratios. Like the hyperscalers and cloud > builders, who can tell companies like Intel and AMD what their product > roadmaps need to look like.
What do not see what they mean by ratio here. Do they mean die ratio between cpu and fpga?
It's one of those things that seem like a good idea, but they just don't work out in practice. FPGA LUTs are just way too slow. You'd have to find a case where doing something on a 3GHz CPU clock running multiple instruction parallel gets outperformed by LUTs that runs at 700MHz (at best). And when you cascade the LUTs, they become slower too.
And that's without solving the problem of closely coupling a CPU pipeline with FPGA logic.
> What do not see what they mean by ratio here. Do they mean die ratio between cpu and fpga?
What they mean is: in something like the Zynq FPGA family, I want a die with 2 CPU cores and 5000K LUTs. The other guy wants 8 CPU cores and 2000K LUTs. It works for narrow applications like signal processing where power efficiency and cost isn't a top concern, but for a hyperscaler, power consumption is a very important metric. As is the cost of paying for a significant part of the silicon die that's sitting there unused.
Imagine an LLM with a new token every 10 nS
GPGPU sucks a lot of air out of the room as well. There aren't many purely computational problems which FPGAs can solve better than a compute-optimized GPU; even though GPUs aren't quite as flexible, they clock a lot faster, they're cheaper, and they're easier to develop for.
So even though I think FPGA could win out from an efficiency perspective for some problems there is just not enough people/companies working on it to win out. It also doesn't help the tools are raging proprietary garbage fires.
Essentially the only time it's gone the other way was when there was a toehold in a critical market, where the benefits were so obvious and profitable that they made up for the difficulty (e.g. graphics).
If, and of course that is a big if, you can repackage a (parallelizable) calculation into FPGA look-up tables and implement multiples of this (e.g. 8 to 80 times) then you can think maybe it's quicker than CPU at 3GHz.
However, you have to include DMA of the data to and fro. It's unlikely to be worth the very extensive effort of integrating two wildly different technologies.
On the other hand, it may not be a complicated calculation but FPGA can do much lower latency and smaller variance in latency (hello high-frequency traders). That is a very narrow niche.
A simple board with CPU and FPGA is the Arduino MKR Vidor 4000: ARM Cortex 32-bit CPU and Intel Cyclone 10 FPGA). Hardware cost: $85. Full suite of development software $1000 or more (although lesser tools are available for free.)
That is exactly the part where having the FPGA next to the CPU helps... You can transparently access the CPU cache via an AXI slave port on the CPU on AMD's MPSoCs at a rate of up to 16 bytes per cycle and you get multiple of those.
Easy: go wide.
Make the FPGA-CPU interface four times wider on the FPGA side than the CPU side. Each tick of the CPU clock reads (or writes) one quarter of the bits.
power pc cores, riscv cores, and by large arm cores
Good enough for a low volume custom solution for which custom silicon is too expensive. Not for a hyperscaler.
In fairness, I never mocked up a true enough implementation in Verilog to get an idea of real world speedup, and even now, I'm not sure exactly what operations you could see real gain with from small reconfigurable fabrics near the CPU. Still, I liked the elegance of having L1-L3+ FPGA's for speeding up operations of increasing levels of complexity, and I figured programmers smarter than me would find creative ways of using the FPGA's with the added instructions.
But most software engineers don't understand the amount of time and compute that would be required to mock up a novel, yet sufficiently complex as to be realistic, CPU. I did indeed have a basic HDL implementation, as is required for most patents in the US (reduction to practice). But to implement it fully enough to understand what kind of performance changes you could get in a modern CPU... it's safe to estimate it as an order of magnitude harder than the most complex pure software project you've ever built alone, and well beyond the resources of an intern who's just doing this in his spare time. Software is a great gig, because it's easy and you can do it all on your laptop.
And I'm sure computer engineers like myself don't appreciate the difficulty of manufacturing true mechanical hardware at companies like Tesla, where a team and a budget would be even more essential to build anything useful.
Anyway, the technical committee thought it was worth patenting, and the idea itself is pretty digestible without it working on production silicon, but I have no idea who comes up with product roadmaps at Intel. It was probably buried in the patent graveyard, or maybe someone qualified actually looked at it and realized it wouldn't work.
If the original author of a patent didn't find it worthwhile to pursue the idea for whatever reason I have very little faith in the idea to begin with. Taken to the extreme it's like saying I can turn lead into gold, and then never actually showing you actually making money that way.
If anything, I'm less convinced than before that you understand how complex a modern CPU is. Perhaps you can re-read my comment; I essentially did what you are suggesting, but that wasn't enough, in my opinion, to glean any information about benchmarks in the real world. It took hundreds of hours. To make it slightly more realistic but still woefully inaccurate using licensable IP's and my own time/money without the company's buy-in seems absurd, even to the most stubborn HN commenter. The fact that we are modifying the L1/L2/L3 structure (perhaps you also didn't read the patent before commenting) makes this essentially re-designing large parts of the CPU from the ground up, at least as far as I understand licensable core IP.
> I would expect you continue working on the apparently golden idea your sitting
Why? No one ever said it was the golden idea. Only that I found it interesting. Apologies if my optimistic curiosity triggered your defenses.
Regardless, your statement makes no sense for other reasons. Perhaps there's a cultural gap (are you not American?). Here are some relevant points about US companies and patents I hope I can help you understand:
1. Big US corporations will patent anything novel, technical, and remotely related to the business as a means of protecting themselves against litigation. Most of these patents only see an initial technical committee, lawyers, and then are never seen again.
2. Inventors have zero individual rights to their patents when patented via company channels (in exchange, the company pays for the lawyers and usually grants a tiny stipend).
3. Interns typically do not influence how research budgets are allocated.
Depends on what AMD does with Xilinx.
Currently the AMD/Xilinx dynamic seems to reverse this: "Depends on what Xilinx does with AMD".
AMD's software roadmap for AI/datacentre leans heavily on Vitis (for software) and AI Engines (as an execution platform). CPUs that integrate AI engines are already shipping (Ryzen AI). It's Xilinx technology, but you should expect it to look more like a GPU accelerator than a traditional LUTs-and-routing FPGA. And, as duskwuff have pointed out, this sucks a lot of the oxygen out of the CPU-with-FPGA design space.
This is incorrect along all 3 dimensions:
1. AMD has its own data-center class GPUs - I don't know how good they are because I don't work on them
2. Vitis is just a brand and will be taken out of the equation before the end of the year.
3. I don't know what execution platform means because AI Engine is one core in a grid of such cores on the chiplets that are on the Phoenix platform (shipped with new Ryzens) and the VCK boards.
> It's Xilinx technology, but you should expect it to look more like a GPU accelerator than a traditional LUTs-and-routing FPGA.
It is correct that there are no LUTs in the fabric but there are "switchboxes" for data traffic (between cores) and you do have do the routing yourself (or rely on the compiler).
That's even more eyebrow raising than an Altera spinoff.
Altera is a good side business, but Falcon Shores is like Intel's consolidated future. If they just let that go... What do they expect? That everyone will just buy Xeon CPUs and IGP laptops forever?
This is not to say Intel will go bankrupt look at the number of quarters AMD spent in red but it really doesn't want to become #2.
A near future with more ARM/RISC-V is very different.
Big chunk of the team from Altera are at AMD now anyway.
Hopefully they finally get back to innovating on the actual FPGA now. I’m so tired of the hardened rubbish and cpu integrated rubbish.
Was there actually a way to access a CPU-integrated FPGA as an "ordinary" user/customer (i.e. not a "special customer")?
Space-Time ran hot, but I don’t see why they would be using more flip-flops than with a conventional FPGA. Also my sources were pretty clear on the cause of their failure.
There are a lot of open FPGA efforts around. I hope someone else gives the idea another shot.
While we are on exotic approaches, Achronix' first generation was using an asynchronous (1 GHz) fabric. I regret I didn't get a chance to play with that before the pivoted to a conventional fabric.
So where the rubber meets the road is how good the tooling is and how well they execute on their tapeout. Xilinx has been much better executing their plans since Intel bought Altera.
What they gain from it? Is there some deal with TSMC behind the scenes?
It seems like TSMC is investing in some Intel's companies IMS and now this
Give my luv to the kitt-ehs ~~~~~ and count the biscuits for me too XD