Intel Completes $16.7B Altera Deal
eweek.com
eweek.com
Combine that with OSS developments by Clifford Wolf and Synflow in synthesis that can be connected to OSS FPGA tools to show even more potential here. Exciting time in HW field.
[1] See: IntelOmniPath-WhitePaper_2015-08-26-Intel-OPA-FINAL.pdf (my copy is paywalled, sorry) [2] http://www.anandtech.com/show/9802/supercomputing-15-intels-... [3] http://www.anandtech.com/show/9702/samsung-950-pro-ssd-revie... [4] https://software.intel.com/en-us/articles/distributed-memory... This is for Fortran, but the same remote Direct-Memory-Access concepts extend over to the new Xeon architecture.
https://news.ycombinator.com/item?id=10805327
Now I mean it even more.
I get the sense there's a hole in my knowledge as to exactly what kinds of limits the structure of FPGAs places on the end result. And more importantly, why.
The flexibility of an FPGA, from routing & configurable logic, takes up a lot of space that otherwise wouldn't be in the system. The blocks themselves are larger and slower than the primitives they simulate. The extra delays from this mean the FPGA can't be as fast as dedicated hardware. If FPGA simulates a CPU or GPU, the real CPU or GPU will always be faster due to optimized logic.
The other issue is power. The FPGA has all kinds of circuits that have to be ready to load up new configuration and simulate something else. Due to dynamic nature, the active parts also use more energy with all the extra circuitry. The result is that FPGA's always use more power than the custom chip.
The last one, cost, comes from the business model. No chips outside of recent GPU's have challenged FPGA's without going bankrupt or being a tiny niche player. Additionally, the EDA tools to make use of them are ridiculously hard to build with Big Two (Altera and Xilinx) investing a ton to get as good as they are. They also give them out cheap to free. So, anyone implementing a FPGA sold at cost will be unlikely to compete with Big Two on tools that make the most of the FPGA. That means anyone using FPGA's will pay high unit prices to line Big Two's pockets for quite some time.
Far as the structure, you have to map the hardware to the structure. Hardware is often done in pieces that connect to each other in a grid. So, the mapping isn't terrible. It's just hard to do efficiently.
Intel is a volume player; this makes FPGAs a bit of a head-scratcher, since in the Grand Scheme, products might get prototyped as FPGAs, but they jump to high-volume, higher-performance ASICs ASAP.
Using FPGAs allows new features and optimizations to be tested and productionized (almost) as quickly as regular software. Upgrading to a newer design is essentially free and instantaneous, compared to expensive and time-consuming process of producing new chips.
So, I was citing them as the opening chapter to the book Intel's about to be writing on how to do FPGA coprocessors. Hopefully. Meanwhile, FPGA's plugging into HyperTransport or an inexpensive NUMA would still be A Good Thing to have. :)
The processors were the FLEXlogic line. They only released a few (looks like 4 total[1]). Here's an announcement for one: https://groups.google.com/forum/#!topic/comp.sys.intel/YBUtO...
0. http://www.embedded.com/electronics-blogs/max-unleashed-and-...
People need to remember that FPGAs are not a magic bullet, especially not for throughput; they're better used for low-latency hardware interaction and things where you need cycle-deterministic behaviour.
Crypto is a far more interesting potential case.
One need error correction to fight disk/OS failures and compression to conserve bandwidth. Both tasks are highly specializable and good candidates for FPGA implementation.
Having tool like FPGA at your disposal makes you look at problems differently.
PHB, how you've grown !
I'd guess more so when sequential part is on top tier CPU and accelerator is on its NOC.
FPGAs have a nice niche for certain applications, and they deserve more prominence, better tooling and widespread use, but they are not magic.
With the right programming model, 50x seems very, very reasonable, not even magical. A typical computer does a decently high amount of context-switching and cache-loading.
But it doesn't work that way for arbitrary algorithms.
I imagine an FPGA in every chipset would go a long way towards writing the software to put them to practical use, though that may also be wishful thinking on my part.
Also, some algorithms can do over 100x. The good ones are just more likely to get less than that so i left it off. I've seen 10-50x in all kinds of papers, though. And those were academics rather than FPGA pro's.
That said, the likes of Intel, AMD, and IBM put so much work into CPU's at most advanced nodes that some sequential algorithms will likely do better on them. Synthesized, amateur FPGA bitstream just can't compete with custom, pro, hard blocks designed for exact purpose.
In practice, most applications require executing plenty of algorithms one after the other, depending on input data, with a lot of glue in between (I/O). FPGAs are terrible at that.
How Intel benefits from owning an FPGA company, I'll confess I don't know. Altera must have some awfully valuable IP that Intel needs.
With FPGA on-chip, you can offload intensive computations to semi-custom blocks. I/O processing, compression, crypto, fast transactions, data mining, you name it.
you can do that today going from ruby/js to c
The kinds of things being discussed here (compute heavy algorithms) are typically already accelerated by using high performance native libraries. Think BLAS or OpenCL or CUDA and the Python bindings in Numpy etc.
FPGAs aren't usually going to be much use in business-logic heavy algorithms (which of course are often written in interpreted languages). The performance limits of these are usually set by IO, and in some circumstances FPGAs may help with parts of that (eg, the aforementioned database indexing).
Realistically, increasing use of the new M.2 interface with fast SSDs will make more real-life difference though.
What's your source for this/why does this happen?
https://en.wikipedia.org/wiki/Out-of-order_execution
https://en.wikipedia.org/wiki/Bonnell_%28microarchitecture%2...
To make a cpu fast you need to shrink them. To the point that the "wires" in the core are so close together that electrons jump from one another. So there is lots of electrical interference. To overcome this voltage needs to be increased. This takes more power and makes the cpu hotter (which requires cooling).
But dont take my word for it. I did some quick Googling. Maybe you can find some better source and explanation:
https://en.wikipedia.org/wiki/Multi-core_processor#Technical...
"For general-purpose processors, much of the motivation for multi-core processors comes from greatly diminished gains in processor performance from increasing the operating frequency. This is due to three primary factors:
- The memory wall; [...]
- The ILP wall; [...]
- The power wall; the trend of consuming exponentially increasing power with each factorial increase of operating frequency. This increase can be mitigated by "shrinking" the processor by using smaller traces for the same logic. The power wall poses manufacturing, system design and deployment problems that have not been justified in the face of the diminished gains in performance due to the memory wall and ILP wall."
Some chatter along those lines:
http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2015/p006...
- Neuromorphic.
- Bye bye Xilinx.
Problem with massive parallelism becomes communication costs and spatial routing. Nothing is free.
More excited about commodity chips with 100s of cores. Rather have something that's easier to program with a faster dev cycle if I'm going to tackle parallelism.