Intel Marrying FPGA, Beefy Broadwell for Open Compute Future
nextplatform.com
nextplatform.com
Intel's Management Engine (which is on all modern Intel chips) acts as a unverifiable second processor with memory and network access running closed proprietary software. Its existence prevents any sort of security from state level actors.
Confused yet?
Most high-end FPGAs aren't "burned" the way you would with a logic array by burning fuses. Configuration data (i.e. the connections between the FPGA's tiny components) is stored in a non-volatile memory and is loaded when the device starts. Altering the design is simply a matter of altering this configuration data, which is easy to do dynamically, seeing how there's a closed piece of code running on a closed module which has access to every bit of it.
The problem ends up being boiled down to problems that we're already aware of: authenticating configuration data (which is akin to the problem of authenticating the OS running on a general-purpose CPU), ensuring that the FPGA's configuration matches that which was programmed in the non-volatile memory an so on.
The configuration data loader is, to the best of my knowledge, a pretty trivial piece at the moment, with the exception of high-end devices for sensible applications (which do include things like encryption, so that the bitstream cannot be retrieved in a useful form). But real-world requirements will soon provide a good excuse for inflating it to a level of complexity where backdoors can be hidden.
It's also important to realize that much of the hardware that ends up on an FPGA isn't really arbitrary data, it's in the form of vendor-supplied IPs that are probably pretty easy to recognize. Implementing your own cryptography hardware is as bad an idea as writing your own cryptography code. It don't think it would be too hard to backdoor a loader that alters the bitstream so that its crypto modules are weak, under specific circumstances.
http://www.xilinx.com/support/documentation/user_guides/ug38...
In case anyone wants a CPU + FPGA combination without this type of security risk, check out Zynq.
Intel ME on the other hand is baked into the hardware, and AFAIK can't be properly switched off, it appears the best you can do is set up a fake Intel MPS Server to point it to:
https://software.intel.com/en-us/forums/intel-business-clien...
With regards to why people aren't talking more about this, maybe fatalism is involved as it seems there is a certain inevitability that whoever controls the silicon also controls to some extent computers. There have been some interesting developments in people trying to create open hardware like lowRISC http://www.lowrisc.org/ and also using ARM so there is some hope in this area.
In short, the overlap between people who know about it and are also security conscious is pretty small. There's also dozens of other things people should be more concerned about in terms of a corporate or state actor gaining unauthorized access to your computer.
- Introduced a long time ago; it's old news.
- It is talked about, but only in certain circles
- Confusion over the relationship between vPro, IME, and AMT, or at least unfamiliarity with those terms making articles/discussions related to them less likely to pop out as interesting topics.
- Some of the people who know about it consider it the normal, boring, status quo, and they react to anyone shocked by it as an alarmist.
[1]: http://hackaday.com/2016/01/22/the-trouble-with-intels-manag...
The x86 platform, like most processors, has a documented instruction set and software loading process. (There are undocumented corners, but the "front door" is open). Whereas historically almost all FPGAs have had fully closed bitstream formats and loading procedures. This necessitates the use of the manufacturer's software which is (a) often terrible and (b) usually restricted to Verilog and VHDL.
If Intel ship a genuinely open set of tools, then all manner of wonderful things could be built on the FPGA, dynamically. That requires being open down to the bitstream level, which also requires that the system is designed so that no bitstream can damage the FPGA.
To me this is most interesting not at the server level but at the ""IoT"" level; if they start making Edison or NUC boards that expose the FPGA to a useful extent.
It also is new and thus not known and used by thousands of RTL coders everyday.
Verilog/SystemVerilog are pretty great at what they do. People are pretty happy with them. The problem is those who get into FPGA (from CS background) expect to write "code" which isn't what you are doing in the hardware world.
There is also SystemC when you want to write test benches, or Bus-Functional-Models. Those is also used pretty extensively in the hardware flow.
I agree that writing code will lead you astray, but not with "people are pretty happy with Verilog". It has a whole load of limitations: no aggregate types, no language-level distinction between synthesisable and unsynthesisable, and quasi-sequentialism that confuses beginners.
Really?
Any HDL language will be confusing to someone who isn't used it because you are defining a set of parallel behaviours and responses to specific stimulus and not a set of instructions.
There can always be improvements but as of today Sys/Verilog is the most popular by far, followed by VHDL. Therefore the tool vendors support those languages which is what I was responding in the original post.
I've been waiting to see this kind of thing for years, ever since I read Adrian Thompson's work on evolution with FPGAs, in which he:
"Evolved a tone discriminator using fewer than 40 programmable logic gates and no clock signal in a FPGA" (slides: https://static.aminer.org/pdf/PDF/000/308/779/an_evolved_cir...)
EDIT: Full paper: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.50....
The field has crawled along pretty slowly since then as far as I can tell.
However, this could be a HUGE thing for computing; developers would finally have a way to create hardware that interacts directly with the physical world in ways we haven't thought of yet. As a small example, Thompson's work revealed circuits that were electromagnetically coupled in novel ways but not connected to the circuit path, yet were required for the circuit to work. Using evolution, in time we should be able to come up with unique solutions to hardware problems not envisaged by human designers.
This is really exciting.
Before the difference between highly optimized code and ok code was maybe 2-3x speedup. Roughly one Moore's law doubling. With heterogeneous computing that is more like 20-30x or more. And Moore's law is dead! This will change a loot in the IT world (More servers are the solution, developer time is more expensive then computer time, etc..). Learn C, down on the metal programming is back, the future is heterogeneous parallel computing.
I would argue that the low level work you are doing should be done in a macro or compiler.
http://www.graphics.stanford.edu/~hanrahan/talks/dsl/dsl1.pd...
http://www.graphics.stanford.edu/~hanrahan/talks/dsl/dsl2.pd...
Pat Hanrahan makes a compelling argument for using special purpose DSLs to construct efficient performant code that takes advantage of heterogeneous hardware.
See the Design of Terra, http://terralang.org/snapl-devito.pdf
I personally really like the idea from the the Halide language, having one language for algo, another for how the computation is done. If something like that could be made general purpose it would be very useful.
>should be done in a macro.. Encourage c programmers to use macros is like encouraging alcoholics to drink :) But I guess you didn't think about pre processor macros.
>...macro
Yeah, I didn't have preprocessor macros in mind. ;*| But wonderful, AST slinging hygienic Macros!
Take a look at http://aparapi.github.io/ it one of the best examples of making OpenCL a first class citizen in Java.
I always imagined that the best use of FPGA's in systems like this would as an I/O coprocessor. If the only way to get to the FPGA is via the CPU, then most (all?) of the benefit is lost.
Kind of wondering if these will work with Intel's Omni-Path, and if so, what the shape of that makes things...
In theory, could be very interesting. Reality though, we get to find out. :)
Plenty of "kernel bypass" and RDMA type functions use shared/user-space memory for "zero-copy" (in reality one copy), operations between NIC and software. If a similar scheme can be used with the FPGA then it would not have too much overhead. I agree, not as direct/efficient as having FPGA serdes I/O go directly to some SPF+/network transceiver, but then you'd also be taking up valuable FPGA gate capacity to run NIC PHY/MAC and standard L2/L3 processing functions that you get from a NIC.
They do exactly this in a programmable PCIe card and custom ASIC.
Buggy compiler implementing a superset of a subset of C89 with totally crazy macro extensions, and bizarre locality properties (e.g manually declaring if a variable is in a register or in ram), bad impossible to decipher (machine generated!) "documentation", 3 (!!) different and incompatible "standard" libraries each implementing different sets of features, 2 of which were written in assembler and inaccessible from C directly and almost no debugging tooling. e.g I had to write my own locking library because there wasn't one. What a nightmare.
The GPU was successful because it had a killer app: Gaming. What's the killer app for the FPGA going to be?
It will be awhile before it shows up in consumer gear, as the use cases are not there yet. Consumers may still benefit as when someone figures out something amazing for it to do, they will get a hardened version of it.
I wonder whether Intel will allow that. Better hardware offloading for various algorithms (SHA, RSA, AES, …) and hypervisor acceleration (VT-x, VT-d, EPT, APICv, GVT, VT-c, SRIOV, …) have been one of the main selling points for new CPU generations. An FPGA would render most of them moot by allowing operators to configure whatever offloading they need without requiring new, expensive Intel chips.
It will also open up the ability for them to separately sell offloading features as IP cores. Also the risk of them having to disable a feature because of an error goes down as they can easily issue an update for it.
That is not quite right. Rather: "Anything reasonably parallel that needs high throughput or low latency and which is not already provided by the CPU or GPU can benefit from an FPGA."
Modern server and desktop CPUs are incredibly fast and offer a lot of parallelism for certain operations. For example, when it comes to floating-point operations, no FPGA has the slightest chance against a desktop CPU (for prices in the same order of magnitude), and even less against a GPU.
Intel Xeon E5-2670 v3: 68 GB/s [1].
NVidia K40: 288 GB/s.
So you're wrong about external memory bandwidth. Regarding internal bandwidth: True, FPGA block RAM has a huge aggregated bandwidth, but it comes with some limitations.
Also, think of all the high-speed tranceivers to feed the data into an FPGA.
For the problems I am working with, FPGAs, even the mid-range ones, are far more suitable than the top GPUs.
What FPGA device are you using? How many memory controllers are on it, what width do they have, and at what frequency are the memory modules operating? What external bandwidth do you actually achieve?
> For the problems I am working with, FPGAs, even the mid-range ones, are far more suitable than the top GPUs.
I believe you, but we're talking about maximum external memory bandwidth, not suitability in general.
The program to reprogram them every few seconds (genetic algorithms?) sounds more interesting.
(I assume the FPGA would need a way to take over the computer bus and access memory? How multiple functional FPGA subunits do that, I leave as an exercise for Intel. :-) )
A dynamic software accelerator, not a specific function one.
(It is probably obvious from the above that the little I know of hardware and FPGAs is a bit old... we almost used keyboards made out of stone. :-) )
The complete compilation process (source code --> netlist --> mapped design --> placed & routed design) can take several hours for large designs; maybe several minutes for small designs. Not suitable for JIT.
No: Physics simulations are typically FLOPS-limited; modern CPUs deliver a lot more floating-point operations per second than FPGA devices at comparable prices.
Unfortunately, simulations often cover a wide range of exponents during a single run (one matrix element might be 5.74293574325e8, while the one next to it might be 3.25356343e-9, and you still want to preserve their precision), and the exponents of the inputs might vary a lot between different runs. You can only use fixed-point if you have a good idea of the exponents of the input numbers, and how those change during the course of the computation. That works well for typical digital signal processing applications, and not so well for generic number crunching libraries.
Even then it breaks down with more complex signal processing algorithms. FPGAs are great for simpler algorithms like FFTs, digital filtering, and motion compensation. They aren't quite as good at more complicated algorithms like edge detection. They really break down when you want to use ML-based algorithms.
More precisely, to put AI on smart phones. Neural network is cheaper on FPGA. You can switch between AIs (voice/image/text recognition, games, etc) quickly and update them over the internet.
If Intel added a fully open, fully specified FPGA to a CPU, that anyone could write tools for, that could change.
The actual problem (as stated in the other comments) is that the tooling is all proprietary (huge and scary) and has received NO love.
Mostly a non-issue. Multi-chip packages have been around for decades. Intel has been putting DRAM chips onto their CPU packages (for their Iris Pro graphics chips) for a few generations now; Core i CPUs also have on-chip PCIE controllers together with their QPI inter-cpu interface, if for some reason the latter can't be made to work with FPGAs.
There are even vendors marketing PCI FPGA cards as "OpenCL FPGA cards": http://www.nallatech.com/solutions/fpga-accelerated-computin...