Intel Xeon processor with FPGA now shipping
fpgaer.wordpress.com
fpgaer.wordpress.com
The initial workload that Intel is targeting is putting Open Virtual Switch, the open source virtual switch, on the FPGA, offloading some switching functions in a network from the CPU where such virtual switch software might reside either inside a server virtualization hypervisor or outside of it but alongside virtual machines and containers. This obviates the need for certain Ethernet switches in the kind of Clos networks used by hyperscalers, cloud builders, telcos, and other service providers and also frees up compute capacity on the CPUs that might otherwise be managing virtual switching. Intel says that by implementing Open Virtual Switch on the FPGA, it can cut the latency on the virtual switch in half, boost the throughput by 3.2X, and crank up the number of VMs hosted on a machine by 2X compared to just running Open Virtual Switch on the Xeon portion of this hybrid chip. That makes a pretty good case for having the FPGA right close to the CPU – provided this chip doesn’t cost too much.
Indeed. What benefits the most from being turned into FPGA circuit are stateless or minimally stateful circuits like media decoders/encoders.
Complex stuff like fancy routers, classifiers, internet protocol inspection/handling will not gain a hundredfold speedup unlike the stuff above. This is why cheap x86 based routers are still a thing.
At the moment, I am involved with one cloud provider in China that bids big on cheap FPGAs and RDMA. AWS and Azures can be defeated in detail.
The plan is following: provide hardware accelerated "building blocks" of any modern dotcom business.
Need memcached? We have it running on RDMA, from a bare metal ASIC, 10 times faster than any x86.
Need transcoding? We have it available on RDMA, from a bare metal ASIC, 10 times cheaper than any AWS instance on buck/megabyte.
Need API proxy for TLS/Gzip with gigabytes per second throughput? We have it running on RDMA, from a box with 4 PCIe accelerators, and AWS has nothing to offer for this use case other than "buy a hundred of top tier high performance instances and put them behind load balancers"
Can you share how this FPGA and RDMA combination works architecturally? What's the software stack composed of? Might you have any links or other resources you could share?
You provision access over in-dc REST api, then rdma capable routers route the appliance to your vm, then you access it through provided libs for linux.
In other words, you don't deal with RDMA yourself at any moment
Pretty narrow still, but the market segment is possibly real.
Buses exposed through the M.2 connector are PCI Express 3.0, Serial ATA (SATA) 3.0 and USB 3.0, which is backward compatible with USB 2.0. As a result, M.2 modules can integrate multiple functions, including the following device classes: Wi-Fi, Bluetooth, satellite navigation, near field communication (NFC), digital radio, Wireless Gigabit Alliance (WiGig), wireless WAN (WWAN), and solid-state drives (SSDs).
And key E is used for wifi. And this card here.
I only know all this because I just bought one, and have been hunting for something useful to put in that slot...
1. Optional M.2 key M 2280 in the main bay via an adapter.
2. M.2 key B 2242, the Lite On T11 works there for an SSD. It's the only consumer M>2 key B+M 2242 PCIe SSD. This generation doesn't support SATA disks in this slot, the T470 did.
3. M.2 key E 2230 as practically all laptops do since Broadwell for wifi.
It's all very confusing and frustrating. So the chances of me finding anything useful to put in that slot are roughly nil, then? I can see no physical or technical reason why my slot should not support this card, except that the card only supports two out of the three possible PCIe-carrying configurations, and I was unlucky. Seems like a poorly designed "standard".
It especially rankles because they devoted two whole PCIe lanes to that damn slot, which would be better put to use unbottlenecking my highly expensive Samsung P981 NVMe drive.
The T480s does not suffer this problem.
The T480s is a new machine.
The galatea from numato is pretty affordable and has 2Gb (256MB) of onboard RAM and comes with a dual ethernet and pmod breakout board from Amazon for $300 USD.
http://www.latticestore.com/products/tabid/417/categoryid/59...
The lattice ECP5 PCIe board comes with I think 128MB DDR3 and dual gigabit ethernet.
And for boards without memory but lots of I/O there is Mesa Electronics: http://www.mesanet.com/ CNC oriented but you can do whatever you want with the FPGA.
You may do some on-line processing like running neural net on the video content, but that's about it. Don't expect anything super exciting from that chip.
Yet, it will bring joy to high frequency traders. The systems there do most of the work in FPGA, including UDP/TCP/IP packets processing and offload some work to CPU (broadcasts about network topology are handled on CPU, for example). They also would like to receive CPU computation results as fast as they can and this chip is exactly that.
By FPGA standards this thing is enormous.
For block RAM in previous generation FPGAs from Altera there were one read and one write paths. To have two read paths you would need to copy block RAM as many times as you need read paths. This means that if you search for block content inside a macroblock with N parallel accesses you would need N copies of macroblock stored.
Tabula's time shifting tech allowed for up to, I believe, twelve paths into the block RAM, six for read and six for write (they were time-scheduled to 1-read-1-write block RAM operating at six times the frequency). I thought this thing would be a road for superscalar FPGA CPUs, but Tabula was closed.
You can imagine not using RAM at all, but then you will spend other resources in FPGA.
These are problems with video compression I see here. I think they are substantial but not unsurmountable and require a balance to solve. It is just me that I saw the balance is not in favour of FPGA.
It is a risky strategy however. Even if they can attain similar performance, which I doubt, programmability remains the big problem for FPGAs. I know Intel is pushing openCL but it simply does not have an ecosystem for software right now, and it remains to be seen if they can even enable much of the feature set of openCL on an FPGA.
Embedded FPGA accelerators are of interest primarily to large server farms and cloud operators, where the cost of development for FPGA acceleration is cheaper than the savings of acceleration-- Microsoft in particular has done a lot of work in this area (for Bing), and probably worked with Intel to define the CPU.
It's also at least plausible that they'd probably have people trained to think like Americans who might think of the same thing in advance.
The answer is rather simple. Xeon Phi wasn't as good as the GPU counterpart.
http://tomforsyth1000.github.io/blog.wiki.html#%5B%5BWhy%20d...
Got a source on this? I've not seen an official mention of what sort of architecture NERSC-9 will be.
> Not sure what Argonne is doing but there’s no reason to think it won’t be GPUs.
It will be a future Intel processor [1].
[1] https://www.hpcwire.com/2017/09/27/us-coalesces-plans-first-...
I agree because there's added levels of complexity with HDLs that SW doesn't have to deal with. What would be nice is if there was a tool that you could declare your problem (functions) and your constraints (throughput, LUTs available) and it would figure out the memory, ALUs and pipe-lining needed to solve the problem.
In my experience, high-level synthesis makes things a lot easier for the programmer, but you still have to be aware of very low-level details (and you still can get bitten by abstractions that you don't fully understand).
If their system accepts "mostly ordinary" code in eg C, a fair amount of stuff would probably work with it, so it could scale beyond educational exploration/tinkering, too.
FPGAs aren't known for their FLOPs though. Sure, the high end ones pack a punch, but compared to GPUs are still extremely expensive.
I am not saying these things don't have applications, but the "usual" computational workloads are not one of them.
In theory, an application with a computation-heavy task could program the FPGA to provide part of that task in hardware (think the hot innermost loop body).
What I am worried about is the infrastructure that is needed to make this happen: Is there even support for this in our compilers? What would support look like?
FPGA offload has been a very active area of compiler research for more than a decade. I like this paper for example https://dri.es/files/fpl05-paper.pdf
Here's one of the few references to that project that I can now find online, a WIRED era piece with all the gushing nerd optimism of the pre-dot-com-bust: https://www.wired.com/1999/02/the-super-duper-hypercomputer/
I suppose this is where Intel is headed now with these first few tentative steps.
This is absurdly far away; it reminds me of the decades of assumption that 4GLs are going to obsolete programmers or the decades of trying to cross-compile C to FPGAs, badly.
It doesn't help that there's a huge infrastructure barrier caused by closed tools. Imagine if Intel brought out a processor with a proprietary instruction set where you were only allowed to use their FORTRAN compiler (no C, let alone anything more modern or JIT) with a per-seat license. That's where FPGA tools are.
We won't see a Cambrian explosion in FPGA tooling until they are made properly open. Building using open-source tools needs to be actively supported by the manufacturer.
There are also conceptual obstacles; FPGAs are sufficiently different that programmers have to re-learn and re-write idioms in order to get usable results. It's as big a jump as going from Javascript to CUDA.
I agree with you about the problems of closed-source - but I feel like the end of Moore's law will encourage creative solutions to performance problems - which FPGAs would at least make technically possible. And, even if it's like the old days when people bought compilers and access to source, people will still do it - since it'll be the best way to get an edge on the competition.
In many applications data access is the bottleneck, and an FPGA coprocessor will not give any improvements in that area.
Using an FPGA to do arithmetic computations (e.g. as in audio/video coding) could provide a benefit, but I fail to see how this improves things much over having an extra CPU core.
The only thing an FPGA does is replace registers by wires, basically, and there's not much to be gained there, I guess.
What has changed here? Are FPGAs faster now or are many more people experimenting, creating sufficient demand for custom hardware at short notice?
FPGAs are a waste of Silicon compared to ASICs, but for fixed applications they are less wasteful than CPUs.