FPGA and Xeon combined in one socket
theregister.co.uk
theregister.co.uk
(Not that shader language or CUDA is that accessible, but you can play with shaders in a browser now. The time to "hello world" or equivalent isn't too bad)
Modelsim is available on Linux as well, but only their more expensive SE product. They charge a premium for Linux platform.
Xilinx and Altera keep trying to convince us that they are software vendors, when in actual fact they are hardware vendors. Rarely is a company good at both.
It reminds me of the days of proprietary vendor compilers where every platform vendor had their own subtly incompatible and/or differently buggy C or C++ compiler. In this way the FPGA development model is a decade or two behind software workflows.
The use cases for FPGAs are a much harder impediment to adoption. Many people get "FPGA boners" when they even hear the word, fancying themselves "chip designers," but practical use cases are much rarer. As evidence, notice they predominate in the military world, where budget is less of an issue than the commercial world.
The technical issue with FPGAs is that they are still one level abstracted from any CPU. They are only valuable in problems where some algorithm or task can be done with specific logic more quickly than the CPU, given 1. the performance hit of reduced real estate and 2. reduced clock speed relative to a CPU and 3. more money than a CPU.
Further diminishing their value is that any function important enough to require an FPGA can more economically get absorbed into the nearest silicon. For example, consider the serial/deserial coding of audio/video codecs. That used to be done in FGPAs, but got moved into a standard bus (SPI) and moved into codecs and CPUs.
Because of this rarity, experienced engineers know that when an FPGA is introduced to the problem in practical reality, it's a temporary solution (most often to make time-to-market). This confers a degree of honor which is why people get so emotionally-aroused about FPGAs.
You can bet though, that if whatever search function Microsoft is running on those FPGAs proves to be useful, it will be soon absorbed into a more economical form, such as an ASIC, or, more likely, additional instructions to the CPU.
Really, installs on 1,600 servers such as this article reports, is not that impressive and certainly only a prototypical rollout.
What chillingeffect is saying (and I agree) is that this logic should be running on the cpu.
And in the case of Catapult, they'll be refactoring the algorithm to represent changes to their search feature matching, which is the majority of what they offloaded to FPGAs.
In general it seems that hardware vs. software fit is orthogonal to concerns such as frequent reconfiguration. Some algorithms simply match hardware well (high concurrency, low-complexity control flow, able to stream through data without complex state). These are often algorithms that do not fit general-purpose CPUs well (cache hierarchy wasted on streaming; lots of control overhead; low core counts relative to FPGA-level parallelism). Some of these algorithms may be for specialized and/or frequently-changing applications such that they should not be burned into an ASIC that will live in a datacenter for 3-5 years.
- 95% throughput increase at same tail latency, - 29% tail latency improvement at same throughput
I would have expected factors or orders of magnitude.
Further ,this increases total cost of ownership by 30%, So the performance improvement is just about 70%.
FPGA is widely used in control circuits for industrial uses, for example.
There's nothing better for prototyping and proof-of-concept work, though. And they'll always have a place in low-volume applications that aren't cost-sensitive.
FPGAs promise the designer 'arbitrary logic' and deliver 'a place others sell into.'
I disagree that FPGA experts "like" the tools they are given, they tolerate them. One of my friends worked at Xilinx for 15 years and understood this all too well. He felt the leading cause of the problem was that the tools group was a P&L center, they needed to turn a profit in order to exist. They got that profit by charging high prices for the tools and high prices for support. His argument was that 'easier' tools cut into support revenue. When I've had high level (E-level, but not C-level) discussions with Xilinx and Altera there has been a lot of acknowledgement about the 'difficulty of getting up to speed' on the tool chain and many free hours of consulting are offered. From a business engagement point of view, making hard to use tools and then "giving away" thousands of dollars of free consulting to the customer to gain their support seems to work well. The customer feels supported, and stops wondering why if the have consultants around for free those consultants wouldn't just make the tools more straight forward to use and available on a wider variety of platforms.
But the biggest thing has always been intellectual property. You buy an STM32F4 and it has an Ethernet Mac on it (using Synopsis IP as evidenced by the note in the documentation), you pay $8 for the microprocessor, work around the bugs, and get it running. If you buy an FPGA, lets say a Spartan 3E, you pay $18 for the chip, and if you want to use that Synopsis Ethernet MAC?[1] $25,000 for the HDL source to add to our project $10,000 if you are ok with just the EDIF output which can be fed into a place-and-route back end. Oh and some royalty if you ship it on a product you are selling.
The various places that have been accumulating 'open' IP such as Open Cores (http://opencores.org/)have been really helpful for this but it really needs a different pricing model I suspect. A lot of HDL is where OS source was back at the turn of the century (locked down and expensive).
[1] I did this particular exercise in 2005 when I was designing a network attached memory device (https://www.google.com/patents/US20060218362) and was appalled at the extortionate pricing.
For a long time I considered soft CPUs a bad idea (the Stretch guys kept trying to sell me on them but since I wasn't really doing things like deep packet inspection I didn't have a good use case, even RAID algorithms on them were better handled by pretty generic DSP type architectures.) However in playing with the Zedboard which has a couple of Cortex A9's attached to the Xilinx fabric I find some interesting things there. If only as a new kind of 'i/o' but that is neither I/O port based nor memory map based (it expresses as memory but it feels different than the memory mapping of old like on the PDP/VAX machines and 68K systems). Could just be nostalgia though.
To me, the entire recent history of computer industry (well, all of it is recent BTW) shows that, if you want your technology to become mass-adopted, you need to make it easier for the little guy to get in the game. The high school kid tinkering with stuff in the parents' basement; the proverbial starving student. That's how x86 crushed RISC; that's how Linux became prominent; that's how Arduino became the most popular micro-con platform (despite more clever things being available).
You make the learning curve nice and gentle, and you draw into your ranks all the unwashed masses out there. In time, out of those ranks the next tech leaders will emerge.
Their counter is of course that they have customers who sweat the $0.25 difference in price. (which I understand but $10,000 in tools and $15,000 in consulting a year is a hundred thousand chips. Which they say "oh at that volume we would wave the tooling cost." And that got me back to your point of "You already have their design win, why give them free tools? Why not give free tools who have yet to commit to your architecture?"
It is a very frustrating conversation to have.
Reconfiguration is also very much needed - customers often experience network and different specific problems that are only fixable in hardware. So we need configurable packet processors (basically top level router in a chip), configurable processing fpga, lots of small fpga etc. And bug fix for one customer is then gradually distributed to everyone. And devices lifespan is long - many customers use 5 or even 10 year old hardware. Performance is enough for them and new functionality is provided for them, partially in way of new fpga firmware.
PS: some devices, especially cable or satellite, are sold in thousands by several companies, so its not like unique hand-made hardware.
PPS: of course any costs to buy toolchain or windows pc in such companies doesn't matter really. Finding talented fpga designers is way harder as far as I understand.
But isn't that because, back then, CPUs simply weren't fast enough for decent video coding? And they still aren't blazing fast for that purpose - compared to GPUs (see GPU-enabled video encoders).
I think the main argument against FPGAs is that it's still a chore to re-purpose them. Sure, it's not as painful as making a new chip from sand, but it's harder than applying code changes to software running in production on CPUs.
If it's a relatively simple algorithm, that changes rarely, where speed is the main bottleneck, that's a very promising scenario for an FPGA.
[0]http://www.anandtech.com/show/7334/a-look-at-alteras-opencl-...
But closed source expensive tools are a problem.
1. I used DigiKey prices from the largest bulk tier. Presumably if you're actually ordering that much you can get them right from the supplier for cheaper.
2. Latency! You can make it almost arbitrarily low on FPGAs.
3. Direct interface with the hardware.
I'd love to see a real cost/benefit analysis by someone with skin in the game because mine was pretty simplistic.
To me, that sounds more like the problem was not a good fit for FPGA hardware.
I'm just speculating, though.
If they're clever enough to make some of those IP cores available to say MATLAB adoption will be faster still.
Nothing sells hardware easier than "do no extra work but spend another couple of grand and see your application speed up significantly"
MATLAB already have MATLAB->HDL, which works very well. We have a team that uses it exclusively for FPGA programming.
If Intel does a good enough job of providing a collection of compute kernels and the surrounding CPU libraries to make using them roughly as "easy" as CUDA then a lot of people will pick that up.
I don't have any hard numbers but I would suspect that there are a great many more people who use MATLAB on a CPU than those who do MATLAB->HDL. So what I'm speculating about is that Intel might support those folks who use MATLAB on a CPU for more general purpose things.
Does that make more sense?
Scattered reports of maniacal cackling amid driving rain and lightning at Chipzilla's lab"
Is this just a Register thing, or do all UK rags use this kind of unprofessional hyperbole? It's literally the most annoying thing in the world.
There is always some kind of British humour on the articles.
Zynq has been out and working in industry for a couple years now.
I've gotten the impression that putting a general purpose CPU in the corner of an FPGA was a pretty standard thing.
One of the things that should differentiate this new effort from Intel is FPGA "direct access to the Xeon cache hierachy and system memory" per "general manager of Intel's data center group, Diane Bryant".
The other thing to know is that HDL languages are mostly the domain of electrical engineers and hence have suffered a lack of any "computer science" in them. The languages and all of the tools are clunky and reminiscent of 1970/1980's style programming when CS and EE diverged. Hence, do not expect to find decent online tutorials or freeware source code available. It's all locked up and proprietary as with all other EE tools.
The best place is to start with a text book, this one (http://www.amazon.com/Fundamentals-Digital-Logic-Verilog-Des...) is a nice introduction to digital design with examples from Verilog.
Personally I prefer VHDL, and this fantastic introduction (http://www.amazon.com/Circuit-Design-VHDL-Volnei-Pedroni/dp/...)
To make either of these useful, you will need a hardware platform and some tools to play with. The DE1/2 is a reasonably priced entry board with plenty of lights, switches and peripherals to play with at a reasonable cost and is well matched with the text books above.
http://www.terasic.com.tw/cgi-bin/page/archive.pl?Language=E...
I also agree that it's stuck in the 70s, comparable to Fortran 77 or ALGOL. Bundling related signals and functionality together to produce something corresponding to an "object" or a "type" is basically impossible. All sorts of errors that could be caught automatically or discouraged by language design aren't. There's a lot of typing, and necessary duplication of effort. IDEs don't help much.
Heavy unit and system testing is fortunately widespread. Because it's the only way to ensure you end up with something that actually works.
I have a back-burner project to design a more modern language that compiles to Verilog which would make this sort of thing much more accessible.
2. Fun fact: the cover of the Fundamentals of Digital Logic book has Chess on it because the author, Zvonko Vranesic, is not only an father in the FPGA/CAD industry, he is also an International Chess master. Also, he's quite good at ping pong for being 76 :(
Any room left on silicon for things like computer vision (well, simple stuff, like recognizing a red ball), or is the whole thing pretty much dedicated to flying the quad?
Also, could you share your design?
http://www.digilentinc.com/Products/Catalog.cfm?NavPath=2,40...
Could you describe why you prefer one over the other?
Perhaps related to your particular application for FPGA?
> and is well matched with the text books above.
Is that likely to be true for the lesser DE0 model too?
It may sound crazy, but I prefer VHDL because it is more verbose. The benefit of the verbosity is precision. with VHDL, you must specify exactly what you want, the syntax doesn't allow for ambiguity. With Verilog, you can let the "compiler" infer some things for you, but you need to think really hard if it will infer the right thing for you. The benefit is that you get to type fewer characters. Since I do not pay (or get paid) per character typed, it is far more important to me to type a little bit more, and get exactly the design I want. When designing circuits, one thing you cannot afford to be is lazy. It is all about precision, because if you get it wrong, there is no step through debugger to help you, and often, it will still work in simulation, but fail to work in real hardware, which means you're stuck.
I can't say for certain with the latest version of the textbook, but IIRC, the version I had many years ago was specifically tied to the DE2. There were a range of exercises to go through with specific problem statements etc. IMHO, the DE1 is equivalent enough that you can change a few pin mappings and get the same result. I cannot say anything about the DE0.
I guess what I'm really trying to say is, study digital logic first, then imagine the circuit you want to build, then write the Verilog that infers that circuit :-)
In the end, the success will boil down to how easy the development is, and how well designed the libraries will be - if the framwork will be capable to automatically reconfigure the hardware to offload CPU-intensive tasks, this has high tech potential for widespread adoption, not just datacenter-wise.
FPGAs are typically used in ASIC development to emulate the ASIC being developed. I've seen boards with 20 FPGAs emulate an ASIC design at <~1/10th of the speed at >>x10 power. While FPGAs are programmable hardware they are far less efficient than custom hardware for various reasons. Naturally ASIC emluation is an application where FPGAs have a very large advantage over software... At volume they're also a lot more expensive and good tools are also very expensive (virtually no mass produced commercial product uses FPGAs). Now obviously if the FPGA is inside the Xeon you're not really paying much more for it (except you lose whatever other function could be crammed in there).
Companies like Microsoft, Facebook, Google have enough servers to make a custom block inside Intel's CPU more attractive than an FPGA in terms of price/power/performance (and they can get that from ARM vendors which is probably scaring Intel).
CPU vendors have spent the last several decades moving more and more applications that used to be in the realm of custom hardware to the realm of software. There are certainly niches of highly parallelizable operations but a lot of general purpose compute is very well served by CPUs (and a lot of it is often memory bandwidth bound, not compute bound). Some of these niches have already been semi-filled through GPUs, special instructions etc.
The FPGA on the Xeon is almost certainly not going to have access to all the same interfaces that either a GPU or the CPU has and is only going to be useful for a relatively narrow range of applications.
I think what's going on here is that as the process size goes down simply cramming more and more cores into the chip makes less and less sense, i.e. things don't scale linearly in general. So the first thing we see is cramming a GPU in there which eventually also doesn't scale (and also isn't really a server thing). Now they basically have extra space and don't really know what to put in it. Also each of the current blocks (GPU, CPU) are so complicated that trying to evolve them is very expensive.
EDIT: Just to explain a little where I'm coming from here. I worked for a startup designing an ASIC where FPGAs were used to validate the ASIC design. I also worked on commercial products that included FPGAs for custom functions where the volume was not high enough to justify an ASIC and the problem couldn't be solved by software. I worked with DSPs, CPUs, various forms of programmable logic, SoCs with lots of different HW blocks etc. over a long long time so I'm trying to share some of my observations... If you think they're absolutely wrong I'd be happy to debate them.
EDIT2: Re-reading what I wrote it may sound like I am saying I am an ASIC designer. I'm not. I'm a software developer who has dabbled in hardware design and has worked in hardware design environments (i.e. the startup I worked for was designing ASICs but I was mostly working on related software).
What if the Intel FPGA did have access to the same resources as a GPU? This isn't inconceivable, it's in the same socket as the CPU.
This gives you the ability to implement specialized algorithms related to compression, encryption, or stream manipulation in a manner that's way more flexible than a GPU can provide, and way more parallel than a CPU can handle.
An FPGA that competes with a modern GPU would probably cost in the neighborhood of US $50,000 per chip.
Being on the same die is better than being on a separate chip but there are some internal interfaces that rely on placement and latency on the die. E.g. it's unlikely that you can add new instructions to a CPU via FPGA or have the FPGA interact with the L1 cache. It's more likely there will be some sort of shared memory and standard peripheral interface to the FPGA (e.g. interrupts, I/O).
EDIT: So to expand on this the FPGA is expected to have a relative high latency to the CPU and a relative low bandwidth to external resources (e.g. if you compare the CPU interface to L1). It's unlikely that the FPGA will have the same cache hierarchy that a CPU core has (size and performance). So it'll be useful where change is expected, the bottleneck is compute, the task is highly parallelizable, the standard instruction set/other blocks aren't very good at, and going for a pure ASIC solution doesn't make sense (either in a separate block or onboard a customer version of the same chip) either because of price/time/volume.
And yes i know that the cpu will be the bottleneck, but it will be the bottleneck anyway.
The other thing that I've seen which may or may not apply to the Intel case is that complexity in chip design can be managed more easily by having blocks that connect to standard interfaces. I.e. if you look inside the Xeon it probably looks like a bunch of different chips that were thrown onto the same die with some standard interconnects. Most of the optimization effort goes inside those blocks, e.g. inside a single core, and it's a lot more difficult to add an FPGA closer to the core vs. just throwing it somewhere else on the chip. That is the number of engineers in Intel who are intimately familiar with the innards of the x86 core design and are capable of making these sorts of changes is probably much much lower than the number who are capable of throwing in some external "block" onto the die and tie it into a standard bus.