*This equation changes if suddenly the pace of semiconductor process miracles starts slowing down
Last I checked, number of cores attached to large memory, on Intel chips, is still going up, and that's what affects throughput on embarassingly parallel jobs, which is what these are.
This is a very interesting phenomenon. It was certainly true from around 1968 to 2005, that custom chips could not survive since the next Intel chip made them irrelevant. But I think that, not only is this going to be changing soon, in fact it has already started to change.
There are credible claims of BLAST (and other life science codes) being ~10x faster on GPUs than CPUs. [1] GPUs are specialized processors (with associated non-standard programming models and compilers) that have overtaken general purpose CPUs in many areas. This has been possible because, for the past 10 years, the "next Intel chip" has NOT been faster (single core clock speed) or exponentially cheaper (cost per core). Intel has been unable to make advances on these fronts due to fundamental engineering factors, such as the breakdown of Dennard scaling around 2005. [2]
GPUs (and FPGAs, and to some extent even conventional CPUs) have been able to continue advancing since 2005 because Moore's Law has been holding (even though the related Dennard scaling law has broken down). However, now we are seeing that even Moore's law is starting to break down. This is evidenced by the increasing delays in the Intel roadmap, with the most recent delay being 10nm process being pushed out to 2017. [3] It is now in question whether silicon process will ever even reach 7nm, and I don't think anyone is willing to bet that Moore's Law can continue on silicon beyond 7nm.
It is at the end of Moore's Law where custom chips become very interesting. This has been covered in 2013, in a presentation by Robert Colwell, who is the Director of DARPA's Microsystems Technology Office. [4] Colwell's thesis is that the End of Moore's law will revive specialized chip design. I find that prospect to be very exciting, not only for chip designers but of course also for software developers (especially compiler and programming language designers, as specialized chips are going to require specialized compilers, programming models and languages).
Of course there is also the other possible future path, within the 10nm to 7nm time frame (i.e. within the next 5 years) that Intel and others will find a way to extend Moore's Law, possibly by finding a viable alternative to silicon substrate. That would also be extremely exciting but somehow I don't think the future will be that simple (DARPA seems to take the view that this simple isn't going to happen, i.e. the Moore's Law exponential must logically end).
[1] https://www.nvidia.com/object/bio_info_life_sciences.html
[2] https://en.wikipedia.org/wiki/Dennard_scaling#Recent_breakdo...
[3] http://www.anandtech.com/show/9447/intel-10nm-and-kaby-lake
[4] http://www.hotchips.org/wp-content/uploads/hc_archives/hc25/...
Moore's law is irrelevant here. It's about the total cost of doing science, and moving sequence analysis to GPUs hasn't really decreased the cost significantly. Note the first paper you linked to is from 2007- neither CPU nor GPU from that time period is relevant today. All the links I see for "GPU HMMER" point to a few marketing pages on the Nvidia site.
Whether BLAST is 10X faster on GPUs (its not) is irrelevant. It's that there aren't interesting problems to be solved by speeding up these kinds of calcuations in terms of single problem latency- what matters is throughput- aligning billions of reads in a short time- and those problems tend to be disk IO bound, not CPU bound.
I have no problem using GPUs- those are relatively easy to program now and we've raised a generation of grad students who can write codes to those platforms. They've proved their way.
It's ASIC and FPGAs which aren't competitive in this area.
"We show that our implementation achieves 5.5x and 5.25x better performance per watt ratios than GPU and CPU implementations, respectively."
This is from Altera and they probably skewed to FPGA side. But still, FPGA can be very competitive if you consider power budget.
Xeon Phi costs $3900: http://www.amazon.com/Intel-Xeon-Phi-7120P-Coprocessor/dp/B0... It consumes up to 300W.
The average kWh in USA is $0.12: http://www.npr.org/sections/money/2011/10/27/141766341/the-p...
So Xeon Phi consumes up to $36 of electricity per hour, $864 per day and electricity start to dominate as quick as in 5 days.
Typical FPGA consume about 10 times as less power as Xeon Phi.
> It's ASIC and FPGAs which aren't competitive in this area.
Out of curiosity, I'm wondering, what about the solutions directly attacking this problem -- i.e., ease-of-programmability & time-to-market?
For instance, I'm thinking of the Altera Software Development Kit (SDK) for OpenCL (AOCL) here -- I don't suppose this would be necessarily worse than "easy to program GPGUs", especially when targeting embarrassingly parallel problems (so, any overheads due to OpenCL model <-> FPGAs impedance mismatch, present due to OpenCL admittedly being originally designed for a very different hardware, could be in fact minimized here)?
In particular, the OpenCL examples don't look particularly complex (speaking as someone with GPGPU background): https://www.altera.com/support/support-resources/design-exam...
In addition, the capabilities present that allow to optimize-around loop-carried dependencies (a _huge_ problem for GPGPUs) like the pipeline parallelism made use of in the HPC examples (like the stateful PRNG; and which makes sense due to the specific nature of FPGA hardware -- more on that in a moment), seem to make this a more attractive platform for a significant set of number-crunching workloads.
This may very well be the right-tool-for-the-right-job decision. There are some very different trade-offs present regarding the kinds of parallelism natural to GPUs vs. FPGAs (admittedly it would be more precise to say "SPMD" instead of "SIMD" in the following, but I don't think it takes away from the key point): "The key difference between kernel execution on GPUs versus FPGAs is how parallelism is handled. GPUs are “single-instruction, multiple-data” (SIMD) devices – groups of processing elements perform the same operation on their own individual work-items. On the other hand, FPGAs exploit pipeline parallelism – different stages of the instructions are applied to different work-items concurrently."
Source: https://www.altera.com/en_US/pdfs/literature/wp/wp-201406-ac...
I don't believe that either kind is universally/strictly "better" than another, so it's all about the use cases -- at least that's how I think about it, perhaps I'm missing some other trade-offs?
Regarding the I/O-bound problems: Isn't this another reason for the attractiveness of high-performance FPGAs -- like, say, Stratix, compared to GPUs? What I'm thinking of is that you can have plenty (relative to GPUs) of very high performance (here, relative to both GPUs -- as well as high-end CPUs) SRAM caches, e.g., QDRII+ SRAM: http://www.cypress.com/products/sync-sram
For instance, one example would be the QDRII+ SRAM options in the block diagram here: http://www.alteraboards.com/product/s5-pcie-hq/
Myself, I'm still unconvinced about the best choice w.r.t. the consistent performance/price ratio maximization (both the device as well as the programmers costs) -- both high-end FPGAs as well as high-end GPUs seem rather on the expensive side, either way (well, and very high-end CPUs too, for that matter).
Completely independently of the above: I'm wondering, what do you think are the reasons for Intel investing in the partnership with Altera and developing its Xeon+FPGA hybrid hardware? I presume there must be something to it, it's a potentially large amount of resources to dedicate for a hardware project.
There's little to gain from making this far faster, either, as the computation costs are still approximately rounding error compared to the cost of the data generation.
It's fun to think about, but there's little to win from this all at the moment. Once sequencing costs drop another 50-100x, it will become more practical.