Domain-Specific Hardware Accelerators
cacm.acm.org
cacm.acm.org
The exceptions were GPUs, where there is enough of a mass market, and maybe super-high-value-added niches, like aerospace, where high unit costs are not a deal-breaker.
The article seems to get there, in the TCO section, but it looks to me like maybe genomics is another niche, where there's enough demand to make the economics work.
I had some hope that maybe fab technology was saturating enough that fab access was getting cheaper, maybe driven by SoCs for phones and all those ARM devices and RPis and stuff, that a new era of awesome ASICs was at hand. This appears not to be the case.
This article suggests mask tapeout costs are under $1 million in older nodes, sometimes well under. If you have an architectural advantage in a problem domain with tens of millions or more in costs, a simple ASIC can be very worthwhile. That architectural advantage might be hard to find, especially when problem domains aren't fixed for long periods of time (e.g. how many ML accelerators only really work well for dense convolutions?), but I suspect too few companies are making custom chips, rather than too many.
https://www.electronicdesign.com/technologies/embedded-revol...
Intel also taught us last year that 4kb of code is all it takes to decode every x86 isa (i.e. 1977-2020) https://github.com/jart/cosmopolitan/blob/d51409c/third_part... Thanks Mark Charney. https://github.com/intelxed/xed
I've lost count of the number of hardware accelerators in the regular expression world that were way faster than some outrageous strawman (e.g. "running libpcre one by one over a thousand patterns") as opposed to doing a little research to find what the state of the art is. Not sure if that's what's happening here, but 15,000x rings some alarm bells.
https://www.illumina.com/products/by-type/informatics-produc...
The analogous thing for regex would be, say, extract a literal factor that probably won't appear in the input and suppress the regex execution if that nice even occurs - but only do it on the h/w implementation. Voila, 1000x speedup - on the "nice" input at least.
It would be interesting to see what the expected limits are for fair comparisons in extremely hardware-acceleration friendly cases. I don't think GPUs, for example, run 15,000x faster than pure CPU based graphics rendering, but am not sure.
If the financial motivation is there, companies will build custom hardware. Companies like Apple and Google (TPUs) are building custom processors for various purposes.
Oh man, this makes me go back to FPGA programming which I did a millions of years ago when i was in college.
We have tried to accelerate different domain specific tasks with FPGAs but the development cost and effort keeps being to much compared to well written and smarter software.
Sure there is a case for some very computational intensive tasks but I think the really interesting part is when you can safely high level program these things.