Baidu Takes FPGA Approach to Accelerating SQL at Scale
nextplatform.com
nextplatform.com
You do have to pass your data through the accelerator to get the processing... which potentially means huge volumes of data moving into this physical processing layer (probably can be done in parallel over a network at high speed) - I would assume this is why shared memory bandwidth was a problem.
This would provide some really interesting options though - imagine feeding data from two disparate databases (say, Oracle and SQL server) in a data flow into this thing - now you have accelerated cross database joins (as long as you can handle the bandwidth and processing on the way in).
There was this post before on HN previously lamenting the state of tools for working with FPGAs, and my related comment wondering if what Baidu has done here was possible:
You get to cheat at that, because the naive implementation loads too much data, while most of it is going to drop off at the filter level.
The ideal TPC-DS schema stores data partitioned by date and maintains a lookup index by item, store and clustered by store's state.
In a database organized like that, neither part of the Query-3
item.i_manufact_id = 436 or dt.d_moy=12
should actually look at every row in the data-set in brute-force fashion - you get to skip massive parts of the data by just inspecting the index or going to the specific date partition and do it much faster than a naive C++ full-scan.
However, if you were limited to a full-load + scan, this would really be killer to have a columnar format which can be handled by an FPGA. The really interesting quote in the blog (for me) is
> the data for the queries is pushed to the accelerator card in columnar format (which is blazingly fast for queries)
I sure could use FPGA filters for columnar data ... even after I've done all my indexing, since I can't index down AND clauses or OR clauses - unlike a traditional CPU, the FPGA will be able to evaluate my entire condition in one sequential flow instead of check + branching.
http://dl.acm.org/citation.cfm?id=1687730
This one appears to be a more recent survey paper by some of those authors:
http://www.morganclaypool.com/doi/abs/10.2200/S00514ED1V01Y2...
As someone else pointed out, this was also one of Netezza's differentiation.
Hard-wiring operators is a minor improvement but really not much because compared to memory access it's not the bottleneck.
Although I focus more on verification these days, I've done ASIC design for about 16 years.
I'd be very interested in working with anyone on figuring out if we can make something better by leveraging hardware.
But I figured I would throw my offer out in case there were people thinking they could solve or enhance something with an FPGA, but didn't have the experience to get started.
I don't see it listed out in the article anywhere. I would guess it's a FOSS database (i.e. either Postgres or MySQL) but also wouldn't be surprised if at their scale they've created something entirely in house.
Taobao http://mysql.taobao.org/index.php?title=首页 360buy http://www.infoq.com/cn/articles/exploration-of-distributed-...
Tencent http://tencentdba.com
https://pgconf.ru/en/2016/89691
> Alibaba has provided a relational database service (RDS) for postgres in our public cloud platform (aliyun.com, the currently biggest public cloud in China). We are also enabling internal applications to use postgres in our other internet business and we can share our experience
http://www.pcworld.com/article/2908692/us-blocks-intel-from-...
and around an year ago some Russian guys got busted for selling the FPGAs to Russia (and those FPGAs were less powerful):
http://kron4.com/2015/03/21/sf-business-owner-arrested-for-a...
> In addition to the CAPI 2.0 coherent links running atop PCI-Express 4.0, there is a further enhanced CAPI protocol that runs atop the 25 Gb/sec Bluelink ports that is much more streamlined and we think is akin to something like NVM-Express for flash running over PCI-Express in that it eliminates a lot of protocol overhead from the PCI-Express bus
http://www.nextplatform.com/2016/08/24/big-blue-aims-sky-pow...
I would also be interested to see if there is a new trend to go back to owning commodity hardware, putting CoreOS+docker on it + a database with a hopefully inexpensive SQL chip. Moving away from VMware, massively expensive servers, massively expensive SANs, proprietary backups..
Would be very interesting to see.
Disclaimer: I work at MapD, a GPU database company. (http://www.mapd.com)
What tools they use to create relational algebra?
I'm not sure what you're trying to refer to here. Are you confusing FPGAs with flash memory?
EDIT: I should note that in the case of using SRAM, the device would only be programmed as long as the device is powered on. You would use EEPROM even on large FPGAs if you have the same design to be flashed to the FPGA between reboots. I personally never use EEPROM for storing the design as I have my FPGAs constantly connected to a computer for programming. My main point is that the FPGA can be programmed a functionally infinite number of times, though if you want the design to remain through reboots through an external memory, the limitation is that external memory, not the FPGA itself.
I'd love to get a real ASIC simulation/emulation platform like Synopsys's Zebu or HAPS... Both use multiple Xilinx FPGAs linked together and easily integrate with the Synopsys EDA flow. Cadence also has the Palladium emulation platform, but those boxes approach the cost of a house in the Bay Area.
My favorite!
"Both use multiple Xilinx FPGAs linked together and easily integrate with the Synopsys EDA flow. Cadence also has the Palladium emulation platform, but those boxes approach the cost of a house in the Bay Area."
Sounds like a business opportunity for someone. I know it could be cheaper and still profitable.
This is only true if the configuration is rewritten to flash every time the device is reconfigured, which would not be typical. Any design which depended on reconfiguring an FPGA frequently would reconfigure the device directly using JTAG or active serial programming, or would switch between multiple configurations in flash.
I personally have several FPGA development boards which I've reconfigured hundreds of times. There is effectively no limit to the number of times a device can be reconfigured; the configuration is stored in SRAM, which has no write limit.
Once a logic structure is in place, it's just a matter of passing signals/data through the FPGA.
That's not an FPGA then. That's a PLA. Even then most of those can be reset using ultra-violet light, a simple process.
No, it's still called "FPGA". The "field-programmable" part refers to them being programmable after they have been manufactured, by someone else than the chip manufacturer. It doesn't matter if you can't reprogram it after it's been soldered onto a board, that's not what "field-programmable" means.
> Even then most of those can be reset using ultra-violet light, a simple process.
Once-programmable FPGAs are typically anti-fuse based. There is no way to reset their programming (we're working with space-grade single-programmable FPGAs, and yes, we're extremely careful about what gets put on them).
HFT firms have been banging on the drum of making FPGA's infinity programmable for a decade and lots of commercial FPGA's can be re-flashed now without worrying about the number of times you do it.
I've seen re-flash counts up to 100,000 times.
Hey put this in the "what has HFT actually done for the world" list of answers:)
> HFT firms have been banging on the drum of making FPGA's infinity programmable for a decade
That's strange, because FPGAs have been infinitely reprogrammable for longer than that.
> lots of commercial FPGA's can be re-flashed now
That's even stranger, because the vast majority of commercial FPGAs are not flash-based, but SRAM-based and can be reprogrammed as often as the registers in your CPU. Besides, HFT firms would be using SRAM-based FPGAs because they are faster and larger than flash-based ones, the advantages of flash-based ones are useless to them.
> "what has HFT actually done for the world"
Well, not this, that's for sure.
The Baidu approach seems very similar.