Xilinx-Samsung SmartSSD Computational Storage Drive Launched
servethehome.com
servethehome.com
At Blekko they had taken this concept to the next logical step and built a storage array out of triply replicated blocks (called 'buckets') that were distributed by their hashid. You could then write templated perl code that operated in parallel over hundreds (or thousands) of buckets and giving a composite result. It always surprised me that IBM didn't care about that system when they acquired Blekko, it was pretty cool. If you implemented it in these Samsung drives it would make for a killer data science appliance. That design almost writes itself.
Also in the storage space, there was the CMU "Active disk" architecture[2] which was supposed to replace RAID. There was a startup spin-off from this work but I cannot recall its name anymore, sigh.
These days it would useful to design a simulator for systems like this and derive a calculus for analyzing their performance with respect to other architectures. Probably a masters thesis and maybe a PhD or two in that work.
[1] Yes MD5 hash collisions are a thing but not for identical length documents (aka an 8K block), and yes NetApp got a patent issued for it.
[2] https://www.pdl.cmu.edu/PDL-FTP/Active/ActiveDisksBerkeley98...
Also, to be fair, IBM wasn't able to do much with any of the companies they acquired in the same timeframe as Blekko. I was working at IBM at that time and witnessed this first hand.
At one time IBM research had a pretty awesome storage group, perhaps they will will have a computational storage fabric offering at some point.
That's batch processing: mapreduce and flume and whatnot. Search is still very much an exercise in getting the queries out to where the index shards live.
http://www.pdsw.org/pdsw-discs17/slides/PDSW-DISCS-Google-Ke...
Asking out of curiosity - isn't this similar to what Venti (from Plan 9) did? Of course, Venti was content-addressed, and in this case I'm guessing this system sat above WAFL (which is definitely not content-addressed).
For a write mostly fabric attached archival device it had some benefits over the SATA based filers (higher density, lower watts/terabyte, less CPU load on the filer head (spread out to the storage retrieval unit which could have many of the hash->block tranalators) etc. I don't believe NetApp ever built a complete one though. Just "too many degrees off their bow" as an engineer I knew would say.
The optimization opportunities are pretty obvious to me. Imagine if SQLite journaling was aware of how long the supercapacitor in the SSD would last, potentially even with real-time monitoring of device variables. You could have your entire WAL sitting in DRAM on the drive as long as it has enough stored energy to flush to NAND upon external power loss.
We have abstraction boundaries for a reason, we give up a small % of performance and in return we can write code once (say SQLite) and use it in many scenarios. For something like SQLite it means it's been around for a long time and had a lot of optimization work done on it (and that probably outweighs the few % gain you might get from such tight integration).
You'd probably get a bigger performance gain from just not using SQL (eg. a DBM).
The bigger challenge I see for implementing something like SQLite on a SmartSSD is that you really don't want your database to exist on just one drive, so you need to figure out how to do HA across multiple SSDs while still offloading most of the computationally expensive database operations to the FPGAs instead of leaving it on the CPU. I think this will condemn SmartSSDs to always working at a slightly lower abstraction layer than what the application really wants.
Something like SQLite might make a decent alternative API to the flash storage layer of the SSD, though. Imagine if the storage controller of your SSD exposed a built "filesystem" that featured robust indexes, transactions, sorting, column families, etc. You could skip talking to the Linux block layer or any POSIX filesystem at all, and your optimized userspace software could directly talk to the storage controller in the SSD instead with a high level software API. This isn't far-fetched; Samsung also has a "Key-Value SSD" on the way that exposes the underlying flash storage using a (surprise!) high-level get/set KV API, for similar reasons.
A design where the controller is this powerful would also allow features like predicate pushdown in the query planner to be implemented. i.e., a `WHERE x > 7` can get pushed into the storage controller, and bad tuples that don't fit the predicate can get excluded/filtered out before getting pushed onto the memory bus. That will save significant processing time and memory bus traffic in aggregate, and it scales with the number of drives (such each drive has its own controller.) Not to mention tricks like inline hardware for sorting, compression, etc.
Outside of fancy SQLite-as-a-filesystem tricks, I suspect the allure of optimizations like predicate pushdown and inline sorting will be very attractive for OLAP systems. Time will tell if these things will stick around, but Xilinx at least seems sure as hell determined to make their way into the datacenter.
(I can't find the video, but his slides are here if you want to go searching - http://sqlite.org/talks/index.html )
In fact, even if it wasn't SQLite but something non standard, I'd be interested in learning it and trying to make it work with my needs.
It is my understanding that FPGA vendors have fought the open source community every step of the way. I would hate to see the future of computing locked up in a new spiffy prison.
I’ve gotten tired of dealing with quirks using SV interfaces in RTL. I’m using structs as a substitute at the moment.
That could get you to a proper prison?
These seem like broad use cases to me. Also consider ETL and database applications. Time series, finance, machine learning, search engines. It seems like the primary benefit is in terms of latency and minimizing data bussed to the main CPU.
Or am I missing sometjing obvious?
Maybe the use-case here is more like transforming the data on the fly. Let's say storing the data compressed, but reading it back uncompressed. This could effectively be transparent to the host CPU, but handled by the FPGA.
The more that I think about it, this data flow sounds significantly more reasonable than asynchronously processing data. Then you could still read / transform / write the new data to the SSD, but you'd limit the main CPU to only sending the read/write IO, instead of the transparent transformation.
SSD already have all provisions for both, and do it. Something like that will genuinely benefit more from a highly optimised ASIC than anything else.
The use case is obviously huge, and you don't see the elephant in the room: money.
Putting all those drives to even a cheapest Xeon around, increases the price n-fold over the price of the flash, unless you talk about multi-terabyte scale SSDs.
The types of computing that can be done in an FPGA require a new style of programming, because in an FPGA, programs get mapped into hardware, such that data flows through the chip, with all of the code running at the same time, instead of in sequence.
There are large differences in the style of programming required to get efficient use out of an FPGA... there are some ways of translating C that work, but they are crude crutches at best.
This is like going back to the early days of computing, when the new machines were expensive, but very fast... and programmer time was relatively cheap. It's going to take a special new breed of programmer to make this really work well. We're at the beginning of a new era.
I hope that makes sense.
Currently one big bottleneck for data processing is moving the data from the drive, over a storage link, into memory, and then back to the drive.
By putting a programmable processor on the drive, you can eliminate that overhead by putting some of your processing algorithms on the storage.
So, for example, if you were running a Hadoop cluster, you could have your common Hadoop algorithms baked into the processor of the drive. Every drive now becomes a Hadoop processing accelerator. Rather than pulling a full dataset into main memory over that very slow link, you run some of the job on each drive, and return only the data you need. Every drive has its own processing power, so the more data you have, the more grunt you have, and you're eliminating the slowest steps.
Because the FPGA is reprogrammable, you can change which algorithms you push down to the drive as you change your workloads over time. Every disk you add becomes a specialized big data, ML, whatever, processing unit.
Other companies in the computational storage area are putting compute resources onto their SSD controller ASICs so that the compute doesn't have a PCIe bottleneck between it and the NAND. But you won't see that kind of design coming out of a Samsung/Xilinx collaboration.
Unrelated: when will Nvidia allow to seamlessly offload Java or another GC based language to the GPU? https://developer.nvidia.com/blog/grcuda-a-polyglot-language... GrCuda seems promising but it would only allow interoperability with Java on the CPU, not offload Java to the GPU, right? Such advances would make gpu computing order of magnitudes more developper friendly and therefore much more mainstream.
This is not a general purpose programming environment, it is more of an data flow / filtering system.
There is no allocation of memory, thus no need for garbage collection.
If AMD can get their act together by seamlessly integrating their GPU, FPGA and CPU with minimum I/O bottleneck it will be a huge boon for the computing industry. People will start doing something that probably unthinkable now but will be very obvious in the near future. Personally I have one application in need of the proper integration of their disparate systems and already started talking to their R&D engineers but see little improvement being implemented.
FYI, AMD has come up with SSD storage and GPU integration before but with limited success [1], but if they can also integrate FPGA together that can probably be a recipe for a great success.
I think AMD and Intel (since they both own GPU/CPU/FPGA technology) really need to come up with open and intuitive design tools for these new systems or just sponsor the work on MLIR and LLHD by LLVM and ETH Zurich, respectively.
[1] https://www.extremetech.com/extreme/232416-amd-announces-new...
Offload-from-GC runtime tools do exist, e.x. Cudafy would translate .Net code into Cuda kernels and handle kernel dispatch. Of course, you were very limited in what constructs and types you could put in kernel functions, but you could write your whole application in C# and accelerate the important blocks.
In practice, a lot of beginner GPU computing has moved to the world of NN training and inference, in which the complexities of GPU offload are entirely wrapped by the libraries you use.
For traditional GPU-accelerated tasks, the limited languages available are not the problem. Decomposing your problem into a form that is amenable to GPU offload can be difficult, and if you're experienced enough to do that well, writing Cuda kernels and dispatch in C++ is not an obstacle. For example, Cudafy meant you didn't need to know Cuda-specific syntax and expressions, but you still had to understand the behavior and limitations of GPUs to write performant code.
Lots of garbage collected languages seamlessly target GPUs. This is typically done at the library level, either within the language ecosystem (eg in Python using Cupy instead of Numpy: https://towardsdatascience.com/heres-how-to-use-cupy-to-make...) or below it (using cuBLAS as your BLAS implementation: https://developer.nvidia.com/cublas)
Java can do this too - something like ArrayFire is reasonably popular: https://developer.nvidia.com/arrayfire