Also, what are the limits like for input data sizes? I've done a little OpenCL, but I've never gone past the GPU RAM size.
Also, what are the limits like for input data sizes? I've done a little OpenCL, but I've never gone past the GPU RAM size.
We could also provide what you're suggesting outside of Hadoop, no problem!
"What are the limits like for input data sizes?" --- There are no limits. You can store your data in AWS, and we will crunch it for you. That said, initially, for IO-bound and disk-bound jobs, ParallelX might not be ideal. This is a problem we are solving as we scale.
Thanks for the feedback! We appreciate it!
I guess I've only delved into GPU stuff for dense matrix math, though, which is a pretty bad fit with Hadoop. Maybe you guys can come up with some other use-cases for them.
Spark is a new computing framework out of Berkeley's AMPLab (https://amplab.cs.berkeley.edu/software/), and it might be an interesting platform to target.
It's being adopted by Twitter, Yahoo, Amazon (http://www.wired.com/wiredenterprise/2013/06/yahoo-amazon-am...), and it's now commercially backed by Databricks (http://databricks.com/), which just recieved funding from Andreessen Horowitz.