I did a quick googling [1] and
> Our algorithm divides the problem into independent `quadrants' ...
> Our results show that our GPU implementation is up to 8x faster when operating on a large number of sequences.
It's still soul crushing. Why did our genome have to be that long :(
BTW, do you have numbers for setups with hundreds of GPUs?
I'm also left wondering about results using stochastic solutions. On how accuracy and problem size relate.
[1] http://ieeexplore.ieee.org/xpl/articleDetails.jsp?reload=tru...