How to accelerate a program using hardware
conorpp.com
conorpp.com
The history of the bitcoin miner has details and us a real-world example of software on x86 ASIC -> FPGA -> Custom ASIC process. It's easy to find the relative performance of the bitcoin miner running on everything from Rasberry PI's to CUDA clusters [1].
Note that the article is using the very flexible DE2-115 and there's lots of interesting trade-offs made to fit a bitcoin miner in only 115,000 gates... iirc, if you have 250k gates, it can run 4x (???) faster due to optimizations during synthesis.
https://github.com/progranism/Open-Source-FPGA-Bitcoin-Miner
Current Performance: 109 MHash/s On a Terasic DE2-115 Development Board [2]
[1] https://github.com/progranism/Open-Source-FPGA-Bitcoin-Miner
[2] https://en.bitcoin.it/wiki/Non-specialized_hardware_comparis...
And why Intel is buying Altera. Stuff like this article will get easier and with even bigger speedups in the near future. Just wait. :)
Also, I wonder if we could bypass the shared main memory, and to turn the pixels on and off directly (by hacking the display driver or whatever).
The pixels exist in a shared memory, do they not?
It's somewhat moving the goal posts IMO. Guess it's ok since it's a student project.
I mean, pixels are just like a memory except that they glow. They hold their value and can be writen a new value in sync with a clock (which is normally 60 Hz for most monitors which is MUCH slower then on chip memory). There could be no performance benefit.