How GPU came to be used for general computation
igoro.com
igoro.com
GPU is also good for anything that requires extensive parallel number crunching, like password cracking.
If you search for "GPU String Matching" you should find some good results.
These guys are using GPUs to convert series of images into 3D models in biology http://vidiowiki.com/watch/m5xbpad/
These guys use GPUs in sound synthesis: http://vidiowiki.com/watch/w9h9mfe/
What can you do with machine learning? Basically everything: "Applications for machine learning include machine perception, computer vision, natural language processing, syntactic pattern recognition, search engines, medical diagnosis, bioinformatics, brain-machine interfaces and cheminformatics, detecting credit card fraud, stock market analysis, classifying DNA sequences, speech and handwriting recognition, object recognition in computer vision, game playing, software engineering, adaptive websites and robot locomotion." (http://en.wikipedia.org/wiki/Machine_learning)
In our lab, we have developed a package called Theano (http://www.deeplearning.net/software/theano/), which allows you to take Python numpy code, adapt it slightly, and then automatically compile the mathematical functions in a optimized function graph which is transformed into C code and compiled to target the CPU or GPU. Which means to say, your matrix mathematics and machine learning algorithms just got a lot faster, with little cost in programmer time.
So add toy programs to that list :)
I suppose this is my first open source project. hurray for me (:
For instance, Nvidia has introduced double-precision support and L1 cache, which has marginal value in traditional graphics. This is going to hurt their profitability on the Fermi chip compared to the simpler ATI alternatives.
I am gonna enjoy watching how all this plays out.
Disclaimer: I am not very experienced in GPGPU field, so my worrying may be proved wrong.
but what you are perhaps missing is that it's ok for gpus to read memory, as long as you have enough threads. they can switch context very quickly, so one set of threads can request memory (hopefully a contiguous chunk) and then drop into the background and let another set of threads do some work (on the same processing unit). this is critical to their efficiency and is very different to a cpu, which instead relies on cache and "sits doing nothing" if it needs to read data from "afar" (obviously there are trade-offs - there's only so much local memory for state, for example).
i worked on a problem that was not as "nice" as you might hope - the memory access was unpredictable to some degree. but i still got a speed up of "tens" on a cheap ($200) graphics card, compared to a meaty xeon. it's more robust than you might expect.
fermi's cache, on the other hand, probably implies more trade-offs. but you can look at it as a necessary step in learning how to find a middle ground between gpus and cpus - which is the next big battle.