All the algorithms requiring training can be optimized using stochastic gradient descent-- which is very effective for large data sets (see http://leon.bottou.org/research/stochastic)
Also, here are some additions for the online learning column:
* Online SVM: http://www.springerlink.com/index/Y8666K76P6R5L467.pdf
* Online gaussian mixture estimation: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.87....
One more thing: why no random forests? Or decision tree ensembles of any sort?