lbfgs is quite common for eg regression w/o l1 penalties.
MPI is not great at even high hundreds of cores; it's too much work to build redundancy / retry / restart / clean failure in. You really need a framework that helps with this.
lbfgs is quite common for eg regression w/o l1 penalties.
MPI is not great at even high hundreds of cores; it's too much work to build redundancy / retry / restart / clean failure in. You really need a framework that helps with this.
I've used MPI on Titan for thousands of cores. It's essentially what MPI was invented for. I also know people who perform QMC simulations using all of the cores on the machine at once using software built upon MPI.
Not necessarily, often derivatives are analytically known in ML.
> MPI is not great at even high hundreds of cores
? You realize that Sequia, which I have run on, has codes that scale to all two million processors.
The focus here is largely on deep neural networks. In this domain, the Hessian cannot be computed and SGD (with minor variants) continues to be the golden standard.