Large-Scale Matrix Factorization on TPUs
ai.googleblog.com
ai.googleblog.com
"Join our vendor-locked implementation" is all I hear. "Trust us by investing into writing software for hardware you can never own".
Maybe I just need to look at them differently, or maybe my perception of their capabilities is just wrong. I do like GPUs, for example. While they are ornery and limited in many ways, they are still flexible enough that you can do enjoyable general-purpose programming on them.
In the last few decades we got really lucky that general purpose CPUs kept getting faster and faster. So many of us could get away with ignoring application specific computer hardware.
Btw, as far as I can tell, TPUs are really just a variant of GPUs with an emphasis on lower precision?
Have a look also at the kind of special purpose hardware people make to mine crypto-currencies. I think the challenge for your teenage hacker is not so much to make them do interesting things, but to design them? (Assuming the tools and simulators become cheap enough?) FPGAs seem like they are cheap enough for a teenager?
GPUs and TPUs have substantially different architectures, it’s not accurate to consider a TPU an a variant of a GPU. See https://dl.acm.org/doi/pdf/10.1145/3360307, particularly section “Contrasting GPU and TPU Architectures”
looking only at the matrix multiply units (and ignoring the hardware multithreading in GPUs that was deliberately not part of the TPU architecture):
TPUv3: 2 cores, 2 128x128 matrix multiply units per core
V100: 80 cores, 8 4x4 matrix multiply units per core
It's almost like a universal law that generic stuff is helpful only to some extent, after which you need specialization. That's why we do PhDs I suppose. And that's why you have cardiologists.
I too am after one of those for my home-security camera system but I've never really looked at what they _are_.
There's only so much of the Internet you can route through a small number of points, before it starts to fail.
No, and what started tipping me off is that those methods are not comparable because CG and Cholesky are only for positive semi-definite matrices, so while it's not mentioned anywhere else, the fact that this method works with these factorization techniques means the matrices that they are looking at fall into a very small subset of special matrices. Even then, the subset is even smaller because CG works sufficiently well here without preconditioning, so the method is for not too ill-conditioned positive semi-definite matrices? And for such papers in matrix factorization, there's always some kind of discussion about numerical stability: what pivoting strategies are used, how the pivoting decreases the numerical error growth rate, with empirical studies based on the conditioning of the matrix. Because none of this is done, and in fact the exact opposite is done by emphasizing the BFloat16 precision, gives me no indication that you can have eigenvalues 1e-15 apart and have this work.
Using low precision arithmetic doesn't necessarily mean you cannot get sufficiently good results, there's a whole body of work in mixed precision algorithms. I'd point to Nick Higham's fantastic work on mixed precision matrix logarithms as a very nice example and read: https://www.maths.manchester.ac.uk/~higham/talks/essam18.pdf https://epubs.siam.org/doi/10.1137/17M1129866. But this paper discusses nothing about precision or error, shrugs it off, and doesn't even empirically measure the error in the 365mil case?
"Large-Scale Factorization of Well-Conditioned Positive Semi-Definite Matrices with TPUs" would be a much clearer title then. All of this doesn't mean it's not useful, it just means it's useful for a very specific class of problems and presenting it as a viable alternative to LU or QR factorization is pretty... out there.
They aren’t presenting it as that. I think you’re just reading it with the wrong lens. “Matrix factorization” in machine learning isn’t the same problem as “matrix factorization” in numerical linear algebra.
In the linear algebra world, you’re trying to essentially keep the matrix, but get it in a different form. In the ML world, you’re trying to learn to predict something.
In general, your loss function isn’t even over all the cells of the matrix (only minimize prediction error over entries where user i actually watched movie j). So the closest numerical linear algebra problem would be factoring a matrix where 99% of the entries are NaN.
Neural nets and DL are just one of the reasons to use TPUs; in the future, they might not even be the preeminent approach. DL is just one part of the much larger world of HPC, it's been very effective, but I expect that as things evolve people will rediscover the value of deep precision.
I disagree with Hooker's hardware lottery thesis for the simple reason that if it was true, she would be able to point to examples of DL-competitive methods which use the same amount of FLOPS to get much superior results (because everyone is so badly neglecting them due to the 'lottery'), or at least show better scaling curves so that they will at some point surpass DL, but at much worse wallclock due to whatever hardware specializations favor DL and penalize those alternatives. This is how DL operated: they ran on CPUs, very slowly compared to GPUs, but they did run, enabling eventual exploitation of GPUs. And that is how they could start a virtuous circle: the success of DL, because it's the right thing, starting from hardware not even remotely designed for DL (like 2010-era GPUs were not) and designed to favor other tasks (classic GPGU stuff) has pulled hardware in its wake. So, what are the DL-like things running on CPU or contemporary GPUs? The only actual example Hooker gives is capsule networks - which I thought were doomed when they were unveiled, a poor attempt at stuff soft attention did better already, and have not impressed anyone in the 5 years since, particularly as larger NNs (whether CNN, Transformer, or MLP) continue to deliver what capsnets promised. Nor have any more compelling examples arisen in the years since. With no examples or comparisons, it boils down to nothing much.
I appreciate your passion but it's clear that outside the world of DL there is a lot of HPC that wants TPUs with higher precision. I don't particularly agree with Sara either, as I have always just moved to whatever resource was most available (IE, finding a better lottery).
https://keg.cs.tsinghua.edu.cn/jietang/publications/WSDM18-Q...
https://www.microsoft.com/en-us/research/publication/netsmf-...
As much as I'm interested in the nuts and bolts of their solutions, I'd also be very interested in the problems that become solvable with quick factorization of matrices with billions of rows.