Furthermore, SVD type of algorithm are orders of magnitude faster than NN to train (which is why they become used to train word embeddings, see Glove).
Furthermore, SVD type of algorithm are orders of magnitude faster than NN to train (which is why they become used to train word embeddings, see Glove).
You're wrong about training times. You can get great results from far less than 117M params. My tiniest music model (similar to https://soundcloud.com/theshawwn/sets/ai-generated-videogame... but not the same one) was around 2M params. Trained it a couple days on my laptop's old CPU. Worked fine.
I think people love to exaggerate their own importance, and in the importance of theory. There's a certain myth of The Smart Hero, where intelligence alone is the decisive factor that saves a project.
Nope. It's determination. Bet on determination every time. And determined people can get very, very far with NNs.
It's a bit like saying it takes forever to grow tomato plants. Well, yes. Yes it does. Welcome to gardening.
ML is problem-space gardening. Yep, it's hyperparameter searches. Yep, it's architecture searches. No, it doesn't take nearly as long as you're implying, because you can quickly narrow the problem space in log(N) steps. (It would be foolish to do a hyperparam search in linear increments.)
Resources have also never been more plentiful. TRC literally gives you 100 TPUs when you sign up. https://jaxtputest.vercel.app/ That means you can (and I did) run 100 training sessions simultaneously. https://www.docdroid.net/faDq8Bu/swarm-training-v01a.pdf
While I’ve seen many data scientists fall into the trap of spending 1 week per iteration even when it’s clear the change produces negligible result. If you stick with the field long enough you’re likely to hit a dog of a problem that should have worked, but ultimately did not.
The areas where ML is advancing fastest also suffer from elusive solutions. Language models and image recognition are only recently tractable problems - the folks trying to advance these fields do not experience magic NN improvements.
1 week is fine for a production training run. But that’s after you’ve locked in your hyperparams on small prototypes. Small prototypes should take no more than two days to give results. If your prototype absolutely must be a large model, you should have many (perhaps even a dozen) attempts training simultaneously.
It’s hard to imagine someone spending an entire week waiting to see whether their latest tweak worked, then tweaking again and waiting another week…
I had models like that, but they were exploratory side projects, not my primary research focus. It’s stuff you try for fun, not for work. That way a week is “whatever” if it doesn’t work out.
But when you’re hunting for a result, no way. Prototype results are measured in “N completed runs per day”. Which means you should have multiple runs going at all times.
It gets tricky to manage, but… I don’t know. Either you’re right and I got incredibly lucky, or determined programmers can do the same thing I did. I’m not that smart.
Try training Imagenet. For some reason this is called a "medium" sized data set.
I do want to point something out though that's a big problem with a lot of your comments. We're seeing the results of what happens when you "just implement a DL pipeline." Lots of people are using these powerful tools without understanding. It's like handing a monkey a machine gun. We're seeing that models don't generalize like people think they do. We're seeing people use models pre-trained on ImageNet or CelebA and think that there is no bias in them. When people are saying you need to learn theory, this is really what they are concerned with.
As more NN methods become viable, some more savvy data scientists complain to me "this NN is just approximating SVD/PCA/POD/etc!" Wonderful, that's explicitly the point! The network we're applying to this problem compares/combines multiple approaches to dimension reduction. The network created a latent space that makes way more semantic sense than just PCA or SVD for this problem (No Free Lunch). It still takes effort and understanding, but the value I've personally gotten over just applying PCA for my problem-sets has been incredible. In fact I'm certain it has made my career. Turns out diagonalizing covariance matrices aren't the only dimension reduction game in town!
It helps to think of each term as an interesting puzzle. For example, SVD. It's fascinating if you dig into it. Most people don't want to, because it feels like work. For me, it's neat understanding ... whatever it is, ha.
I think it's finding the basis eigenvectors in a higher dimensional space, which basically just means that e.g. the eigenvectors of a cube are the X, Y, Z axes you're used to. If you skew it along the X axis, the Y axis bends a bit, along with the cube.
The eigenvectors form a shape that, when you find the volume of it, is the area of the resulting form. So the determinant of a cube's SVD is the volume of the cube.
In higher dimensional spaces, it's the same thing, except it's called "eigenvectors" (named after Sir Eigen of Eigenmadethisup) because mathematicians have reasons for using complicated language, some of which is valid. But as you see from me muddling through this, the underlying concepts are all small simple pieces that fit together.
Or I was nowhere close to the explanation of SVD. But it was close to something interesting, since it leads to the question of "What's the SVD of a sphere? How about a point cloud?" It was easy to figure out for a cube. Not so easy when it's an arbitrary shape. "And why is it useful?" Because it gives a lot of hints about what that object is. In the optimal case, in StyleGAN for example, the SVD can even be the basis vectors like "smile", "age", and so on. (You know in Faceapp how you can drag the "Age" slider and make yourself look older or younger? That's a basis vector in higher dimensional space. It's orthogonal -- more or less -- to "smile", because if you drag the "smile" basis vector around, it doesn't cause you to age older or younger. Except it's not quite orthogonal, because it's a higher dimensional weird-ass shape and therefore can't be orthogonal, so sometimes when you make someone older their hair turns grey even though "blonde hair" is orthogonal to "age" in theory.)
Yada, yada. Rinse and repeat and dive in for a couple years. You'll find it's fun once you jump in.
P.S. All the people reading this that feel offended like "No, you really must start with theory; you can't possibly learn anything if you don't know what you're doing," you better read this: http://thecodist.com/article/the_programming_steamroller_wai...
That steamroller is coming for you. Once the legions of javascript programmers realize that hey, I can do DL just like an ML researcher, you're gonna be doomed. Because a 17yo JS programmer has roughly 10x as much determination as even I can muster these days, let alone someone who clings to the idea that theory is the only path forward.
It’ll democratise but it’s not there yet.
It wouldn't be "approximating" anything. An optimal one layer linear NN "autoencoder" is PCA. There are other learning algorithms for PCA than gradient descent, but the infrastructure for learning NN:s with big data sets makes it painless.
As soon as you add activations and layers, you're improving on SVD/PCA. For dimensionality reduction, it means the "manifold" is more complicated than just a linear projection.
You're expanding the space of realizable functions, which is an improvement in a specific sense, but not in all senses! The SVD, since it is better understood theorist theoretically, is a more straightforward problem to solve robustly. There are fewer hyperparameters (like learning rate) to choose, and you aren't left wondering whether your solution is at a bad local minimum.
I think it's wrong to think that it's an obvious improvement.
I'm assuming by "ahead" you mean score the best on some arbitrary metric?
Doing this approach you will not come out "ahead" when your model runs into problems in production and you need to diagnosis and get it performing as soon as possible.
You also won't come out "ahead" in 5 years when a magical matrix of weights is the core of your predictive modeling efforts and no one on the team really understands exactly what was being done but the output of that model is crucial to another team... except it doesn't seem to be working.
You also won't come out "ahead" when an event like the pandemic strikes, and the assumed distribution of all of your data no longer holds but you still have to keep your teams performance up for the month.
I've been in the top 100 globally on kaggle so I understand pushing a metric higher, I've also worked for 10+ years doing DS work where a minuscule bump in a metric is not worth all the associated cost of a complex model. There have been at least three times where I have walked into a company that was running some mess of a neural network and replaced it with a one or two parameter model (that took seconds not days to train) and gotten practically identical performance to a NN that was causing massive headaches to maintain, train and debug.
One of the most important principles of engineering that seems frighteningly lost on many data scientists I've met is that simplicity is the goal, and we only take on complexity when absolutely necessary.
There are certainly cases were NN are the best path and you should let a million parameters do the thinking for you, but there cases are relatively few and getting solid performance here requires serious experience and skill.
You can do similarity search and all the sorts of things you do for word embeddings on embeddings generated for other scenarios.