I just wanted to blast these other applications, because I think people get this idea that AI has to be AI for anything interesting to happen... but there are really niche applications where people don't think these tools are experimental. And what you describe may already be happening, Geoff Hinton's critique of modern deep nets seems to be a call to get more biological. (Thinking of capsules nets).
I have a neural net onboard my phone which automatically detects songs offline and tells me what they are. Is that semantically 'quantitative math' and not machine learning?
>I have a neural net onboard my phone which automatically detects songs offline and tells me what they are.
MP3 uses something called psycho acoustics, which is a quantitative model on human perception, which is used to eliminate frequencies that can't be heard based on this model.
Your neural network doesn't tell you what features make songs distinct, it's not a quantitative model at all, but a black box heuristic on what the important features are superficially. If actual mathematicians worked on this problem, I guarantee you they'd do a better job, and their models would work on a commadore64, with real time training. Moreover it would tell you things like who is singing, if it's a live performance, which concert it was.
No, this is wrong.
Some of the most brilliant people in the world have been working on image recognition, voice recognition etc. and AI is crushing all of their work.
"Your neural network doesn't tell you what features make songs distinct, it's not a quantitative model at all" - it doesn't matter at all if our objective is detecting the song. Neither does the mp3 compression algorithm.
This is very true. I take my stronger statements back, MAINSTREAM mathematicians attempting this problem are all wrong, and have been wrong for 50 years. But you do need the right theory, and the right math that realizes this theory.
"AI" is superficially beating the work in computer vision. Computer vision is complete bogus. The gabor filters, fourier transfroms etc. are all wrong conceptually. The known methods do abysmally on basic tasks like object recognition, texture segmentation etc. But they keep trying it.
I would take this one step further: computer vision, audio and NLP researchers have been stuck in a rut for the past 50 years. DL is beating THEIR math, but this is because of data and computation speed, not because of any insights. But DL is also wrong, and giving you an illusion of progress. Both of these things are doomed to go the way of GOFAI.
I can go into great detail and carefully explain why MAINSTREAM contemporary ideas in math for vision, audition and language are completely wrong, and have been wrong for 50 years. What is the right model? Like I mentioned before, the right ideas are emerging, neural networks will dominate, just not DL.
This is a No True Scotsman. Actual mathematicians did work on this problem, training the neural network to achieve it's target task of identifying songs using minimal power and storage consumption - which works.
Collecting observations aka data.
> deriving the mathematical laws that govern what you see
Fitting a model.
> Your neural network doesn't tell you what features make songs distinct
It literally learns better features that you could ever come up with by hand. This is why CNNs do better in computer vision that hand engineered filters.
> I guarantee you they'd do a better job, and their models would work on a commadore64, with real time training.
LOL if you think that a room full of people can listen to TBs of audio data, decide what mathematical functions when combined together are better descriptors of that data than a DL model learning its features.
You don't have the slightest clue what you're talking about.
It's analogous to a human being able to identify songs by remembering the chorus, just that the NN uses it's own features for both the memory and offline perception.
>In 2017 we launched Now Playing on the Pixel 2, using deep neural networks to bring low-power, always-on music recognition to mobile devices. In developing Now Playing, our goal was to create a small, efficient music recognizer which requires a very small fingerprint for each track in the database, allowing music recognition to be run entirely on-device without an internet connection.
https://ai.googleblog.com/2018/09/googles-next-generation-mu...
Why did you pick a neural network? What mathematical properties does a neural network have that makes it appealing to this problem? How were the networks trained? Back propagation? It doesn't converge, and worse learning weights for a new batch can cause you to forget previous batches. This isn't a desirable property of neural networks or back propagation. You probably had a lot of heuristics on top, fine. How do you know that the weights you ended up with will always work in practise? Given an arbitrary track, you can encode it? What about growing the database? Does the neural network get updated for new songs, or do you use the same neural network to fingerprint new songs and update the data base?
Here's how I would have done it:
A song file is just a sequence of amplitudes. I would do some kind of an interpolation of piece-wise trig function. Trig functions have very desirable properties: they are continuous everywhere, and infinitely differentiable. Moreover, a sine basis decomposition will be able to reconstruct the original signal very well. This is great, because now you can use theories from DSP and fourier analysis. So we take the entire song, do a continuous time discrete cosine transform, in a block size of 32. Now you compute the square norm of all feature vectors, sort them, eliminate the vectors that are within 1e-3 radius (they are too similar to each other, there's not point in keeping them) and only store the top 25% of feature vectors by the square norm. The 25% cut off threshold and 1e-3 radius of similarity are heuristics, and adjustable parameters.
Now you have a database. For a new song, repeat the procedure, and get a feature vector for every 32 interval. There are probably theories in DSP you can use to get a better similarity measure, but for now, we'll just use the L2 norm of the difference. Do a nearest neighbour search in your data base for all feature vectors, and rank the results based on hits. I can run all of this on a computer from 2000s which are crappier than modern phones, and have the entire backend run on equally crappy hardware too. All parts of what I'm doing are fully deterministic, updating the DB is incredibly fast, CTDCT is super fast, there are no questions of convergence, no need for training. You can probably increase the accuracy and speed by doing some DSP and doing the nearest neighbour search based on different voice, bass, instrumental etc. features.
In practise how would it compare to your neural network? No idea, but I imagine it should be very competitive. The big benefits are that you have only 3 parameters (radius of similarity, cut off threshold and block size). This seems very easy to bench mark against, it should take like a week to implement. I'm not sure about the compression of the finger print however. Not sure how much space 1000000 songs will take (probably 25% since that was our cutoff). You can probably borrow psycho acoustics to make a better data base, and get a better compressed representation. Another alternative would be to down sample the song to 64kbps before hand.
This was a paper from Shazam from 2003. This is essentially what I proposed, there is no training. Shazam works pretty well. It's not even going into the mathematical consideration I went into.
>You'll never be able to develop features with the heuristic methods you described that will work as well as the features learned by a neural net.
False.
Deep nets are here to stay. They're just not magic bullets that solve all problems equally well, especially those when training data is minimal.
What on earth are you talking about?
>Deep nets are here to stay.
Maybe in silicon valley for consumer products in things like snapchat and siri. They won't work for industrial problems.
What do you think machine learning is, if not “quantitative math”? Deep learning is just linear algebra and calculus, and things like random forests are even simpler mathematically.
AI is in a hype bubble right now surely, but it's a 'very real' thing that's going to infiltrate a lot of areas.
Being state of the art doesn't imply that these things will solve these problems. In ML terms, how do you know that NN/AI isn't a local maxima that we need to jump out of? All NLP systems are joke. Sure replace Watson with DL, might perform better on Jeopardy. But in real conversations? Forget it.
I wouldn't bet on these things. NN will win, but not the back propagation, ReLu, sigmoid or whatever pseudo science that is the current buzzword. There is 50 years worth of understanding in actual neuroscience and cognitive modelling that no one has paid attention to, and new design principles are emerging that will influence mathematics.
It's the best performing tool we have for NLP, image recognition, etc. Is it a local maxima? Probably. But it's out there solving real problems nonetheless. We'll capture all the gains we can and then move on after.
I suggest you are misinformed about the state of AI.
AI is currently ahead of all other approaches in many fields.
It's lead to quite a number of practical advances and breakthroughs.
The 'best examples' are those that I described, but there are many more.
Your comments indicate I think some ignorance on the issue - I think I see the point you are trying to make but I also submit that you're not aware of what AI is doing today.
'Self driving cars' would be impossible without AI today, for example. The vision systems depend on AI it's a breakthrough without which we simply wouldn't have the tech.
>And what were the equivalents of DL/ML for physics before calculus?
This is a good question. Before Newtonian calculus, and the laws of gravitation, people were building very complicated conic models (ie. eclipses, parabolas etc.) to get better and better prediction of planetary motion. A lot of parametric math came out of this, with many sophisticated models getting better and better, giving these astronomers an illusion of progress. However, Newton's insight was that motion is connected to mass, and this insight was the basis of how to derive the laws of motion, which gave us the laws of gravitation (F = (Gm1m2/r^2)). This insight eliminated the previous Keplarian models of motion, because you were now able to predict the motion of arbitrary rigid bodies using very simple math (we teach this in highschool). Ofcourse, Newtonian motion has its limitations that's why we have quantum physics and Einstien's relativity theory. But for practical technological applications, Newtonian physics on its own gets you incredibly far.
Where is ML/DL? It would be akin to Keplarian elliptical motion. More realistically however, it's closer to aether theory of light, and will go the way of GOFAI. This stuff isn't grounded in modelling any scientific observation. Moreover, they are mathematically useless. Back propagation doesn't converge, and why should you fit your data to an arbitrary mathematical structure? In practise, DL/ML doesn't work at all, you will be much more successful by modelling your problem mathematically. For example, consider an automobile manufacturer, which has all kinds of moving parts in their planes. They typically model each part mathematically (ie. gear x under goes exponential time decay), and imply their parameters using rigours test data. Then you use some sort of an empirical statistical model to predict the failure.
I've seen deep learning companies come and fall flat on their face trying to beat the accuracy of these deterministic systems. Those guys needed a lot of data, and GPUs. I'm not even criticizing the fact that DL is a black box. It's worse, it's inferior to everything out there on every metric imaginable. These mathematical models in contrast have been in production for decades, with yearly updates, and they run in real time with little historical data, they are fully understandable and they beat every method we know of.
This isn't the first time multi layer perceptrons gained hype. They didn't work in the 80s, or 90s or the 2000s, they don't work now. The math behind DL is the same that we had in the 80s, they just called it multi layer perceptron. None of the ideas in modern ML/DL are new, all these ideas like reinforcement learning, GANs etc.
2. Likewise Newtonian physics is also an approximation: it does not fare well near relativistic speeds or high gravity. But at least we have models which seem to be accurate to many decimal places today. Who knows what the future may hold.
3. Not all useful problems can be represented by simple equations, but they can be computed analytically (e.g. N-body problem).
4. Ultimately DL is popular because it works better than anything else in some very specific domains like speech recognition and image recognition. It is overapplied I'll admit, but if you can do better then feel free to publish a paper.