Computer Vision Research: The deep “depression”
linkedin.com
linkedin.com
I last dabbled in image processing research around 2011. Probably most of the papers i read during the previous 5 years were small little epsilon papers that added no real value. I did some work in other fields and noticed a similar trend there. I always attributed it to the trend of PhDs being pumped through the system in ever greater numbers and the need for researchers to publish a paper every few months.
My major professor diluted the paper and added other content consistent with the previous method. Not just adding prior art to the introduction, but changing the meat of the paper so that it didn't seem like a departure.
He assured me that this would make it easier to publish, and publishing was all that mattered. There were no bonus points for publishing a novel technique, and there would certainly be extra work having to deal with referees.
I'm very glad to be out of that environment now. I noped out of academia and happily dealing with corporate B.S.
Smart people doing remarkable things don't seem to have a place in our society anymore, neither in academia nor in the private economy. Sure there are exceptions.
Microsoft Research, Google X and of course a handful of universities that actually work as they are supposed to, but for most of my ex colleague's these weren't real options as noone ever instilled the courage in them to find their way there, or survive the competition.
It's strange that most of them are building CRUD Software, Writing Shaders for Game Engines or work in marketing, instead of pushing us towards breakthroughs in CV and with that AI.
1. Resources. Time and money. These research ventures might take years and sometimes need some funding in addition to your own salary. It is basically borderline impossible for most developers to contribute anything useful in an environment where they need to worry about paying their bills for this month.
2. Know-how. Right now, most industries have advanced so much that you need very specialized knowledge and a metric fuckton of math/stats to contribute to any particular field in any significant way. A developer with a BS in CS can most often not even understand the papers being currently published due to the high math/specialized required knowledge.
In addition to that, not every publicly funded school provide these publicly funded software packages freely for public use, hoping by keeping them private to exploit their "business value" somehow. At least in this part of Europe.
I think smart people doing remarkable things have greater visibility than they have at any point in the recent past due to the internet and platforms like YouTube, etc. It is easier than ever to share and discover knowledge now than it ever has been in the past..
The biggest problem I see is that there is limited compensation for participating in these endeavors. I think this is an issue from the standpoint that some research is very costly or requires resources that the average person cannot obtain. There is still a great number of discoveries that can be made by people working in labs they have made wherever they found space.
[1]: http://www.andreykurenkov.com/writing/a-brief-history-of-neu...
He was a terrible professor! Yes, publish or perish is real, but that's such a terrible, pessimistic attitude to have. Its like he had given up on trying to actually do research. Tell him to go become an adjunct.
My professor during my Master's adventure was insistent on publishing a lot. He didn't mind if it was little deltas. He burned the phrase "publish or perish" into my brain. But had I figured out something novel, he would have absolutely supported putting it out there. That would have been the whole point of doing this work! Deltas help you survive, but novel ideas are what you should strive for.
No, the reason for this is that all the easy stuff has already been discovered. 20 years ago it was still kind of easy for a single PhD student to make enormous progress in his field, but after a dozen of PhD generstions there is just not much left which can be discovered by a si gle student.
Just look at physics, they had to build a multi-billion dollar research facility below the ground to advance our knowledge. Or the satelites for discovering gravity waves, ... All of this can not be done in the classic PhD model where a single PhD student works on a topic.
tldr: After 150+ years of science there is almost nothing left which can be discovered by a single PhD student
With 150+ years of science lots of new kinds of questions emerged that hardly anyone went deeply into. Just begin working there and make deep contributions. To make a few things clear:
* You probably won't get any academic recognition for this or probably no research funding agency will be willing to fund your research (far too experimental).
* It is quite possible that despite you being really talented your research into this will come to nothing. That is a prize one has to pay for the possibility of doing a deep contribution.
If you are looking for ideas where to start, just look around yourself and from what you see try to create a deep general theory (the more abstract and general the better IMHO) which has lots of predictive power and either allows you to formalize a theory where you can prove theorems about (my personally prefered way, since I'm mathematician) or has strong falsifiable predictions that one can in principle do experiments on.
> I mean actually you could do physics this way, instead of studying things like balls rolling down frictionless planes, which can't happen in nature, if you took a ton of video tapes of what's happening outside my office window, let's say, you know, leaves flying and various things, and you did an extensive analysis of them, you would get some kind of prediction of what's likely to happen next, certainly way better than anybody in the physics department could do. Well that's a notion of success which is I think novel, I don't know of anything like it in the history of science. [from the linked transcript]
that is good experimental data and the role of theorists here is to look how that performance achieved and why. For example, one can reasonably suspect that there is a good reason why the kernels in a well trained image recognition deep learning net do look like receptive fields of neurons in visual cortex. I'm pretty sure that there is some kind of statistical optimality in that, something similar to like normal distribution is maximum entropy distribution for a given standard variation. The same way i'd guess Gabor of neuron receptive field is something like maximum entropy on the set of all possible edges or something like this. The point here is that the great success of deep learning generates a lot of very good data for theorists to consume. You can do only so much theory without good experimental data, and in the decades before the availability of computing power (and resulting success of deep learning) there wasn't that much of the computer vision theory advances to speak about, really.
>leaves flying and various things
Newton did that for 20 years. With great success.
It also goes without saying that the phrase "statistically optimal" is meaningless in this specific context. You can claim they are a part of minimizing the cost function, but, again, you have to be very careful about the chicken and egg problem, because humans are the ones who manually craft the cost function.
https://courses.cs.washington.edu/courses/cse528/11sp/Olshau...
and do we know why? Usually it would mean some optimality. It should be relatively simple math here (back at the time at our University it would be given to a student as a thesis project and couple months later we'd have it), and that would give us 2 things - insight into biological visual cortex (which we suppose follows some optimality too and know we would have a very good candidate for the one) as well as to allow to start some primary convolutional layers with the (optimal set of) Gabors instead of going through learning them. Actually some of the best results i saw 15-20 years ago were produced by the simulation of visual cortex through such construction. And now image trained deep learning nets converge to the same.
what are you talking about? where did i say such a thing?
...
For example, LSD-SLAM: http://vision.in.tum.de/research/vslam/lsdslam
Deep learning / ML approaches certainly have a place, and they're getting a lot of attention right now, but the computer vision domain is about a lot more than segmentation and classification.
Maybe in coming years we'll see some more breakthroughs from the ML side on encoding priors - for example, teaching a network about projective geometry is a lot worse than just structuring it in a way that it 'knows' what projective geometry is. This could result in a closer collaboration between the two fields.
Well, the article is arguing deep learning has taken most of the "mindshare", the attention of most researchers. If great results are coming from other parts of the field, that would be a reason to be concerned.
Yet there are still people out there working against the tide trying to find a 'unified' theory of what's going on, granted with limited success or support. Some argued at the time that the problem is simply too difficult to tackle with our statistical 'tricks' and computing power. It's somewhat disheartening that folks are still grumbling about the same things 10 years after I left.
The fundamental problem is that funding is given to those who promise the best outcome ("device that can recognize cancer") rather than the truth ("Where is the data located in an HBM").
Now, engineering work isn't bad, but today's university still has relics from a previous generation, like research papers. Hence, we're left with a bunch of research papers with little scientific content. The only fix I can think of is to offer useful alternatives to the PhD and prefer or mandate other markers of achievement like patents instead of research papers.
It's the difference between giving a "brute force" computer proof like the four color theorem than try and come up with new theory where it's just a result from it.
Most folks I know are desperate to do actual science, experimental or theoretical. Instead, they optimizing some procedure/protocol.
if someone wants to do science, there is no one stopping them from doing science.
if someone comes to someone with their hands out, then theres going to be strings attached.
The copyright laws for scientific articles are. I just link to Aaron Swartz' Guerilla Open Access Manifesto: https://archive.org/details/GuerillaOpenAccessManifesto https://archive.org/stream/GuerillaOpenAccessManifesto/Goamj... https://archive.org/download/GuerillaOpenAccessManifesto/Goa...
However, in practice that isn't an issue:
1) At least in this domain, all publications are de facto open access, as in, if you just google the name of a paper in a random citation in 99% cases you will get a non-paywalled full text version - if not from the actual place of publication, then on arxiv, author's home page, etc. It's not totally appropriate as there could be differences, but it's definitely enough to say "there is no one stopping them from doing science".
2) If you do actually need access to the university library databases for paywalled articles, then just go to the library. If you want to do science, there are options. Most people simply have or get some kind of university/college affiliation. If you don't, in many places you can still use the university library to access the data without the paywalls. If not, then you often can (depending on your country) "join" university to audit a single course, which would get you that affiliation and access to their infrastructure. I'm getting to more and more obscure scenarios, but even then there are options - the publications are accessible (though at some times not conveniently enough) and that is not a serious obstacle to doing science; it's still far less effort than actually reading and understanding these papers.
3) If you do need something that's really not available to you, just email the author. As a rule, people write articles because they want people to read them, use them and cite them. My advisor has a bunch of papers that he received that way in pre-internet time when that involved expensive mailing over the ocean. The only realistic case where an author wouldn't send a preprint version to you is because you're either rude or haven't taken the five seconds to click on the link in their homepage to get that paper.
Except they are learning this stuff prior to their graduate work...so I don't know where the author is coming from here. All of our Computer Vision people are very familiar with all of those topics - especially complex geometries and topology.
I would still like to hear what the author of the post recommends as a course of action- Maybe he can write a followup post that provides these details to clarify this.
But the point of the article is that there is more to computer vision. Stereo, optical flow, geometry, and physics can only be aided by deep learning so much.
Another point not mentioned is the computational power required for deep learning. Consider programming the physics for a ball rolling down an incline. You could use (1) the math itself, vs. (2) a neural net. It's clear that the direct math approach could be orders of magnitude faster than neural nets. I wouldn't be surprised if directly coding the physics would achieve 1,000x the performance of a neural net.
Edit: Although what they seem to describe is replacing GPs with neural networks in Bayesian optimization which is supposedly more efficient.
Since the point of Bayesian optimization is to limit the number of times you have to evaluate a new set of hyperparameters, I am not sure how useful it is to "be able to scale" (i.e. even if maintaining the GP is O(n^3) with the number of evaluations, the costly part should be to evaluate the hyperparameters in the first place) but I haven't read the paper so they may show some high dimensional hyperparameter cases where performing a lot of evaluations pays off.
If training CNNs does it faster/better than having a better understanding of Stereopsis/etc... then who cares?
Deep learning undeniably works well. It works so well that it has almost completely taken over the field, sucking the oxygen out of other parts of the field. This is a little disconcerting since no one really understands how much of deep learning's success is due to a) the algorithms themselves vs. b) the massive amount of data and computational power being thrown at vision problems.
There is something to be said for some "counter-cyclical" efforts to encourage people to keep exploring approaches that aren't just slightly different network topologies, or least to ensure that they're not totally forgotten. That's the message I took away from the article.
In a way, you could take this as a perfect example of the authors' point: we might have had a lot of these advances sooner if we hadn't let interest in neural networks fall completely off a cliff.
Think about it: Do we really, in the US of A or anywhere in the 1st world countries, have a lack of resources? In the 21st century? With hardly anyone actually working on the basics the population needs, like food, because our productivity is through the roof?
And "money" only is an issue when humans decide it is - it is an entirely made-up concept represented only in bits in computers (never mind that little bit of printed money, which can be produced at will too).
We do pay a huge amount of people in both private and public sector to do useless work. See dilbert.com cartoons and their popularity, and the results of research into the mood of working people, a huge amount of them thinking their work is useless to society or even just their own company.
The feeling (and reality) that we all have to work hard to get by is mostly an artificially created system outcome. With the productivity levels we have achieved, and the tools and knowledge at our disposal, we should have plenty of room for experiments.
Yes, I just say national debt. As long as this is not paid back (or we are working seriously to decrease it) we can't claim that we have an abundance of money - quite the opposite. Even more: As long as we aren't seriously working to decrease the national debt we are near to delayed filing of insolvency, which is not a good idea as any founder can tell you.
And no, debt does NOT have to be paid back. It can be rolled over indefinitely, or, if demands on "payments" really are an issue just go bankrupt. NOT A JOKE. It happened A LOT. I myself - and I'm not that old - have seen three currencies in my country in my life (Germany), my grandparents more. Germany is a horrible 3rd world place... no?
What are people going to do when all they get is "sorry, no more payments on the old debt. Are they going to emigrate to Alpha Centauri? Are they going to refuse to live from now on? No, they will grumble and get back to doing what they've been doing all the time. As long as the government is strong etc. Plenty of examples - see Germany. No that doesn't lead to Argentina - only Argentina leads to Argentina. A countries industry, culture, people, government don't suddenly stop working just because the currency is changed and old debts forgiven. Money does not exist, it's a made-up concept.
> we can't claim that we have an abundance of money - quite the opposite.
Uhmm... what should I tell someone like you? Except to repeat:
Money does not exist, it's a made-up concept.
Guess what happened after the financial crisis - we invented a trillion dollars. Out of nowhere. Suddenly it was there.
Also, the attitude that ANNs are "all we need" was there even before deep learning delivered all the recent SOTA results. I remember Patrick Winston commenting on that in one of his lectures in 2010(!).
One does science also to understand why something works (i.e. what is the theory behind it). Deep learning gives something that (perhaps) works, but doesn't give an any clue to the question why it works - on what theory does the neural net internally compute in its black box.
Based on this criterion one can even argue that deep learning papers should not be considered as science as long as the authors of the deep learning paper make no really hard attempt to present at least a partly understanding of the internal theory that the trained neuronal network executes (i.e. a very detailed interpretation of the weights with ideally falsifiable predictions). This is hard work, I know, but isn't that nearly anything in science?
The reproducability problem should be solvable if the authors would simply release the training data and source code that trains the network (as yould scientific practise would require). Or is there anything else that is necessary to assure reproducability (I'm not that deep into CV papers)?
Even on this very website... I feel there is an immense bias in favor of anything "deep" and "neural" when it comes to AI. Can't recall any recent AI papers that made it to the front page without having those two words in the title.
And please, don't tell me there is nothing interesting going on in the field outside of deep learning. Even when an approach doesn't beat SOTA in terms of error rates, it can still contain valuable ideas or have interesting properties.
It still is. Any system using structured light will have big problems when confronted with sunlight (windows, outdoors etc...).
I'm not confident there is a solution within the structured light domain to this as the beacons will (more than likely) never overcome sun intensity. We're doubling down on passive systems and reference maps.
Remember that throughout the eighties/nineties ANN classifier performance was as underwhelming as other approaches or even worse at many tasks. Now it reached a local optimum and all other approaches are being discarded.
You better separate between using lots of data (perfectly OK, in my opinion) and the number of parameters that the model uses (say, number of weights in the neural network) and where you better have good explanation for the existence of any variable and why it has this concrete weight and no other and why it is even necessary to introduce (consider any parameter that you have to introduce as some kind of physical constant - physicists invest lots of time to explain/reduce the number, so should CV researchers).