Engineering is the bottleneck in deep learning research
blog.dennybritz.com
blog.dennybritz.com
Also I learned to communicate with other teams and exchange with ideas proved by practices - this really helps.
To improve, I would suggest to publish whole setup, all the parameters used, the programming code, and publish either all the data or a reference to a large free data set (no MNIST anymore in papers, please).
If you have a breakthrough in transfer learning then you will be able to very effectively demonstrate it with MNIST.
The race to the bottom is essentially over, but that doesn't mean MNIST can't be used to demonstrate learning.
Regarding setup and parameters. I hope AI researchers move toward something like pachyderm [https://pachyderm.io] -- providing a single docker image to completely replicate their work. However, I sincerely doubt that will happen. As "open" as research is the details are almost always obfuscated to prevent competition with the spin-out company (or other researchers).
But sadly, IMO at the amateur-level, TensorFlow considered harmful. I have repeatedly observed novices blow the thing up by starting from one of many of its amazing and fantastic teaching examples. It's not a question of the TensowFlow API, but rather of the engineering quality of its underlying engine, which kind of sucks. Nothing ruins an enthusiastic data scientist's day like a cryptic seg fault for no apparent reason whatsoever.
And I know they're working on it, but fer cryin' out loud, the API is great, and Google has the bottomless pockets to do a lot better than this. It's been over a year and I still see people throw their hands up in frustration trying to make use of the thing. Of course, Google has never been a customer-driven company, but if we don't want an AI Fall, methinks this needs to be fixed.
As far as the TensorFlow API is concerned. This may be a tradeoff between speed and robustness. In order to have every operation checked every time would certainly slow down the code for general use. Probably better how-to / setup / use guides are a better solution for this (unless it is a flat out bug).
The paper made it seem like they had been using a standard PCFG parser (which circulated in the research community at the time) to achieve their results. It turned out they hadn't and instead had written a custom one and in fact their results were not reproducible using the standard parser.
What was meant to be a timesaver in terms of engineering (using a standard parser instead of writing your own) turned out to be a massive time sink. It also turned out that by using a custom parser they had unintentionally diverted from a vanilla PCFG (probabilistic context free grammar), or in other words, some implementation details had led to a departure from the assumed underlying theoretical model.
I've been to 100+ colloquia in the physics dept at Cornell and I have never been at one that I felt was a waste of time or that the person should not belong there.
The CS department colloquium is a different story: yes I got to see Geoff Hinton before he became a celebrity but maybe half of the talks are awful.
I have sat through too many presentations obsessing on HW-level perf/W especially w/r to Deep Learning ASIC wannabes. Just writing one's code in C/C++ (and doing it well) guarantees at least a 2x improvement over Java and a 10-100x improvement over Python. I won't even bring up the computational coup that is CUDA.
But hey, let's base a mobile phone OS on Java and block low-level access to its GPU, that's a fantastic idea, right?
See also many experiences with data scientist and CS primadonnas dismissing low-level coding as "ops." I liken this to the Eloi dismissing the Morlocks as "the help."
This wasn't a noticeable issue until about a decade ago. But it is now and it continues to get worse IMO. The "programming language" bit is just one of its symptoms when the root cause is ignorance of practical computer architecture.
That said, the mentality of throwing all big data problems at Hadoop clusters with 4 year-old GPUs and flaky 10 gB interconnect (many of which could be solved faster on one.big.modern.and.cheaper.machine(tm)) is working wonders for my Amazon stock so maybe I should just shut up and get rich?
I had to select features from multiple papers in order to try and select the best ones with classification results to prove it.
A few problems included:
- Incomplete/unavailable datasets (404 on some copyright pictures)
- Features consisted on Math formulas and text descriptions (no code whatsoever)
- Classifier names only (which framework did you use? parameter values?)
In the end i couldn't contribute as well, got instructions to save my work in a private repo despite being funded by an EU academical scholarship.
As a researcher, I expect 50-90% of my time to be slogging through organizational and preparatory work.
Is it? Maybe the paperwork is important for other members of the care team and the surgeon is the only one who is familiar enough with the surgery to fill out the forms. And, you can't reasonably be doing surgery round the clock.
We've had a few machine learning experts working here for a couple of years, but recently brought in a software engineer with a passion for machine learning. He was able to, within a few months, streamline the data acquisition pipeline to the point where we could iterate on a new models in about 30 minutes, down from days. He accomplished this not just with better data but by building efficient in-memory data structures. It saves literally days of time per iteration because of disk I/O.
Before his work the training data versus the data we used in production had minor differences. Each new release required intensive manual verification to make sure that our model worked. Now we have much more certainty that the two match up.
Looking down on engineering problems is like a famous architect looking down on structural engineers. You're not gonna have a very good skyscraper if your foundation is shaky and ad-hoc.
So that you could link your code to a dataset, have it automatically run, and show the result...
Not sure if it is worth the time...
A general dataset pool in OpenAI would be nice. Kaggle has quite a few just basic datasets (MNIST etc) for evaluation.
i thought i came up with a good set of hyper parameters using aws gpu instances (python 2.7). i wanted to visualize some of the outputs so I copied the code to my machine and ran under python 3.5 (windows) and only got 57% accuracy. these swings in accuracy are huge