OpenAI Baselines
blog.openai.com
blog.openai.com
I ran into this recently by accident when writing a simple RL example. With two weight matrices to learn, the first weight matrix was given correct gradients, the last weight matrix was only supplied with partial information. Surprise, still works, and I only discovered the bug _after_ submitting to OpenAI's Gym with quite reasonable results. I've seen similar issues in the past such as accidentally leaving a part of the network frozen (i.e. it was randomly initialized and never changed) yet the model still happily went along with it.
This is good and bad. Bad in that it makes errors difficult to catch. Good in that, if you had a reason for freezing part of the network (maybe transfer learning etc) your model will learn to happily use it, even if that "information" is more akin to noise.
Regarding reproducibility, most papers I've gone to reproduce take far longer than expected and usually involve deducing / requesting additional information from the lead authors. Even minor issues, such as how the loss is calculated (loss / (batch * timestep) vs loss / batch) can confuse substantially and given they seem "insignificant" and that there are space constraints in papers, they are rarely written down.
Worst I have seen recently was a state of the art published result where the paper was accepted to a conference yet they didn't include a single hyperparameter for their best performing model - and no code. There is near zero ability to reproduce that given the authors spent a small nuclear reactor worth of compute performing grid search to get the optimal hyperparameters.
tldr There are reproducibility issues all the way up the stack, from gradient descent working against you to minor omissions in the papers to full fledged omissions that are still accepted by the community.
For instance, I suspect a large fraction of claimed "better than baseline" results in papers are actually a result of well-tuned proposed model and not-very-well-tuned baseline. Papers usually say something along the lines of "we tune these hyperaparameters with crossvalidation", but you can be very thorough or not very thorough in this process and achieve very different results. Or if the author is (intentionally or unintentionally) devious, they could adjust the hyperparameter search ranges in a way that doesn't cover the optimal region, or is way too broad.
All of these problems combined are the reason that I pay much less attention to numbers in the tables of papers. Papers are, unfortunately, more of a source of cool ideas and inspiration for stuff to put into your toolbox and maybe try yourself in your next project.
A year or two may not sound like much but for deep learning but you might as well be comparing Heron's engine[1] to a modern combustible engine. The most cited attention paper, Bahdanau et al. 2014, was only two and a half years ago - and their results weren't all that strong initially anyway. There are so many minor auxiliary improvements / newly found domain expertise that can add up to a substantially improved result (initializations, regularization techniques, hyper parameter choices that work well for that given task / dataset, ...) even for almost the exact same model.
As an example of this, I implemented a Keras SNLI model[2] as a teaching aid for friends only to find my simplest baseline outperformed a number of fairly recent papers. This is not necessarily an issue at all with those papers - many of the ideas and techniques they propose have gone on to be used in other tasks and models - but it does indicate how these additional best practices / more time on the baseline can really accumulate over only months. Keras is a brilliant example of that as it's "best practices included" - when you use an LSTM it'll default to strong weight initializations etc.
I'm vaguely hoping that in the future there can be workshops at conferences titled "Pushing the baselines" where we take some time to see what new techniques can apply to old baselines and/or tune baselines with the explicit goal of getting them to regain their edge. There's at least one at ICML on reproducibility[3] which I'm looking forward to :)
[1]: https://en.wikipedia.org/wiki/Aeolipile
[2]: https://github.com/Smerity/keras_snli
[3]: https://sites.google.com/view/icml-reproducibility-workshop/...
Recently I had to implement gradient calculations by hand recently (writing custom CUDA code) and had a pretty terrible time. Mixing the complications of CUDA code with my iffy manual differentiations and floating point silliness can drive you a little bonkers. I ended up implementing a slow automatic differentiated version and compared resulting outputs and gradients to help work through my bugs.
Here's hoping that Tensorflow's XLA and other JIT style CUDA compilers/optimizers will make much of this obsolete in the near future.
For those not familiar, the overhead for calling a CUDA kernel can be insanely high, especially when you're just doing an elementwise operation such as an add. Given your neural network likely has many many of these, wrapping many of these into one small piece of custom CUDA can result in substantial speed increases. Unfortunately there's not really any automatic way of doing that yet. We're stuck in the days of either writing manual assembly or being fine with suboptimal compiled C.
[1]: https://www.tensorflow.org/versions/r0.11/api_docs/python/te...
[2]: https://github.com/pytorch/pytorch/blob/master/torch/autogra...
Right now we hand write all of our own gradients as well.
The overhead can come from a ton of different places. This is why we wrote workspaces: http://deeplearning4j.org/workspaces
Allocation reduction and op grouping are only a few things you can do.
it is a computer algorithm, so by definition it is trivial to reproduce results, you just run the program again.
> I ran into this recently by accident when writing a simple RL example. With two weight matrices to learn, the first weight matrix was given correct gradients, the last weight matrix was only supplied with partial information. Surprise, still works, and I only discovered the bug _after_ submitting to OpenAI's Gym with quite reasonable results.
so you want to say that you coded a bug, but you don't have a method of testing whether you have a bug. So you didn't code a bug. And if you didn't code a bug, you can't reproduce a bug.
So yes, reproducing a bug is difficult when you have no means of determining whether or not you have a bug...
...maybe you should look into choosing a means of determining whether or not you have a bug.
Not really. Deep learning is still quite a lot of dark voodoo where random initialization and data shuffling can matter significantly. People also adapt hyperparameters manually during training, stop early with no clear metric, and don't share their code for preprocessing the data or even the exact architecture of the network.
It's certainly better than in other fields, but it's not trivial.
If I have a function f and an element of the domain x, then the value y = f(x) does not change. This is not dark voodoo.
For my Master's thesis in AI (12 years ago, so before most of this open stuff) I compared an existing Genetic Algorithm, described in a published paper, against my improvement. My improvement was significantly better.
However, I relied on the prose description of the original algorithm. The original paper (cited many, many times) didn't even have pseudocode, let alone source code.
For my paper, I included pseudocode of both the original algorithm and my improved algorithm. But we still didn't have established practices for how to make source code available to readers, in such a way that they'd be archived long term.
Is there an established way now?
Seems to be gaining popularity.
Your broader point is spot on however. My general hope is that people are wary when their results are strong^, ensuring that you don't get a good result via "cheating", so the majority of bugs are likely to harm performance. If a result is also not reproducible (i.e. "cheating" bug) it won't be used and built on - but if a result is bugged but reproducible (i.e. bug where performance was lower than it should be) then we can still move the field forward even in spite of these issues.
^ When I achieved state of the art for a task - especially given it was a huge jump in accuracy for a relatively small model compared to the previous state of the art - I spent many days sitting there double checking I hadn't accidentally cheated ;)
I'd like to imagine it's how effective peer review could be if given sufficient motivation ;)
As you note, though, most papers don't get anywhere in the same magnitude of focus, and others which do may still be entirely unreproducible anyway :(
This is a historically backed insight. If you're interested in a good critique of the decompose-by-function-then-combine-later approach, I recommend "Intelligence without Representation" from Rodney Brooks http://www.scs.ryerson.ca/aferworn/courses/CPS607/CLASSES/su...
Hmm... I can agree that vision and NLP could be seen as "applications", from one point of view. But I can see another position where each simply represents a different aspect of underlying cognition. Language in particular, seems to be closely tied up in how we (humans) think. And without proposing a strong version of the Sapir-Whorf hypotheses, I can't help but believe that a lot of human cognition is carried out in our primary language. Now to be fair, this belief comes from not much more than obsessively trying to introspect on my own thinking and "observe" my own mental processes.
In any case, it leads me to suspect that building generally intelligent AI's will be tightly bound up with understanding how language works in the brain and the extent to which there is a "mentalese" and how - if at all - a language like English (or Mandarin or Tamil, whatever) maps with "mentalese". Vision also seems crucial to the way humans learn, given our status as embodied agents that learn from the environment using site, smell, sound, kinesthetic awareness, proprioception, etc.
Quite likely I'm wrong, but I have a hunch that building a truly intelligent agent may well require creating an agent that can see, hear, touch, smell, balance, etc. At least to the extent that humans serve as our model of how to construct intelligence.
On the other hand, as the old saying goes "we didn't build flying machines by creating mechanical birds with flapping wings". :-)
When we "work on NLP" it looks something like https://blog.openai.com/learning-to-communicate/ or http://www.foldl.me/2016/situated-language-learning/, as an emergent phenomenon in the service of a greater objective - not an end but a means to an end. It does not look like doing sentiment analysis or machine translation.
When we "work on Computer Vision" it looks like including a ConvNet into the agent architecture or a robot, it doesn't look like trying to achieve higher object detection scores.
Ah, sounds like we are on the same page then.
Creative and inventive thought is very picturesque and non-linear: https://www.amazon.com/Psychology-Invention-Mathematical-Fie...