Open sourcing Sonnet – a new library for constructing neural networks
deepmind.com
deepmind.com
Looks like there are some new layers for special kinds of attention RNNs, word embeddings, alternate implementations of spatial transformers, and so on. They also have another Batch Norm implementation that of course requires tons of fiddling to work properly, a classic tf staple :-)
As a machine learning environment, tf is so complex that different research groups have to define their own best practices for it. Using tf is like learning C++ where everyone learns a slightly different, mutually incompatible, but broadly overlapping subset of the language. We're just seeing a glimpse into DeepMind's specialized tooling along with reference implementations of the operations they use in their work.
This will be really useful for researchers who want to mess with deepmind's ideas/papers, but I'm a bit relieved that there isn't anything claimed to be fundamentally paradigm-shifting about this release.
I've been working with TF in detail for several months now and it is a very nicely designed framework. It is quite low-level, as you pointed out, (and so can be a challenge to get used to, and why you see helper wrappers/libs like this and keras popping up) but ML models and methods are becoming increasingly complex and TF provides exactly the foundation and flexibility needed to implement this stuff while still being able to slap together some basic layers if you need to. It also helps that the engineering is very solid (e.g., distributed models and datasets) and most of the performance kinks by now have been worked out.
For those maybe putting off working with TF because of its steepish learning curve, I'd strongly suggest you dive in. I learned last year in the usual method (docs + google), but the new OReilly books are the first good ones I've seen if that's your style.
The analogy with C++ is a good one, and historically appropriate given how scarce GPU resources are now.
So you can still do whatever with Keras by combining it with lower level TF operations, it's not really keeping you from doing anything. But more complex layer types and training might have to be rolled by hand, at which point it can be easier to just do it all at the lower level TF ops. This comes up a lot when implementing newer papers where they are proposing new training methods or Operations where you really need the flexibility to get in there and manipulate the ops directly.
My impression is that it's an ecosystem that's developing incredibly rapidly with all the associated growing pains. Even for such a simple-on-the-surface task as feeding very large datasets into the graphs, there are many different ways, some deprecated already, some about to be deprecated by functionality not in 1.0, etc.
It's also quite new, but the distributed execution, while awesome, still requires a huge amount of hand coding and tuning the machinery, reminding me quite a lot of just rolling the damn thing in MPI yourself. I'm very excited to see where distributed tf will be in a year or two, but it's a chore today.
I'm really bullish on their Cloud ML product, a tool to auto-distribute the execution graphs and associated data ingestion, but the documentation and examples for that are a bit all over the place at the moment, and require writing models to conform to the newish Experiment API, which is also a bit underdocumented and underexampled.
My frustration is likely only because of how awesome tensorflow actually is. It's new, but so promising that I want to use it for everything, regardless of the development roadmap!
# 1. Define our computation as some op
def useful_op(input_a, input_b,
use_clipping=True, remove_nans=False, solve_agi='maybe'):
# ...Sad face.
Is Python 2.7 requirement really a big deal for mostly using other people's code (adding a bit of your own application layer code)?
I use Anaconda and just have different environments for 2.7 and 3.x, TensorFlow or spaCy, or NLTK, etc.
Really easy to switch to the environment I need.
The 3.x is catching up, but still not fast enough as the community wants it to. But 2.7 is here for a reason, and shaming people that used 2.7 to push 3.x adoption, in my opinion is a bad practice.
If either their library (or a dependency) depends on something which binds to C, it could take some effort to make it python 3 compatible.
(FWIW, I use 2 or 3 depending on the OS or libraries that I am dealing with)
Hopefully they improve the docs as they go along; I could follow along because I work in the area, but they're pretty researchy, and I would imagine most would have difficulty parsing them unless they're regularly reading papers in this area.
I'm looking forward to this - DeepMind has a frustrating track record of not open sourcing their code. While it's their prerogative and I know of many valid reasons for doing so, the many other researchers who publish their (sometimes terribly hacked together) models have been incredibly helpful in verifying that their work, well, works.
EDIT: Also, what about TF-Slim?
Keras, though, still really incredible for usability.