Deep Learning Systems
dlsyscourse.org
dlsyscourse.org
> “keys”, “queries”, “values”, in one of the least-meaningful semantic designations we have in deep learning
And in the context of LSTMs:
> throwing in some other names, like “forget gate”, “input gate”, “output gate” for good measure
This makes me feel more confident about actually understanding these topics. Before, I was totally misled by the awkward terminology.
But I think it is good to have memorable names that one can use to talk about the concepts verbally.
This is probably just an unfortunate situation, due to progressive understanding. Pointing this out in the slides gave me a sense of relief.
Personally, I wouldn't put names to every minor part of an algorithm or formula that was discovered to work empirically. But then again, I haven't discovered anything, and the authors of the respective papers certainly deserve some credit for their inventions!
These are legit:
cell_h = cell_(h-1) * forget_gate + tanh(linear(input_h)) * input_gate
out_h = cell_h * output_gate
see? forget_gate masks input by multiplying with numbers in [0, 1], input_gate controls the external input, output_gate controls the output of course
In most Deep Learning courses, the implementation is left to TAs and neither recorded nor made available. This course is an exception. Another bright exception is NYU Deep Learning course [0] by Yann LeCun and Alfredo Canziani. In that course, too, all recitations ("Practica") are recorded and made available. And Canziani is a great teacher.
Seems like he really cares. I looked him up and I guess he was a student of Andrew Ng (the legendary ML lecturer!!) so it makes sense.
Deep learning methods are so computationally intensive, many advances have come through new algorithms and optimization methods.