118 karma · joined November 22, 2019
If you file an issue here, I think it would work to update things:
Interestingly, there is a cuBLAS 13.1 whl on PyPI, not sure what that does.
For example, Alex Graves's (great! with attention) 2013 paper "Sequence Generation with Recurrent Neural Networks" has this line:
One difficulty when training LSTM with the full gradient is that the derivatives sometimes become excessively large, leading to numerical problems. To prevent this, all the experiments in this paper clipped the derivative of the loss with respect to the network inputs to the LSTM layers (before the sigmoid and tanh functions are applied) to lie within a predefined range.
with this footnote:
In fact this technique was used in all my previous papers on LSTM, and in my publicly available LSTM code, but I forgot to mention it anywhere—mea culpa.
That said, backpropagation seems important enough to me that I once did a specialized videocourse just about PyTorch (1.x) autograd.
In Thunder[1], a PyTorch to Python JIT compiler for optimizing DL models, we are maintaining a bytecode interpreter covering 3.10-3.12 (and 3.13 soon) for our jit. That allows to run Python code while re-directing arbitrary function calls and operations but is quite a bit slower than CPython.
While the bytecode changes (and sometimes it is a back-and-forth for example in the call handling), it seems totally good once you embrace that there will be differences between Python versions.
What has been a large change is the new zero cost (in the happy path) exception handling, but I can totally why Python did that change to that from setting up try-block frames.
I will say that I was happy not to support Python <= 3.9 as changes were a lot more involved there (the bytecode format itself etc.).
Of course, working on this has also means knowing otherwise useless Python trivia afterwards. One of my favorites is how this works:
l = [1, 2, 3]
l[-1] += l.pop()
print(l)
1. https://github.com/Lightning-AI/lightning-thunder/My intuition would be that you get more orthogonal directions to the gradient (of previous samples) if you have larger model.
If we follow the ordinary chain rule (for a single coordinate if you want) through the edges of the computational (DAG) graph, we get the right thing in each step.
The only other rule you need is that "if you use one variable several times in a calculation (i.e. several edges from(fw)/to(bw) the same node), you need to add the gradients computed for each", but IMHO that is pretty basic and intuitive, too. (So if you plug in z for both x and y into f(x, y), you have d/dz f(z, z) = f_x(z, z) + f_y(z, z), where the subscript indicates partial derivative.)
To me this seems both mathematically simpler than mixing the two into a "more than chain rule" thing and closer to what is actually going on algorithmically in a given implementation (the one I'm most familiar with is probably PyTorch's).
A more proper solution could be to chain the generator expressions: with
lines = ["a,b", "c,d"]
you could do ((a,b) for a, b in (l.split(',') for l in lines))https://karpathy.github.io/2019/04/25/recipe/
but the general idea of "get something that can overfit first" is probably pretty good.
In my experience getting the data right is probably the most underappreciated thing. Karpathy has data as step one, but in my experience, also data representation and sampling strategy does quite the miracle.
In Part II of our book we do an end-to-end project including e.g. a moment where nothing works until we crop around "regions of interest" to balance the per-pixel classes in the training data for the UNet. This has been something I have pasted into the PyTorch forums every now and then, too.
Then (as far as I know), in contrast to generation, training is done on the entire output of the transformer (so all tokens of the full input) rather than serially token-by-token (in the RNN days, this was called teacher-forcing), so that may give you a significant boost in the tokens per second rate over generation.
A long while ago, I wrote a little tutorial[0] on quantizing a speech commands network to the Raspberry. I used that to control lights directly and also for wake word detection.
More recently, I found that I can just use more classic VAD because my uses typically don't suffer if I turn on/off the microphone. My main goal is to not get out the mobile phone for information. That reduces the processing when I turn on the radio...
Not high-end as your solution, but nice enough for my purposes.
[0]. https://devblog.pytorchlightning.ai/applying-quantization-to...
Also JAX[1], PyTorch[2] come with JIT compilation specifically aimed at GPU kernels "fusing" multiple higher-level operation
And NumPy/Scipy (also) uses Pythran[3], an AOT compiler not too unsimilar to Numba.
[0] https://github.com/faster-cpython/cpython [1] https://pytorch.org/docs/stable/generated/torch.compile.html [2] https://jax.readthedocs.io/ [3] https://pythran.readthedocs.io/
I think some useful classification criteria would be - does it replace running code in Python (either own interpreter or compiler), vs does it speed up certain bits, - does it aim to faithfully implement Python or does it intentionally diverge in the semantics, - does it provide low-level semantics (where numba, pythran shine) or higher-level (e.g. what PyTorch, JAX do) - target architectures (CPU, GPU offloading, ...)
1. https://www.newyorker.com/tech/annals-of-technology/chatgpt-...