Building an LLM from Scratch: Automatic Differentiation (2023)
bclarkson-code.github.io
bclarkson-code.github.io
(see also andre karpathys zero to hero nn series on youtube as well its very good and similar to this work)
https://arena-ch1-transformers.streamlit.app/%5B1.1%5D_Trans...
Edit - it is. Not to talk down on the series. I’m sure it’s good, but it is actually “LLM with PyTorch”.
Edit - I looked again and I was actually not correct. He does ultimately use frameworks, but gives some early talk about how those function under the hood.
[1] https://github.com/cafaxo/Llama2.jl/tree/master/src/training
I quite like Jeremy's approach: https://nbviewer.org/github/fastai/fastbook/blob/master/17_f...
It shows a very simple "Pythonic" approach to assemble gradient of a composition of functions from the gradients of the components.
But obviously the author already thought of that. The source repo has a great motto: "It don't go fast but it do be goin'" [1]
I love the idea of the project and I'm curious to see what the endgame runtime will be.
i realize the cost and time to train may be prohibitive and that quality on general english might be very limited, but is the code itself available ?