This is excellent. Thanks for sharing. It's always good to go back to the fundamentals. There's another resource that is also quite good: https://jaykmody.com/blog/gpt-from-scratch/
Your resource is really bad.
"We'll then load the trained GPT-2 model weights released by OpenAI into our implementation and generate some text."
What a bad take. That resource is awesome. Sure, it is about inference, not training, but why is that a bad thing?
The GPT from scratch post explains, from the ground up, ground being numpy, what calculations take place inside a GPT model.