For such a simple change to the softmax it wouldn't take long to verify. It's really embarrassing to not do that before publishing.
For such a simple change to the softmax it wouldn't take long to verify. It's really embarrassing to not do that before publishing.
To validate the idea the author has, it would be required to train a LLM from zero. If the author is right, you would get similar results to the current generation of LLMs, but with (a lot) less space required for the intermediate layers.
The time to achieve that is still measured in kilo- to mega-dollars, why is it wrong to put that idea in the open to substantially criticize or adopt?
And yes I do disregard his research effort. There are hundreds of well-justified and well-researched "clever tricks" for improving Transformers, and almost all of them don't work. I'll believe it when I see the results.
Train a Transformer based model with and without the modified Softmax (Suggestions: GPT-2 or nanoGPT)
Measure performance - I'd probably start with Perplexity and see if there is any difference (we'd expect little difference).
Quantize both models with different quantization strategies.
Measure the perplexity of the quantized models of different sizes. We'd expect the performance to drop off quicker for the non-modified model than the modified one if this is working.
In any case, that was an lmgtfy-level question. Here's what I found: https://til.simonwillison.net/llms/training-nanogpt-on-my-bl...
I shall try that soon.
I did a writeup like this. (Not as nicely as Simon though) where I modal.com (cloud GPU, containers, quick starts, free $30/m spend) to use their GPUs (e.g. T4, A100).
https://martincapodici.com/2023/07/15/no-local-gpu-no-proble...
T4 I think was good enough for the job, not much need for the A100.
Since this post I am working on an easy way to do this with a script called lob.py that requires no code changes to the nanoGPT repo (or whatever repo you are using) and runs in modal.com. The script exists but gets refined as I use it. Once it is battle tested a bit more I will do a post.
(It is named lob.py as it "lobs the code over to the server" where lob is UK slang for throw)
Watch this space.
BERT 109M, testing perplexity
OPT 125M, testing perplexity
ViT 22M, testing on ImageNet top-1.
I think there might be some curse of the auto-didact here, hinging on the meaning of publish: it would be embarrassing if he was capital-P publishing, as in a scientific paper.
The blog goes to great lengths to point out it is _not_ capital-P publishing.