1,717 karma · joined June 6, 2014
- it implements a real word-level LLM instead of a character-level LLM
- after pretraining also shows how to load pretrained weights
- instruction-finetune that LLM after pretraining
- code the alignment process for the instruction-finetuned LLM
- also show how to finetune the LLM for classification tasks
- the book it overall has a lots of figures. For Chapter 3, there are 26 figures alone :)
The video looks awesome though. I think it's probably a great complementary resource to get a good solid intro because it's just 2 hours. I think reading the book will probably be more like 10 times that time investment.
Now, the secondary goal is, of course, also to help people with building their own LLMs if they need to. The book will code the whole pipeline, including pretraining and finetuning, but I will also show how to load pretrained weights because I don't think it's feasible to pretrain an LLM from a financial perspective. We are coding everything from scratch in this book using GPT-2-like LLM (so that we can load the weights for models ranging from 124M that run on a laptop to the 1558M that runs on a small GPU). In practice, you probably want to use a framework like HF transformers or axolotl, but I hope this from-scratch approach will demystify the process so that these frameworks are less of a black box.
1) ... be theoretically a tad slower if you add the LoRA values dynamically during the forward pass (however, this is also an advantage if you want to keep a separate small weight set per customer, for example; you run only one large base model and can apply the different LoRA weights per customer on the fly)
2) ... have the exact same performance as the base model if you merge the LoRA weights back with the base model.
PS: All winners of the NeurIPS 2023 LLM Efficiency Challenge (finetuning the "best" LLM in 24h on 1 GPU) used LoRA or QLoRA (quantized LoRA).
E.g.
class MultilayerPerceptron(nn.Module):
def __init__(self, num_features, num_hidden_1, num_hidden_2, num_classes):
super().__init__()
self.layers = nn.Sequential(
nn.Linear(num_features, num_hidden_1),
nn.ReLU(),
nn.Linear(num_hidden_1, num_hidden_2),
nn.ReLU(),
nn.Linear(num_hidden_2, num_classes)
)
def forward(self, x):
x = self.layers(x)
return x
model = MultilayerPerceptron(
num_features=num_features,
num_hidden_1=num_hidden_1,
num_hidden_2=num_hidden_2,
num_classes=num_classes
)
model.layers[0] = LinearWithLoRA(model.layers[0], rank=4, alpha=1)
model.layers[2] = LinearWithLoRA(model.layers[2], rank=4, alpha=1)
model.layers[4] = LinearWithLoRA(model.layers[4], rank=4, alpha=1)I hope they mean 200 "million" daily active users. Otherwise it would indeed be quite bleak lol
What I meant to say here was 500B domain-specific tokens. Maybe domain-specific is not the right word here, but tokens related to the problems that the LLM aims to solve.
EDIT: Updated the text to be more clear.
I'd say that it's probably a mix of all of the above (incl some distillation).
100%! LOL. I was traveling and typing this on a mobile device. Must have been some weird autocorrect/autocomplete. Strange. And I didn't even notice. Thanks!
Not sure what happened there. Someone must have changed it! So weird! And I agree that the current title is a bit awkward and less representative.