HNHacker News
TopNewBestAskShowJobs

rasbt

1,717 karma · joined June 6, 2014

AI researcher and statistics professor
submissionscomments
rasbt··on Implementing a ChatGPT-like LLM from scratch, step by step
Yeah, I don't think creating educational materials makes sense from an economical perspective, but it's one of my hobbies that gives me joy for some reason :). Hah, and 'insane amount of work' is probably right -- lots of sacrifices to carve out that necessary time.
rasbt··on Implementing a ChatGPT-like LLM from scratch, step by step
Sorry, in that case I would rather recommend a dedicated RL book. The RL part in LLMs will be very specific to LLMs, and I will only cover what's absolutely relevant in terms of background info. I do have a longish intro chapter on RL in my other general ML/DL book (https://github.com/rasbt/machine-learning-book/tree/main/ch1...) but like others said, I would recommend a dedicated RL book in your case.
rasbt··on Implementing a ChatGPT-like LLM from scratch, step by step
I added notes to the Jupyter notebooks, I hope they are also readable as standalone from the repo.
rasbt··on Implementing a ChatGPT-like LLM from scratch, step by step
That was pretty smooth. They reached out whether I was interested in writing a book for them (probably because of my other writings online), I mentioned what kind I book I want to write, submitted a proposal, and they liked that idea :)
rasbt··on Implementing a ChatGPT-like LLM from scratch, step by step
Lol ok, otherwise it would probably be not very readable due to the verbosity. The book shows how to implement LayerNorm, Softmax, Linear layers, GeLU etc without using the pre-packaged torch versions though.
rasbt··on Implementing a ChatGPT-like LLM from scratch, step by step
Haven't fully watched this but from a brief skimming, here are some differences that the book has:

- it implements a real word-level LLM instead of a character-level LLM

- after pretraining also shows how to load pretrained weights

- instruction-finetune that LLM after pretraining

- code the alignment process for the instruction-finetuned LLM

- also show how to finetune the LLM for classification tasks

- the book it overall has a lots of figures. For Chapter 3, there are 26 figures alone :)

The video looks awesome though. I think it's probably a great complementary resource to get a good solid intro because it's just 2 hours. I think reading the book will probably be more like 10 times that time investment.

rasbt··on Implementing a ChatGPT-like LLM from scratch, step by step
It's in progress still. I have most of the code working, but it's not organized into the chapter structure, yet. I am planning to add a new chapter every ~month (I wish I could do this faster, but I also have some other commitments). Chapter 4 will be either uploaded by the end of this weekend or by the end of next weekend.
rasbt··on Implementing a ChatGPT-like LLM from scratch, step by step
I'd say my primary motivation is an educational goal, i.e., helping people understand how LLMs work by building one. LLMs are an important topic, and there are lots of hand-wavy videos and articles out there -- I think if one codes an LLM from the ground up, it will clarify lots of concepts.

Now, the secondary goal is, of course, also to help people with building their own LLMs if they need to. The book will code the whole pipeline, including pretraining and finetuning, but I will also show how to load pretrained weights because I don't think it's feasible to pretrain an LLM from a financial perspective. We are coding everything from scratch in this book using GPT-2-like LLM (so that we can load the weights for models ranging from 124M that run on a laptop to the 1558M that runs on a small GPU). In practice, you probably want to use a framework like HF transformers or axolotl, but I hope this from-scratch approach will demystify the process so that these frameworks are less of a black box.

rasbt··on Implementing a ChatGPT-like LLM from scratch, step by step
Glad to hear and thanks for the support. Chapter 3 should be in the MEAP soonish (submitted the draft last week). Will also upload my code for chapter 4 to GitHub soonish, in the next couple of days, just have to type up the notes.
rasbt··on Implementing a ChatGPT-like LLM from scratch, step by step
It kind of is, but it's also kind of motivating :)
rasbt··on LoRA from scratch: implementation for LLM finetuning
During training, it's more efficient than full finetuning because you only update a fraction of the parameters via backprop. During inference, it can ...

1) ... be theoretically a tad slower if you add the LoRA values dynamically during the forward pass (however, this is also an advantage if you want to keep a separate small weight set per customer, for example; you run only one large base model and can apply the different LoRA weights per customer on the fly)

2) ... have the exact same performance as the base model if you merge the LoRA weights back with the base model.

rasbt··on LoRA from scratch: implementation for LLM finetuning
I think the main use case remains behavior changes: instruction finetuning, finetuning for classification, etc. Knowledge addition to the weights is best done via pretraining. Or, if you have an external database or documentation that you want to query during the generation, RAG as you mention.

PS: All winners of the NeurIPS 2023 LLM Efficiency Challenge (finetuning the "best" LLM in 24h on 1 GPU) used LoRA or QLoRA (quantized LoRA).

rasbt··on LoRA from scratch: implementation for LLM finetuning
Yeah, the LoRA part is from scratch. The LLM backbone in this example is not, this is to provide a concrete example. But you could apply the exact same LoRA from scratch code to a pure PyTorch model if you wanted to:

E.g.

    class MultilayerPerceptron(nn.Module):

        def __init__(self, num_features, num_hidden_1, num_hidden_2, num_classes):
            super().__init__()

            self.layers = nn.Sequential(
                nn.Linear(num_features, num_hidden_1),
                nn.ReLU(),
                nn.Linear(num_hidden_1, num_hidden_2),
                nn.ReLU(),
                nn.Linear(num_hidden_2, num_classes)
            )

        def forward(self, x):
            x = self.layers(x)
            return x

    model = MultilayerPerceptron(
        num_features=num_features,
        num_hidden_1=num_hidden_1,
        num_hidden_2=num_hidden_2, 
        num_classes=num_classes
    )

    model.layers[0] = LinearWithLoRA(model.layers[0], rank=4, alpha=1)
    model.layers[2] = LinearWithLoRA(model.layers[2], rank=4, alpha=1)
    model.layers[4] = LinearWithLoRA(model.layers[4], rank=4, alpha=1)
rasbt··on LoRA from scratch: implementation for LLM finetuning
Hah, yeah that's LoRA as in Low-Rank Adaptation :P
rasbt··on Coding Self-Attention, Multi-Head Attention, Cross-Attention, Causal-Attention
Yes, totally agree. These implementations are meant for educational purposes. You could in theory use them to train a model though (GPT-2 also had a from-scratch implementation if I recall correctly). In practice, you probably want to use FlashAttention though. You use it through `torch.nn.functional.scaled_dot_product_attention` etc.
rasbt··on Atlassian Acquires Loom
It's fascinating that it's a 1B business. I thought it was just uploading screen recordings to the cloud (basically UI around uploading. Like macOS QuickTime + YouTube private video upload)
rasbt··on Twitter / X is losing daily active users. CEO Linda Yaccarino confirmed it
> When Yaccarino was first asked about user metrics during the interview, she seemingly wanted to move away from that particular conversation, saying that X had between 200 and 250 daily active users.

I hope they mean 200 "million" daily active users. Otherwise it would indeed be quite bleak lol

rasbt··on Optimizing LLMs from a Dataset Perspective
I think it could potentially make the model smarter, but it's up to how you collect the data to train the reward models. Currently, companies & papers that use RLHF focus on "safety" rankings, for example. But you could potentially collect labels "smartness" or "correctness" instead and train the the reward model one these. (And then use that reward model to finetune the LLM you want to improve.)
rasbt··on Optimizing LLMs from a Dataset Perspective
RLHF is a popular candidate, but the focus is more on "helpfulness" and "safety" -- I don't think it necessarily improves LLMs on reasoning benchmarks
rasbt··on Understanding Llama 2 and the New Code Llama LLMs
Yes, when I remember correctly, they said they didn't release the 34B Llama 2 model yet because they haven't had a chance for "red teaming" that one, where with "red teaming" they mean something along the lines of identifying and exploiting vulnerabilities
rasbt··on Understanding Llama 2 and the New Code Llama LLMs
Haven't seen that one, yet. Thanks for sharing!
rasbt··on Understanding Llama 2 and the New Code Llama LLMs
Good catch. Above that paragraph, I wrote that the Code Llama models were initialized with the Llama 2 weights, which makes this contradictory, indeed.

What I meant to say here was 500B domain-specific tokens. Maybe domain-specific is not the right word here, but tokens related to the problems that the LLM aims to solve.

EDIT: Updated the text to be more clear.

rasbt··on Understanding Llama 2 and the New Code Llama LLMs
Interesting, I thought GPT-3.5 was considered GPT-3 + InstructGPT-style RLHF on a large scale, whereas GPT-4 is considered to be an MoE model.
rasbt··on Understanding Llama 2 and the New Code Llama LLMs
I think so too. But in general, it could also be due to other reasons: faster hardware, lower timeout for batched inference, optimizations like flash attention and flash attention 2, quantization, ...

I'd say that it's probably a mix of all of the above (incl some distillation).

rasbt··on Why the original transformer figure is wrong, and some other tidbits about LLMs
Agreed, compared to other architectures, transformers are actually quite straight-forward. The complicated part comes more from training it in distributed setups, making the data loading and tensor parallelism work due to the large size etc. Like the vanilla architecture is simple, but the practical implementation for large-scale training can be a bit complicated.
rasbt··on Why the original transformer figure is wrong, and some other tidbits about LLMs
> Misspelling "Attention is All Your Need" twice in one paragraph makes for a rough start to the linked post.

100%! LOL. I was traveling and typing this on a mobile device. Must have been some weird autocorrect/autocomplete. Strange. And I didn't even notice. Thanks!

rasbt··on Why the original transformer figure is wrong, and some other tidbits about LLMs
So weird, I posted it with almost the original title (only slightly abbreviated to make it fit: "Why the Original Transformer Figure Is Wrong, and Some Interesting Tidbits About LLMs".

Not sure what happened there. Someone must have changed it! So weird! And I agree that the current title is a bit awkward and less representative.

rasbt··on Show HN: A fully open-source (Apache 2.0)implementation of llama
I think some businesses and people are worried about using GPL code in their code bases because that's incompatible with their own licenses.
rasbt··on Show HN: A fully open-source (Apache 2.0)implementation of llama
Not sure, but I think the point was that if you have something in GPL license (like the code in this case) it's open source, but that doesn't mean you can use that for your business application. That's because GPL requires you open sourcing all derivative work and most businesses don't want to/can't do that.
rasbt··on Show HN: A fully open-source (Apache 2.0)implementation of llama
I guess that means time to fire up a few GPUs later today and get some weights! We should have a weight exchange platform for that maybe, haha.
← PreviousPage 2 of 3Next →