HNHacker News
TopNewBestAskShowJobs

gpjt

1,810 karma · joined January 12, 2009

https://www.gilesthomas.com/
submissionscomments
gpjt··on LLM from scratch, part 28 – training a base model from scratch on an RTX 3090
OK, early indicators support both you and Gemini quite strongly re: batch size. On my (somewhat ad-hoc) test dataset, I get losses like this:

  * OpenAI medium weights: 3.231
  * OpenAI small weights: 3.500
  * My locally trained model, FineWeb Chinchilla, batch size 6: 3.944
  * My locally trained model, FineWeb-Edu Chinchilla, batch size 6: 4.167
  * My locally trained model, FineWeb-Edu double Chinchilla, batch size 6: 4.135
  * My cloud trained model, FineWeb Chinchilla, batch size 13 \* 8 = 104: 3.674
That last one was trained on an 8x A100 machine with 40 GiB per GPU, with the same code as before, just converted to DDP. It certainly looks like the much larger batch size has improved the model significantly.

I'll be trying on larger machines. No gradient accumulation yet, but it's certainly looking like a valuable lever to pull for local training runs (and, I suspect, might also be useful on "small" cloud machines like the one I used -- will have to see what things look like with the bigger mini-batches I can squeeze onto 80 GiB and 160 GiB GPUs).

gpjt··on LLM from scratch, part 28 – training a base model from scratch on an RTX 3090
OP here -- agreed! I tried to summarise (at least to my current level of knowledge) those 12-18 hours here: https://www.gilesthomas.com/2025/09/maths-for-llms
gpjt··on LLM from scratch, part 28 – training a base model from scratch on an RTX 3090
OP here: one thing that surprised me in this experiment was that the model trained on the more curated FineWeb-Edu dataset was worse than the one trained on FineWeb. That is very counterintuitive to me.
gpjt··on LLM from scratch, part 28 – training a base model from scratch on an RTX 3090
OP here -- thanks! I'm in the process of doing some trains using the same code plus DDP on big Lambda Labs machines, and (within the bounds of what I can afford) will hopefully have some interesting results about all of those shortly.
gpjt··on Man unexpectedly cured of HIV after stem cell transplant
How much of that low survival rate is due to the condition they received the transplant, though? Conceivably a patient with "just" HIV might do better than one with eg. leukemia and HIV.

That said, IIUC the whole stem cell transplant procedure is unpleasant enough that it still might not be worth it.

gpjt··on A cryptography research body held an election and they can't decrypt the results
Thanks for the reminder of a brilliant IT crowd moment!
gpjt··on OpenAI researcher announced GPT-5 math breakthrough that never happened
To be fair to the OpenAI team, if read in context the situation is at worst ambiguous.

The deleted tweet that the article is about said "GPT-5 just found solutions to 10 (!) previously unsolved Erdös problems, and made progress on 11 others. These have all been open for decades." If it had been posted stand-alone then I would certainly agree that it was misleading, but it was not.

It was a quote-tweet of this: https://x.com/MarkSellke/status/1979226538059931886?t=OigN6t..., where the author is saying he's "pushing further on this".

The "this" in question is what this second tweet is in turn quote-tweeting: https://x.com/SebastienBubeck/status/1977181716457701775?t=T... -- where the author says "gpt5-pro is superhuman at literature search: [...] it just solved Erdos Problem #339 (listed as open in the official database erdosproblems.com/forum/thread/3…) by realizing that it had actually been solved 20 years ago"

So, reading the thread in order, you get

  * SebastienBubeck: "GPT-5 is really good at literature search, it 'solved' an apparently-open problem by finding an existing solution"
  * MarkSellke: "Now it's done ten more"
  * kevinweil: "Look at this cool stuff we've done!"
I think the problem here is the way quote-tweets work -- you only see the quoted post and not anything that it in turn is quoting. Kevin Weil had the two previous quotes in his context when he did his post and didn't consider the fact that readers would only see the first level, so wouldn't have Sebastien Bubek's post in mind when they read his.

That seems like an easy mistake to entirely honestly make, and I think the pile-on is a little unfair.

gpjt··on Why do LLMs freak out over the seahorse emoji?
This is a great post on many levels, but what struck me as particularly clever was the use of lm_head to decode the outputs of earlier layers. That linear layer is only trained to decode the output of the last layer, so intuitively it might only be able to do that -- the embedding spaces used between earlier layers might be different and "incompatible". It's really interesting that that is not the case.
gpjt··on The maths you need to start understanding LLMs
Post author here. I agree 100%! The post is the basic maths for people digging in to how LLMs work under the hood -- I wrote a separate one for non-techies who just want to know what they are, at https://www.gilesthomas.com/2025/08/what-ai-chatbots-are-doi...
gpjt··on The maths you need to start understanding LLMs
Check the first link in the parent comment, it's a link to the book.
gpjt··on Tell HN: I just made a first ever dollar on my SaaS
Congrats! Amazing feeling, isn't it :-)
gpjt··on Ask HN: Go deep into AI/LLMs or just use them as tools?
This, 100%. A full-stack engineer will likely have at least a solid understanding of the HTTP protocol, HTTPS, WebSockets, the interface layer between the frontend server and their chosen Web webdev stack, and so on. Then a more general understanding of networking protocols, TCP vs UDP, DNS, routing, etc. In general, you need to have a solid understanding of the layer below where you're working, some understanding of the layer below that, and so on, less and less detail needed for each layer down.

(That's not to say that you shouldn't bother with learning more -- more knowledge is always good -- or that the OP specifically only knows that. It's more a sensible minimum.)

My own "curriculum" for that has been Jeremy Howard's Fast AI course and Sebastian Raschka's book "build an LLM from scratch". Still working through it, but once I'm done I think I'll be solid on your point 2 above. My guess is that I'll want to learn more, but that's out of interest more than because I think its necessary.

gpjt··on Writing an LLM from scratch, part 13 – attention heads are dumb
As the author of the original post above, let me say that if that's word salad, it's a Michelin star salad. Just the right mix of lettuce and tomato, and the dressing is spot on :-)

Seriously, though, differentiable hash tables is an awesome way to look at them, I wish I'd heard it before.

gpjt··on Writing an LLM from scratch, part 13 – attention heads are dumb
Author of the post here -- I'm being careful not to do that. My posts are more about filling in the gaps; they're covering the things that aren't mentioned. The book's target audience is, I think, people with a bit more background knowledge about the inner workings of AI than I have, so I'm having to play catch-up a bit.
gpjt··on Writing an LLM from scratch, part 13 – attention heads are dumb
Author here: I endorse this comment ;-) That's definitely the route I've optimised for for reading the series.
gpjt··on Gandi March 9, 2025 incident postmortem
Another one leaving for Porkbun here.
gpjt··on Why are there no thunderstorms in the UK?
Huh, I was thinking the same thing, and was wondering whether it was just moving to London. Could be both, I suppose.
gpjt··on Show HN: GS-Calc – A modern spreadsheet with Python integration
As co-founder of Resolver Systems -- we tried but ultimately failed to take on Excel with a Python-enabled equivalent back in 2007 -- and current employee at Anaconda (providing Python in Excel) I really do hope you get this one to work. Excel is a mess, Python is better, and someone surely will eventually be able to fix the former with the latter. Let's hope it's you :-)
gpjt··on NoProp: Training neural networks without back-propagation or forward-propagation
Interesting. As you say, that certainly makes sense for mammala. But I'd be interested in knowing what mechanisms you might conjecture for birds, where pretty much all foetal development happens inside the egg, separated from the mother -- or fish, or octopuses.
gpjt··on NoProp: Training neural networks without back-propagation or forward-propagation
Presumably that is limited by the gig or so of information in our DNA, though?
gpjt··on Ask HN: Are you afraid to travel to US to tech conferences?
It does depend on where you enter the country, at least in my experience (UK citizen, have visited a lot both on visa waivers and prior to that on visas, since the 80s).

Every time I've flown into Austin, TX, they've been super-nice. DC likewise. NYC/Newark are brusque but not nasty. San Francisco are scary. Boston on the one time I flew there was just horrendous, though that might have just been one agent who was having a bad day.

gpjt··on Writing an LLM from scratch, part 10 – dropout
My intuition is very undeveloped on this, but it makes some kind of sense to me that dropout would make convergence slower, because you're ignoring a bunch of parameters in every batch. The goal seems to be to get a better, more general model by trading off some training time.

The Llama thing is interesting, though!

gpjt··on Writing an LLM from scratch, part 10 – dropout
OP here -- I'm new at this, but I don't think so. A zero output from a neuron still contributes to the output. Taking a silly toy case, imagine a network with one input, one neuron, one output. You're training to it match data where whatever the input is, it outputs 1 -- that is, your target state would have the weight set to zero and the bias set to 1. If it was initialised with the weight zero but the bias also zero, then when you pushed your test set through you'd get zero outputs, but the error would be non-zero and there would be adjustments to propagate back.

I could well be misunderstanding you, though!

gpjt··on Writing an LLM from scratch, part 10 – dropout
OP here -- that's the one! Highly recommended.
gpjt··on Revealed: How the UK tech secretary uses ChatGPT for policy advice
I've found it very useful with organisational problems too. I had a complex issue to work through recently and tried working with ChatGPT, Claude and Grok 3 to optimise letters to some of the people involved to try to get things solved in the way I felt was fairest for all involved (anonymised, of course, with no memory on ChatGPT). One neat trick was to export the letter from one chat, then to start a fresh one, say that I am <role of recipient> and provide the letter and ask what it thinks -- basically red-teaming it.

The process of doing that clarified my thoughts and arguments so much that I never wound up having to send them -- I'd already made the right points on Slack and in meetings, and a compromise was achieved.

gpjt··on Writing an LLM from scratch, part 8 – trainable self-attention
I have 100% set myself a "no side quests" rule while going through the book so that I don't do that. I've had... patchy success with that, but I think I'm doing pretty well apart from the week I spent getting LaTeX rendering working on my blog so that I could do the pretty maths in that post.

What I'm doing is building up a list of things to dig into in depth once I've finished the book. Kind of like a treat to encourage me to push forward when I'm working through a bit that's tough to understand.

gpjt··on Writing an LLM from scratch, part 8 – trainable self-attention
OP here: my posts are kind of reading notes for the book, so I don't normally copy the code from there -- so you wouldn't have seen the packages used. So far there's been tiktoken for tokenization (Raschka shows how to write a simple tokenizer and explains the workings of the byte-pair tokenization that he recommends, though) and PyTorch for CUDA-acceleratable matrix maths, automated differentiation for gradient descent, and so on.
gpjt··on Writing an LLM from scratch, part 8 – trainable self-attention
I do see your point!

But at the end of the day, it depends on where you want to spend your time. "Build an LLM from scratch" is over 300 pages -- and they are very dense pages. My blog post covers fewer than 10 of them (though TBF they are the hardest pages). Adding on tokenizers in depth from scratch would add on 100 or so more. Adding on efficient-enough matrix multiplication to do anything would add on a few hundred more, and doing it in CUDA would probably be a couple of thousand. Now add on automated differentiation to work out the gradients for training -- a few thousand more? Optimizers for the training -- even more than that, perhaps.

You have to draw the line somewhere, as otherwise (as you suggest) the "from scratch" book has to start "go out and get some really clean sand" so that you can start fabbing your own chips. I think that tiktoken and PyTorch are a solid choice for that line, as it means that the book is manageable in size and gives you enough of an overview of the underlying stuff to be able to work out what you want to dig into next.

gpjt··on Writing an LLM from scratch, part 8 – trainable self-attention
OP here: that is an excellent point! Of the eight read-throughs, four were on Friday night, then I had dreams involving some kind of TRON-like vector spaces, and at brunch on Saturday things seemed to start gelling. The four extra read-throughs were to crystallise that intuition to a level that I felt I could start writing it up. I'm 100% sure the sleep was what built some kind of intuition at a pre-linguistic stage that I could build on.
gpjt··on The benefits of learning in public
Oh God, I just worked out the joke. Very funny, Kendrick :P
← PreviousPage 2 of 9Next →