HNHacker News
TopNewBestAskShowJobs

gpjt

1,810 karma · joined January 12, 2009

https://www.gilesthomas.com/
submissionscomments
gpjt··on Blog about things you don't understand yet
Hmm. Great post, not so sure about the title.

I'd say (and I think this is actually compatible with the article itself, if not the title) that a better rule is to write about things that you've only just started understanding. The OP is, I think, correct that you might write a better introduction than an expert for whom it's all obvious if you've just been through the struggle of learning it, because the journey is fresh in your mind.

The risk of sounding more knowledgeable than you actually are is a real problem, though, and I'm glad he mentions it. It's really tricky to strike the right balance between making it clear that the topic is new to you, and hedging so much that you sound like ChatGPT on a bad day.

gpjt··on Midlife Vascular Risk Burden and Dementia-Free Survival Years
I must admit that I found it confusing that they lumped people with unmanaged hypertension in with those who were on meds to manage it. Those feel like separate groups. Does anyone here know more about this area, and why the combined group might make sense?
gpjt··on Ask HN: What are you working on? (August 2026)
I'm adding MoE support to the GPT-2 code from Sebastian Raschka's "Build a Large Language Model (From Scratch)". Just the minimal changes. Planning to do (or at least start!) a from-scratch base model training run on my own hardware before the end of the month.
gpjt··on I use AI on this blog
Lovely bit of inadvertent editorialising by the submission form there :-) Title should be (as it is on the page) "How I use AI on this blog".
gpjt··on Ask HN: Anthropic banned me from using Claude Code and I don't know what to do
In the EU, at least one of the problems is regulatory/tax for digital services. They need to charge you the VAT rate for the country where you are based. For that, they need two pieces of evidence about your location. There are various things they can use for that -- telephone number, IP address, card billing address, and so on. If they can collect two that indicate the same country, they're safe -- but if all of them point in different directions then they could get in trouble during a tax audit.

Of course, for larger transactions you'd expect that a human in the loop could work with you to get the right info so that they would be covered. But I guess for Microsoft, their definition of "larger" might be more than a few grand...

gpjt··on 10Gb/s Ethernet: switching to a Broadcom SFP+ module
OP here -- yes indeed. I ever do a new series of posts on upgrading my network to anything faster, the first one will probably be titled something like "25Gb/s Ethernet: how much it costs to rip out your CAT-6A and replace it with fibre".
gpjt··on 10Gb/s Ethernet: switching to a Broadcom SFP+ module
OP here -- yes, 100%! Within my study it's DACs all the way. I only use 10GBASE-T where it's the only option: through the wall cabling, and from the connector on my ISP's crappy router to the mini-PC that acts as the real router.
gpjt··on 10Gb/s Ethernet: switching to a Broadcom SFP+ module
Thanks for making the comment! It saved me having to bolt a USB fan over the old, misbehaving SFP+ module or something similarly silly...
gpjt··on Statement on US government directive to suspend access to Fable 5 and Mythos 5
Doesn't say anything about citizenship though. There are plenty of US residents who are not citizens. And a lot of people abroad appear to use US billing address credit cards -- in my last company we had hundreds of people with the same US billing address who appeared to be managing Africa-focused businesses and used IPs that matched that.
gpjt··on Ask HN: High school student – is learning programming still worthwhile?
Excellent points. It's also worth noting that many people don't wind up working in the field they studied at university. CS grads have probably been the exception in recent years, because the industry has been booming, but the two most successful entrepreneurs I know studied philosophy and art history. A friend who is very senior in recruitment studied economics.
gpjt··on Ask HN: High school student – is learning programming still worthwhile?
That is an excellent point so long as you don't take it too far! Lisp/Haskell/Erlang yes. INTERCAL/Brainfuck less so...
gpjt··on Scientists say they've reversed brain aging in mice with a nasal spray
Hmmm. "Comforting haze" seems dubious. My grandmother had Alzheimer's and at least from the outside it seemed like a bad LSD trip that never ended. She didn't understand what was happening and was scared.

Sample size of one and anecdotal of course, but...

gpjt··on Scientists say they've reversed brain aging in mice with a nasal spray
Love the verbification!
gpjt··on GPT Guesses Between 1 and 100
I was thinking the same. It's a simple idea, heavily over-explained. The code is similar, massively overengineered for such a simple test.
gpjt··on Cooling copper plates could slash data center energy use by 90%
I guess if it's on the inside of a water-cooling loop, you should be OK if the water is pure enough. I don't know how hard that "enough" would be, though.
gpjt··on Eden AI – European Alternative to OpenRouter
Yup, agreed -- it's amazing how close they are getting! I was just wondering if there was some true frontier non-US model that I'd missed.
gpjt··on At least 10 people tied to sensitive US research have died or disappeared
Thanks for the link -- I read that when it was published, then the other day I wanted to send it to someone but I'd forgotten where it was.
gpjt··on Eden AI – European Alternative to OpenRouter
Which are not? There are Chinese models that are only months behind, which is impressive -- but they are still behind.
gpjt··on Ask HN: Who is using OpenClaw?
I have it running on a Proxmox VM. It basically just sends me summaries -- it reads a bunch of RSS feeds I pointed it to (news, tech, etc) and gives me a daily summary, along with an image of the day based on that. It also sends me recommendations for times to go for a run based on the weather and my Strava activity, daily recommendations for stargazing (what's visible and when, weather, etc), and a couple of daily reminders for things I tend to forget.

I'm using Claude as the model, though, so it's smart but pricey. Should configure it to use different models for different things, but it's trickier than I would have expected to do that.

gpjt··on Sam Altman's Coworkers Say He Can Barely Code and Misunderstands Basic Concepts
Wait, that's it? Seven paragraphs, all short? Two quotes, one from some anonymous MS exec? Is the site sending some minimal version of the article to me because I'm using Brave, or is this the lowest-content article I've seen in weeks (and I'm on Twitter)?
gpjt··on Squirrel seen 'vaping' in London park
It's a bit of a double edged sword. As someone who smoked and found it impossible to quit for decades, I'm very happy to have been able to switch to (reuasable) vaping. It's probably added years to my life expectancy.

OTOH the upsurge in nicotine use amongst young people feels suboptimal, and disposable vapes are a scourge.

gpjt··on Squirrel seen 'vaping' in London park
Looked to me like it was trying to work out whether it was edible. Sensible behaviour for an animal in a world where unfamiliar edible things appear from time to time.
gpjt··on After ruining a treasured water resource, Iran is drying up
Not sure that every browser advertises English, but mine certainly does. However, as I'm in Portugal, many websites ignore what my browser says and send me to translated versions, I assume based on my IP. That causes problems because the translations are often quite bad, and they do it with redirects to PT URLs so I can't share links with people who don't speak the language.
gpjt··on LLM from scratch, part 28 – training a base model from scratch on an RTX 3090
Awesome, thanks! I'm still doing trains on the big machines right now (hopefully will write up over xmas) but I think once I've worked out the sweet spot for memgatokens per dollar for this model, it's time to start tweaking the other controls -- LR and cosine variation of it, as you said, and also dropout, bias, weight tying, and definitely gradient clipping (which should at least get better bang for the buck from time/$ spent). I'll leave it to Google to follow up Chinchilla with a "best batch size across a thousand trained models" paper ;-)
gpjt··on How, and why, I invented OnlyFans. In 2004
I think the punctuation makes it clear -- imagine "How I invented Facebook. In 2001." The full stop in the middle of the sentence breaks it and makes you realise he's speaking figuratively.
gpjt··on The Average Founder Ages 6 Months Each Year
100%, I think there were weeks when I aged a year...
gpjt··on LLM from scratch, part 28 – training a base model from scratch on an RTX 3090
Thanks re: gradient accumulation, I'm glad to hear my intuition was right!

As part of the upcoming post I'm running the DDP train on A100s with 40 GiB and 80 GiB, H100s with 80 GiB, and B200s with 160 GiB, so I'll have at least three loss vs. batch size points to plot. So that might be interesting.

I guess a full test would be to train at various batch sizes on the 160 GiB machine and plot the resulting loss. That would be very expensive as a hobby project (the bs=64 train cost a bit more than $40 excluding overhead) so I won't do it.

But perhaps a shorter train would still be of value? That is, train for 300M tokens for a tenth of the cost and see where the loss landed? The problem with that would be if the impact of batch sizes varied with the length of the train, eg. if batch size 64 was better than 512 for short trains but weaker at longer ones.

gpjt··on LLM from scratch, part 28 – training a base model from scratch on an RTX 3090
Hmm, interesting. With a batch size of 512 (8x B200s with 160 GiB each) I get worse results! Maybe there's a sweet spot somewhere in between.
gpjt··on LLM from scratch, part 28 – training a base model from scratch on an RTX 3090
Exactly! If I can get it down to an hour or two (seems very plausible on an 8x H200 with 160 GiB VRAM per GPU, though those are almost never available on Lambda Labs), I'll do the experiments with dropout and the other possible causes of issues, then see if I can bake that all into a new train on the RTX 3090 and confirm it repros there. Looks like I'll definitely need gradient accumulation there.

I assume the zero_grad would need to go in the same if block?

gpjt··on LLM from scratch, part 28 – training a base model from scratch on an RTX 3090
OP here -- with a 112M model you should be able to get something worth playing with using 2.24B tokens. The Chinchilla heuristic is tokens = 20 x parameters. Obviously you cam get a better result by grinding through more tokens, but it will be very slow progress. It's worth noting that Andrej Karpathy is using the 20x thing for his nanochat project.

I try to explain the Chinchilla paper in the post, but your favourite AI should be able to explain it well, and has the benefit that you can ask follow-up questions.

Page 1 of 9Next →