HNHacker News
TopNewBestAskShowJobs

gpjt

1,810 karma · joined January 12, 2009

https://www.gilesthomas.com/
submissionscomments

Extending Raschka's GPT-2: an MoE trained from scratch on an RTX 3090

gilesthomas.com·1 pts·gpjt·
0

Putting my Jax-trained models on the Hugging Face Hub

gilesthomas.com·2 pts·gpjt·
0

Why do OpenAI's GPT-2 weights beat mine? Part four: digging into dropout

gilesthomas.com·1 pts·gpjt·
0

First patient to undergo live AI-assisted brain surgery has tumour removed

bbc.com·6 pts·gpjt·
0

Adding diagrams to my static site generator with D2

gilesthomas.com·3 pts·gpjt·
0

Use the built-in GELU, don't roll your own

gilesthomas.com·1 pts·gpjt·
0

A Quick(ish) Chinchilla Check

gilesthomas.com·1 pts·gpjt·
0

I use AI on this blog

gilesthomas.com·1 pts·gpjt·
1

Why do OpenAI's GPT-2 weights beat mine? Part three: testing overtraining

gilesthomas.com·2 pts·gpjt·
0

Why do OpenAI's GPT-2 weights beat mine? Part two: the bugfix

gilesthomas.com·8 pts·gpjt·
0

Why do OpenAI's GPT-2 weights beat mine?

gilesthomas.com·4 pts·gpjt·
0

Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090

gilesthomas.com·16 pts·gpjt·
0

Building intuition about LLM parameter counts

gilesthomas.com·2 pts·gpjt·
0

Poppy the training box, part 1: the beginnings

gilesthomas.com·3 pts·gpjt·
0

From bigrams to GPT-2, one component at a time (in Jax)

gilesthomas.com·1 pts·gpjt·
0

Building a Jax training loop for an LLM training run

gilesthomas.com·2 pts·gpjt·
0

Thoughts on Role Confusion

gilesthomas.com·3 pts·gpjt·
0

Flax debugging: making a hash of things

gilesthomas.com·2 pts·gpjt·
0

10Gb/s Ethernet: switching to a Broadcom SFP+ module

gilesthomas.com·195 pts·gpjt·
170

Jax: Commitment Issues

gilesthomas.com·4 pts·gpjt·
0

Jax Back Ends and Devices

gilesthomas.com·2 pts·gpjt·
0

Using Safetensors with Flax

gilesthomas.com·2 pts·gpjt·
0

First Looking into Jax

gilesthomas.com·3 pts·gpjt·
0

10Gb/s Ethernet: using mini-heatsinks with a 10GBASE-T SFP+ module

gilesthomas.com·3 pts·gpjt·
0

10Gb/s Ethernet: what I did to get it working in my home

gilesthomas.com·232 pts·gpjt·
177

10Gb Ethernet: what I had to (re)learn

gilesthomas.com·1 pts·gpjt·
1

LLM from scratch, part 33 – what I learned from the appendices

gilesthomas.com·5 pts·gpjt·
0

LLM from scratch (32l) – Interventions: updated instruction fine-tuning results

gilesthomas.com·1 pts·gpjt·
0

How an LLM becomes more coherent as we train it

gilesthomas.com·3 pts·gpjt·
0

LLM from scratch, part 32k – Interventions: gradient accumulation

gilesthomas.com·2 pts·gpjt·
0
Page 1 of 4Next →