I were 17, I'd learn how to build LLMs from scratch
twitter.com
twitter.com
Learning to hack something together in high school using the latest technology (vacuum tubes, radios, microprocessors, web/javascript) has been a common theme in the tech world for generations. With LLMs and online tutorials, this isn't even a difficult suggestion. Do people think learning new tech is somehow wasted effort?
I can't recall or point out exactly when, but there is a stark before/after moment where the opinions of anything pg went from "Interesting and maybe true in some ways" to what we see today, lots of knee-jerk reactions and hardly any comments about the actual content.
Hazarding a guess, I think the moment Altman became the CEO and later during COVID, the sentiment seemed to have been shifting towards what we see today. But this is all based on hazy memory, rather than looking at the data. I'm sure there is a blog post waiting to be written about analyzing the sentiment of comments to PGs articles on HN, and you'll see a shift somewhere.
Hard disagree. This submission is still being highly upvoted, while another recent post[1] on the harms caused by Graham’s fellows[2], with a fairly tame comment section, has been flagged. That is a constant on HN. It’s not a fluke, it’s as predictable as the sunrise and getting more pronounced.
I’m sure we’re both biased in our perceptions. Mine is that HN in general (certainly more than any other website) used to worship[3] everything he wrote, together with others like Musk, until things started to really go to shit and many eyes have been opened to the effects of the unfettered greed of rich tech guys out of touch with reality.[4]
[1]: https://news.ycombinator.com/item?id=49411762
[2]: A better English word is escaping me.
[3]: That word I choose hyperbolically but deliberately. It definitely was not “interesting and maybe true in some ways”, it was much more hardcore than that.
[4]: That is not “knee-jerk” but a slow realisation still ongoing.
The sycophancy on HN is starting to break down because there is a higher proportion of users sceptical towards the outputs of the VC and wider investment world than ones who believe they're potential beneficiaries of it.
Tech industry people are becoming less interested in HN as a warm handshake into the startup world because, frequently, they're disgusted by it. And this reflects on the sentiments people post on PG's articles.
Increasingly if those at YC want the same kind of low-bar praise they got before, they will need to get it from machines.
This might be hard to disentangle from the tech correction layoffs of 2023 and 2024, but my perception is that most of the spite comes from Tech Outsiders opposed to jaded Developers.
Don't many of the commercial ones prevent you from using them to build LLMs?
I would say the reason for the negativity is not because it's a bad idea for a project, or that doing projects in general is a bad idea (it's not!), it's because it's a very specific thing that is not for everyone. The best thing about computing is the low barriers to entry. You can basically work on anything that takes your fancy. So those who are interested in ML will be drawn to learn about LLMs. They don't need anyone to tell them to do it. Telling everyone to do it reminds me of the "just learn to code" stuff of a decade ago. No, please don't, please find something you enjoy.
The above is, after all, the whole genesis of the word 'hacker'. We should celebrate that.
I tried to modify the embedding output of bert to make it generate box embeddings instead of point ones. At the time I had access to university provided A100 gpus but even with all that a training run took half a day. Models these days I don't think I can train it in any reasonable time with that much compute.
(As a TML person, I'm obviously biased, but I couldn't resist because of "tinkering").
TBF it's hard to imagine a real architecture change that wouldn't require a ton of compute, but you could certainly fine tune and play with different recipes, loss functions, etc. And Claude can carry you a lot of the way through doing this.
One fun task is to invent a tool and then train a small model to use it. You could export that small model and run it locally for free forever to do your thing. I think this is what a lot of Software Engineering will look like later.
There are a lot of other high level abstractions here to look at. Prime Intellect has one.
The other thing to play with is self-hosting small models, but IMO most of the interesting stuff is actually related to multi-gpu or multi-node inference so there's not necessarily a ton to learn here.
So, you build one from scratch.
The best analogy is strip mining (big labs) vs cave exploration (solitary/small teams). I think this is how science progresses at the boundaries by smart/curious/hardworking individuals because depth is a requisite for finding the right questions and then the answer. It is not for everyone and it does not always work. But you learn a ton even if it doesn't pan out to be a big breakthrough.
https://ravinkumar.com/GenAiGuidebook/book_intro.html
This guidebook covers pretraining, post training (SFT, RL) and a couple other topics. And others authors have also written books that fit on single node reasonable hardware.
If you want to start with a pretrained base I built Gemma 270m and released it last year. This fits on a raspberry pi.
https://developers.googleblog.com/en/introducing-gemma-3-270...
The fundamentals of AI don't require industrial amounts of large scale. Think of it like this, when I was learning how a plane worked when I was a kid I didn't build a 747 at home, I started with scale sized model planes. Same idea here.
And FWIW I'm a staff researcher at Deepmind (opinions here are my own) so I want to specifically encourage all people out there, you can learn a lot about how these LLMs work at home, for (mostly free), using resources like colabs or spot pricing on accelerator providers. There's many great resources out there and I encourage anyone willing to learn to go for it!
No. But funnily enough that is a promise by some of the AI cretins and their boosters. Oh yeah best case scenario you learn how to build LLMs for us. We’ll employ you. And then ultimately that just becomes training data for the LLMs to do it themselves.
But why are people cynical? they ask.
Looking back, it was an ideal start to a 40 year career as a software engineer. And of course currently I'm reading "Why Machines Learn: The Elegant Maths Behind Modern AI" because the urge to understand the machines I want to master and control hasn't gone away.
Good. It took a long while for ICs to be paid more than 'management' in the US. It's still a hard-fought battle here.
// a flood of coders in the '10s-early '20s whose only creed was FYPM and "what's the minimum
I wonder what possibly could have happened in that timeframe?
https://www.theguardian.com/technology/2014/apr/24/apple-goo... https://en.wikipedia.org/wiki/High-Tech_Employee_Antitrust_L...
Why the everliving fuck has everyone stopped building tools?
Sadly, I attribute this to a decline in having “good will” towards others, and an increase in cynicism, prejudice, and meanness.
It would be a good idea for young people to deeply know how these programs work. Not so that they can spend their career building them, but so that they can approach the next class of problems we'll all start trying to solve, with intuition all the way down to the weights and underlying mathematics. And also, to develop a healthy intuition of when "Just LLM it" will not be the right choice.
"Build an OS" wasn't a common university project because we were all expected to go out and work on Windows, but because understanding the bare-metal firmware for a computer helps you deeply understand how to intuit building for a whole class of problems.
Sure, but it was for a specific degree with a syllabus that taught you the foundational knowledge. It was not expected from the law students to learn how to build one.
IMO this is still relevant, everything surrounding the LLMs needs such a vast infrastructure that I don't know if I would find it more useful to learn the maths behind ML than CS
Well it is framed as quite specific advice.
(I'm done with mining PG tweets for meaning)
It would be a good idea for _everyone in the industry_ to deeply know how LLM training, inference and "agents" work, not least because it removes the ability of shysters to bamboozle with bullshit.
But, as much as a good idea it is for the young to understand this, it's the elderly who will be really taken advantage of if they do not keep up - just look at Facebook for good examples of why.
But the load-bearing assumption is that intuition has to be human-shaped intuition. Humans can't intuit thousands of orthogonal directions because we project everything down into a 3D metaphor and hope it holds. That's a fact about our hardware, not about the systems.
And the reason why is the most interesting part: nothing requires the compression step. A model or an agent can operate over the actual objects, holding thousands of runs and ablations in context and noticing regularities in the native dimensionality, without translating them into a picture of a ball rolling down a hill. No bottleneck at "can you visualize it."
So the narrower claim: it's not that intuition here is impossible full-stop, it's that human intuition is unreliable. Your post-hoc rationalization point is evidence for that, not against it. The story exists because a person needs something to hold in their head. Drop that requirement and the failure mode goes with it.
Linking to a Twitter thread about training a small LLM to be good at Wordle is not an example of what I'm talking about. It might well be a useful task but it doesn't allow us to understand deeply what's happening.
Now with llms we need way more that 6 dimensions so we can start thinking of assigning matrices to each 3d point for example. That allows us to increase the dimension from 3 to 3 + whatever the matrix dimension is.
We can visualize the matrices instead of having numbers as having colors for each entry, so they can be a sort of cube with each vowel being a different color.
Now you can start to visually intuit about how these massively high dimensional spaces can be formed of these colored matrices that can react to some input training data.
That's a start of an idea for intuiting things that might seem impossible to have an intuition about. I think visualization is a great way to start.
visualizign 100 dim matrix will not tell you how llm work. so what you even talking about.
Depending on your knowledge of math, I recommend starting with linear algebra, building an understanding of the equations and try to visualize more and more complex systems, then study llms to see how you can apply your linear algebra intuition to your understanding of llms.
VTK is a great toolkit for visualizing complex systems. 3 blue one brown on YouTube has other visuals that might help you.
It takes time but it's possible. Good luck.
> Visualizing large dimensional matrices is a way of starting to develop intuition about llms.
one last time before i disengage. how do you know this and what intuitions have you personally developed.
I've spent a great deal of time studying llms and NN architectures in general. This allows me to intuit things about them and that intuitive understanding for my comes mainly from visualizations.
Are you able to visualize much about llms? I ask because if not, starting by learning to visualize high dimensional complex systems is what I would personally recommend.
Check out this video for a really nice visualization into NN architecture https://youtu.be/QgH9sr7G13Q?si=iTn1FiFsYCZCJKE_
Like a cat who knows how to teleport so rapidly around the room that he becames a blur. All the while taking into account stuff falling around he knocked down.
As you say, there's a lot of value to unlock with understanding the generic principles, and I would add specific application.
Lots of people are building fantastic Tools or pulling down million dollar salaries without groking the precise representation of a single weight.
I'll borrow my response from a far greater applied mathematician than myself: https://www.youtube.com/watch?v=_oNgyUAEv0Q
Most software engineers do not have a deep understanding of CPU architectures. In fact they probably don’t even have a shallow understanding and get around just fine. How many of them are looking up the instruction set for the CPUs they deploy their CRUD app to in EC2?
Basically every part of the original transformer was replaced with something more efficient or better:
LayerNorm -> RMSNorm
Sinusoidal position encoding -> RoPE
MHA -> GQA
ReLU -> GELU
What this means is that there is ample opportunity to improve on what we’ve done thus far.
Yes you might progress the field, but will you really understand why? You can make up an explanation and anthropomorphise it with a few contrived diagrams and everyone will cheer!
Lot's of neat stuff to learn.
From a biology perspective - we have a good understanding of the carbon atom and how forces influence it. We really don't know much about a cell.
Absolutely none of that mattered in the long run because capital markets think environmental destruction is a small price to pay as long as some people get rich.
What kids learn about personal cooling tech and what not won't matter when the supply of the raw materials required to make them is reserved for some VC-funded trillion-dollar 'startup' trying to synthesize an undetectable chemical weapon to get past the bureaucratic red tape of international treaties.
Some people will be delighted. Most are going to have a very hard time adjusting.
YC and adjacent isn't it.
I didn't do it because it was useful to me in a practical sense. It's because LLMs are fascinating and I want to know how they work. From that perspective it's been a great experience. I have afirm grasp of the basics. This makes it much easier to understand frontier concepts like compressed latent attention. I can follow the field and understand it.
Not sure I would have got as much out of it at seventeen. I have a lot of background and experience which made it much easier to learn. I wasn't struggling with the linear algebra or with python. I already knew pytorch and neural networks. That helped a lot and I covered these tutorials fast and could skip over large sections. A few evenings and the odd weekend day over a couple of months was enough for me.
For seventeen year olds the tutorials are good enough to make it possible to learn this but it would have taken a lot longer to understand. On the other hand I would have learned a lot more. I think I would have learned a lot of valuable stuff.
However I also think 17 year old me was studying for his A levels and probably this was right choice in terms of maximising future opportunities. I'm not sure I think learning about LLMs instead is sensible. Indeed it might be bad advice. But I can absolutely agree with the sentiment.I think 17 year old me would have wanted to do this too.
I would suggest starting with Andrej Karpathy's YouTube video: https://youtu.be/kCc8FmEb1nY?is=oiDsrBYJg_MUUmoD
This video is excellent. I'm a huge fan. Also the video is zero commitment and instantly available which makes it a good way to check you are interested.
The book by Sebastian Raschka is slightly less accessible but very reasonably priced and the experience of working through a book is a lot nicer than skipping back and forth in a video (for me). Sebastian's blog posts on recent architectures are absolutely great too.
http://languagemodelbuilder.com teaches you (in a few hours to days) how to build an LLM from scratch. It's entirely free, without accounts, and without data collection.
Thanks a ton for building this.
Putting this in contrast with programming, I learned coding when I was 8, and it was incredibly stimulating to learn because you can quickly iterate and there were thousands of books and YouTube tutorials that dumb everything down and teach you fundamentals. All you needed was a $300 computer, and you can learn nearly anything you want, without being gatekept from this or that because you don't have enough vRAM / an sm_100 GPU.
There are other types of models like diffusion models right now that are showing more efficiency and have a higher ceiling for improvement. Understanding math and fundamentals are more important.
I fail to see how PG suggesting that a 17 year old, today, should study something... is in any way a parallel to formalizing it as part of the curriculum.
I would, in fact, expect PG to recommend something completely different in 10 years.
The whole point of hacking is to understand the world around you, today
But of course, 10 years ago this wasn't obvious.
Or more reasonably, the same old thing : use Linux, hack a little, why not learn programming basics. But learn to own your technology, fight against centralization of technology. The same old RMS story.
IDK where the tech industry is going, if there will be jobs anymore or not, but what I'm sure (and what have been the case for the last 10-15 years anyway) is that for most tech jobs, having good technical knowledge beyond the basics is pretty useless and will probably not be recognized.
If you can, stay a computer geek if that's your thing, but don't make it your career choice, the Eldorado is behind us.
> why not learn programming basics
Do you think implementing an LLM from scratch would not require programming!?
A generic project based goal like this is the best approach to self learning, because you fight your way to it, picking up what you need to know on the way. And, once it's all working, and you get your first meaningful sign of success (in this case, some tokens out), after your long path, you fucking celebrate!
I really hope we get another hiring boom like in 2020 when I decided to study CS. Otherwise my career will be very rough. I love it and can't imagine doing anything else.
who would have guessed when first world were enjoying their heyday as they were colonizing the world, and teaching everyone and their mother to learn English.
Kinda sad now you have to compete with serfs from thirls world. chu chu chu. so cute.
shouldn’t have robbed others then. a richer world for you comes at poorer world for someone. we are playing infinite game in finite world.
Young people are not responsible for the crimes of their ancestors centuries ago.
The reason remote worker can out compete you is because of this currency arbitrage you designed. Your pennies goes a long way in this so called third world. Your one day dinner at nice place is a months salary for family of 4.
Seriously, if you can’t compete globally, then you don’t deserve your privilege of being born in rich world.
Stop labelling other humans as serf when your entire wealth was built on shaky foundations.
You == not you the person, but the government and laws.
The fix would be obviously to control and limit foreign workers by regulating companies.
You should really read up on disruptive innovation. Many of those who you call "serfs" who'll "take 1/3rd of my pay" will tomorrow establish themselves, justify their presence, and move up the wage/income scale, while you'll be left complaining.
I am from and live in a third-world country and have spent some time in a first-world one. I've seen the ETH Zurich tag open doors that I didn't even know existed. If you aren't leveraging your passport, your education and your network to reach a place where you don't need to worry about "serfs", then you're doing something seriously wrong. What you have, compared to the competition you denigrate, is invaluable, making best use of it is your job.
Bad advice. You’re going to need family connections and loyal people you can bet your life on to survive if the rest of what you’re saying is remotely true.
We never imagined that we'd live to see Snow Crash depict a utopia by comparison to what we got.
What? You want me to tell kids not to be ambitious?
Watching Karpathy's "Zero to Hero" won't make you AI specialist. It will improve your understanding of the basics.
That said, I a good starting point for a 17yo is reading about perceptrons[0], then the basics of neural networks[1] (ex. 3-layer perceptron) then writing a program to train a 3-layer perceptron and classifying the MNIST dataset[2] - a dataset of characters.
This can anywhere between a day and a week and you will demystify the basics of neural networks and work your way forward with more advanced contemporary concepts.
Fun fact: any multi-layer perceptron neural net can be reduced to a 3-layer perceptron network.
[0] https://en.wikipedia.org/wiki/Perceptron [1] http://geeksforgeeks.org/deep-learning/neural-networks-a-beg... [2] https://www.kaggle.com/datasets/hojjatk/mnist-dataset/data
I’d probably say something like: do something you enjoy and seems like it might be useful, but accept that the pace of change may mean that whatever you study ends up being irrelevant.
Whatever solution there ends up being to this, it’s not going to be one that an individual 17 year old can implement. We’re past the point where individual good and bad choices matter that much to economic outcomes.
Telling other people's children what to do is easy and basically doesn't have any downside to being wrong. With your own children things are a bit different.
So: what are people here with school age children telling their own kids about the future? If their kids ask, what kind of careers would they encourage them to pursue, assuming they have the skills and interest?
When I was last in the Bay Area, maybe about a decade ago the bookshops were full of titles like "Python for Preschoolers" (I exaggerate, but only slightly). Clearly at the time a lot of people working in tech thought that cultivating an interest in programming was going to be the path to being a successful (by some metric) adult. Is that still the case?
In this regard, Healthcare is a polar opposite. It's pretty hard as a male nurse to not "accidentally" become a home wrecker.
Being a super rich and an unhappy workaholic, or a super-impressive engineer who wakes up one day at 45 and realizes they regret wasting half their life (I ran into way too many of these) is a much worse fate than "not being rich from your startup" and working a relatively regular job while feeling fulfilled and happy by more than just work.
Especially in the US, which is uniquely bad at this and encourages people to work themselves to death, mental health and work life balance are much more valuable things for 17 year olds to focus on than finding good startup ideas.
In case you think i'm being a bit dramatic, let's look at the state of 17 year old mental health in the heart of Silicon Valley:
"The City of Palo Alto and the Palo Alto Unified School District approved a funded contract to place 24/7 human security guards and monitors at all four local Caltrain grade crossings, including the Churchill Avenue crossing directly adjacent to Palo Alto High School."
(in case it's not obvious, it's because of suicides by high school students)
The 17 year olds do not need advice on better startups, and this situation will never get better if we focus our advice on how to be better at work instead of how to be better at life. This will require redirecting the conversations.
I'd move to the middle of nowhere and work multiple jobs on a farm and in construction. Learn how to grow food, and build things. Meet the farmer's daughter, and marry her. Then, buy my own land, grow my own food, and build my own things.
I started investing in farms, have 50 pigs and 100+ chickens now. We are planning to grow to 100 pigs and 2000 chickens in a year. We will start growing Shiitake mushrooms in a few months too.
Nothing wrong with learning the theory and understanding the papers. Getting to that point you’ll have to get your fundamentals down. Might be an interesting exercise.
But as a future? I guess we’ll see. I suspect the next financial apocalypse will determine if there is one. Another AI Winter that may outlast all others so far.
The few exceptional individuals will innovate what the millions will benefit from.
At 17, the mother of Isaac Newton removed Isaac from school and tried to make him a farmer. We all know that wasn't his destiny.
You could also walk out your front door and get hit be a meteor tomorrow.
On an Amiga, I took various public domain text documents from cover disks and counted the probability of the next word given the previous word. Then spat out random sequences of words from it and printed them out. It was called "Splurge". Basically a very very simple single layer statistical language model.
Some of the sentences were randomly not bad sentences, which seemed amazing at the time!
That kind of thing (and Core Wars and Tierra etc) did lead me to getting a job at an artificial life startup at the end of the decade. But that was in turn about 10/15 years too early (no GPUs).
There's some lesson from this about timing, but honestly I've gained the most as a person when I did something that was fun, ethical and gained an audience. A tricky combination.
None of us know what the future of work, education, or AI is going to look like. But your best bet is to become a life-long learner. Be it LLMs, musical instruments, physics, or business.
Learning AI isn't like learning HTML in the 90s then expecting to get a job at a tech company building websites. You can't just "learn how to build LLMs" and expect a frontier lab to hire you so I'd argue this is rather bad advise.
Additionally, unlike web development in the 90s you cant really do anything interesting yourself... All of the interesting/useful stuff will require huge amounts of compute and data so there isn't even much point in learning to start your own thing either.
As someone whose built many of NNs from scratch (hand written code, long before the days of LLMs), it's more or less useless knowledge if I wanted to work in a frontier lab or do anything interesting in the field.
I also think anyone thinking about going into a field which is basically a crossover of CompSci and Maths is absolutely insane right now. Even if you think there is a place for CompSci and Maths post LLMs, there's almost no chance anything you learn today will be relevant to the skills required in say 5-10 years.
Moreover I am not sure it is even good advice? Would you advise a 17 y.o. to learn how transistors work or how to code (i.e. is LLM training the right level in the stack)? LLM training, a discipline where relevant work is already out of reach for 99.999% of budgets really as essential as this post implies?
Personally I don't see the problem, as long as you're aware there is survivorship bias involved here.
What's the alternative really, seek advice from unsuccessful people? That seems worse :)
Personally I do both, read about what worked for people, also read about what didn't work for people, then ignore both and do whatever the fuck I want.
Intuitively, I would guess that they have a better grasp of what made them fail than successful people have of what made them succeed.
Nonetheless, there are many successful people I would gladly listen to for advice, though they are often successful in a different meaning than what venture capitalists would use (e.g. parents with great kids, managing to keep a healthy work-life balance, happiness, and maybe even having time to spend on some cool hobby project -- you are heros!)
What bucket should I put this advice in?
But some advice is less specific than other advice though. Some things are always stupid, and some things are always smart, if you look at the context of our world and society. I find myself pursuing these "fundamental truths" with great interest lately, especially now that the world is changing so quickly.
That’s what you get from listening to “successful people”. You get to learn about all the things they tried that failed, then the things that did work on that 24th try, which was successful.
The “survivorship bias” people always seem to assume that the “survivor” lucked into his fortune on his first try ever, so he can’t have learned anything, so we don’t have to listen to him. But that’s seldom the case.
I’ve written about this before:
There are successful people that failed many times before being successful. Somehow these days, previous failures and persistence seems to be ignored and they focus just on the luck you got on the 20th try.
Seek advice from the averagely successful people, since that is statistically what you're most likely to be.
I know who the 17 years old is closest to.
If you can't do that - not sure anyone can help you.
In 2000 (his era), it would have been really smart to study the source of Linux or Apache. Would have paid dividends over decades. Cuz that knowledge was so rare. The number of people hacking on LLMs now dwarfs the number of people hacking on web servers 30 years ago, by several orders of magnitude.
And if you turn back the clock even more, I mean just even having access to a computer, let alone owning one, would have put you at a massive advantage.
I don't know what to call it. The pioneers should be respected obviously, but at the same time you need to understand that for them, the game wasn't nearly as played out as it is now.
I just don't think you can afford to be dicking around with LLMs like you could afford to dick around with random Linux distros 20 years ago. Too many people willing to do it for free these days.
You don't wanna end up being the 2030 equivalent of a certain SNES emulator developer, or maintainer of a package manager for jailbroken iPhones, I mean the list goes on and on. Being a hacker doesn't automatically give you a path to being rich, or even making a decent living. It hasn't been that way for a while.
(Edit: And learn how honest business works)
I think the last one was seeing a skilled electronics repairman do surgery on a CT machine controller.
>"Someone asked what I'd do if I were 17. I'd learn how to build LLMs from scratch..."
That's funny, Paul Graham, because if I were 17 again,
I'd learn how to program in LISP.
https://www.paulgraham.com/rootsoflisp.html
https://www.paulgraham.com/iflisp.html
https://www.paulgraham.com/hundred.html
(And/or other LISP derived languages... Clojure, Scheme, Racket, TinyScheme, etc.)
I guess "the grass is always greener..." as that old expression, that old "chestnut", goes... :-)
I first studied them in 2011, and I was like what? Just a bunch of partial derivatives?
I keep looking at AI to check if now it's something else but it keeps being gradient descent.
Okay, it's great that you can perform miracles using gradient descent but that doesn't make it captivating in any way.
I'm much older and less wise now, but I still afforded myself the opportunity to follow karpathy's tutorials to build a LLM from scratch. Got to play with a few ideas. Seen similar ideas turn up in frontier model work, which is quite gratifying.
There are so many ideas to try.
Currently playing with autoencoders that takes A and B and produce latents A', B', and C'. Reconstruction of A is from A' and C', B is from B' and C'
The idea is if C' can be made to improve both outputs, it must store as much information as it can about what is common to both inputs.
Also many people/kids don't have access to proper "productive" systems anymore, since the whole computing and electronics industry shifted to make "consumer"-devices like smartphones or laptops made for netflix, gaming and spotify.
Breaking the barrier to build a custom system, install linux (or developer tools for Windows, MacOS) is already a complex AND costly task. It was just way simpler in the late 90s and 00s to get something working.
There is very little reason for humans to get all too engrossed in this type of work now, today, with the hope of being good enough at it to command a high salary in 3-5 years. AI can already do it incredibly well, and they can do it persistently and doggedly 24 hours a day.
And younger people I meet don't even own laptops. I had a genz/millennial cusp friend who wrote all her college papers on her iphone.
The reality is that an incredibly small minority of companies in the world do any real training or optimisation. It's unnecessary and inefficient for most purposes unless you are fully dedicated to being an LLM company, and still then it's a struggle. Those few that do train, they spend most of their budget on compute and have relatively small teams.
Getting experience in this field requires having access to very expensive hardware to begin with. And the skills will be quite hard to convert into any real value for someone, leading to a decent income, unless you have a ton of funding from patient investors, or you have decent contacts in Bay Area networks to get hired at the right place.
With all due respect, paulg is in somewhat of a bubble, this is not congruent with the global situation.
In reality the problem is that it gets blasted out of the water by a much worse architecture trained on 10000x the infrastructure. And while I'm sure the freshly brought in ML student came up with a 10%, even 30% better architecture, it just doesn't matter. (and never mind that even OpenAI hasn't really solved a voice model yet. Try it. It can probably match 2026-quality call centers, but it's no substitute for an actually empowered human)
... and yet, if you look at what hyperscalers are getting paid for ... comfortably more than half the income is training. Which makes no sense on so many levels.
e.g. https://valueaddvc.com/blog/inference-chips-vs-training-chip... (I get it, not great first source, but st
There are plenty of areas were we need people to do this for insurances, banks etc.
AI/ML exists on many levels.
The only jobs that he found he was highly over qualified or paid very little.
In any case, it doesn't look like there's this crazy rush to hire all ML talent, even the one that understand the math and technology deeply.
Maybe people simply don't want math PhDs but something else? Since 1-2 years ago I started doing consulting/freelancing in the ML space, but more on the infrastructure, deployments and similar stuff, as a general purpose developer, and I have a waiting list of clients interested in more work, some of them even trying to recruit me to work for them full-time as well. I'm based in continental Europe, fwiw.
You’re missing the point. Understanding how Unity works fundamentally makes you a better Unity dev.
It is viable as a toy project, but there are vanishingly few career opportunities.
Good luck with that approach when trying to toy around with models and their training/inference.
There is also a lot of math basics missing that a 12 year old may be able to grasp, but I would bet they are at least 13 by the time the knowledge is deep enough to understand what operations are happening.
[Blake Ross] worked as an intern at Netscape at the age of 16 ... Ross became disenchanted with the browser he was working on and the direction given to it by America Online, which had recently purchased Netscape. Ross and Hyatt envisioned a smaller, easy-to-use browser that could have mass appeal, and Firefox was born from that idea ... in 2003 all of Mozilla's resources were devoted to the Firefox and Thunderbird projects. Released in November 2004, when Ross was 19, Firefox quickly grabbed market share ... with 100 million downloads in less than a year
https://en.wikipedia.org/wiki/Blake_RossBrowsers from scratch are multi-year projects for multiple people. Even just skinning and minor tweaks to modern browsers is a deep well for one person.
Building a browser in the early web was actually a very achievable goal for exactly the reasons it isn’t now. There was not JS. No CSS. No SVG. In fact very few widely supported image formats (and graphical browsers weren’t around in the earliest days of the web anyway). TLS didn’t exist. HTML only had a subset of methods. And even POST was usually just managed by CGI/BIN calling an external process, often written in C++ or Perl.
It was a simpler time.
That’s like saying someone who fabricates cars doesn’t have the skills to drive them. Perhaps not, but they’re very well placed to pick it up quickly. They’ve also shown they can do something far more challenging, which is actually better than hiring for narrow immediate skills.
In other words: I’d hire that candidate in a heartbeat.
Oh hey thank you for that. It really helps. Hope you have your rug pulled from under you today too.
— signed, a career changer trying his best.
But that was way back in the early 1970's and all I had to work with was a mainframe.
Well the mainframe itself wasn't bad, the real show-stopper was that I didn't own the computer outright, no strings attached, no debt, etc.
>I'd probably try to make an LLM that I could use on some specific problem.
I thought so too back then, still do so I guess this is one of those things that could stand the test of time. I always wanted to start with something a lot simpler than a Moon mission myself. At 17 I already had a significant breakthrough in the chem labs and it was from alternatives to a single processing step plus everything that descended from that, rather than trying to tackle a much more complex detailed multi-step synthesis. I was only 17 but I was not trying to be a slouch, I don't think pg was either at that age but his advice is not for just anybody. I couldn't have done it if I hadn't made major progress since being 16, and it really emphasized at the time how much maturity can make a difference. My imagination ran wild as I extrapolated :)
In a reply from LeCun to pg:
>>I'll figure out a set of methods and architectures beyond LLMs that can quickly learn to perform physical tasks as efficiently as humans and animals. That last item is also what I would if I were 30, 40, 50, or 66 years old
I see no reason to stop at 66 either ;)
But I figured that people owning more computer power than I could ever afford were going to be doing something like this as soon as they could, without having to wait for something like an LLM to arrive before getting peoples' attention.
It did seem like things were going to take longer than you expect, so it's pretty good to have a lifetime of concentrating on the specialized natural science domain expertise, focused now for 50 full years on how it would combine if AI ever got good enough.
Both the natural science and the AI need to be a major cut above, I still see dramatic room for improvement in my own work. If I'm going to have to rely on "other peoples' AI" then that natural science component is going to have to pull a lot of weight to keep up with the kind of computers that only rich-as-hell high-rollers have access to.
I just finished fine tuning Gemma e2b for local code completion on my local machine.
This comment just reinforces what the post actually means. We need people that are LLM natives, computing solves itself with time and with scale adjustments
For anyone who wants to dork around there is https://github.com/rasbt/LLMs-from-scratch which is something amazing that I think anyone who wants to engineer things around LLMs should at least blast through and read.
And so I think the idea is more to understand tomorrow ... from first principles.
In the late 80's, as a teenager, I learned x86 assembly and C because that was the only way to squeeze out enough juice from my shitty CGA (and later VGA) card to programm the games/graphics that interested me.
I haven't written assembly in years.
But whatever I did in my career: it helped me and gave me an edge over my peers to have a foundation that is very close to the metal.
Current AI can automate significant amounts of grunt work in programming and math. It's good at running web searches and writing summaries. There are a few other niches where it is currently successful. But other than that, many corporate AI projects are spectacular failures.
So just given what we have in hand, assuming no further breakthroughs, then we're maybe looking at AI being somewhat bigger than the Internet. Which would make it a revolutionary technology, sure.
But to get from "a revolutionary technology" to "the substrate the future runs on", then you need to assume more breakthroughs: long-context operation over weeks or months, displacing human workers 100% instead of 75%, and the ability to directly economically compete with actual humans. And people are investing literal trillions of dollars to make that future come true, without really thinking through what truly competitive-with-human AI would actually mean. We might be looking at massive job loss, centralization of power, fully automated "companies" with no humans dominating markets, and other dystopian scenarios.
And in those worlds, it's unclear that being good at CUDA and matrix math will be all that helpful, careerwise. The AIs are already pretty good at that stuff. Data scientists get paid OK when they actually get hired, but it's not everything college students were promised in the 2010s, either.
We can't yet build a fully-general competitor for the human mind. But we're getting closer. And if we ever do build one, the consequences will be really weird in any number of ways. So I worry about visions of the future that assume AI keeps improving significantly, but that also assume it still somehow remains a "normal" technology that doesn't, for example, render most humans fundamentally uncompetitive.
We don't know this. So many people are simply claiming this confidently, and a lot of them are betting their careers on it, but nobody has a crystal ball. Whenever someone tells you confidently, and without any doubt or qualifications, that something "is the future," be skeptical.
I remember when the Segway was definitely going to change urban planning worldwide.
AI is a great solution in niche areas but generally doesn't make much money. All the large companies are in the negative.
The steam engine was less of a bubble and was much more revolutionary and had a greater impact.
Incidentally, the skills for the lowest levels of LLMs aren't that far removed from those needed for mobile telephony, in that both are based on maths, computation and information theory.
What most people want from mobile technology is for it to work, not too expensively, and for it to get out of their way.
What most people want out of AI is for no leader to emerge and wield supremacy against the rest of us. People are afraid of it in ways they weren't afraid of mobile, so they're more willing to work together against whoever is in the lead.
Its more like an arms race and less like a utility. The disadvantage I face when my competition has better mobile coverage and bandwidth is minor. The disadvantage I face when my competition has better intelligence on tap is much more significant.
No.
That’s what people like us on HN want. The people out in “Greater Userland” just want the black box to answer their questions. They could care less who is behind it. They don’t yet attach their black box to Amazon or Microsoft etc. And most won’t care enough to be inconvenienced even when they do make the connection. (As your competition argument implies.)
Heck, a lot haven’t even made the connection between the black box that gives them answers and data centers. They think, “ ChatGPT good” and at the same time think “data centers bad”.
And the job postings are often ridiculous. I recently was an AMD job advert in Germany for an ML Kernel Engineer, not Senior mind you. The requirements went something like
> Masters Degree required with strong preference for a PhD with peer reviewed articles in {journals_list} > 10+ years of experience in C/C++ > GPU programming experience required > 10 more ridiculous lines
No idea how a teenager self teaching himself LLMs is supposed to even get a shot...
1. I can make turn a stone and water into a delicious soup
"17 year olds, learn to build an LLM from scratch"
2. This soup would be more delicious if we add a few carrots. Does anyone have carrots
"Increase your chance of success by getting a Masters degree"
3. How about potatoes?
"And get a PHD"
4. What about some salt?
"And publish some peer reviewed articles in {journals_list}
5. We should also add beef
"Now work in the industry for 10 years"
6. See, this soup is delicious, and I made it all with a stone
"See, you're rich, and it's all because you learned LLMs as a 17 year old"
Finetuning model is cheap and incredibly useful for deployment. You don't need to pre-train a frontier llm from scratch to make useful models.
There is tons of domains where you and fine-tune llms and deploy them for value in companies and for your own entrepreneurship ambitions. I have made this a big part of my career for the last few years and now I'm working on finetuning models for starting my own companies.
I feel like the "ALWAYS HAS BEEN" meme is apropos here.
-- 3Blue 1Brown's Neural Network Series: https://www.youtube.com/playlist?list=PLZHQObOWTQDNU6R1_6700...
-- Karpathy on LLMs: https://www.youtube.com/watch?v=7xTGNNLPyMI
-- Stanford CS336: https://www.youtube.com/watch?v=JuoVZkPBiKk
Then do this hands-on:
-- Karpathy's zero-to-hero: https://karpathy.ai/zero-to-hero.html
It similar to understanding how a very basic CPU works. Just because I'm not going to work at intel or nvidia or whatever optimizing the hell out of a chip, it doesn't mean I just throw my hands up and think "magic" - the basic architecture isn't that difficult, and the value of knowing it is astronomical for anyone writing software.
Building an LLM from scratch has a hard split between a tutorial project you can complete in a weekend (that's useless for actual usage) and then a solid 1km high brick wall if you want to create anything actually useful from scratch.
Modified open models have a very active community around them, without the need to look much further than Hugging Face.
Benjamin: Yes, sir.
Mr. McGuire: Are you listening?
Benjamin: Yes, I am.
Mr. McGuire: Plastics.
Benjamin: Exactly how do you mean?
Mr. McGuire: There's a great future in plastics. Think about it. Will you think about it?
---
I love this scene because it so perfectly captures what it's like to be young and given advice, however well-meaning, by an older generation living in a world that no longer exists for the young. And it's ambiguous and trite enough to be essentially useless even if the underlying idea isn't terrible.
Good times.
I wonder if you can dig that out of the historical reddit database. I'd like to see that again. I love how everything is recorded now.
For the AI part, I recently started working through this CMU course on Modern AI which teaches you how to build an LLM from scratch. It was posted on HN a few months ago and has been great: https://modernaicourse.org/
If anyone has other recommendations for beginners, please share them!
Long-term there's probably a lot more potential, and much bigger markets, in robotics than in LLMs, but it will really take the long term to get there.
There's also this little-known concept called learning things for learning's sake and not always trying to capitalize on it.
Is it something that's going to be fundamental for future technologies? I always plan to learn, but end up in a death spiral feeling like I'll invest so much time and energy only for the puck to have moved somewhere completely different.
Would gladly accept any advice :)
I mean first that is already what plenty of 17yo are actually doing, because that is what they do at school or in parascholar activities. There are already countless of such tutorials where you can do that in an afternoon.
The pointless part though is precisely why Amazon and others are hunting for rare books, all the low hanging fruits have been picked already so just training a bigger model will simply mean burning more energy and money. Sure training a small one for the basic principle is a great pedagogical thing, training another one, medium, then maybe a large one, is also good in term of learning the process and architecture, but one should not expect it to be useful out of that context.
Pure players are precisely doing everything they can to corner the market by making their own scale unreachable by others. Smaller players with access to lesser infrastructure are thus betting on different market, e.g. embedded systems.
17yos should definitely build their (L)LMs from scratch and whatever bigger model they can train for free, or for cheap, but they should not expect that to bring them any riches.
I'm sure they'll get right on powering up their computer from the hamster wheel, Paul.
My heart goes to all the kids out there that didn't get the fair shake let alone fair access to tech that gets these condescending "learn to code/learn to LLM" bootstrappy talks from rich pricks that don't know what life really can be like for a lot of American kids out there.
2. LLM from 0 to Hero, and nanoGPT by Andrej Karpathy
He's seeing a future for models running on everyday hardware just capable enough to do what the use case requires.
Agi is the academia solution Software is the practical solution
I spent a days reviewing the lecture notes for CS336: Language Modeling from Scratch - and then trained a nanoGPT-esque model in PyTorch.
I'd recommend trying it for those who are curious. Computational bottlenecks become much more intuitive when you've looked at the overall process.
That's why. That future is uncertain. So why gamble your future on something that's popular at the moment for something that could change completely a year later?
I don’t see why learning how LLMs work is a bad project for a 17 year old.
Optimizing your entire career and the next decade+ of your life on LLMs? Yeah, probably not ideal. It’s almost always a bad idea to make long term decisions based on current trendy things.
And since everyone is using this topic to give their ideal advice to 17 year olds, my advice as a mid-30s guy: seriously consider becoming highly skilled at a specific thing, and don’t be scared off by the idea that it’ll take 5-10-15 years to get there.
When you’re 17-25, the timescale of a decade seems infinite. But it’s really not, and a decade spent “exploring and keeping your options open” sometimes just ends up with you being pretty decent but not amazing at a lot of random things.
Sometimes I wish I had just become a carpenter, chef, electrician, etc. – a specific skill set that leads to mastery over time, rather than the endless exciting-new-thing hamster wheel of working in tech.
And this is a problem with modern society today: expecting 17 year-olds to know what they want to do professionally for the rest of their lives, and to focus heavily on professional development aligned to that.
By that age, I think it's not uncommon for individuals to have interests and perhaps even dreams, but a well-defined career focus that serves as the foundation of an actionable skills development plan? Nah. That just isn't common and I'd argue not desirable. 17 year-olds should be exploring their interests, enjoying early adulthood, learning valuable lessons in the social realm, etc. Not training themselves to become compliant little worker bees.
(Besides the obvious nano gpt)
Even worse when they ask themselves.
As to the substance of your comment, there's nothing stopping a 17-year-old from writing poetry, learning guitar, and also learning about how LLMs work, if they're so inspired. Indeed I'm sure pg would encourage it, and that he did the equivalent of all those things himself when he was young. There's plenty of evidence of that in his essays: he wrote short stories, published a scandalous school newspaper, played soccer, and studied fine art and philosophy. See:
I assume the people at HN are intelligent enough to understand this without me having to explain the most obvious banalities to them. The point is why you would do anything at 17:
> if they're so inspired
You said it yourself. The Twitter post, to me, reads like career advice or something in that vein. One has plenty of time to run the wheel like a lab rat later, but at that very point in life you have possibility to find the thing that truly inspires you, that is its own motivation.
I was incredibly lucky to have had a father who worked that out for himself then impressed it on me. The people I feel sad for are those who don’t have people around to make them feel it’s ok to try things that seem impossibly advanced for your age or social status.
Wealth inequality is a huge factor in our ability to excel.
"Why would you reinvent the wheel rather than making it better?"
"To learn how wheels are made"
Banger reply.
My first read was "if you could redo your life from age 17 what would you do". To which "figure out how to make an LLM" would be an insane answer.
I think his advice here is maybe... a year or two too late to be good advice. I can't foresee the job market for machine learning experts being better than it is now in 5-6 years time when said hypothetical 17 year old would be most ready to start career hunting. Either the bubble is gonna pop and the market will be flooded with laid off AI talent.
Even if I'm wrong about there being a bubble at all, I still think that in 5 years time the tech will just have matured to the point of diminishing returns on refining existing architectures. Plateauing until some PHD comes up with something as ground breaking as attention.
I’m sure I’ll get torn to pieces for this but it’s frustrating to continually witness people treat a single person’s prose as the Word of God.
If you are 17, go be yourself, whatever that is, in whatever way you want that to be, but do it so authentically and fully. Be unapologetic about what you love and what motivates you, and pursue that with passion and commitment.
Owner of Golf Club Company says I should dedicate my life to golf lmfao.
Then vote for someone who will make the debts go away?
I mean, have you seen the options for people graduating right now? How people are behaving?
Or forget the data, look at how the story of the new future technology is being told. The people making it recognize that it has the potential to put swathes of white collar workers out of jobs, and they are openly talking/warning/PR-ing about it.
People in tech and SV, the places which have a underlying culture of near delusional optimism, are talking about trying to avoid being part of "the permanent underclass".
Gambling is up, and prediction markets are being treated as financial investments. Wall street bets is a thing, and outright speculative investments are the hope people have to get ahead.
This is happening in the USA, forget the weaker or smaller economies.
When people see the future as one massive zero sum game, with no way to win by building, then they are going to change how they plan their future.
As opposed to vote for someone who gives even more power and money to the oligarchs? You bet your last dollar that people will choose the former over the latter.
He capitalizes on greed and hype but with a soft, sober and thoughtful voice so as to lull you with rationalism and now 20 years of his “disruption” has mostly ruined modern society and a whole generation of techies have been led astray into trying to “change the world” is the world of today (minus the magic technology really any better than 20 years ago?)
- Good for him and his Tech Bros, bad for the rest of society
the lessons have svg diagrams, ai chat inline (google docs), etc
apologies if this comes off as slop, but its been working great for me.
There really isn’t a very strong correlation between tech industry hiring strength and AI as of yet. Various studies that are out there haven’t even witnessed AI workflows contributing more than modest gains in software engineering efficiency. I.e., being able to write code 20-40% faster isn’t a seismic shift in the industry where everyone is getting laid off tomorrow and we’re all replaced by software.
Even with the questions surrounding the current job market, it’s still an incredibly good ROI career compared to so many other jobs out there.
For example, in my local area you can get a job as a registered nurse working nights in the emergency room and only make ~$115k.
I make almost double that telling an LLM what to do from my house in my pajamas during the day with less time spent in university.
Even if tech roles lose half their salary to automation pressure it’s still a really good gig.
This right here is why nobody is shedding tears for the massive employment crisis in tech.
You make double that shilling ai slopware while they work nights saving lives.
I would tell a 17 year old that the world will always need nurses, same can’t be said for guys sitting in their pajamas burning tokens.
What is happening now in tech has been a long time coming, and it can’t happen fast enough.
I don’t make the rules for how much each profession is able to make in compensation.
The world will always need nurses, but that doesn’t mean that it’s a fantastic career to get into if you have a neutral career preference and your primary consideration is university tuition ROI, expected compensation, work schedule, and day-to-day physical exertion.
My point isn’t to debate the virtues of each career, I am intending to stick to objective aspects of them.
In that sense, telling today’s kids that there’s no future in tech careers just because there’s a short term hiring slump is extremely premature. I certainly wouldn’t tell a kid who is passionate about tech to avoid the field just because the unemployment rate is currently 7%.
https://www.investopedia.com/bachelor-s-degrees-with-the-bes...
I'd rather simply write another mnist implementation and check if I really like all that AI stuff at first place. Even then, before going into mature-on-the-way-to-dying tech (LLMs) I'd rather focus on fundamentals - good ols linear models, regressions, stat etc.
But sure, make the kids even more depressed by telling them they need to learn how to build an LLM so they can get a job working themselves to death to make Paul and friends rich.
I told my much younger brother when he was 12 what programming was and it'd be a great career. He looked into it and within months was writing CLI games. Eventually releasing his own unity 3d game on steam as a teen.
Eventually he got into CS and did really well because none of it was scary and new. He parlayed that into role at Meta out of university.
My point being, 17 year olds have time to learn new skills and guidance can go a long way.
I don't think you can really call yourself a developer unless you at least have an idea how to build a more complex software project like a compiler, and maybe have built a toy one either at uni or for fun.
It's not clear how long this LLM age of AI will last (to be replaced by something better), but nowadays any developer should at least understand the basics of ANNs, and more than just the "hello world" of a cat vs dog CNN. An LLM/Transformer is maybe the equivalent of a compiler in that regard - something that we all use and is complex enough to present a bit of a challenge. You should at least understand the basics of how an LLM is built, and maybe building a toy LLM will/should become the new Comp. Sci. degree toy compiler replacement.