Let's Build the GPT Tokenizer [video]
youtube.com
youtube.com
No metaphors trying explain "complex" ideas, making them scary and seem overly complex. Instead, hands on implementations with analogy explainers where you can actually understand the ideas and see how simple it is.
Steeper learning curve at first but it is much more satisfying and you actually earn the ability to reason about this stuff instead of writing over the top influencer BS.
Definitely recommend watching those videos and doing the exercises, if you have any interest in how LLMs work.
Then I abandoned it in 2021. This year, it struck me that LLMs would be great to infer business insights from the schema. I could create reports and dashboards automatically, surface critical action points straight from the schema/data and users chatting with the app.
So for the last couple weeks, I have been building it, running test on LLMs (CodeLlama, Zephyr, Mistral, Llama 2, Claude and ChatGPT). The results are quite good. There is a lot of tech that I need to handle: schema analysis, SQL or API calls, and the whole UI. But without LLMs, there was no clear way for me to infer business insights from schema + user chats.
To me, this is not a niche anymore now that I have found a problem I wanted to tackle already.
A good example is llangchain
2013: LDA with MALLET
2015: spaCy
2018: BERT
2023: GPT-4
2024: every person is an NLP expert in four lines of LangChain code
If you had to pick, building a project using off the shelf tech would better prepare you to work your first AI engineering job. However, the knowledge in these videos could help you land that first job, and is a useful base for concepts that aren't going away any time soon.
Also, please let us know if you figure out the secret. I would love to also switch from generalist backend to ML/AI.
To fix it, I'd throw a few gigabytes of synthetic data in the training mix before fine tuning that included the alphabets of all the relevant languages, things like.
A is an upper case a
a is a lower case A
the sequence of numbers is 0 1 2 3 4 5 6 7 8 9 10 11 12
0 + 1 = 1
1 + 1 = 2
etc.It still amazes me that Word2Vec is as useful as it is, let alone LLMs. The structure inherent in language really does convey far more meaning that we assume. We're like fish, not being aware of water, when we use language.
It's exactly like lexers for compilers. This parsing strategy coupled with the decision to then map the results into an embedding space of arbitrary dimensionality is why these models don't work and cannot be said to understand language. They cannot reliably handle fundamental aspects of meaning. They aren't equipped for it.
They're pretty good at coming up with well-formed sentences of English, though. They ought to be given the excessive amounts of data they've seen.
And while it's a core feature, it's a fairly robust one, while you can get some targeted improvements, the default option(s) are good enough and you won't improve much over them.
That's my first point. In 10 years we have word2vec, GloVe, GPT-2 and... tiktoken. lol. It's as if directional, numeric magnitudes in an embedding space of arbitrary dimensionality have magically captured or will magically capture the nuances and expressivity of language. Optimization techniques and new strategies for domain adaption are what matters, particularly for mobile devices, on-device ASR and short-form videos.
I don't think robust is a good characterization of clusters of semantic attributes in space or a distributional semantics of language. I'd say crude and without understanding are more accurate descriptions. Capturing semantic properties sometimes is not the same thing as having a semantics.
By targeted improvements you must be referring to domain adaptation and by the default option you must be referring to attention over BPE tokens? You can move directional quantities around in directional quantity space all day. If it results in expected behavior for your application that you weren't getting before that's great. If that's all you want to get out of these models then indeed there's nothing to do here. I'm not after improvements so much as I'm after something that works.
If you don't care about tokenization and use any of the reasonable default options without caring about them, and if you're doing a proper pre-training on non-tiny quantities of data, then the next few layers of whatever neural architecture you have on top of these tokens will generally be able to learn to compensate for any drawbacks in your tokenization, perhaps at some computation overhead - e.g. perhaps you could have had one less layer or smaller layers if you had the best tokenization possible, and edging out that computation cost improvement is pretty much the only thing you can hope to get out of having a better tokenizer.
Tokenization, the gateway to word embeddings, is a means to an end. I'm not suggesting that better tokens are needed or that BPE tokens should be replaced with something else. I'm suggesting that aiming for a distributional semantics is setting the bar pretty low and that there are better places to end up than These Things Are Over Here And Those Things Are Over There Let's Combine Them And See What Happens. I'm expressing disbelief that these representations have been taken at face value and that there has been practically no discussion of applying alternative formalisms which may be more expressive.
Modeling language in a latent space only makes sense for certain aspects of language and certain kinds of analyses. Crucially, you have to have meaningful primitives to begin with. This line of thinking that an understanding of language and an understanding of the world is somehow going to emerge from mapping character spans onto a latent space and combining them with dot product attention is pretty half baked. These systems remain in Firth Mode™.
There are books from oreilly and paid MOOC courses that are just padded with lots of unnecessary text or silly "concept definition" quizzes to make them seem worth the price.
And there are excellent free YT video lectures, free books or blog posts.
Andrej's YT videos are one great example. https://course.fast.ai is another.
Disclosure: Professor Malan is a friend of mine, but I was a fan of CS50 long before that!
- beej's networking guide is the best thing for network layer stuff https://beej.us/guide/
- explained from first principles great too https://explained-from-first-principles.com/
- pintos from Stanford https://web.stanford.edu/class/cs140/projects/pintos/pintos_...
Andreas Kling. OS hacking: Making the system boot with 256MB RAM https://www.youtube.com/watch?v=rapB5s0W5uk
MIT 6.006 Introduction to Algorithms, Spring 2020 https://www.youtube.com/playlist?list=PLUl4u3cNGP63EdVPNLG3T...
MIT 6.824: Distributed Systems https://www.youtube.com/@6.824
MIT 6.172 Performance Engineering of Software Systems, Fall 2018 https://www.youtube.com/playlist?list=PLUl4u3cNGP63VIBQVWguX...
CalTech cs124 Operating Systems https://duckduckgo.com/?t=ffab&q=caltech+cs124&ia=web
try searching here at HN for recommendations https://hn.algolia.com
Fellow hackers might also enjoy:
I like the book better than the online course.
There's also a tremendous amount of extremely low quality YouTube and blog content.
But from my limited sample size, the best free content is better than the best paid content.
If the web page /content is too polished, they're most likely optimizing for wooing users.
Unlike a lot of the examples I gave in the sibling comments. Where the optimization is only on the love for the topic being discussed
There's an inverse correlation with the glossiness of the content as well.
This is probably due to survivorship bias. Sites that have poor content and poor visual appeal (glossiness) never get on your radar.i.e. Berkson's Paradox: https://en.wikipedia.org/wiki/Berkson%27s_paradox
I'm not sure if the crew of the Nostromo would agree ;)