GPT in 500 Lines of SQL
explainextended.com
explainextended.com
I'd been inspired by the makemore lecture series[0]. At the 1hr mark or so, he switches from counting to using a nn, which is about as far as I've gotten. Breaking it down into a relational model is actually a really great exercise.
def generate(prompt: str) -> str:
# Transforms a string into a list of tokens.
tokens = tokenize(prompt) # tokenize(prompt: str) -> list[int]
while True:
# Runs the algorithm.
# Returns tokens' probabilities: a list of 50257 floats, adding up to 1.
candidates = gpt2(tokens) # gpt2(tokens: list[int]) -> list[float]
# Selects the next token from the list of candidates
next_token = select_next_token(candidates)
# select_next_token(candidates: list[float]) -> int
# Append it to the list of tokens
tokens.append(next_token)
# Decide if we want to stop generating.
# It can be token counter, timeout, stopword or something else.
if should_stop_generating():
break
# Transform the list of tokens into a string
completion = detokenize(tokens) # detokenize(tokens: list[int]) -> str
return completion
because that looks a lot like a state machine implementing Shlemiel the painter's algorithm which throws doubt on the intrinsic compute cost of the generative exercise.A GPT in 60 Lines of NumPy - https://news.ycombinator.com/item?id=34726115 - February 2023 (146 comments)
I actually plan to do what you describe after I do the embeddings video (but only for a “toy” neural net as a proof-of-concept introduction to gradient descent).
No. You've got to have a solid background in computer science to even start to understand fully this article.
Even the title itself is not accessible to 99% of humans.
Is there any simplistic blog posts / training courses which go through how they work, or expose a toy engine in python or similar that? All the training I’ve seen so far seems oriented at how to use the platforms rather than how they actually work.
Particularly [0], [1], and [2]
[0] http://jalammar.github.io/illustrated-transformer/
[1] http://jalammar.github.io/illustrated-gpt2/
[2] https://jalammar.github.io/visualizing-neural-machine-transl...
I did make the mistake though of clicking "+ expand source", and after seeing the (remarkable) abomination I can sympathize with ChatGPT's "SQL is not suitable for implementing large language model..." :)
That is not true. See ByT5, for example.
> As an illustration, let's take the word "PostgreSQL". If we were to encode it (convert to an array of numbers) using Unicode, we would get 10 numbers that could potentially be from 1 to 149186. It means that our neural network would need to store a matrix with 149186 rows in it and perform a number of calculations on 10 rows from this matrix.
What the author calls alphabet here, is typically called vocabulary. And you can just use UTF-8 bytes as your vocabulary, so you end up with 256 tokens, not 149186. That is what ByT5 does.
But from my experience that cannot be true - it has to learn somehow. There is an easy example to make. Tell it something that happened today and contradicts the past (I used to test this with the Qatar World Cup), and then ask questions that are affected by that event, and it replied correctly. How is that possible? How a simple sentence (the information I provide) changes the probabilites for next token by that far?
1. The trained knowledge included in the parameters of the model
2. The context of the conversation
The 'learning' you are experiencing here is due to the conversation context retaining the new facts. Historically the context windows were very short and as the conversation continues it would quickly forget the new facts.
More recently context windows have grown to rather massive lengths.
For real though, and knowing this is a leading question, the author has near-on 15 years of blog posts showing complex problems being solved in SQL. Is their brain bigger than yours and mine? Maybe a little bit. Do they have a ton of experience doing things like this? Most definitely.
It is easy to get overwhelmed when looking at the end result.
To learn you should probably run one block at a time to understand what each piece does. (In "normal" languages those pieces would be an isolated function)
E.g. The Sultan’s Riddle in SQL https://explainextended.com/2016/12/31/happy-new-year-8/
We might end up with more regularised language, and a more consistent model of the world, but that would come at the expense of accuracy and faithfulness (two things which are already lacking).