Each turn of the conversation, the entire history of the conversation is first tokenized. Each token is like a short ID for a very (multi-thousand-element) vector called an embedding. An LLM is built around putting all these vectors together into a matrix, and then multiplying that matrix by another bunch of matrices, anll of which results in another embedding vector that represents the next token in the conversation. This vector is then added to the end of the input matrix (aka the “context window”), and the process starts again. This continues until the token coming out of the LLM is a special “end of message” token.
A human designed the network of matrix multiplications that make up the model, but the matrices all started out filled with zeroes. The training process is what puts nonzero values in these matrices.
The values are such that if you look at how each column of the input matrix (which, recall, is an encoding of each token of the conversation history) gets multiplied as it works its way through the network, you can think of the coefficients as the model’s understanding of a “concept”. Notably, because modern LLMs are built around the concept of a “transformer” which looks at every word of the context simultaneously, the “concept” can involve coefficients that act on multiple (usually nearby) tokens. So the model may have internalized the concept of “empathetic” as “a cluster of nearby tokens whose values, when multiplied by M1[2492, 59272] * M2[592827, 7394] * M3[93732, 429474] * ..., are all similar to each other.”
As might be apparent, this means that the values for each token embedding are also highly sensitive to the coefficients buried deep within the LLM’s matrices. These values are the result of the training process too. The whole thing gets trained at once on an enormous body of text, some of which explicitly explains concepts like “empathy,” which will enable the model to associate empathetic clusters of words with other clusters of words that define empathy. But if you removed all the definitions and explanations of empathy from the training set, the model would still probably form a (weaker) concept of “empathy” that it would have a harder time explaining. It might not even know the word “empathetic,” but it would have some internal structure that causes token clusters like “I’m sorry for your loss” and “I know what you’re going through” to have similar values across some of their elements.
Using LLMs for medical research is all about encoding drug data rather than words and trying to extract the “concepts” that the training process has formed. The hope is that the model has internalized a novel insight from seeing way more data than a human could consume, much less reason about.