What are embeddings?
vickiboykis.com
vickiboykis.com
The magic is creating the embedding, embeddings themselves are nothing special.
[0] https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...
My current belief is that a hundred thousand+ dimensional latent space, as in SOTA LLMs, seems to have enough dimensions to reduce a large array of cognitive skills into a proximity search.
Combine the two embeddings into a new vector space and BAM, you've invented "embedding2word".
Curiously, this also applies to humans: we learn words first, spelling later; we almost always think in terms of words, subwords and phrases, and we're really fast at it - while anything to do with spelling seems to demand much more focus, and some degree of conscious hand-holding.
As for encoding sequence, I'm curious how that happens and need to find relevant papers, but I imagine there might be some "n-gram dimensions" in the vector, as in one value for "I'm a first token in a sequence", one value for "I'm a second token in a sequence", etc., which would encode the "occurs before"/"occurs after" relationships using few dimensions, leaving the rest for more interesting relationships.
The feature itself is something I wanted to play with too, as it's kind of an obvious thing to want. I mean, these models execute a pipeline:
[text] -> [tokens] -> {[embeddings] -> [inference] -> [embeddings]} -> [tokens] -> [text]
Where the part in { ... } may or may not be implemented as a single step (i.e. all three parts interleaved).
Now, apparently all the magic[0] of transformer models sits in the latent space and is invoked by the { ... } bit. We also know for sure that you can make the pipeline look like this:
[text] -> [tokens] -> {[embeddings] -> [inference]} -> [embeddings]
So with the two things in mind, it's kind of obvious you'd also want a pipe that looks like:
[embeddings] -> [inference] -> [embeddings] (and optionally -> [tokens] -> [text])
for the sole purpose of messing around and exploring the latent space itself.
I'm very much not up to date with the whole space, so I might be missing something, but I'd thought that poking around the latent space would be getting a lot more attention than it seems to be getting.
(EDIT: replaced < ... > with { ... } for readability.)
--
[0] Not the "how do transformers tick" details, but the "how the hell are they this good" / "GPT-4 is uncanny valley" / "could this thing be actually thinking?" kind of magic.
There’s been no real effort to especially expose it because that is what you get by default. Even OpenAI has an “embed” endpoint. So you’re not going to see a huge push for it the same way you won’t see a push for reasoning about websites “in the HTML”. :)
- Get embeddings for e.g. "blue" and "red", or "the sky is blue" and "galaxy redshift";
- Average them, resulting a vector that's bound to not be expressible with tokens alone;
- Input that to the same model I got the embeddings from, and see what comes out.
If by "embeddingendpoint" you mean OpenAI, they provide that for a specialized model, and (AFAIK) you can only get embeddings out (for the purpose of comparing various vectors yourself). They have no API endpoint for an LLM that can take those embeddings as input.
So, this has been possible already for quite a long time.
Each of these layer types have different computational costs for training and inference, and encode different inductive biases, which may be more or less appropriate to a given problem.
https://unzip.dev/0x014-vector-databases/
It also explains how to use them IRL and why.
Vicki mentions this survey I wrote some time ago: https://towardsdatascience.com/milvus-pinecone-vespa-weaviat...
Hoping it'll be useful as well!
But even if I did consider it, a link called Get PDF wouldn't make me think it's a book. PDFs can contain books but they usually don't.
Bear with me on this question as I am still learning...
From a quick skim I couldn't find it. It might be there? But with a transformer, at which layer do you consider the vector to be the embedding? Or are there multiple choices?
For example you have a positional and token embedding - so they are of course embeddings, but simple ones. King - Man + Woman = Queen type thing.
But I imagine you might want to consider the output of the 1st ? 2nd? Nth? transformer on the decoder side as an embedding of sorts.
For example the sentence "attention is all you" has gone through a couple of layers you now have a vector that understands those words pretty well. But if you go too far through I imagine you get closer to the "prediction" and so lose the meaning and just get something closer to "need".
But if possible to get, that sentence embedding (by which I mean N tokens up to context size) might be useful for search indexing of long text.
For example: the Flutter example is introduced, then not mentioned for a while, then mentioned again with a specific problem, then ignored again, then used again with some other different problems. I keep getting lost about what we’re actually trying to achieve with it, and sometimes it feels like it’s entirely unnecessary to refer to it.
There are other times where we jump from pretty simple explanations to an intense matrix decomposition. A lot of jargon also gets used that, though previously explained, is used really densely. I feel like keeping the wording simpler would still get the point across.
Other than that, I’m finding it quite useful, and I’m only on page 30/75. Thank you for writing it :)
(my experience: CS degree, self-studied some AI stuff for the past half a year)
They're just locality-sensitive hash functions.
lsh: similarity(x, y) < thresh => E[P(d(lsh(x), lsh(y)) < 1)] > 1-eps for some eps
For a model that learns a good representation, and some suitable distance function to measure distances between embeddings (e.g. Euclidean) model: similarity(x, y) < thresh => E[P(d(emb(x), emb(y)) < r] > 1-eps
for some cutoff distance r in the vector spaceFarting out a quick one line answer is a boring comment
OP might be down voted because, even though the answer might be technically correct, it still completely fails to explain the concept (i.e., it's a useless explanation) and does so with a hint of arrogance and demeaning to readers.
Yours too, by the way.
It's like someone asks what is temperature and some smartass answers it's a scalar quantity that's mapped from a physical state into a semi-infinite continuous set. Great to know.
I found "a monad is just a monoid in the category of endofunctors" the most useful explanation of monads. I didn't know what an endofunctor—or a moniod—was when I read it, but it gave me some hooks for research. Other explanations using analogies just left me confused, and nothing further I could look into.
Though that said, I revised that explanation after the others did not help me, so maybe I had absorbed enough context.
It's like if someone wrote a book called "What is consciousness?" and just provided an annotated list of neurotransmitters. That approach leaves a lot out.