So, this has been possible already for quite a long time.
Each of these layer types have different computational costs for training and inference, and encode different inductive biases, which may be more or less appropriate to a given problem.
The feature itself is something I wanted to play with too, as it's kind of an obvious thing to want. I mean, these models execute a pipeline:
[text] -> [tokens] -> {[embeddings] -> [inference] -> [embeddings]} -> [tokens] -> [text]
Where the part in { ... } may or may not be implemented as a single step (i.e. all three parts interleaved).
Now, apparently all the magic[0] of transformer models sits in the latent space and is invoked by the { ... } bit. We also know for sure that you can make the pipeline look like this:
[text] -> [tokens] -> {[embeddings] -> [inference]} -> [embeddings]
So with the two things in mind, it's kind of obvious you'd also want a pipe that looks like:
[embeddings] -> [inference] -> [embeddings] (and optionally -> [tokens] -> [text])
for the sole purpose of messing around and exploring the latent space itself.
I'm very much not up to date with the whole space, so I might be missing something, but I'd thought that poking around the latent space would be getting a lot more attention than it seems to be getting.
(EDIT: replaced < ... > with { ... } for readability.)
--
[0] Not the "how do transformers tick" details, but the "how the hell are they this good" / "GPT-4 is uncanny valley" / "could this thing be actually thinking?" kind of magic.
There’s been no real effort to especially expose it because that is what you get by default. Even OpenAI has an “embed” endpoint. So you’re not going to see a huge push for it the same way you won’t see a push for reasoning about websites “in the HTML”. :)
- Get embeddings for e.g. "blue" and "red", or "the sky is blue" and "galaxy redshift";
- Average them, resulting a vector that's bound to not be expressible with tokens alone;
- Input that to the same model I got the embeddings from, and see what comes out.
If by "embeddingendpoint" you mean OpenAI, they provide that for a specialized model, and (AFAIK) you can only get embeddings out (for the purpose of comparing various vectors yourself). They have no API endpoint for an LLM that can take those embeddings as input.
Combine the two embeddings into a new vector space and BAM, you've invented "embedding2word".
Curiously, this also applies to humans: we learn words first, spelling later; we almost always think in terms of words, subwords and phrases, and we're really fast at it - while anything to do with spelling seems to demand much more focus, and some degree of conscious hand-holding.
As for encoding sequence, I'm curious how that happens and need to find relevant papers, but I imagine there might be some "n-gram dimensions" in the vector, as in one value for "I'm a first token in a sequence", one value for "I'm a second token in a sequence", etc., which would encode the "occurs before"/"occurs after" relationships using few dimensions, leaving the rest for more interesting relationships.