> embedding vectors you've calculated from the code? If so, those are likely quite easily reversible
I don't think embeddings are generally reversible... you're usually projecting onto a lower dimensional space, and therefore losing information.
I don't think embeddings are generally reversible... you're usually projecting onto a lower dimensional space, and therefore losing information.
> We train our model to decode text embeddings from two state-of-the-art embedding models, and also show that our model can recover important personal information (full names) from a dataset of clinical notes.
https://arxiv.org/pdf/2310.06816.pdf
There's certainly information loss, but there is also a lot of information still present.
“a multi-step method that iteratively corrects and re-embeds text is able to recover 92% of 32-token text inputs exactly”.
In theory if you know the model being used you could reverse them.