Putting this into perspective in the age of LLMs:
The bigger the model the better the embedding, so one could take the middle activations of large language models and use that as embedding but using smaller models is often good enough and more performant.