The whole point of embeddings and tokens are that they
are a compressed version of text, a lower dimensionality. now, how low depends on performance, lower amount of vectors=more lossy (usually).
https://huggingface.co/spaces/mteb/leaderboardYou can train your own with very very compressed, i mean you could even go down to each token=just 2 float numbers. It will train, but it will be terrible, because it can essentially only capture distance.
Prompting a good LLM to summarize the context is probably funnily enough the best way of actually "compressing" context