I was using embeddings to group articles by topic, and hit a specific issue. Say I had 10 articles about 3 topics, and articles are either dry or very casual in tone.
I found clustering by topic was hard, because tone dimensions ( whatever they were ) seemed to dominate.
How can you pull apart the embeddings? Maybe use an LLM to extract a topic, and then cluster by extracted topic?
In the end I found it easier to just ask an LLM to group articles by topic.