I run into the same problem. I made a RAG system and imported 15 years of reddit and HN content (my own message logs) + 2 years of LLM chats, totaling about 80MB of text. I can use it to retrieve fragments but there is duplication, and almost all duplicate fragments have slightly different approaches. How do I merge all of them, and how do I get a deduplicated taxonomy? I got about 290K unique keywords extracted from the text, it doesn't fit into LLMs.
I am gravitating towards building graphs of ideas, and having a way to generate unique (non duplicate) new ideas while ingesting new text.