As an LLM novice, can someone explain what these "your document" apps are doing? My understanding is that GPT-4 doesn't support fine-tuning, and 50MB is too large to add to the prompt (which would be too expensive anyway).
Since GPT can use things from his context arbitrarily ,does it solve the hallucination issue, even for ebooks?
Use SebtenceTransformers in python to write to the database (PineconeDB) and then do the same for queries. Use the results as context.
When you query something like "What is this research about?" is it able to use data from all chunks?
I feel that it's inevitable that OpenAI et al. will be able to handle large PDF documents eventually. But until then I'm sure there's a lot of value of in this kind of pre-processing/chunking.
You are right GPT-4 doesn't support fine-tuning but, I think (in general) people might be misunderstanding what fine-tuning does.