Show HN: Unriddle – Create your own GPT-4 on top of any document
unriddle.ai
unriddle.ai
I asked for whitelist access, didn't get it yet and now I won't get it because my country banned openai an openai closed my pro account.
You are right GPT-4 doesn't support fine-tuning but, I think (in general) people might be misunderstanding what fine-tuning does.
When you query something like "What is this research about?" is it able to use data from all chunks?
I feel that it's inevitable that OpenAI et al. will be able to handle large PDF documents eventually. But until then I'm sure there's a lot of value of in this kind of pre-processing/chunking.
Use SebtenceTransformers in python to write to the database (PineconeDB) and then do the same for queries. Use the results as context.
Since GPT can use things from his context arbitrarily ,does it solve the hallucination issue, even for ebooks?
I just tried this one on the NBA CBA, which I would think is an ideal use case, and it didn't answer a single question correctly. Hopefully we find a more clever way to use LLMs in conjunction with a knowledge-base.
"What's a mid-level exception?" (Initially stated the info wasn't in the provided context, but upon asking again, it got it correctly)
"How long does a team have to wait after receiving a player in a trade before trading the player again?" (It's 60 days under certain circumstances, it told me 6 months)
"What's a traded player exception?" (Claims isn't in the document)
"How are traded player exceptions created?" (Answers correctly)
So after playing with it a bit more I think I understand the sorts of questions it does and doesn't like to answer, but still, I would think fuzzy matches for section titles should work.
Potential issue might be that chunks just serve to activate massive knowledge of GPT4 and not actually used as a basis for an answer. For example, GPT4 has surely seen Dune in its training corpus and could be answering from memory.
There are lots of solutions using embeddings which basically boil down to a text search and picking some number of sentences / paragraphs around the search results to construct a "context", which may work with a limited set of technical / non-fiction documents but is otherwise of narrow use.
But I expect this kind of querying will be much better as the context windows for LLMs increase
But those tools are severely limited on maximum amount of pages. This tool has max 300 pages. ChatPDF (or how it was named) has limit on 2000 pages in paid version. So I will need to wait until those models are more available.
I couldn’t find any information about pricing? This is going to be a free tool?
Do you maintain the costs from ADs? Do you charge users?
Thank you
Can it ingest multiple PDFs in the same 'context', or would I have to assemble it all into one (under the 50mb limit)?
What is it using to
Can we (provoke you to) set the model temperature in the conversation to either minimize the hallucinating or increase the 'conjecture/BS/marketing-claims' factor?
I've also found in my area that it'll happily hallucinate stuff -- after all, it has zero actual understanding, it just predicts the most likely filler in the given context.
Tamping that down and just getting it to cut out the BS/overconfidence response patterns and reply with "I know X and I don't know Y" would be incredibly useful.
When we get back an "IDK", we can probe in a different way, but falsely thinking that we know something when we are actually still ignorant is worse than just knowing we've not yet got an answer.