Oracle of Zotero: LLM QA of Your Research Library
github.com
github.com
I've spent quite a lot of time in the medical/scientific literature space. With regards to LLMs, specifically RAG, how the data is chunked is quite important. With that, I have a couple projects that might be beneficial additions.
paperetl (https://github.com/neuml/paperetl) - supports parsing arXiv, PubMed and integrates with GROBID to handle parsing metadata and text from arbitrary papers.
paperai (https://github.com/neuml/paperai) - builds embeddings databases of medical/scientific papers. Supports LLM prompting, semantic workflows and vector search. Built with txtai (https://github.com/neuml/txtai).
While arbitrary chunking/splitting can work, I've found that integrating parsing that has knowledge of medical/scientific paper structure increases the overall accuracy and experience of downstream applications.
it would accelerate research so much if LLM accuracy increased on biomedical papers.
very much agreed on the potential to extract signal from paper structures.
two questions if you don't mind:
1. did you post a summary of your chunking analysis somewhere? i'm curious which method maximized accuracy, and which sentence-overlap methods were most effective.
2. do you think general tokenization methods limit LLMs on scientific/biomedical papers?
> 1. did you post a summary of your chunking analysis somewhere? i'm curious which method maximized accuracy, and which sentence-overlap methods were most effective.
Good idea on this but nothing posted. In general, grouping by sections of a paper has worked best (i.e. methods, conclusions, results etc). GROBID is helpful with arbitrary papers.
> 2. do you think general tokenization methods limit LLMs on scientific/biomedical papers?
Possibly. For vectorization, specifically with medical, I do have this model (https://huggingface.co/NeuML/pubmedbert-base-embeddings) which is a fine-tuned sentence embeddings model using this base model (https://huggingface.co/microsoft/BiomedNLP-BiomedBERT-base-u...). The base model does have a custom vocabulary.
In terms of LLMs, I've found that this model (https://huggingface.co/Open-Orca/Mistral-7B-OpenOrca) works well but haven't experimented with domain specific LLMs.
RAG is good for semantic search, but really we need something that works at a knowledge/understanding level as opposed to data/information level.
txtai (included with paperai) has the ability to build semantic graphs (https://neuml.hashnode.dev/introducing-the-semantic-graph).
I agree that RAG is just one part of the equation. But the tools are available if one wanted to build their own complex multi-agent workflow.
The agents are likely very narrow and specific, they do one very very specific task. Then the workflow is a DAG chaining their work together.
In neo4j, for example, relations tend to have natural language names. (The cat BELONGS_TO the human)
So LLMs appear to be apt at making those queries
I guess, this could be an application of the agent model. I've seen multiple LLMs recently trained specifically on LateX parsing. One model would recognize from the parsed PDF garbage that there is probably an equation there and call a different want to parse it.
At the moment it returns runtime error.
Edit: It's because of missing open-api key, https://huggingface.co/spaces/whitead/paper-qa/blob/main/app...
I know that in theory arxiv, being a pre-print server, shoulnd't give any credibility but practically that is the case and it still is a good quality/bs filter compared to e.g. Medium articles.
There are preprints of articles since then published (which have the same credibility as the peer-reviewed article), articles form mates (which are obviously great), and the rest, which might be interesting but not a solid source on its own.
It seems to be working as intended, to be fair. ArXiv has precious little ways of improving the accuracy of the preprints.
The feature I find most useful is the table automation which I use for literature review, since it lets me run the same QA prompts on a collection of documents all at once.