Citation Needed – Wikimedia Foundation's Experimental LLM/RAG Chrome Extension
chromewebstore.google.com
chromewebstore.google.com
You can read more about it at https://meta.wikimedia.org/wiki/Future_Audiences/Experiment:...
Disclaimer: I worked on this.
I’m heavily into browser extension development! I’ve done some insane things with them.
I’ve built about 8 browser extensions in the last 6 months, most of them have thousands of users and one of them had half a million users
Currently building an LLM powered design assistant extension.
If you’d like to chat im reachable by email at “hello[at]papillonsoftware[dot]dev”
Do you have any concerns about your ability to properly maintain so many extensions?
For instance, one of them reveals salaries on job seeker sites, and is feature complete as it does what it’s meant to do bug free and fast
Disclosure: author on that work.
If not, what extra work is needed to bring it to that level?
A key engineering challenge will be speed ... when you're navigating a document you want a fast response time.
[0] https://diff.wikimedia.org/2022/01/19/the-wikipedia-library-...
Wikipedia works because we can update it in real time in response to changes. LLMs that need to constantly recrawl every time a page on the internet is updated, and that properly contextualize the content of that page, is a huge ask. Because at that point, it stops being an LLM and starts being a very energy-hungry search engine.
I also have my doubts on whether it is possible to implement efficiently (or at all). I suspect that just yanking in the article and all the sources is non-feasible, and any smaller chunking would be missing too much context. Plus LLM logical capabilities are questionable too, so I don't know how well the comparison would work...
https://gitlab.wikimedia.org/repos/future-audiences/citation...
edit: the readme build instructions are incomplete and i don't think hotreload works. use `npm run build-dev` to get a working build.
it's not obvious to me what prevents this from being a firefox extension as well - it might be the sidebar/sidepanel api differences, but i haven't played with those much
In academia, Wikepedia citations are generally a no no. One reason is their unrelaibity (the author is citing a source that they themselves can edit). More importantly, Wikepedia may be a good place to find primary sources, but in itself it is a secondary source.
Wikipedia is good for finding secondary source, and then primary sources by following the links.
[1] https://en.wikipedia.org/wiki/Wikipedia:Primary_Secondary_an...
There is plenty of reasons why wikipedia is an inapropriate source to cite most of the time in academia, but that surely is not one of them.
Acedemics cite their own papers or other sources they have editorial control over all the time.
"Wikipedia articles should be based on reliable, published secondary sources, and to a lesser extent, on tertiary sources and primary sources. Secondary or tertiary sources are needed to establish the topic's notability and avoid novel interpretations of primary sources."
https://en.wikipedia.org/wiki/Wikipedia:No_original_research
> "Citation Needed is an experimental feature developed in 2024"
- The provided passages do not contain any information about a feature called 'Citation Needed' being developed in 2024.
- Wikipedia
- Discouragement in education
- not be relied
I know I'm not using it for intended purposes but it seemed funny.
The second part does seem more problematic, but still, as essentially a yes/no question, it should be significantly less likely to hallucinate/confabulate than for other tasks.
Also, if you look into their “wrong” example closer, it is a bit misleading, as both sources are correct. Joe Biden was 29 on election day, but 30 when he was sworn in. Understanding this requires more context than the LLM was apparently provided.
I am not sure if when people say this they just don't have experience building with LLMs or they do and have experience that would make for a very popular and interesting research paper.
I disagree. LLMs that have been "grounded" in text that has been injected into their context (RAG style) can absolutely stolñ hallucinate.
They are less likely to, but that's not the same as saying they "won't hallucinate" at all.
I've spotted this myself. It's not uncommon for example for Google Gemini to include a citation link which, when followed, doesn't support the "facts" it reported.
Furthermore, if you think about how most RAG implementations work you'll spot plenty of potential for hallucination. What if the text that was pulled into the context was part of a longer paragraph that started "The following is a common misconception:" - but that prefix was omitted from the extract?