Wikipedia search-by-vibes through millions of pages offline
leebutterman.com
leebutterman.com
Unfortunately, I tried describing a few terms across philosophy and psychology and for all of them, the entry I was aiming for was only around the ~20th rank. (Far more popular but less accurate items were populated above it -- e.g. no matter what I typed trying to define a specific modality of psychotherapy, "psychotherapy" was always the #1 result.)
In contrast, I've used ChatGPT to identify the names of certain niche subfields when I couldn't remember what they were called, and it was right every time.
I love the idea of an AI service specifically designed to identify the names of things from descriptions. But I don't think restricting it to Wikipedia (or Wikipedia page titles) is the right approach, and it seems like general-purpose LLM's are doing a great job.
Still, as a proof of concept and as something you can run locally in the browser, this is extremely cool.
Wikipedia is a great demo dataset, and I’m definitely up for adding more datasets. Specifically, just like iPhoto lets you search “mountain” and you can get pictures with mountains, might be cool to search with some multi modal models like CLIP on various datasets
Whereas when I type the same query into google, I get exactly what I expected. Which is kind of disappointing, I was hoping to find out about some weird looking monkeys that I didn't know about.
For instance, Reddit data dump (https://academictorrents.com/details/7c0645c94321311bb05bd87...), filter for Wikipedia links, include a context of the thread, combine that with the contents of the article
I don’t keep any analytics on the page about what people find and don’t find, so I haven’t set myself up to improve the search results :/
One trick that might be helpful is to embed only the defining (usually the first) sentence or paragraph of the Wikipedia article, rather than the whole document -- not clear to me which portion you're using now.
My own site, OneLook, has had a similar feature (https://onelook.com/thesaurus/) since '03 that lets you find words and concepts by description. It was a pure reverse-dictionary search back when I started, but over the past two decades I've explored word embeddings, then sentence embeddings, and more recently LLMs. Nowadays it uses GPT to generate some guesses for inputs that it can't answer itself.
LLMs are so much better than earlier methods at this task, it's taken some of the wind out of my sails on improving this aspect of OneLook. I frequently hear from people for whom reverse-definition lookups are the main reason they use ChatGPT!
However, there is a recent paper that actually does try and do this: "Retrieving Texts based on Abstract Descriptions" (Ravfogel et al., 2023) https://arxiv.org/abs/2305.12517.
They give many examples of searching by vibes: "an architect designing a building", "a company which is part of another company", "a book that influenced the development of a genre", etc. etc. Their embeddings apparently facilitate this type of search much better. Would be interesting to retry the offline Wikipedia search from the linked post with this new type of embeddings.
I think we may be doing awful things to Lee Butterman's bandwidth bill.
Although I know from experience it's really difficuly to assess search result quality by hand, you can be very close to something great and return far worse matches than this does.
I searched "pointy building in Paris", and got :
Tourism in Paris, Bourse de commerce (Paris), Grands Projets of François Mitterrand, List of tallest buildings and structures in the Paris region, List of tourist attractions in Paris, Palais des congrès de Paris, Landmarks in Paris, Palais de la Bourse, Lyon, Outline of Paris, Architecture of Paris
no mention of the most famous pointy building in Paris...
Maybe sentence embedding of the entire article is not the best thing for this kind of application.
I just checked the article, and of the 19 times the word "building" appears, it's mostly a verb, followed by "Chrysler Building"
Unless there's some other famously pointy building I'm not thinking of.
I was mostly too excited to trot-out the thing I remembered about why most towers are not buildings.
Without this signal, a lot of useful information is ignored and the result doesn't feel as magical.
Still impressive, fascinating demo
* "The wizard in The Lord of the Rings": No Gandalf or Saruman, only books about LOTR and such.
* "Protagonist of Scorsese's Taxi Driver": No Travis Bickle.
* "A person that plants trees for a living": Somehow a gardener isn't on the list.
* "Curly-haired painter on TV": No Bob Ross anywhere.
* "Unusually shaped modern art museum in Spain": Bilbao does show up as number 4, but none of the others are unusually shaped.
* "Dog shaped like a sausage": Surely a dachshund should be in the top results.
Looks like the embedding model used (all-minilm-l6-v2) currently ranks 35th on the hugging face leaderboard [0]. I'd love to try with other models if anyone wants to +1 this demo :). This feels like a nice dataset to build intuition around embeddings used for RAG etc.
I think it's worth it.
[0]https://www.google.com/search?q=%C3%A9corch%C3%A9+site%3Aen....
What's the number means near words in the results? For example "book 70k" what's 70k refer to? The number of discussions on Wikipedia about it? #s of Edits? Or articles that mentioned it?
"https://en.wikipedia.org/wiki/" + encodeURIComponent(title)
Here you have it. The major feature is in the title: you can hardly call it a Wikipedia search engine if you can’t access the articles.