Embedding Archives: Millions of Wikipedia Article Embeddings in Many Languages
txt.cohere.com
txt.cohere.com
Would love to see similar but embedded with a more open representation model or even sent2vec
If you want to query for a search term, you can use a trial API key which is free to use for prototyping. The embedding model itself is not open source, though. [co-author of the post here]
For the headings, I mean the Wikipedia section headings (which isn't always a paragraphs, my mistake).
In both cases the data can be used like to classify/visualize Show HNs in your linked post.
Edit, going to [1] the datasets are labelled 2022-12.
It's an embeddings database of Wikipedia abstracts with page view data integrated to enable filtering pages based on popularity in addition to similarity.