HNHacker News
TopNewBestAskShowJobs

ikuyamada

375 karma · joined April 30, 2013

submissionscomments
ikuyamada··on Wikipedia2Vec: Optimized Implementation for Learning Embeddings from Wikipedia
My past paper describes an entity linking method based on Wikipedia2Vec: https://arxiv.org/abs/1601.01343

You need to extract entity names using an NER software (e.g., SpaCy, Stanford NER), and resolve the names to knowledge base entities using the entity linking method.

ikuyamada··on Wikipedia2Vec: Optimized Implementation for Learning Embeddings from Wikipedia
Unlike Word2vec, this tool learns embeddings of entities (i.e., entries in Wikipedia) as well as words. And although the model implemented in this tool is based on Word2vec's skip-gram model, it is extended using two submodels (Wikipedia link graph model and anchor context model). Please refer to the documentation for details: https://wikipedia2vec.github.io/wikipedia2vec/intro/
ikuyamada··on Wikipedia2Vec: Optimized Implementation for Learning Embeddings from Wikipedia
You can see the visualization of the embedding vector space here: http://projector.tensorflow.org/?config=https://wikipedia2ve...

I recommend to use T-SNE instead of PCA, which can be selected by the button at the bottom left.

ikuyamada··on Wikipedia2Vec: Optimized Implementation for Learning Embeddings from Wikipedia
Author here. More broadly, embedding is a mapping from objects (e.g., words and entities) to vectors of real numbers. And as described in the mcxlog's comment, an entity refers to an entry in Wikipedia in this paper.
ikuyamada··on Show HN: Wikipedia2Vec – A tool for learning embeddings of words and entities
Thank you for your feedback! I am also interested in conducting experiments on extrinsic tasks such as text classification. In addition to word embeddings, Wikipedia2Vec also contains entity embeddings which are likely beneficial for these tasks, so I would like to design a model that uses both the word embeddings and entity embeddings.
ikuyamada··on Show HN: Wikipedia2Vec – A tool for learning embeddings of words and entities
Thanks :) 1) I think learning entity embeddings using the Doc2Vec (paragraph vector) model is an interesting idea, but we did not test it. 2) This tool was initially developed to address the entity linking task. Mapping words and entities into a same vector space enables to model the contextual information that is useful for entity linking. For details, please refer to this paper: Joint Learning of the Embedding of Words and Entities for Named Entity Disambiguation: https://arxiv.org/pdf/1601.01343.pdf
ikuyamada··on Show HN: Wikipedia2Vec – A tool for learning embeddings of words and entities
Regarding word embedding algorithm, I am interested in supporting other models that uses subword information (e.g., Fasttext). Further, there have been proposed various recent models to learn entity representations from KB, and I plan to work on them.
ikuyamada··on Show HN: Wikipedia2Vec – A tool for learning embeddings of words and entities
We did not add Fasttext to our benchmarks because of a minor technical issue but we will work on it. Further, to conduct a fair comparison with ELMo, I think it is needed to use extrinsic tasks such as question answering and textual entailment.
ikuyamada··on Show HN: Wikipedia2Vec – A tool for learning embeddings of words and entities
The current code is written specifically for Wikipedia. However, its algorithm is portable for knowledge bases that contains articles and their entity annotations.
ikuyamada··on Show HN: Wikipedia2Vec – A tool for learning embeddings of words and entities
What kind of output did you mean? Wikipedia2Vec learns embeddings of entities which have links from other articles more than min-entity-count times.
ikuyamada··on Show HN: Wikipedia2Vec – A tool for learning embeddings of words and entities
Please note that similar to other approaches (e.g., node2vec), Wikipedia2Vec learns embeddings for Wikipedia entities in addition to embeddings for words.
ikuyamada··on Show HN: Linkify, auto-linking keywords for easier search on mobile apps
When clicking on the link, Linkify displays a small widget that contains links to typical search sites such as Wikipedia, Google, Twitter, etc.
ikuyamada··on Json-wikipedia
In my understanding, DBpedia is a project for extracting data mainly from Wikipedia infoboxes (not whole data dumps) by collaboratively creating rules for converting data into cleaner schema that enables to perform a query such as SPARQL. I think this is a project that directly converts wikipedia dump xml to JSON for easier manipulation, which differs from DBpedia.
ikuyamada··on PiCrawler: A distributed web crawler using PiCloud
I am also a fan of using PiCloud as a platform.

Using PiCloud for implementing crawler requires some hacks, so I decided to create this as a separate package.