Show HN: Generate a quiz from a Wikipedia page
github.com
github.com
DBpedia is big as well.
WikiData is several magnitudes smaller than Freebase and DBpedia at the moment, but has an active healthy community.
What's the best way to run an offline copy of WikiData these days? Does it run on MySQL (or Postgres) like all other Wikimedia properties?
This isn't what happened. Freebase is still available to download and is slowly (with Google support) being migrated to WikiData[1].
There are some other pretty large Knowledge Graphs around. ConceptNet and Probase/MS Concept Graph are two that are worth looking at.
[1] https://static.googleusercontent.com/media/research.google.c...
The supported Virtuoso database is quite esoteric ( https://en.wikipedia.org/wiki/Virtuoso_Universal_Server ).
Has one succeeded writing a script to import the data to Postgres, MySQL or Lucene?
You need a RDF/Graph database (unless you are up for a lot of re-engineering)
I bet you're going to get some offers from making this post on HN. Make sure you have your contact info in your profile. :)
Regarding your codebase: clear and to-the-point code, well commented, and helpful commit messages. Including a `requirements.txt` is a plus.
Good job, keep it up!
# splits a Wikipedia section into sentences # and then chunks/tokenizes each sentence
If I had interviewed the author, I would have asked him what's the purpose of commenting like that.
Imagine using the time that spent on "re"-visiting "re"-opened bugs that are vague enough to be argued about on writing code that doesn't need these "re"s in the first place.
I contend that that might be a difficult place to get to especially because it's a team effort as well, but I feel it's more productive and less stressful to work like that.
To each his own I guess.
# Iterate through article's sections
for section in self.page.sections:*Mostly related to: whitespace / spacing, indentation, mixing single quotes and double quotes, magic numbers, naming conventions
One potential improvement is to remove the common parts of answer and question (as in your Triumph example)
I used nltk (natural language toolkit), which takes care of most of the hard work. It tokenizes whatever text you pass it, and even assigns each word a part-of-speech (noun, adjective, etc).
The grammar is where I tinkered the most. You can see I have 3 grammar rules set up (NUMBER, LOCATION, PROPER). nltk will go through the tokenized words and see if any sequences of words match any of the rules.If it finds a match it groups/chunks those words together into a phrase with the tag you've specified (ie. LOCATION).
As for the rules themselves, they're very easy to write once you understand the syntax. For example, let's look at my PROPER rule, {<NNP|NNPS><NNP|NNPS>+}
Everything in the {} is the rule. The tags inside of the <> are the parts-of-speech assigned by nltk. Translating the rule literally would be: match any sequence that has: [an NNP or an NNPS] followed by one or more of [an NNP or an NNPS]. In other words, any sequence of two or more NNP or NNPS words.
Let me know if you plan on continuing it. I'd love to collaborate.
[Jar'Edo Wens]
> Correct!
It's very buggy though... I get more invalid questions than good ones, haha
<class 'TypeError'>, TypeError("a bytes-like object is required, not 'str'",),
Please report back if that worked or not with an Issue on the repo, so I can follow up with a fix. Thanks!
Unfortunately, I'm getting a 500 error on every request.
What did I do wrong?
I solved it by downloading 'averaged_perceptron_tagger' from nltk.
>>>import nltk
>>>nltk.download('averaged_perceptron_tagger')