Topic Suggestions for Millions of Repositories
githubengineering.com
githubengineering.com
I chose to look at this one as the README is quite descriptive and offers examples, and it is reasonably well structured.
It's a Go HTML sanitizer.
The suggested topics included "go" but none of the others I have now tagged it with.
It did include things like "html-element" and "data-uri". I can see how it did this from word prevalence, but these were far too specific to examples and documentation, and did not describe the project.
It feels as if the word counted should be weighted towards the early part of the README, perhaps no farther in than the 3rd heading.
Has anyone else implemented something like this before? War stories?
Caveat: Unlike Github we actually do have user-generated labels to start with, so our modeling problem has one-step less complexity than theirs but are otherwise the same.
(TL;DR If you have gold standard data to start with and no time to read academic articles, try fastText supervised learning. If you have time to read academic articles, start with a lit search.)
Can I ask what your site is?
Moz have a good write up of their Keyphrase Extract pipeline here. https://moz.com/devblog/machine-learning-approach-to-keyword...
I would be happy to discuss this with you more if you would like.