Predicting Stack Overflow Tags with Google’s Cloud AI
stackoverflow.blog
stackoverflow.blog
No NLP approach I’ve tried has been able to predict question score based on content better than my baseline of “choose the mean”. (I’ve tried random forest on bag of words, AWD-LSTM, and Google AutoML so far).
The author of this post tried score prediction as well and pivoted to tag prediction after she couldn’t find anything that worked well: https://twitter.com/srobtweets/status/1125860523377979398?s=...
It’s so crazy to me that which posts get popular might just be random. And makes me wonder about the correlation between post content and popularity on other social sites like HN and reddit.
Additionally, you found average (all up) to be a better predictor than average per user or per category?
Apologies for grilling you here, I should frankly dig in myself, but if you happen to feel like indulging me it's much appreciated :)
A lot of questions that should have a very low score get fixed by people other than the original author. This may make your modeling more difficult because not only are there multiple authors, but the expected score of a question is not time-invariant. In fact, the final score of a question is a path dependent composite score of potentially several intermediate formulations.
The language model (viewable at https://stackroboflow.com ) scores much higher accuracy predicting the next word on Stack Overflow than other datasets like IMDb.
But on the IMDb dataset this approach led to state of the art sentiment analysis results by pulling out the encoder and adding a custom head whereas on Stack Overflow it didn’t grok anything about the score.
Now all I know is I haven’t seen evidence that they aren’t.
I could see how there could be some element of luck in whether a post gets attention before getting buried. It depends who’s online when it’s posted and whether it piques their interest, and many other random factors that don’t have anything to do with the content of the post like how quickly other posts come in afterwards to bury yours, and how any automated algorithms decide to surface it.
(I added tags to the language model later.)
Generally very short questions score badly because they will be stuff like "How do I make a video sharing website" but a very similar question like "How do I copy the current line in vim" will score well as long as its not a duplicate.
Both of those questions look similar when you just look the words and sentence structure. To know the difference you have to understand what a video sharing website is and know what kind of task copying a line in vim is so you can know which one is a reasonable question and which one is likely to help more future readers.
And (theoretically) duplicates are deleted by the moderators.
For my side project: As we received emails from developers asking for clarification/help with the APIs, the system would provide relevant documentation URLs so that anyone could pick up an inquiry and brush up on the API (and have a handy link if the the docs could be leveraged in the response).
Both Emma and Sara are in our Developer Relations (aka Dev Rel) organization, like Kelsey Hightower and Felipe Hoffa.
They’re explicitly not in Sales, and their job is focused on explaining and demonstrating stuff to Developers. They go to meetups, give invited talks, write blog posts, and so on.
Sorry if that doesn’t come across as clear. They work for Google, are paid by Google for that work, but aren’t measured by revenue or anything.
DevRel is measured on how many people they reach, but even that only vaguely. Marketing events, by contrast, are more often measured by pipeline. And you can be certain that onboarding (Professional Services) and Sales/Solution engineers are measured on revenue and time to revenue.
Otherwise you're just playing "fastest gunslinger in the west" trying to answer generic stuff before anyone rather than sniping questions you genuinely find interesting and can teach you.
I guess tags and searching (when dupe-closing) are more essential if you primarily use SO to post answers.
The same could probably done with search terms saved to your profile, but the tags are a much more organized alternative.
And yes, Google is worse than useless, it doesn't even search for what I type in, and when I force it to do that, it returns SEO spam websites, bot-farmed-content, clickbait and advertising ahead of useful content. Or instead of useful content.
As for asking, I suppose if my question has never been asked before by someone else, I'd rather not know the answer.
So even if tags are not your thing then whatever question you see on StackOverflow will be fully tagged up. If the software doesn't suggest some when asking some keen person will add them in.
If you are interested in a particularly obscure software package that does not have its own StackOverflow site then you will fond the square bracket search option (for the tags) is a good way of finding out what is new in that niche and what cool features or tips you can borrow for your own project.
At a guess the unrelated questions - 'top network questions' - are more likely to get clicks than the SO search box. So, to answer your question, the answer is 'no'. Aside from the one DDG user on HN that maintains a gopher site and the really computer-phobic uncle that uses Bing! with Windows XP the whole English speaking world is using Google.
China is different.
So tags probably do not work well on a popular topic, but outside of that, they can be useful.
It's certainly common for e-commerce. Especially when you can search by specific product attributes, shipping options, etc.
Thats why they have the requirement for a tag "Could someone be an expert in this tag?"
https://www.kaggle.com/c/facebook-recruiting-iii-keyword-ext...
I'm in a waiting room right now, but is anyone interested in summarising or commenting on the differences/ gains/ losses between the performance and aspects of comparing the two? :)
https://github.com/lettergram/sentence-classification
It uses Keras and goes through everything from encodings, to the way various networks function, to hyperparameter tuning.
To access the dataset you need to be logged in with a Google account. Details here: https://github.com/GoogleCloudPlatform/ai-platform-text-clas...