Machine Learning and Link Spam: My Brush With Insanity
seomoz.org
seomoz.org
My general strategy is to invest into training set curation and evaluation. I also use quick scatter scatter plots to check I can seperate the training sets into classes easy. If its not easy to do by eye then the machine is not magic and probably can't either. If I can't, then its time to rethink the representation.
The author correctly underlines the importance of training set, but also equally critical to have the right representation (the features). If you project your data into the right space then pretty much any ML algorithm will be able to learn on it. i.e. its more about what data you put in, rather than the processor. k-means and decision trees FTW
EDIT: Oh and maybe very relevant is the importance of data normalization. Naive Bayes classifiers require features to be conditionally independent. So you have to PCA or ICA your data first (or both), otherwise features get "counted" twice. E.g. every wedding related word counting equally toward spam catagorization. PCA realizes which variables are highly correlated and projects them into a common measure of "weddingness". Very easy with skilean its preprocessing.Scaler() and turn whitening on.
There are papers[1][2] that outline possible benefits of SVM-based spam filtering. Unfortunately, SVMs are still in their infancy and not many people know how to implement and use them. I do think the are the future, however.
[1] http://trec.nist.gov/pubs/trec15/papers/hit.spam.final.final...
[2] http://classes.soe.ucsc.edu/cmps290c/Spring12/lect/14/007886...
SVMs have been around for almost two decades now, which is an eternity in the ML world, rather than infancy.
SVMs don't require the problem set to be linearly separable.
Please note that there's a myriad of robust, scalable SVM implementations -- SVMlight, HeroSVM, LIBSVM, liblinear... (the latter two also have wrappers in scikit-learn, a Python library mentioned in the OP).
Like I said, I haven't worked (much) with SVMs but they really do seem like the future. Unfortunately, they are difficult to work with.
And finally, the first workable SVM algorithm was proposed in 1995 (not nearly the two decades you claim) and implemented several years later. SVMs are very much still in their infancy -- especially considering that not much work is being done on SVMs since ANNs are much (much) easier to work with.
Geoff Hinton's claims that DBN's outperform traditional (shallow) ANN's, random forests and SVM's for any of the tasks he has tried them for (including classifying textual documents), though I have not benchmarked how well they work myself (not yet anyway!).
I have only worked with text classification methods where I chose the features myself. As I understand it, a deep network still has (like a 'traditional'/non-deep ANN) a fixed number of inputs in its input layer, i.e. one would have to process each input text somehow before feeding it into the network (to make the input sizes equal). Is there a usual way to do this without doing feature-extraction?
So the parameter would be the number of slots (== number of input units of the deep NN).
And the transformation of the text into bag-of-words / n-grams would not be considered feature-engineering - or at least only 'low level feature engineering' - the higher level features will be learned by the deep network.
I guess one could go lower level still and even do away with bag-of-words / n-grams : limit the text size to e.g. 20000 characters, represent each character value with a numerical value (e.g. its ASCII code point when dealing with mainly English texts) and then simply feed this vector of codepoints to the input layer of the deep network. Given enough input data, it should learn location-invariant representations like bag-of-words / n-grams (or even better ones) itself, right?
[0] https://www.youtube.com/watch?v=AyzOUbkUf3M
[1] For a general overview of the study, here are its slides (pptx): http://users.cs.uoi.gr/~gtzortzi/docs/publications/Deep%20Be...
[2] The same study in full (pdf): http://users.cs.uoi.gr/~gtzortzi/docs/publications/Deep%20Be...
Unless you have a baseline of "what was here first" and "exactly when every website went live with what links" like Google does (because they have been indexing websites since the dawn of time as far as the internet and linking is concerned. Heck, there wasn't even backlink spamming prior to Google because Google was the first search engine to rank by number of backlinks!)... you're going to have a really tough time determining what spam is and what it isn't.
Which 1% do you decide to focus in on?
It does however make it very difficult to gauge who is a legitimate linker and who is not.
All I'm really saying is if you're starting now with all of the years of random linkspam backscatter is that you are in for a rough ride.
As an aside, you have use Ahrefs.com to get pretty decent tracking of when links appeared since it started (I think ~18 months ago or so). Given that the rate of spammy pages is increasing extremely fast and old spam pages are dying off, I imagine that in the not too distant future you'll be able to get decent link history for many sites.
Oops, I just clicked Reply to your comment.
Sentence = "This sentence is semantically and syntactically valid."
P(Sentence) = log(p(START,START,This)) + log(p(START,This,sentence)) + log(p(This,sentence,is)) + log(p(sentence,is,semantically)) + log(p(is,semantically,and)) + log(p(semantically,and,syntactically)) + log(p(and,syntactically,valid)) + log(p(syntactically,valid,.)) + log(p(valid,.,STOP)) + log(p(.,STOP,STOP))
where START and STOP are special symbols that aid in determining the proximity of a word to the beginning and end of a sentence.
If your training set fails to sufficiently generalize, you could use Bayesian inference to estimate the likelihood that the sentence is spam. Under this framework, you'd be calculating the posterior probability of the sentence being spam given the observed sequence of n-grams, which combines (i) the inherent likelihood that any sequence of words is spam and (ii) the compatibility of an observed sequence with (i), which is proportional to the impact it has on (i).
[1] http://storage.googleapis.com/books/ngrams/books/datasetsv2....
Then, run the model over your data and start playing whack-a-mole (and refining the model).
You can also turn off links in comment bodies and the URL field of the comment form to try and prevent scrapers from even finding you a worthy target. Won't help identify spam though.
Finally, centralized spam identification systems like Akismet work really damn well because they are watching the whole site network at once and can use those heuristics to identify spammers rather than the actual spam content itself.
And it is going to fail A LOT.
Do this instead:
1. Contact a company that has a searchengine and therefore access to all your links. ( http://samuru.com ) springs to mind.
2. Do keyword extraction of those pages. Assume that anything that doesn't have any of the keywords of the page that is being linked to is a Bad link.
3. The ones that remain Google the keywords you extracted. (like 10 of the words) if the linking page doesn't appear in the top 50 results it is probably a Bad neighbor according to Google.
This method doesn't require NTLK, or Grammar checking. You can do it your self, and you are using Google to tell you if the site is on the Bad Neighbor list so you don't have to guess.
One of the most linked to page on the internet is the download page for Adobe Reader. It is definitely not spam but millions of those links aren't going to have "the keyword" on the page, so by your logic are bad links. This is an extreme example, but it is not an uncommon scenario.
Furthermore, if you have millions of backlinks, it becomes quite difficult to scrape Google (but you can use services like Authority Labs).
You don't have to scrape Google they have an API for Search that is about $10 per 1000 calls at volume.
I have done this for BILLIONs of links.
Do you have a link to the API, please? Thanks!
If they are buying those links then finding the bad ones is as easy as contacting the people they cut checks to.
Sorry, what do you mean by that?
Then I collected the HTML from those pages and turned it into a feature vector, then tried to learn if a page would have less or more than 5 votes. My prediction rate was as good as random. Fail.
Google does this since Panda. They machine-predict if a page would have success or not with the people and use that as a ranking factor. It makes SEO into a holistic art - you need to think of everything.