Rolled out a machine learning model and trained it on the database. 99% of them vanished.
Next day, the machine didn't work and success rate was around 5%.
Found out, they have learned the trick and now using symbols from different languages to make it look like English.
Trained again, success rate went up again.
Next hour, success rate fallen.
This time, they mixed their content with other valid content of our own blogging platform. They would use content from our own blog or other people posts and mix it to fool the machine learning.
Trained it again and was success.
Once a while such content appear and machine model fails to catch them.
It only takes couple of minutes to mark the bad posts and have the model get trained and redeployed and then boom, bad content is gone.
The text extraction, slicing through good content and bad content, finding out symbols vs sane alphabet and many other thing was at first challenging, but overall pretty excited to make it happen.
Through this we didn't use any platform to do the job, the whole thing was built by ourselves, little bit of Tensorflow, Keras, Scikit-learn and some other spices.
Worth noting, it was all text and no images or videos. Once we got hit with that we'll deal with it.
edit: Here's the training code that made the initial work https://gist.github.com/Alir3z4/6b26353928633f7db59f40f71c8f... it's pretty basic stuff. Later changed to cover more edge cases and it got even simpler and easier. Contrary to the belief, the better it got, the simpler it became :shrug