The system has a compound set of machine learning models and parsers to finally extract from the government official public news (http://in.gov.br) documents with the following structured info:
- who was/will hired and fired (PERSON entity)
- which job role it will/did have. (JOB entity)
- when will happens¹ (DATE entity)
Each entity is extracted individually using a custom trained NER and each sentence is passed to the Relation Extraction system, which is built using SVM. Features are concatenated word vectors² compound by the slices of the text in the form (entity1, entity2, before, between, after).
The system is being alive for almost two years. It produces great results. Just did need retrain the entity recognizer twice in all that time (built using spacy which uses a averaged perceptron).
The SVM part (Relation Extraction) was not retrained since the first day deployed and it still works gracefully :D
¹this info is on the text as natural lang, sometimes is different from the post date ²gensim.Word2Vec custom model trained on this corpus.