In general I used a custom word vector model with a fixed dimension of n=100. As you should now, this feature transformer only maps word to vectors, but I have a more complex structure than just words to work on. The model works sentence wise and preparsing of text is made, let's say I have the following sentence:
"Hire Bill Gates as Front-End Developer at 05/05/2020"
Suppose the NER extracts the following entities from it:
- Bill Gates (PERSON)
- Front-End Developer (JOB)
- 05/05/2020 (DATE)
In my domain problem, I need the relation PERSON-JOB-DATE, which can be decomposed in two binary relations: PERSON-JOB and JOB-DATE. Each binary relation it's a model by itself with the following class outcomes: invalid, hiring, firing. If two binary relations has the same job and outcome classes, I build the triple PERSON-JOB-DATE.
At feature engineering level, which is you asked for, I build a tuple of local attention from entities perspective based on slices of the text: (entity1, entity2, before, between, after). Each part it's transformed into a vector by using average word vector and finally each part it's concatenated in a final vector that will be used by SVM. Since word vector dimension I choose in my experiments was n=100, my model will have n=500 features.
An example about the slices for PERSON-JOB it will be:
("Bill Gates", "Front-End Developer", "Hire", "as", "at 05/05/2020")
1. Each element of the tuple is tokenized
2. With tokens available, use word vector model to transform each one.
3. For each element of the tuple take the average of the vectors.
4. Concatenate each averaged vector into a final vector.
Using that structure the model is biased strongly by how the sentence is written with the words around of the entities. For my domain where the sentences are very regular it worked well. I am not sure if will work in more general domain, like social media.
I hope this clarify your question in some level.