I love Spacy, but I was recently trying to embed documents with it and it seems that it uses word-vector averaging for that. Is it possible to do sentence or paragraph embedding like sentenceBert?
Document similarity is a tricky thing to get into the main API though, which is organised around the idea of sending a Doc through a sequence of steps, each of which adds or updates annotations. The signature for document similarity is fundamentally different. In the end we decided to not try to shoe-horn it in. Not everything needs to be one function call. So document relation predictions (including document similarity) wouldn't be a standard pipeline component.