Looks great!
Dumb question: How does Weaviate know that "Scandinavian" is close to "Finnish" ? The source not having "Scandinavian" at all. If their vectors are close, then the "vectorization" is quite standard for any text, and also per language?
Dumb question: How does Weaviate know that "Scandinavian" is close to "Finnish" ? The source not having "Scandinavian" at all. If their vectors are close, then the "vectorization" is quite standard for any text, and also per language?
EDIT: Just realized I didn't answer the second part of your question. Yes, the models are language-specific, but there are also multilingual models that work across a large no. of different languages.
[1] Sentence-BERT: https://sbert.net
[2] Weaviate Customizer with Out-of-the-box models: https://www.semi.technology/developers/weaviate/current/gett...