Love this demo but as others noted it's really easy to find queries where it performs poorly (e.g. typos).
Looks like the embedding model used (all-minilm-l6-v2) currently ranks 35th on the hugging face leaderboard [0]. I'd love to try with other models if anyone wants to +1 this demo :). This feels like a nice dataset to build intuition around embeddings used for RAG etc.