Encoding both word and image embeddings into the same index then doing ANN on that index might also work. See this example of text-to-image retrieval: https://paperswithcode.com/task/texture-image-retrieval
Maybe someone from Dropbox can add more color to this and explain what other options they considered.
As it stands, I still can't find what I need in Dropbox. And never could. From reading this article I'd think searching for a basic keyword like "dog" or "ship" or "runner" would yield some results from my tens of thousands of photos, yet I get nothing (nothing relevant, at least).
Edit: On second reading, this is only available to Dropbox Pro and Business users. I hope they roll this out to other paying users soon.
Encoding words and images into the same space and doing ANN is kind of what the current system is, if you look at it right. The ANN is framed in terms of similarity rather than distance -- and is approximate because of the sparseness approximation. But the big difference from the papers you linked is what we use as the encodings: not the traditional penultimate layer of a network, but classifier scores for images and projected word vectors for text. This gives us a space with semantically meaningful dimensions, which lets us build the system without a large multimodal training set; our text and image models are independently trained on different datasets.
1. Did you look at CLIP? it provides a common (to images & text) embedding.
2. Do your models need specialized training (vs. open models)?
Assuming you are searching just a single account, there are unlikely to be more than 1 million images in a typical account.
A simple linear scan of all of those feature vectors should be a simple matter of milliseconds.