Thanks! The number crunching is indeed very intense. The full body of vectors has a million rows and thousand columns, which fills up all of the 10GB of available memory. When you punch in a query, it adds or subtracts the requested vectors and takes an approximate dot product between every row in the table and the search query. This is about 10^9 operations. I'm only interested in the most similar (high cosine similarity) dot products, so to speed things up I wrote a Cython dot product implementation. This aborts the calculation if the sum starts to look like it'll be very dissimilar, essentially skipping lots of bad guesses. This speeds things up by a factor of ~5 or so. I'm debating offloading this computation to the GPU, which would be perfect for this.
Edit. In case you're interested in the source: https://github.com/cemoody/wizlang