Thanks for pointing the memory optimizations out. The Google model trained on 100 billion words unzips as a 3.4 GB binary, which often is a non-starter on smaller servers, while these vectors indeed come in at less than 100 MB. Mainly this is because only the most common 100k of the 3m vectors distributed are in here, but also because 32 bit floats for each of 300 elements per vector are converted to 16 bit floats, with virtually no loss in precision. (Research suggests we may be able to compress even further, perhaps to 5 bits of entropy per vector element.) In practice, across a wide variety of words and phrases, this pipeline can quickly derive new vectors by web searching, and thanks to the OP's HDBSCAN, can cluster them quite well. Typically it's the quality of keywords found online that ends up most limiting the quality of maps derived here, while there's probably room for breakthroughs in the keyword extraction algorithm put together here (which relies on word2vec indexes as a proxy for the idf in tf-idf). In any case if greater precision is needed - like when I wanted to map ~50 really great scientists who deserve all the credit here - one trick that can help is to increase the number of websites to scan per unknown word (e.g. from 10 to 50) in the research_keywords function.