Fast word vectors with little memory usage in Python
github.com
github.com
We also have a paper on Magnitude we will be presenting at EMNLP 2018 and ELMo support is coming soon!
How would ELMo work if a neural network needs to be run?
Developers are more familiar with SQLite, it's bundled with Python, it is easy to add to nearly every language, and it has specialized features we use like the Full Text Search module for out of vocabulary lookups.
* reddit (Storage and retrieval of Word Embeddings in various databases): https://www.reddit.com/r/LanguageTechnology/comments/9cf0zk/...
* ... links to this GitHub repo: https://github.com/martialblog/word_embedding_storage
EDIT: sorry it wasn't a PR but an issue
However, newer formats, such as the fastText binary format, store the embedding matrix contiguously. In such formats you can just memory-map the embedding matrix and you only have to load the vocabulary into memory. This is even simpler than the approach described here and in the Delft README, you don't have any serialization/deserialization overhead [1], you can let the OS decide how much to cache in memory, and you have one dependency less (mmap is in POSIX [2]).
[1] Of course, if you have a system with different endianness, you have to do byte swapping.
[2] http://pubs.opengroup.org/onlinepubs/7908799/xsh/mmap.html
By default, LMDB Embeddings uses pickle to serialize the vectors to bytes (optimized and pickled with the highest available protocol). However, it is very easy to use an alternative approach such as msgpack
So, they use serialization/deserialization.
https://radimrehurek.com/gensim/models/keyedvectors.html
The low memory footprint in LMDB is interesting. There were also some exciting papers recently that used compression (vector quantization) [0]. We're still evaluating whether to add those to Gensim [1].
[0] word2bits, https://twitter.com/RadimRehurek/status/976015607941517312
[1] word2bits PR, https://github.com/RaRe-Technologies/gensim/pull/2011#issuec...
You can replace a list of 500K words with 50K ngrams, and it also works on unseen words and agglutinative languages such as German. It's interesting that it can both join together frequent words or split into pieces infrequent words, depending on the distribution of characters. Another advantage is that the ngram embedding size is much smaller, thus making it easy to deploy on resource constrained systems such as mobile phones.
Neural Machine Translation of Rare Words with Subword Units
https://arxiv.org/abs/1508.07909a
A Python library for BPE ngrams: sentencepiece
Edit: this work uses a different byte-layout + parallel reader that heaves the word vecs into memory as compressed byte arrays. Load time is seconds (haven’t benchmarked with current ssds). Memory footprint is on the order of the size of your word vecs (memory is cheap for me, but could easily be extended to support mem-mapping if memory resources are scarce).
Thanks
It actually utilizes an LRU cache (configurable cache size argument on the constructor) so you can utilize an in-memory LRU cache as you query words off-disk using SQLite indexed disk-seeks. And you're right, due to Zipf-ian properties of most text, you can see gains even with a small in-memory cache size :).
https://www.influxdata.com/blog/benchmarking-leveldb-vs-rock...
I think rocks gives better read+ write performance together whereas lmdb is heavily skewed to read performance in our experience -- which is mirrored by the benchmark.
http://www.lmdb.tech/bench/ondisk/
Also the InfluxDB blog post is missing quite a lot of discussion that originally occurred here https://disqus.com/home/discussion/influxdb/benchmarking_lev...