I hacked in support for critbits from this repository:
https://github.com/jgehring/critbit89I modified it to use something like the pool allocator that TFA's hash table implementation used (testing later with and without showed only small performance differences).
I couldn't replicate the dataset from TFA, since it wasn't provided, so I downloaded the java corpus and just took the first 3M symbols in it. Relative performance of the 3 benchmarks from the article were similar.
The result was that it was much more space-efficient than the trie implementation from TFA (only slightly larger than the hash table), but about 1.8x slower than the trie implementation.
A quick -pg run showed 98% of the CPU time in cb_tree_insert which isn't useful for determining why this was so slow, since it's the monolithic function that does all of the insert work other than allocation.