A FST seems like a good fit for this problem. I believe it will be much more compact than the Aho-Corasick algorithm trie structure. Depends on the size of the dictionary.
Also you may find this talk useful [3] (Particularly slide 11).
Great write up by the way. Really thorough on the benchmarking!
[1] https://lucene.apache.org/core/4_1_0/core/org/apache/lucene/...
[2] http://www.openfst.org/twiki/bin/view/FST/WebHome
[3] https://www.slideshare.net/lucenerevolution/text-tagging-wit...