The only point of reference between the two papers I see is the CoVe models, which you guys beat pretty handily, but the ELMo model also beats the CoVe model handily, just on different datasets, so not clear how they stack up.
Any chance you could do some more direct comparisons? You do say that it's a more complex architecture, and the tokenized char convolution stuff is a bit of a pain to do, but if that actually helps, it's not that bad to do once.
From an engineering perspective, not changing the LM weights is kind of nice because then you can train multiple separate models on top of the embeddings without needing to retrain everything (and deal with the associated "noise" when retraining models) and it gives some nice modularity. It would be nice to know how much of the performance is lost from having embeddings that can be shared across a lot of tasks.
Random note: it seems like in Table 7, you have bolded "Freez + discr + stlr" in the IMDb column which has a value of 5.00, whereas "Full + discr" has a value of 4.57, and so should probably be the bolded number.