Nice work accelerating convolutional models! It might be better to see (or cite papers about) the trade-off how model performance (accuracy, etc) changes w.r.t. how it is quantized.
With learning-based program optimizer, we can competitive performance on benchmark models and significant boost on emerging models against TensorRT(int8).