From experience I know that models using RNNs have trouble training with FP16 precision. The common solution is to do training in FP32 and inference in FP16. To make this happen you often have to implement custom code (e.g. using Tensorflow or Keras as a meta framework)