Not everything in the compute pipeline is going to be converted to fp16 operations. Anytime you are doing accumulation or exponentials, you would have to have it in fp32.
There was a good talk from NVIDIA at last years GTC: http://on-demand.gputechconf.com/gtc/2018/presentation/s8923...
Here is another relevant blog post: https://devblogs.nvidia.com/mixed-precision-training-deep-ne...
EDIT: Also not everything in the training loop is a matrix multiplication where tensor cores are useful.