Just spent a few minutes this morning spinning up a H100 model server and trying an FP8 quantized version (including kv cache quantization) to fit it on 2 H100s -- speed and quality looking promising.
I'm excited to see if the better instruction following benchmarks improves function calling / agentic capabilities.