I'm curious whether actual inference workloads actually push to 600W (and not 350W) and what the last 250W get you. Rare is the (generic gpu) workload where I get >5%, some rare light inference benchmarks up to 10%...
They don't usually max out for single prompts, but concurrency in the form of multiple users and/or multiple agents will make them draw full power.
Yes that was the second part of my comment. I've been benchmarking many things AI or not on those boards and while one can sometimes get to 600W I haven't seen more than 10% gain for those last 250W. If someone has a workload that gets more from those watts I'd be interested. For now on anything I run (full-capacity, continuous, batched...) capping at 350W seems better in bang-for-bucks when including energy costs (from the GPU + added cooling).