we got this running on 4 RTX Pro 6000's and for single request we're getting around 250 tok/s we can support about 48 concurrent requests we're seeing around 2400 agg tok/s peaking around 24-31 concurrent users. Model performance feels like gpt 5.4 - mostly using it with pi agent. the only thing i'm missing with this model is vision and i see some folks have done some work like
https://huggingface.co/webbrain-one/DeepSeek-V4-Flash-0731-V... but have not yet tried it out.