In my benchmark Deepseek-v4-flash did much better than Qwen 3.8 27B at reverse engineering.
"Did much better" is a bit misleading w/o that context and 1 hour time limit -- your benchmark design heavily favors V4 Flash. From results on your page V4 Flash processed 1-1.5M tokens an hour, while Q3.8 27B was failed before even reaching 200K tokens.
By the way, how are you running V4 Flash on single Spark? Was it quantized?
https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-One-DGX-Spark