> This project evaluates local language models running on a single NVIDIA DGX Spark.
"Did much better" is a bit misleading w/o that context and 1 hour time limit -- your benchmark design heavily favors V4 Flash. From results on your page V4 Flash processed 1-1.5M tokens an hour, while Q3.8 27B was failed before even reaching 200K tokens.
By the way, how are you running V4 Flash on single Spark? Was it quantized?