Wrote this up in a bit more detail on my blog, including some thoughts on what value the pelican benchmark can still provide here: https://simonwillison.net/2026/Jul/16/kimi-k3/
Try setting reasoning levels yourself manually. We see in the benchmarks that one of the graphs shows low, mid, max, so its clearly there.
I had the same issue with GLM 5.2 only offering high/max.
By playing around with openai compatible protocol, and setting the reasoning level from none, low ... high, xhigh and testing a flawed logic test.
It was easy to see that GLM had all the different reasoning levels. Low was like one line, medium did a few, high started to really expand, xhigh was a page or 2, max was MAX.
Very sure that you can force K3 into using less reasoning.