Benchmarking open-weight models for security research
dualuse.dev
dualuse.dev
Qwen 3.6 35B A3B also exceeded my expectations. It's surprisingly performant, even though the previous generation wasn't even able to use the testing harness.
(Tbd on Kimi K2.6; the eval is still running.)