I know it's only on a single benchmark, but I dont understand how it can be so bad...
I know it's only on a single benchmark, but I dont understand how it can be so bad...
HN is the last bastion of serious inquiry these days. But its not immune as OPs comment proves.
The models not availble on copilot were tested through opencode (max reasoning) and deepseek v4 was tested through Cline (with max reasoning too).
I really like this benchmarking. Have you evaluated the judge benchmark somehow? I'd love to setup my own similar benchmark.
I haven't evaluated the judge benchmark. You have everything needed in the repo to do so though, so be my guest. It took me a bit of time to put all this together and won't have much more time to dedicate to it before a couple of weeks.
BTW, if you explore the repo, sorry for all the French files...
Your prompt is extremely slim yet you score it on a bunch of features.
The eval prompt is quite extensive: https://github.com/guilamu/llms-wordpress-plugin-benchmark/b...
I personally develop with very detailed spec, and I don’t want nothing more and nothing less compared to the spec.
I found 5.4/5.5 much better at following spec while Opus makes some things up, which aligns with your benchmark but that makes 5.4/5.5 better for me while worse for you.
What strike me as very strange though is that 0 model were able to just use the search input already present in GravitYForms forms list page and all created a second input.
Also, I know it's not in the prompt, but adding a ctrl+f shortcut to a search input? Is that that crazy? I don't know.