Choosing an AI model: one prompt, 11 models, different results
netlify.com
netlify.com
I'm asking because I personally only use AI with specific and detailed instructions, building my projects piece-by-piece.
I mostly don't look at the low level code and some of it I don't understand as much as I'd like, but I very much give much more technical instructions than a simple, two sentence prompt.
So, it seems to me this "oneshot from a simple prompt" eval is fairly meaningless when it comes to model evaluation itself, as it is in no way representative of real world application.
This would be more something for "vibe coders", people with little to no programming background wanting a website?
It appears that they're mostly testing the ability to make business tasks autonomous, with ~20% associated to development tasks (ssh here, install this, etc.), but not actual programming.
On many of my tests, there was no difference in the result between the smaller and bigger model, but there was a big difference in speed and price.
Much more interesting is providing a million tokens of meaningful input and getting 1000 tokens out (high level critique of a detailed design doc, finding a subtle bug in a big codebase, etc).
I'm interested to know how the classic AI "purple preference" emerged (organically?) and if the beige wave came out of specific training efforts to combat it?
To your point on development work (the code itself), I was talking to some friends on the Google Chrome team about any research understanding the model's preferences around framework ergonomics and abilities to properly implement core web standards for given tasks. I think that would be super fascinating.
You basically have to explicitly say not to use those colors/themes.
1. Long form task based examinations like this that test the ability of the model+harness to remain on task, tool calling, overall effectiveness and taste.
2. More direct 1:1 and qualitative comparisons that you might get with a tool like https://evvl.ai/ - which also uses OpenRouter and does similar one off model comparisons (or lets you use it as a MCP from your dev env to be like: "take the prompt from this loop and try it against these other models")
It's still a work in progress but preliminary results reveal that Sol is able to reproduce 70%-90% of Fable's performance. This is a very meaningful result for me because code review is what I use AI for.
Yes, the most valuable benchmarks and evaluations you can write are those that resemble your work.
The evaluations are extremely hard to write and test.
And yes, virtually all benchmarks are E2E one shots, they do not reflect multi turn processes or how most people interact with LLMs.
Which is why every Opus after 4.6 looks better on benchmarks, but is hard to work with interactively.
As a non-designer trying to build something solo, I think value's in the higher-tier models being able to brainstorm and infer a variety of design directions, as well as dissect the "this looks off" comments that would frustrate human designers.
Some of the frontend design skills try to bridge the gap, but the better models perform far better as is, and often better without one of those frontend design skills trying to ram their own workflows in.
I've been super impressed with one shot AI images and designs in the past, but have never been able to adjust a design without things going off the rails.
No, you are not wrong. If you now crack the "What is?" in a generalizable way, there is very good money in that.
If that's the entire prompt, it's quite depressing how much alike these all look. I appreciate some of the details from the Opus 5 version, but I can't help but strongly feel the AI vibes emanating from that design.
On the other hand, if you want something different with LLMs, all it takes is a few more words of creative flair in the prompt.
> wheelchair
You really took "cripple the model" to heart!
The similarities are beyond coincidence, to the point I'll be scrapping Claude's version of the redesign.
I think websites have always looked mostly alike. It's sorta always been a thing. Reminds me of "Corporate Memphis" (https://en.wikipedia.org/wiki/Corporate_Memphis)
It's alright though, because some people are OK with middle of the road (Wordpress, Boostrap, Squarespace templates, now AI). And others are willing to either pay a developer to get involved or put in the extra effort to differentiate themselves.
My university's Principles of WebDev course 10 years ago had a similar assignment and the results all ended up looking like that too.
>Our default skills also include some UI design guidance, mainly to avoid known gotchas (e.g., the now-dreaded purple AI slop) and get the model to reason about the visual identity appropriate for the user’s ask. But beyond that, each model is free to go build what it thinks we’ll want.
If all you want is "opening hours, the address, a short menu and a photo" there are easier and cheaper ways to do that.
Me, I'm saying that, and I've skipped going to coffee shops and restaurants because they don't have a website, just a fucking Facebook page. I don't use Meta products and can't see their page if I'm not logged into an account I don't have, so I do what the business owner intended: I go fuck myself and get coffee somewhere else.
And why is it so wrong to go to a place to drink coffee without researching it in the first place?
Since all coffee stores are not the same, I look at the website of the coffee shop I am considering visiting, and either I choose it or I don't.
This is the reason the world is slowly becoming a boring ass place and the internet basically consists of 3 websites that are all trying to suck your soul dry.
If I ran a business, a website that stands out from the crowd and provides a good experience would be in my top 3 things to do.
And believe it or not, pre-ground coffee has a risk of gluten cross contamination. That's not a problem with most coffee shops though as long as I get an espresso instead of drip coffee
In my benchmarks, I started insisting on having at least 5 runs.
This also makes me very suspicious of some of the posted benchmarks that compare models. If the results don't say how many benchmark runs were performed and what the variance was, there is really no basis for comparison. You are just guessing.
This means that generic benchmarks and evals are sort of passé. There's no need to have them write "make a coffee shop page" or whatever. Instead, simply focus on the real problem you have and evaluate whether they solve that problem. You want to move along the price-performance frontier alone and the SOTA models are decidedly at the top of the performance curve but very expensive so they serve very well to produce the gold standard and to judge.
The change from prior to today is that benchmaxxing and per-token pricing allowing for evaluation point to the same direction: do not proxy results. Instead deploy the highest end model you have as a judge for others until you have statistical confidence in discrimination and move along the task-specific frontier.
And even for the tasks that are amenable to having LLM judges assess them, there's still a huge benefit in looking at traces yourself and labeling them. It's time-consuming, yes, but you'll learn a lot about the ways an agent fails in your particular domain, it gives you more reliable golden datasets, you have a mechanism to evaluate your judges, etc.
I haven't done this for coding yet (honestly can't really figure out the best approach), but for in-app evals, I built a whole interface to review traces. My app sends up 10k - 30k input tokens on the initial request, between fixed prompt, schema, and dynamic context, so just reviewing it is a huge pain. My interface converts all that raw input data to human-readable objects, highlighting important things and tying together inputs vs outputs for the request. Makes reviewing it, grading it, tagging it, etc. much easier. And then once I've done that for a batch of requests, the passing requests have their inputs and outputs frozen as a golden dataset that I can run other models against. And there the LLM judges are great, because they're very adept at looking at an input prompt, a given model response, and a (human-approved) ground truth, and pointing out any significant differences between the output and the ground truth. Much easier than actually producing the ground truth.
I highly recommend Hamel Husain's writings on evals.
I highly doubt that you have solved it. Writing proper evals and benchmarks for real-world scenarios is far from trivial.
Benchmarking an agent essentially means freezing, at the very minimum:
- the model
- the model's configuration (e.g. effort, permissions, provider)
- the dataset (e.g. a git repository at a specific sha)
- the code running the agent itself (you can build your own harness, trivial, but you still need to ship it as a single executable, froze in time. benchmarking against closed source runtime like claude code is quite useless, they change too frequently and in ways you cannot directly inspect).
- the tools at agent's disposal. Even a slightly different implementation of tool X (e.g. grep or readfile or sed) has an impact. In general this implies also freezing a very specific container image. In my personal benchmarks I provide a specific list of tools that come with the executable, there's no possibility of interacting with the outside world besides the provided apis, the agent bundles its own tools.
And even then: there's significant noise coming from the LLM providers themselves which noticeably change the models behaviour, I don't know whether this is because they optimize some settings or change the inference over time, etc.
And, last but not least, the output of LLMs is non deterministic.
Also, the LLM as judge presents essentially the same non-deterministic problems, has to be benchmarked itself thoroughly, and writing quality rubrics or "golden answers/outputs" is just difficult. One of the metrics I consistently try to emphasize is to avoid the "shotgun vomit dump" of information. So answers that get right to the point in plain terms avoiding dumps of information filled of jargon on top of the user are rated differently.
In short: its far from trivial to benchmark models on real-world agentic work taken from your personal or professional projects.
And even creating the test cases themselves is hard. No: you cannot take the output of some "sota" and use it as gold standard. This is a very crap approach. It's the sloppiest solution to the problem, in the very sense of slop: plausible, average, lacking any creativity or out of the box thinking, the things that make the real difference in complex software development.
The very point of creating these benchmarks is to find which configuration/model/tools/harness (skills/mcps/documentation/agents.md, etc) works better.
And it only works if you create these benchmarks yourself from genuinely difficult non-trivial work and find a solution that is better by most metrics implementation-wise, albeit you could settle on the implementation solving a series of cases and edge cases.
Some designs use the screen estate so ineffectively that only the title and a big, boring, generic graphic is shown on a phone. Users have to scroll all the way down to see the content and find what they need. Better designs show the menu, navigation points, and meaningful, aesthetic graphics. Other designs, such as Gemini 3.6, were quite sophisticated but not optimized for traffic and would not load on a 3G connection. However, a simple static website should load instantly on a mobile connection.
That said, even a simple web page has many requirements, so expecting a turnkey, ready-made design if the user is not guiding the process is not realistic. Thus, I believe the best choice nowadays is a model with good design skills that understands and adheres to an iterative design process, offering a good initial design as a starting point but also prompting the user to provide guidance and feedback. As the design process runs through many cycles, the initial cost should be modest. But more importantly, the model should understand its own design, be able to explain its choices so it can converse with the user using concrete elements in a accurate language to guide the process. IMO, this is still missing from even the frontier models today.
When there are implicit boundaries, negotiable tradeoffs, taste, whatever, in the mix, then the differences in the model capabilities become way more interesting.
There's no consistency over multiple runs of the same prompt on the same model.
Also, for the purpose of music lyric analysis Qwen 4B is laughably bad. Like it's going out of its way to be extremely wrong, misunderstand the prompt. When it does correctly understand what I asked for it ALWAYS tells me that the mood is Angry. Sometimesiit just gives back all the lyrics. Sometimes it claims that it doesn't have the list of moods or the lyrics and tells me I should look them up on the Internet first.
All with the same prompt every time.
Are models being overtuned for coding tasks?
I haven't tried it but I think this can be replicated with a system prompt.
I remember the codex system prompt contains something like, "Do not consider a task complete until you have verified the result."
Although I've been running the new GPT models in a custom harness and they do that anyway now, without being prompted. So I think that prompt was for a previous generation.
Take from that what you will - and maybe the AUE is more desirable - but Sol & K3 are genuine game changers for existing codebases. Anthropic currently have an antagonism problem which has worked for them in the past, but not anymore I think, as other models have become as competent.
Come up with an arbitrary test, let a bunch of LLMs work on it. Make some very subjective judgement about the result...
Kimi, GLM, and Deepseek all absolutely run away with the "Can I quickly read the menu and find the address" challenge.
Most of the rest of the pages are stylistic, but hard to parse.
If I were in a car on a mobile phone trying to find the address of the place to meet a friend for coffee... I don't want a bunch of fluff and stylistic design that makes it hard to parse the information on the site.
How so? Surely they can just steal such generic graphics off existing web sites.
LLM