Result: https://www.dropbox.com/scl/fi/8b03yu5v58w0o5he1zayh/pelican...
These are tough benchmarks to trial reasoning by having it _write_ an SVG file by hand and understanding how it's to be written to achieve this. Even a professional would struggle with that! It's _not_ a benchmark to give an AI the best tools to actually do this.
Promoting a pelican riding a bicycle makes a decent image there.
I would not hire a blind artist or a deaf musician.
Like asking you to draw a 2D projection of 4D sphere intersected with a 4D torus or something.
Most the non-math design work of applied engineering AFAIK falls under the umbrella that's tested with the pelican riding the bicycle. You have to make a mental model and then turn it into applicable instructions.
Program code/SVG markup/parametric CAD instructions don't really differ in that aspect.
Ergo no you can't just say throw a bicycle into an LLM and a parametric model drops out into solidworks, then a machine makes it. And everyone buys it. That is the hope really isn't it? You end up with a useless shitty bike with a shit pelican on it.
The biggest problem we have in the LLM space is the fact that no one really knows any of the proposed use cases enough and neither does anyone being told that it works for the use cases.
You too, Monet. Scram.
It's a fun way to deflate the hype. Sure, your new LLM may have cost XX million to train and beat all the others on the benchmarks, but when you ask it to draw a pelican on a bicycle it still outputs total junk.
https://chatgpt.com/share/684582a0-03cc-8006-b5b5-de51e5cd89...
A similar test would be if you asked for the pelican on a bicycle through a series of LOGO instructions.
My CV had a stupid cliché, "committed to quality", which they correctly picked up on — "What do you mean?" one of them asked me, directly.
I thought this meant I was focussed on being the best. He didn't like this answer.
His example, blurred by 20 years of my imperfect human memory, was to ask me which is better: a Porsche, or a go-kart. Now, obviously (or I wouldn't be saying this), Porsche was a trick answer. Less obviously is that both were trick answers, because their point was that the question was under-specified — quality is the match between the product and what the user actually wants, so if the user is a 10 year old who physically isn't big enough to sit in a real car's driver's seat and just wants to rush down a hill or along a track, none of "quality" stuff that makes a Porsche a Porsche is of any relevance at all, but what does matter is the stuff that makes a go-kart into a go-kart… one of which is the affordability.
LLMs are go-karts of the mind. Sometimes that's all you need.
Go kart or porsche is irrelevant.
That's the point.
The market for go-karts does not support Porche.
If you bring a Porche sales team to a go-kart race, nobody will be interested.
Porche doesn't care about this market. It goes both ways: this market doesn't care about Porche, either.