Pelicans on a bicycle
simonwillison.net
simonwillison.net
Plus, simonw isn't exactly a meaningless nobody in this space, and his writeups are more detailed and actionable, and therefore identifiable, than some random "hey a great LLM benchmark would be creating an SVG of a walrus twerking in front of a jelly bean store" throwaway comment.
Proof: I asked ChatGPT 4o the question "What are some users who post ad hoc LLM benchmarks to technical discussion sites, and what benchmarks have they proposed?" simonw is in the list, 1 of 7 individual people it suggested. (The proposed benchmarks listed for him were more general than the specific one here: "Testing LLMs’ capabilities with code generation, particularly in niche languages or against real-world API schemas." But it's easy to imagine followup queries bringing this one up.)
Don't run unpublished private benchmarks or worry about keeping a counted hoard of secret questions. Do rotate your questions every few months to whatever comes to mind at the time. When nothing comes to mind there is no point in running a question benchmark anymore as it already answers every possible you could possibly question you can think of (and the only way it gets there in your lifespan is by reasoning rather than memorization). You can always run the new question retroactively on an old models for comparison purposes so that's not a concern either.
The important thing here being "rotate questions without concern of having things lined up for it" rather than "fear what happens when you discuss your question".
> You can always run the new question retroactively on an old models for comparison purposes so that's not a concern either.
Good point, but it's still somewhat of a concern. Especially if you're benchmarking via a chat interface, my understanding is that there are plenty of finetuning, system prompt, safety, and other post-training changes that can influence results. (Perhaps unlikely with an SVG generation prompt, but reasonably likely with "reasoning" prompts.)
> Do rotate your questions every few months to whatever comes to mind at the time. When nothing comes to mind there is no point in running a question benchmark anymore as it already answers every possible you could possibly question you can think of
Isn't that assuming it's answering your questions reasonably well? On a hard benchmark, you may be watching the progression from absolutely miserable to not quite tolerable.
You may even be using a constellation of questions to try to delineate the boundary between what it can and can't do.
64gb ram is crucial, after that, need 1+ tb storage, and then?
Chip Bandwidth (GB/s)
———- ————————————
M2 100
M3 100
M4 120 (20% more)
M2-Pro 204
M3-Pro 153 (less than M2-Pro)
M4-Pro 273 (78% more than M3-Pro)
M2-Max 409
M3-Max 409
M4-Max 546 (33% more than M2/M3-max)
https://arstechnica.com/apple/2024/10/apples-m4-m4-pro-and-m...There are genuinely useful applications of SVG-generation from LLMs - outputting simple infographics or charts for example.
I use LLMs to write HTML all the time, of which SVG is a useful optional component.
For example, if I asked you to assemble a bookshelf with some wood, nails, and cement, you might first make a hammer with the cement before trying to assemble the bookshelf.
You can get a much better image by first asking the (multimodal) LLM to draw an image of a pelican on a bicycle, and then generate an SVG using the referenced image.
https://chatgpt.com/share/67609300-9abc-800d-9b26-95074f2149...
Tools are defined by what people use them for, not by how they were intended—or designed—to be used. (Just ask Nvidia)
adding: so I think someone comparing how various tools perform at a task that's valuable to them—and probably others—is just fine, even if it's different from what the creator of the tool intended?
[1]: https://github.com/simonw/pelican-bicycle/blob/main/README.m...
FLUX, for all its benefits, is really annoying in this aspect as you cannot really provide negative guidance without complicated hacks [1]. Therefore, "orignal" ideas that go against popular wisdom or mainstream culture will be very hard to prompt for. This is, in my opinion, an important caveat of ML "art". ML "art" is already considered useless slop because of how much it conforms to biases built in the weights. The inability to tune down adherence to mainstreaming norms is therefore all the more problematic.
[1] https://www.reddit.com/r/StableDiffusion/comments/1estj69/
https://web.archive.org/web/20240419001426/https://www.wired...
So the fact that the AI models screw this up so badly is understandable. Sure, they screw up in ways that humans wouldn't, such as the beak backwards in one of the pictures (pointy end toward the bird!) because they don't know or care about something every human would know: What a beak is for and what it looks like in general. Or for that matter the biodynamics of how a pelican's long, spindly legs could, in fact, work a pair pedals. But ask me to draw a pelican from memory, and have a good laugh (if you're better at it than me) because to me, they're just kind of a peripheral vision, pink abstraction, not something I focus on understanding. And that's what they are to the AI model too.
are there pink pelicans, or are you thinking of flamingoes?
> a pelican's long, spindly legs
and when you then also described it as pink, it started to make sense and I too understood that you must be thinking of flamingos :^)
Mo link, sorry, but on youtube, GCN asked pro riders to draw a bicycle...none could.
Whereas a pro rider can probably tell you all about the biomechanics of how to optimally interact with the bike, the right foods to eat and how much to sleep and when. But the actual wrenching around with them? That's the pro mechanic's job.
I also suspect it strongly correlates with knowing the term "diamond-frame". In addition to bicycle-repairers probably knowing the term, it's also used among people who like/know other frame styles--in my case recumbent bicycles.
But, what about this workflow: given prompt, LLM generates two SVG outputs. Both are rendered by an SVG renderer, and then we combine the two into one image, one on the left and the other on the right. We then ask a visual LLM (could be the same LLM or could be a different one) to tell us whether the left half or right half of the image is a better response to the prompt. Now we've got preferences which can be used to fine-tune the LLM using DPO. And you could iteratively repeat the process – as the LLM is fine-tuned it may produce even better outputs which then produces new preferences for further fine-tuning.
Would be interesting to see what kinds of results it might produce in practice.
I would not be surprised to see that the LLM generating SVG does a UNO reverse, and makes use of little colored squares in a grid to draw a “vector image” where each of the squares represents individual pixels :p
maintains website collecting SVG files of pelicans on bicycles
Remember that these are basically one-shot. Very different to how you or I would solve the problem (get a circle up on the screen, have a look at it, make some changes, add some wings, tweak the dimensions, etc.). We would go through hundreds or thousands of feedback cycles before we got something half-decent -- in this situation the model only gets one attempt.
If I had the latest version of Illustrator then I would consider seeing how well its image generation does, but I do not because it has a lot of exciting new bugs that break my normal workflow. I believe that under the hood that works by feeding your text prompt to a bitmap image generator and running the same old autotrace on it, which results in some pretty messy and hard-to-edit shapes.
Latest Claude does a suspiciously good job…