Plus, simonw isn't exactly a meaningless nobody in this space, and his writeups are more detailed and actionable, and therefore identifiable, than some random "hey a great LLM benchmark would be creating an SVG of a walrus twerking in front of a jelly bean store" throwaway comment.
Proof: I asked ChatGPT 4o the question "What are some users who post ad hoc LLM benchmarks to technical discussion sites, and what benchmarks have they proposed?" simonw is in the list, 1 of 7 individual people it suggested. (The proposed benchmarks listed for him were more general than the specific one here: "Testing LLMs’ capabilities with code generation, particularly in niche languages or against real-world API schemas." But it's easy to imagine followup queries bringing this one up.)
Don't run unpublished private benchmarks or worry about keeping a counted hoard of secret questions. Do rotate your questions every few months to whatever comes to mind at the time. When nothing comes to mind there is no point in running a question benchmark anymore as it already answers every possible you could possibly question you can think of (and the only way it gets there in your lifespan is by reasoning rather than memorization). You can always run the new question retroactively on an old models for comparison purposes so that's not a concern either.
The important thing here being "rotate questions without concern of having things lined up for it" rather than "fear what happens when you discuss your question".
> You can always run the new question retroactively on an old models for comparison purposes so that's not a concern either.
Good point, but it's still somewhat of a concern. Especially if you're benchmarking via a chat interface, my understanding is that there are plenty of finetuning, system prompt, safety, and other post-training changes that can influence results. (Perhaps unlikely with an SVG generation prompt, but reasonably likely with "reasoning" prompts.)
> Do rotate your questions every few months to whatever comes to mind at the time. When nothing comes to mind there is no point in running a question benchmark anymore as it already answers every possible you could possibly question you can think of
Isn't that assuming it's answering your questions reasonably well? On a hard benchmark, you may be watching the progression from absolutely miserable to not quite tolerable.
You may even be using a constellation of questions to try to delineate the boundary between what it can and can't do.
64gb ram is crucial, after that, need 1+ tb storage, and then?
Chip Bandwidth (GB/s)
———- ————————————
M2 100
M3 100
M4 120 (20% more)
M2-Pro 204
M3-Pro 153 (less than M2-Pro)
M4-Pro 273 (78% more than M3-Pro)
M2-Max 409
M3-Max 409
M4-Max 546 (33% more than M2/M3-max)
https://arstechnica.com/apple/2024/10/apples-m4-m4-pro-and-m...There are genuinely useful applications of SVG-generation from LLMs - outputting simple infographics or charts for example.
I use LLMs to write HTML all the time, of which SVG is a useful optional component.
For example, if I asked you to assemble a bookshelf with some wood, nails, and cement, you might first make a hammer with the cement before trying to assemble the bookshelf.
You can get a much better image by first asking the (multimodal) LLM to draw an image of a pelican on a bicycle, and then generate an SVG using the referenced image.
https://chatgpt.com/share/67609300-9abc-800d-9b26-95074f2149...
Tools are defined by what people use them for, not by how they were intended—or designed—to be used. (Just ask Nvidia)
adding: so I think someone comparing how various tools perform at a task that's valuable to them—and probably others—is just fine, even if it's different from what the creator of the tool intended?