I'd be interested for any LLM emitting any kind of text-to-picture instructions to get results that are beyond a kindergartner-cardboard-cutout levels of art.
Claude 3.5 Sonnet is in second place: https://github.com/simonw/pelican-bicycle?tab=readme-ov-file...
This is what o1-pro yielded: https://gist.github.com/carbocation/8d780ad4c3312693ca9a43c6...
My personal test has been "A horse eating apples next to a tree" but the deliberate absurdity of your example is a much more useful test.
Do you know if this is a recognized technique that people use to study LLMs?
https://int19h.org/chatgpt/lakeside/index.html
One interesting thing that I found out while doing this is that if you ask GPT-4 to produce SVG suitable for use in HTML, it will often just generate base64-encoded data: URIs directly. Which do contain valid SVG inside as requested.
The most significant part I took away is that when safety "alignment" was done the ability plummeted. So that really makes me wonder how much better these models would be if they weren't lobotomized to prevent them from saying bad words.
Isn't it just like any kind of conversion or translation? Ie. a relationship mapping between diffrent domains and just as much parroting "known" paths between parts of different domains?
If "sun" is associated with "round", "up high", "yellow","heat" in english that will map to those things in SVG or in whatever bizarre format you throw at with relatively isomorphic paths existing there just knitted together as a different metamorphosis or cluster of nodes.
On a tangent it's interesting what constitutes the heaviest nodes in the data, how shared is "yellow" or "up high" between different domains, and what is above and below them hierarchically weight-wise. Is there a heaviest "thing in the entire dataset"?
If you dump a heatmap of a description of the sun and an SVG of a sun - of the neuron / axon like cloud of data in some model - would it look similar in some way?
I don’t think it reflects any understanding. But to go from screenshot to conceptually accurate and working code was impressive.
https://gist.github.com/uschen/38fc65fa7e43f5765a584c6cd24e1...
Copied SVG from gist into figma, added dark gray #444444 background, exported as PNG 1x.