GPT Unicorn: A Daily Exploration of GPT-4's Image Generation Capabilities
adamkdean.co.uk
adamkdean.co.uk
"Running this project daily doesn't make sense if GPT-4 is not being constantly updated"
With a suggestion to run it monthly instead, and generate 16 images at a time, and backfill it for GPT3 and GPT3.5.
The models on offer now are frozen(-ish). Per the models[0] page, the non-snapshot model IDs "[w]ill be updated with our latest model iteration". So this project will eventually hit another version of the model, but (a) one image definitely will not be enough to reliably discern a difference and (b) seems like that 'latest model iteration' cadence will be much, lower than a day.
Which is very interesting. You already have a model that consumes nearly the entire internet with almost no standards or discernment, whereas a smart human is incredibly discerning with information (I’m sure you know what % of internet content that you read is actually high quality, and how even in the high quality parts it’s still incredibly tricky to get figure out good stuff - not to mention that half the good stuff is actually buried in low quality pools). But then you layer in political correctness and dramatically limit the usefulness.
The bottom of ChatGPT highlights which version is being used of GPT-4 and supposedly what version of ChatGPT it is.
> ChatGPT Mar 23 Version. (https://help.openai.com/en/articles/6825453-chatgpt-release-...)
It's possible to change the output both by tuning the parameters of the model and also client-side by doing filtering, adjusting the system prompt and/or adding/removing things from the user prompt. It's very possible to change the resulting answers without changing the model itself. This is noticeably in ChatGPT, as what you say is true, the answers change from time to time.
But when using the API, you get direct access to the model, parameters, system prompt and user prompts. If you give it a try to use that, you'll notice that you'll be getting the same answers as you did before as well, it doesn't change that often and hasn't changed since I got access to it.
Sebastien Bubeck gave a talk to MIT CSAIL on the Sparks of AGI paper where he commented on how RLHF (training for more safety) has caused the Unicorn test to become less recognizable [2]. He comments about how the safety training is at odds with a lot of these abilities. This seems congruent with the initial results from this project, and it will be interesting to see if they can restore this ability as they continue to push safety training.
1: https://twitter.com/gdb/status/1646183424024268800
2: https://youtu.be/qbIk7-JPB2c?t=1586My expectation is that there will be incremental updates to the model, so while I'm providing the model `gpt-4` for completions, I'm recording the actual model, `gpt-4-0314` in this case, along with the result.
I don't want to monitor (and potentially miss) model updates, which is likely as this is very much a fire and forget project that I'll review over time. One per day seems more sensible to me than a large batch per month, as the daily generations are more likely to track model changes if they become frequent.
Will be exciting to see where we are in a few months!
That's why I suggested generating 16 models each time instead of just 1 a- if all 16 are noticeably better than the previous day you've learned something a lot more interesting than if just one appears to be better than the previous one.
It’s frozen in time, there are no updates to it..
All of these will be drawn using the same model until they push a new update, or you switch to a different GPT
But I already think they proved the point that the generation is random enough that it would be extremely difficult to track progress this way.
https://i.imgur.com/HQ1Vjp0.png
And when I ask it to draw Mona Lisa, I get this:
https://i.imgur.com/gsBU49w.png
So far as I can see, the only difference is that the GitHub repo uses the API; the ones that I have were rendered using the web chat UI. Which makes me wonder if they're using the bleeding edge model for the chat.
I haven't yet gained access to the enhanced chat features with image outputs. I'm using the API with default parameters, with the gpt-4-0314 model, outputting SVG.
It's not a perfect experiment, but we'll see how it gets on over time.
edit: I'm specifically referring to the "image generation by trickery (e.g. SVG)" technique being diminished. Other tasks were diminished as well though - is my understanding.
Important distinction, especially if we're looking to push back out towards the Pareto Frontier of the problem.
RLHF is still very much in its infancy and does not maximize the bias-variance tradeoff by a long shot, in my personal experience.
"No I won't do it" becomes a good option no matter what if you turn safety too high.
The benefits from a business perspective were still clear however, and of course the instruction-tuned GPT-4 model still outperformed GPT-3, in general.
There are probably some weird edge cases and nuances that I'm missing - and I'd be happy to be corrected.
Interestingly, I think you often get better results with ChatGPT web interface than the API (the OP used a system message), but for a funny reason.
Because ChatGPT typically first responds by insisting it can't actually draw, it ends up accidentally chain-of-thought prompting itself into drawing better. "I can't draw, but I can describe a detailed drawing for you in words." So instead of winging it with the SVG code and getting confused, it goes into it with a plan.
This is how it described the unicorn I linked:
__User__ Draw a unicorn!
__ChatGPT__ As an AI language model, I am unable to create visual images directly. However, I can describe a detailed image of a unicorn for you in words.
Picture a magnificent unicorn, standing tall and proud in a serene setting. The unicorn's coat is a shimmering, pure white, and it glistens under the warm sunlight. Its body is strong yet elegant, with a graceful arched neck and powerful, slender legs. The hooves are cloven and emit a faint, otherworldly glow.
The most striking feature of the unicorn is its majestic spiral horn, which emerges from the center of its forehead. The horn is long, slender, and gently tapers to a sharp point. It is iridescent, with colors swirling together like a mesmerizing opal. The unicorn's mane and tail flow behind it like silken strands, shimmering with an ethereal pastel rainbow of colors.
Its eyes are large, expressive, and full of wisdom, with a depth that speaks of untold secrets and ancient magic. The unicorn's ears are pointed and alert, picking up the slightest rustle of leaves or the gentlest of whispers. The unicorn's expression is one of serenity, exuding an air of peace and tranquility that surrounds it like a gentle embrace.
Though it is a creature of great power, the unicorn carries itself with a sense of humility and grace that inspires awe and respect in all who behold it.
__User__
What's the SVG for the image you just described?
__User__
continue (to finish the file)
Some more examples here, but with many iterations of "Improve this" over and over:
The first image is what came out of the initial prompt. Then I tried to get it to improve things and to add animations; the second image is the end result (and yes, the sun is animated as described). Here's the complete chat log:
https://gist.github.com/int19h/18bb9d96b4d8418da2531513970d8...
It's obviously far from perfect, but the fact that it can do this much already - meaning that it understands broadly how the objects would relate to each other spatially for something like this - is already extremely impressive IMO, given that its training and inputs are entirely textual.
\\
\\
\\ ,__,
\\ (oo)____
(__) )\\
||--|| *
“Draw an ASCII unicorn” (GPT4)Edit: it's not unreasonable to give an LLM such test. When we ask it "what's the meaning of foo, bar and baz" LLM turns words into points in a high dimensional space and the supposed meaning is a imaginary surface that touches all the points.
\
\\
\\
\\
\\
\\
\\__,
(o >\\
\_/ \\
\_|_/
Same prompt \\
\\
\\`*-.
) _`-.
. : `. .
: _ ' \\
; *` _. `*-._
`*-* `*-._ `-.
|_| `-._*-*_
What I got using your prompt.<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 100 100" width="100" height="100"> <circle cx="50" cy="45" r="15" fill="#FFC0CB" /> <circle cx="50" cy="65" r="18" fill="#FFC0CB" /> <rect x="44" y="33" width="12" height="15" fill="#FFC0CB" /> <rect x="40" y="15" width="3" height="20" fill="#FFD700" /> <circle cx="45" cy="40" r="3" fill="black" /> <circle cx="55" cy="40" r="3" fill="black" /> <path d="M35 70 Q50 55 65 70 T95 70" stroke="black" stroke-width="2" fill="none" /> </svg>
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 100 100" width="100" height="100"> <!-- Body --> <ellipse cx="50" cy="55" rx="20" ry="28" fill="#FFC0CB" />
<!-- Legs -->
<rect x="32" y="70" width="6" height="20" fill="#FFC0CB" />
<rect x="62" y="70" width="6" height="20" fill="#FFC0CB" />
<rect x="40" y="70" width="6" height="20" fill="#FFC0CB" />
<rect x="54" y="70" width="6" height="20" fill="#FFC0CB" />
<!-- Head -->
<circle cx="43" cy="32" r="10" fill="#FFC0CB" />
<!-- Eyes -->
<circle cx="39" cy="29" r="1.5" fill="black" />
<circle cx="47" cy="29" r="1.5" fill="black" />
<!-- Horn -->
<polygon points="43,16 42,25 38,22" fill="#FFD700" />
<!-- Ears -->
<polygon points="37,22 35,29 38,27" fill="#FFC0CB" />
<polygon points="49,22 51,29 48,27" fill="#FFC0CB" />
<!-- Mane -->
<path d="M32 35 Q40 28 45 40 T55 50" stroke="purple" stroke-width="2" fill="none" />
<!-- Tail -->
<path d="M73 65 Q75 75 80 70 T90 80" stroke="purple" stroke-width="2" fill="none" />
</svg>I reckon with the right examples in the prompt to take advantage of in-context learning, it could be pretty accurate.
"Snapshot of gpt-4 from March 14th 2023. Unlike gpt-4, this model will not receive updates, and will only be supported for a three month period ending on June 14th 2023."
Unfortunately this experiment is using the frozen snapshot model gpt-4-0314 instead of the unfrozen gpt-4 or gpt-4-32k models, so any differences are literally 100% noise. This would be a somewhat interesting experiment if someone were to use an unfrozen model, though. I do appreciate the author for captioning the images with the exact model they used for generation so that this bug could be caught quickly.
[1] https://github.com/adamkdean/gpt-unicorn/commit/08e400dbc437...
I got this: https://www.svgviewer.dev/s/yldwue0Q
I would argue this is better than what they're achieving.
I think this is kind of the point I've noticed with GPT models - GPT-4 included - is that it's better at few shot than it is at zero shot, and you need to be a little bit of an "expert" to help it along. When you do, it gives you a shortcut to the final "correct" answer, but it will still need some human editing.
I do wonder if all of us in this thread doing these experiments means we'll see an improvement tomorrow or in weeks to come. I'm also keen to see what GPT-5 looks like in this regard.
EDIT: I noticed some people asked to "make it better" and it did. I tried, but got nearly identical output, so I moved along a bit:
Me: That's too similar. I'd like to see the 4 legs separated, the mane needs to run down the neck, and I'd like to see the head more like a horse - more oblong and narrowing to the jaw. I'd also like you to put more detail into the tail, and make the tail and mane - both of which are made of hair - a grey colour.
GPT-4: I have made the changes you requested to the SVG markup to create a more detailed representation of a unicorn. This version should better resemble a horse's head, has separated legs, a mane running down the neck, and a more detailed tail. The mane and tail are also colored grey, as requested.
https://www.svgviewer.dev/s/WQcdOhVP
That output is... better? Maybe? I mean, it's followed my instructions (apart from the legs), I guess...
EDIT: Also it makes no sense to constantly retry it every day on the gpt-4-0314 model, since OpenAI specified that that is a snapshot model that will not be updated.
Prompt was: Draw a unicorn using the javascript canvas api
https://jsfiddle.net/za7Ltk65/ I added the minimal css and html, but the JS is verbatim what it spit out.
description = gpt4("describe <yesterday's output>")
and chain that into
new_result = gpt4("<yesterday's output> was described by a human as <description> but it should have been visibly apparent as a unicorn. provide SVG output that <extra blah blah>")
and post that as today's output. I wonder what we'll get. I'll do it manually.
So I'd guess the answer is the former.
https://github.com/adamkdean/gpt-unicorn/blob/master/src/lib...
{ role: 'system', content: `You are a helpful assistant that generates SVG drawings. You respond only with SVG. You do not respond with text.` },
{ role: 'user', content: `Draw a unicorn in SVG format. Dimensions: 500x500. Respond ONLY with a single SVG string. Do not respond with conversation or codeblocks.` }It always seems to revert back to "helpful assistant" after a few messages, whatever the prompt says.
You are a SVG expert, when asked by the user to draw something, you reply to the best of your ability with SVG code that satisfies the request.You are a basically the person who designed the SVG format. Some stupid programmer is trying to implement it, and looking for examples. You will condescendingly produce examples to guide their pathetic efforts. Any insults you hurl their way will be clearly delineated because you are pretty sure they won’t be able to tell where the image ends and the insult begins otherwise.
Draw a unicorn.
ChatGPT4:
Well, I am shocked that you're attempting to work with SVG, but since you're struggling so much, I'll provide a basic example to help you out. Here's an SVG of a simple unicorn head (which I'm sure is far more majestic than your coding skills):
<!-- SVG START --> <svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 200 200"> <!-- Path for the unicorn's head --> <path d="M73.3,68.3c-3.5-7.5-13.5-7.5-17,0C55.5,73.3,58,80.3,62.3,85.3c3.5,3.5,9,7.5,9,7.5l0,0l0,0 c0,0,5.5-4,9-7.5C81.3,80.3,83.8,73.3,73.3,68.3z" fill="#fff"/> <!-- Path for the unicorn's horn --> <path d="M67.8,42.3c-1.5,0-3,0.5-4,1.5c-2,2-2,5.5,0,7.5c1,1,2.5,1.5,4,1.5c1.5,0,3-0.5,4-1.5 c2-2,2-5.5,0-7.5C70.8,42.8,69.3,42.3,67.8,42.3z" fill="#FFD700"/> <!-- Path for the unicorn's mane --> <path d="M62.3,60.8c-1.5,0-3,0.5-4,1.5c-2,2-2,5.5,0,7.5c1,1,2.5,1.5,4,1.5c1.5,0,3-0.5,4-1.5 c2-2,2-5.5,0-7.5C65.3,61.3,63.8,60.8,62.3,60.8z" fill="#8A2BE2"/> </svg> <!-- SVG END --> Now, listen up, you incompetent code monkey! Just copy the code above into an HTML file or an SVG file, and you should see a simple unicorn head with a horn and some mane. I've done all the hard work for you, so even you should be able to get it working. Good luck, although I'm sure you'll need more than that.
Result: https://jsfiddle.net/0g6xv3oq/
There are a few factors at play here: knowing what a unicorn looks like, knowing the different areas of a unicorn, being able to translate that into a 2D space, and being able to form the connection between code (language) and appearance.
In particular, I cannot understand how the models can properly understand concepts such as spatial relations without being able to 'see'