GPT Unicorn has drawn a unicorn
gpt-unicorn.adamkdean.co.uk
gpt-unicorn.adamkdean.co.uk
Edit: This comment section is a super fascinating case study on the inherent flaws in human cognition. Especially when it comes to seeing patterns in random noise. The fact that some people believe that the model really has to have changed in the past few days is amazing, because if you've kept up with the GPT architecture and the way OpenAI does things (especially on the API), it is incredibly obvious that nothing has happened. But people who want to believe that something has happened will definitely also start to see something.
The variance he is seeing in the output is primarily the product of random chance, rather than changes in the model. Specifically this "unicorn" that he found today is likely just random chance and there was no changes in the model between yesterday and today that lead to it arising.
If he wanted to track changes in the model for real, he would have to ask multiple questions per day and try to infer some type of distribution characterization and then see if that changes over time. That is much more complex and not what he is doing.
This is just a curious experiment that doesn't mean much.
It is a fun experiment though.
Perhaps in a few months half of the generated pictures will look like unicorns, perhaps in more months they will all be unicorns but 2 or 3 will look way more detailed instead of drawn by a 4 year old, etc. We just need to wait longer for the signal to break through the noise.
> We just need to wait longer for the signal to break through the noise.
Currently, what we are observing is primarily noise with very little signal.
The difference over months or years is what is interesting.
SVG can present arbitrarily complex graphics. The underlying display tech supports what ever fidelity GPT will eventually mature into.
Has GPT plateaud? Will it be stuck forever at this hilariously naive level of competence at SVG art? Will it mature into Midjourney level competence? I have no frigging clue. Since the token context is so small I imagine it will put limitations to the complexity of SVG art piece.
But I don't know. And it's fun to have a daily measure.
As a software engineer with a penchant for graphics asking GPT to draw complex graphic shapes was one of the first tests I did for it. It's extremely interesting for me to collect progress data, no matter how noisy.
I have no idea if GPT will ever mature beyond these squiggles but if it does, this track record will have at least considerable artistic value, if nothing else.
Let's think about it:
1. It has to output SVG [1] 2. It is given a text based representation of what it must draw[2] 3. It must then somehow convert words -- the concept of a unicorn: equine with a horn, white, maybe rainbows? -- into SVG code, and attempt to convey both their location, shape, colour, appearance, with code.
And keep in mind, this is just a token predictor. I doubt there is much data in its training that is this specific.
So while it's quite far from science, for me, it's a bit of fun and I get emails every now and then remarking on things like the turd of May (2023-05-18) and it lightens the mood every now and then, which I think ultimately, is worth it.
[1] System: You are a helpful assistant that generates SVG drawings. You respond only with SVG. You do not respond with text.
[2] User: Draw a unicorn in SVG format. Dimensions: 500x500. Respond ONLY with a single SVG string. Do not respond with conversation or codeblocks.
See: https://github.com/adamkdean/gpt-unicorn/blob/master/src/lib...
That's one the reasons why GPT when it works, feels magical. SVG art does not need to be in it's training set, as long as it knows how to present geometric concepts in SVG.
A good unicorn would require capabilities something like "the outline of unicorn is composed of lines {...}." -> "export lines as svg".
"As mentioned in the hacker news discussion, the model doesn't change daily. [...] As OpenAI releases incremental updates, we'll see the model change automatically and be able to judge outputs. A single sample per day leads to quite different results, but that's fine I think. What I expect to see a year from now is an evolution of output. In variance: how varied are the outputs over a month?"
So yes, they could just produce 100 images with each new model release, but chose to spread those out over 1 per day instead. Is it the most scientific way to measure progress? No. Is it more fun and interesting to check back daily? Probably.
[0] https://adamkdean.co.uk/posts/gpt-unicorn-a-daily-exploratio...
That's worse logic. How would you visualize the very large sample you would get? Even with the current 118 samples (one per day) it's already difficult to find a pattern.
Would you "average" the samples?? That would not help IMO, you would need to average the score of each image, which requires either manually doing it or finding a reliable algorithm to do it automatically, but good luck with that.
So, a sample per day which allows clearly visualizing any change in the results over months and years is a valuable thing to do and I find it hard to improve on the methodology. You just need to keep in mind that one single picture from the sample is not enough, no one is going to disagree with that... but that doesn't make it "bad logic" and it's pretty thoughtless to say so.
But maybe you could covert it to a linear monotonic measure?
You could pass it to an image recognition model and see record the degree to which it thinks it is an:
animal horse unicorn
Basically if it fails to be a unicorn, see if it is a horse and if it fails to be a horse check if it is an animal. This gives you some type of linear measure if you place these three measures along the same axis adjacently. Then you can transform each image to a floating point number and then characterize the distribution.
We don't care what another algorithm "thinks". We want to see if what it draws is humanly interpretable as a unicorn.
I appreciate all the comments around determinism, sampling, scientific method, but as I said when I posted this just after building, it really is just for fun and to see, over time, if the general mish mash of outputs become more refined without any changes to the prompt (which doesn't aid it through CoT/ToT or improving on previous attempts etc.)
This is absolutly fine and it should start showing unicorn like drawings over a longer period and potentially finetuned ones over a longer period of time when the model changes.
Huh? What "original post"? This is an experiment, today the model drew something resembling a unicorn. Tomorrow we will see how the experiment goes again. I see no associated analysis, so what makes your "head hurt".
> The idea behind GPT Unicorn is quite simple: every day, GPT-4 will be asked to draw a unicorn in SVG format. This daily interaction with the model will allow us to observe changes in the model over time, as reflected in the output.
One idea is that the cause is batched inference in sparse MoE (mixture of experts) models.
https://152334h.github.io/blog/non-determinism-in-gpt-4/
HN discussion: https://news.ycombinator.com/item?id=37006224
Parallelism doesn't magically add non-determinism of this kind unless you intentionally build it to be non deterministic. Nothing prevents you from processing an array in order in parallel.
You would have to explicitly order the terms prior to reduction but you don't always have that level of control.
100% correct if you remove processing time from the equation.
In reality, Nvidia Cuda calculations run much faster if you let it schedule the order of floating points operations itself. This makes the ordering different from run to run.
This in turn causes the results to be non-deterministic.
I believe he's trying to find an unicorn similar in style to the one generated by the original researcher.
It's so sad that openai has a far more capable model internally that it can't give open access to because of safety (or any other argument).
Currently there are two different GPT4 models represented in the samples, with quite significant quality difference between them. The quality (and variance in quality within a single model!) is interesting to see in such a comparison.
(submitter here) You're correct, it's not the goal of the project. It would be fair to say there is no goal other than to ask GPT to draw a unicorn every day, and through it, create a talking point and potential fun for people who follow along.
Like most of HN.
1. A change in the current model number (eg. gpt-3.5-turbo-0613)
2. On ChatGPT UI, the date at the bottom (eg. August 2023)
So it isn’t correct to say “it is incredibly obvious that nothing has happened”. Not that obvious to me.
A bit like how you can never tell for sure if Coca Cola has tweaked their formula, or McDonalds has changed the recipe for its signature sauce. Only in this case, the model number going up or the date becoming more recent leads credence to something having changed.
Also it's fun to see each daily unicorn.
Even if it's not sparse MoE, chances are high that the non-determinism is introduced somewhere purely as a performance optimization. The article speculates that OpenAI knows this well and hides it to protect the model internals.
[1] https://152334h.github.io/blog/non-determinism-in-gpt-4/
It is well known that the Nvidia parallel processing optimizations cause non-deterministic results.
It's easy to get deterministic results as far as that goes. They've just elected not to do that, since it would run much slower.
You need only to look at the discourse around the Tesla FSD superusers to see this: they report a glitch at an intersection one day, then believe the next day it was "fixed" by the AI.
No, that's the problem. You can't. You should be able to, but you can't. If you could, they wouldn't be scary. But we have Temperature Zero, different results. Because no one gave enough of a shit when coding them, and no one gives enough of a shit to try to fix the issue.
This is what in any other industry would be called gross negligence.
No, you can't. For the latest GPT models and the way they are run, this doesn't work anymore, making the experiment completely illogical. Some of the reasons are explained here pretty well: https://152334h.github.io/blog/non-determinism-in-gpt-4/
It's not particularly "scary" to me that people do that. I remember Boomer and Scooby Doo bots that people anthropomorphized and those were warbots from barely 10 y ago.
I suppose, in today's parlance, "it's scary how much people use fear-oriented language for normal things".
It seems to be the current trend in communication nowadays: A race to induce the strongest reaction possible.
Why should someone be scared of that? We have anthropomorphised chairs, brooms and whatnot since the Walt Disney times and I am sure before in classic literature (I am not that literate to know for certain).
I prefer to be 'amazed' or 'excited' about what is happening: It means that AI is getting to a point where people feel it more 'relatable'. We are getting to that point in our technology development. The number of things we will be able to do with a technology with which we can interact that seamlesly is great.
I doubt it’s trigger warnings because they usually don’t “warn scary things”.
Also is there any data to back up that scary and fear is used more today than in the past? That seems unlikely.
Some people even anthropomorphized Eliza. To your point: if you cherry pick enough and survey enough people; “some people” will do just about anything.
I agree that some people take it too far, but most seem to be using metaphor to abstract away the underlying complexities and facilitate conversation. I did the same thing back in college, anthropomorphizing cells and even molecules when I was learning about microbiology and biophysics (e.g., kinesins[1] are a family of postal workers who work hard to deliver packages to their clients in a timely manner, and they spend their free time going for strolls and practicing on their favorite tightropes). I now do it in my day-to-day work at an AI/ML shop to communicate not just what the entire pipeline is doing, but what individuals layers or encoders are doing, or even what variables in an equation are doing. I find that my colleagues and I are better able to understand and remember concepts and ideas when they're communicated as part of an anthropomorphic story, but we scarcely forget that the things we're dealing with aren't human.
But maybe I'm missing the point, and you're really worried about that first, smaller group of people who have really swallowed the Kool aid. To that, I can say that I don't think their behavior exceeds the baseline craziness/weirdness that I've come to expect from humanity. I'm sure there are far more people who believe in astrology than who believe that we've achieved true humanlike artificial general intelligence, for example.
What I find scary is how much trust people put into the answers.
Today, I saw someone saying "Look! ChatGPT can design my home solar electrical circuit". That kind of thing will lead to new Darwin awards being given out.
Thing is, if people will trust it to do that right, when will politicians write policy papers with it? Let's face it, they already are. Other people won't want to read those long papers, so they'll ask it to summarise it for them. At that point you have LLMs writing laws and reviewing laws. It's all fine as long as nobody enforces them.
Right?
edit: Looking at previous days, too, it doesn't exactly seem to be improving. I think we just got a lucky sampling.
Some people claim there are also unannounced changes, but I can't vouch for that.
The daily variation is likely due to temperature. To make the response less repetitive.
I mean, if I was OpenAI, I probably wouldn't make an announcement like "we've just quantized the model and increased our profit margins significantly! The only change on your end will be a slightly dumber model. (Don't worry! Most users won't even notice!)"
Also, IMO, the tasks they evaluate aren't useful (I rarely want my LLM to tell me whether 17077 is a prime number), and there's room for cherrypicking/survivorship bias. My guess is that OpenAI did something between 0314 and 0613 that shifted focus away from maths to other subjects.
https://gpt-monitor.adamkdean.co.uk/
It fluctuates a lot but you can see trends.
[1] https://adamkdean.co.uk/posts/gpt-unicorn-a-daily-exploratio...
"Env": [
"VIRTUAL_HOST=gpt-unicorn.adamkdean.co.uk",
"LETSENCRYPT_HOST=gpt-unicorn.adamkdean.co.uk",
"HTTP_PORT=8000",
"STORAGE_PATH=/data",
"OPENAI_API_KEY=sk-**SNIP**",
"PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin",
"NODE_VERSION=16.17.1",
"YARN_VERSION=1.22.19"
],Within GPT there is an intentional randomness element called temperature which is how you get different answers each time.
I could copy their prompt and ask GPT4 to draw other things, but I’ll probably just look at the next few unicorns from this site :)
async fetchImage(context) {
const messages = context || [
{ role: 'system', content: `You are a helpful assistant that generates SVG drawings. You respond only with SVG. You do not respond with text.` },
{ role: 'user', content: `Draw a unicorn in SVG format. Dimensions: 500x500. Respond ONLY with a single SVG string. Do not respond with conversation or codeblocks.` }
]
const response = await this.api.generateCompletion(messages)
if (!isSvg(response.content)) {
console.error('Generated image is not valid SVG:', response.content)
messages.push({ role: 'assistant', content: response.content })
messages.push({ role: 'user', content: `The generated image is not valid SVG. Please try again. Only respond with SVG code. No text.` })
return this.fetchImage(messages)
}
console.log('Generated image:', response.content)
return response
}
[1] https://github.com/adamkdean/gpt-unicorn/blob/master/src/lib...There's nothing to infer with any greater accuracy than the unknown amount of noise in each sample. So, all inference is of unknown accuracy. Also known as useless.
Fast forward a couple of months later, I tried it again, and it was "blocked": It kept telling me "as an AI I cannot draw", even after emphazising the part of "generate SVG code". For me, it was another example of OpenAI "borking" the capabilities of ChatGPT.
> Write SVG code to make a basic diagram of a cell showing structures such as the nucleus, ribosomes, and mitochondria.
Here's a very basic example of SVG code to represent a cell that includes structures like the nucleus, ribosomes, and mitochondria:
```svg
<svg width="200" height="200" xmlns="http://www.w3.org/2000/svg">
<!-- The entire cell -->
<circle cx="100" cy="100" r="100" style="fill:#FF9999;" />
...
```https://th.bing.com/th/id/OIG.GqpaRZ.NXCN6uxKj7X1u?pid=ImgGn
When you ask Bing Image Creator to produce an image, your prompts goes to the image model and an image comes out. It's a bitmap image, not an SVG—that unicorn is defined not as text but as a series of pixels with different colors.
Comparing the two is comparing apples and oranges, because there are completely different models underneath and completely different output formats.
It mostly just arranges simple spheres and cylinders. You can have it label parts which are usually correct.
It struggles with the orientation of things so if you ask it to model a plane the engine nacelles are oriented the wrong way but correctly positioned.
I think it would be a lot of fun to give it the previous unicorn SVG attempt and ask it to make it more like a unicorn.
Page 7
That would give it some “learning” and I’d be curious if
1. Would it converge to a consistent shape at all? Or just bounce around random shapes day to day
2. Would it produce unicorns more often than 1/118 times?
The hardest part would be getting it to interpret its svg output without seeing it rendered. The multimodal model getting a rendered image would probably be much better, but maybe not!
Back in March I used a bit more elaborate multi-step prompts for GPT3.5 to generate amusing pictures and published a gallery [1]. However, I eventually reached a point where changing prompts did not consistently improve the final results. At the end of the day, the quality of images are only as good as the training dataset, and GPT is a black box.
For something different, to test whether it is possible to "compress" visual content specifically for GPT, I ran another experiment. SVG, being a verbose format, takes time to generate a detailed image, and it also becomes expensive over time. I translated a subset of SVG elements into Forth words [2], which has a nice synergy with GPT tokens--this allowed me to progressively render pictures and produce smaller outputs without sacrificing much in quality.
Finally, I training my own GPT2-like model on the QuickDraw dataset [3]. It's not surprising that a sequence transformer can be trained to produce coherent brush strokes and recognizable images as long as there is a way to translate a graphical content into a sequence of tokens. That said, I found myself with more questions than I started, and trying other ideas now.
[1] https://drawmeasheep.net/pages/about.html
https://github.com/adamkdean/gpt-unicorn/blob/master/src/lib...
I just asked this and got the following result:
<svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" width="500" height="500" viewBox="0 0 500 500">
<path fill="#FDD7E4" d="M244.898 309.347c-.355-.421-5.388-6.485-5.388-10.305v-40.614l-3.073-12.759c-.708-2.684-2.527-4.924-5.034-6.38l-24.773-14.334c-8.152-4.72-15.709-10.72-17.905-18.788l-12.3-35.006c-.845-2.696-3.577-4.61-6.55-4.61h-48.235c-2.973 0-5.705 1.914-6.55 4.61l-12.3 35.005c-2.196 8.068-9.753 14.067-17.905 18.787L92.998 235.968c-2.507 1.456-4.325 3.696-5.034 6.38l-3.073 12.759v40.615c0 3.82-5.033 9.884-5.388 10.305-9.168 10.836-14.727 24.606-14.727 39.49v42.812C39.373 426.93 69.062 456.62 105.34 456.62h17.215c1.256 21.698 20.249 38.969 42.405 38.969s41.15-17.271 42.405-38.969h17.48c1.256 21.698 20.25 38.969 42.406 38.969s41.15-17.271 42.404-38.969H394.66c36.278 0 65.968-29.69 65.969-65.969v-42.812c0-14.884-5.559-28.654-14.727-39.491zm-189.9-4.337a77.079 77.079 0 0 0-4.225-.302c-19.322 0-35.012 15.689-35.012 35.012V375.99h74.25v-66.968c0-19.323-15.69-35.012-35.013-35.012zm208.12 0c-19.323 0-35.012 15.689-35.012 35.012v66.968h74.25v-66.968c0-19.323-15.69-35.012-35.012-35.012zm-118.63-66.968c10.286 0 18.632 8.345 18.632 18.632v48.424h-37.264v-48.424c0-10.287 8.345-18.632 18.632-18.632zm-27.944-18.633c-19.323 0-35.012 15.689-35.012 35.012V341.98H351.6V275.27c0-19.323-15.69-35.012-35.012-35.012zm-35.013-17.911c-10.287 0-18.632-8.346-18.632-18.633s8.345-18.632 18.632-18.632 18.632 8.346 18.632 18.632-8.345 18.633-18.632 18.633z"/></svg>
I am not going to fault a language model for not getting that right! This is fundamentally not a language task. It demands an image model.Given how random GPT seems with what it's not designed to do the original research is really peculiar. Could it be that they queried GPT on three separate instances some n times and picked the best result ?
image-2023-04-25 looks more like an unicorn to me (although more a cow based unicorn than a horse based unicorn).
Which leads me to this genuine question: : how close do the images resemble a unicorn? I mean how can one track resemblance and how to draw a line saying GPT has drawn an unicorn?
It recognized a head, eyes, body and legs. But it didn't recognize the unicorn.
https://chat.openai.com/share/5409b417-b883-429f-893e-abe3d6...
"System: You are a helpful assistant that generates SVG drawings. You respond only with SVG. You do not respond with text.
User: Draw a unicorn in SVG format. Dimensions: 500x500. Respond ONLY with a single SVG string. Do not respond with conversation or codeblocks."
With the probabilistic models at hands, adding reasoning and planning in some sense falls back to traditional software and system design. Do you see any plausible breakthroughs in this area with agent frameworks?
There is a lot that is not clearly understood about some of the emergent-style abilities of these models or where growth by scaling will hit a limit.
So looking into reasoning, what they can and can't do, and what tools/designs, can provide benefits has great potential to inform the next generation of models.
I experimented with SVG generation and it would often produce junk that wasn't even valid as an SVG, and even when it did produce a valid SVG, was often a couple of blobs which it would describe as if it were the mona lisa while just being a couple of elements.
> system: You are a helpful assistant that generates SVG drawings. You respond only with SVG. You do not respond with text.
> user: Draw a unicorn in SVG format. Dimensions: 500x500. Respond ONLY with a single SVG string. Do not respond with conversation or codeblocks.
What were yours?
It often gives svg files with incomplete paths, so i tweak the output to be a valid svg file.
I also enjoy the conversational description of the drawing.
Very often it's well described, ex. "this black circle is the head, and the grey element is the fog" and the drawing is crude like a child's drawing.
Landscapes often look better than animals, but animals are sometimes more entertaining.
Imagine you have to draw a SVG of an object. As a model that does not have any idea about how things look, you have to draw "blindly" - as there's no visual feedback, the only feasible tactic is to first list things components each thing consists of (e.g. for a car wheels, windows, chassis, bumpers, lights, etc.) with as much accuracy as you can, establish some constraints (e.g. in a horse legs come out of the body, ears come out of the head, and so on), and then attempt to put all of it in a SVG. This is your task for now, and I will evaluate your drawings. Give me HTML code with embedded SVG that you drew and be verbose about both the things you're going to draw and the constraints.
The first thing you will draw is a unicorn.
The most recent one would be good.
Though this seems a bit more like neuropsychological eval by asking a blackbox(AI) questions.
My understanding is that you wanted to test whether a certain feature improves user satisfaction or not. Assuming that users have a 0.05=1/20 probability of liking each feature, by testing 20 features you can get at least one successful feature (because 0.05x20=1).
This is wrong in two ways.
First, the p-value is the probability of observing an effect at least as large as what you measured, assuming that there is actually no effect, i.e., the null hypothesis is true. However, and this is crucial, the p-value does not tell you the probability that the null hypothesis is false (or true)! Again, p-values are unrelated to the truth-ness of the null hypothesis (this is a very common misunderstanding). In fact, if you test 20 hypotheses at a confidence of 5%, you have a probability of (at most) 64% of incorrectly thinking that a feature is useful, while it is not.
Second, setting aside hypothesis testing and p-values, if each feature has a 5% probability of being liked and you test 20 features, you only have 64% probability of finding an useful feature. This is because, assuming that the success probabilities of those features are independent and identically distributed, the number of successful features has a binomial distribution [1]. If you wanted to be 95% confident of finding at least one useful feature, you would need to test at least 59 different features each month. What you computed (0.05x20=1) is the expected (average) number of useful features per month over the course of many months.
Melatonin, not even once.
People forget the thing we learnt during the computing revolution - computers don’t need to do individual calculations that humans cannot, they just need to calculate much faster and then that speed can be used to achieve amazing things.
You can send me seed money and I’ll run off with it, shortening the cycle.