"Running this project daily doesn't make sense if GPT-4 is not being constantly updated"
With a suggestion to run it monthly instead, and generate 16 images at a time, and backfill it for GPT3 and GPT3.5.
"Running this project daily doesn't make sense if GPT-4 is not being constantly updated"
With a suggestion to run it monthly instead, and generate 16 images at a time, and backfill it for GPT3 and GPT3.5.
The models on offer now are frozen(-ish). Per the models[0] page, the non-snapshot model IDs "[w]ill be updated with our latest model iteration". So this project will eventually hit another version of the model, but (a) one image definitely will not be enough to reliably discern a difference and (b) seems like that 'latest model iteration' cadence will be much, lower than a day.
Which is very interesting. You already have a model that consumes nearly the entire internet with almost no standards or discernment, whereas a smart human is incredibly discerning with information (I’m sure you know what % of internet content that you read is actually high quality, and how even in the high quality parts it’s still incredibly tricky to get figure out good stuff - not to mention that half the good stuff is actually buried in low quality pools). But then you layer in political correctness and dramatically limit the usefulness.
The bottom of ChatGPT highlights which version is being used of GPT-4 and supposedly what version of ChatGPT it is.
> ChatGPT Mar 23 Version. (https://help.openai.com/en/articles/6825453-chatgpt-release-...)
It's possible to change the output both by tuning the parameters of the model and also client-side by doing filtering, adjusting the system prompt and/or adding/removing things from the user prompt. It's very possible to change the resulting answers without changing the model itself. This is noticeably in ChatGPT, as what you say is true, the answers change from time to time.
But when using the API, you get direct access to the model, parameters, system prompt and user prompts. If you give it a try to use that, you'll notice that you'll be getting the same answers as you did before as well, it doesn't change that often and hasn't changed since I got access to it.
Sebastien Bubeck gave a talk to MIT CSAIL on the Sparks of AGI paper where he commented on how RLHF (training for more safety) has caused the Unicorn test to become less recognizable [2]. He comments about how the safety training is at odds with a lot of these abilities. This seems congruent with the initial results from this project, and it will be interesting to see if they can restore this ability as they continue to push safety training.
1: https://twitter.com/gdb/status/1646183424024268800
2: https://youtu.be/qbIk7-JPB2c?t=1586My expectation is that there will be incremental updates to the model, so while I'm providing the model `gpt-4` for completions, I'm recording the actual model, `gpt-4-0314` in this case, along with the result.
I don't want to monitor (and potentially miss) model updates, which is likely as this is very much a fire and forget project that I'll review over time. One per day seems more sensible to me than a large batch per month, as the daily generations are more likely to track model changes if they become frequent.
Will be exciting to see where we are in a few months!
That's why I suggested generating 16 models each time instead of just 1 a- if all 16 are noticeably better than the previous day you've learned something a lot more interesting than if just one appears to be better than the previous one.