Hidden Changes in GPT-4, Uncovered
dmicz.github.io
dmicz.github.io
We know it can copy parts of the context and the system prompt while a special part of the context isn't immune to being copied.
You can test it yourself by adding random strings to the system prompt, you can consistently have the model copy them over. Is that not enough to have a reasonable belief that the model can copy system prompt instructions in the web interface?
Those that experiment with local models like myself can demonstrate to you that leaking the system prompt is not difficult at all.
It's some strange kind of neurosis to harbor such incorrect and strong beliefs on matters you have zero expertise in.
and because we know it well enough to be sure that it's not smart and devious enough (yet) to conspire like this? (nor has any clear reason to)
1. ChatGPT reliably produces the same output across several different instances, and other people have independently found identical versions of this text with different prompts.[1][2] This typically wouldn't happen with a hallucination, which may change with each time the model is prompted.
2. The instructions accurately describe the capabilities and restrictions of ChatGPT's function calls. When making custom GPTs, the browser tool and DALL-E image generation tool are options, and the "system prompt" given by the custom GPTs reflect whichever tools you've selected.
3. ChatGPT has reliably changed with the changes noticed in the "system prompt". I document a recent change made to the prompt on my post from yesterday.[3] On the older instances of ChatGPT I have, the model will suggest it has no idea what the "guardian_tool" is or make any attempt at stopping a discussion of U.S. elections, while newly made instances of ChatGPT will discuss the "guardian_tool". The "system prompt" repeated then must give at least some sense of what updates are being made under the hood.
I certainly think we should still take the idea that this is the "system prompt" with a grain of salt. There is a good discussion about this here: https://news.ycombinator.com/item?id=37879077
[1] https://www.reddit.com/r/ChatGPT/comments/18494zo/what_are_t...
[2] https://medium.com/@dan_43009/what-we-can-learn-from-openai-...
[3] https://dmicz.github.io/machine-learning/chatgpt-election-up...
Of course, it could be that GPT-4 has been instructed to lie about its prompt, but failing that, you should expect any answer that stays the same across multiple wordings and prompting methods to be accurate.
An accurate answer is often driven by a concrete and highly confident fact in the training dataset (e.g. structured data fact, like a birth date from Wikipedia etc.).
The hallucinations are derived facts of (hopefully) low confidence. Nondeterminism is more common if you have low scores. Only a few facts can take high score (in a usable system), while many can take a low score -- then numeric instability can make a mess.
I'm not very familiar with LLMs, but I do have experience with the traditional ML models and content understanding production system. But, LLMs are not far from them.
Especially since (I'm fairly sure? though this is apparently not trivial to find today) it doesn't use temperature=0.
1. DALL-E generation metadata: 1. A stylized chat bubble integrated with an abstract brain design, representing AI and chat. 2. A sleek, modern avatar that looks like a digital assistant, with elements like a clipboard or list to represent organization. 3. A microchip or circuit board pattern in the shape of a speech balloon, merging AI technology and chat. 4. A robotic hand holding a pen or stylus, positioned over a notepad, symbolizing AI's role in organizing chats. 5. An organizer folder or agenda book with digital or futuristic motifs for AI integration. 6. A light bulb with chat bubbles around it, showing the concept of generating and organizing AI conversations. Seed 8893527578
2. DALL·E displayed 1 images. The images are already plainly visible, so don't repeat the descriptions in detail. Do not list download links as they are available in the ChatGPT UI already. The user may download the images by clicking on them, but do not mention anything about downloading to the user.
I give those prompts when I want to confirm or verify the input prompt is what I expect so, that's kind of the point.
I've never gotten back a result that was different than my recall and then I can verify it with a previous conversation stored.
Are you worried that the system will gaslight you into believing you gave a prompt that you didn't?
To me personally it's crazy how many people think that we would be better off without any kind of copyright protection. Copyright solves many real world problems and protects people against having a company profit off their work... but as soon as AI is involved so many people start to advocate for throwing it away.
Also, it's a subtle difference, but copyright is not intended to solve the problem of companies profiting off of artist's works, it is intended to promote the progress of science and useful arts. It attempts to do this by giving creators limited exclusive rights.
Even scientists are tired of the predatory and rent seeking behaviour of the publishers they have fallen prey to and are looking for any way out.
This is not promoting progress this is the opposite of it
Because they made it, it wouldn't exist without them, and others value it. If this data wasn't objectively valuable, we wouldn't be having this discussion.
It looks like to me that many companies want to use the new generative tools, and many others want it not to impact their stake in the copyright system. I’m pretty sure they will both come to a compromise which will leave most users without any benefits, either from reduced copyrights or from availability of generative tools. It’s what would make both powerful parties satisfied (if not happy), and will impact the status quo the least.
Say, for instance, that they instituted a mostly mandatory licensing scheme, so that an individual artist had no choice but to allow use of their art as input when creating generative tools. People using art in this way have to pay a rather high licensing fee, but it is not paid to the artist, but to some sort of central copyright office. Huge copyright holders can also pay an exorbitantly high fee (to the same recipient) to opt out of licensing. Win-win-win; Existing copyright holders keep their existing copyrights, only large-ish actors can create new generative tools, new political positions and institutions are created with lots of money flowing in. Of course, artists then get screwed by being co-opted by generative tools which they can never afford to create themselves, and the general public get robbed both of the opportunity of using and creating new generative tools, and of any less restrictive copyright law.
(I agree that it would be terrible if they began enforcing other copyrighted content and for training purposes, because it would lead to centralisation)
They're the worst, eg they will notoriously come after you if you play public domain music as well.
How did portrait artists feel when photography was gaining popularity? Should we have let them control the industry so that if we want to record a memory of a person we must have them stand or sit for hours while someone draws them?
etc.
I feel bad for copy editors and people who write corporate blog posts or design logos or come up with ad jingles, but their niche is gone now and they need to adapt.
Thanks for being respectful and cordial though.
To me, it's a kind of cowardice that people like you shrug your shoulders at and sigh and say "that's just the way things are". You can say that's just how the markets work. I don't have to respect you for it.
Things change -- people's jobs will be different. It isn't going to mean artists will stop making art or machines will make everything bland, it is just a new tool that will change industries and make things easier for people to do well and thus make more art. Some people won't be able to live well doing the same thing they do now, but what they do now wasn't what they would have been doing if they were in their grandparents time.
They showed some examples in the past and showed that society adapted.
We could try and improve our society and systems to have a safety net, education that allows us to adapt to rapidly changing technologies, etc sure but that's a whole discussion in itself.
If you're angry that independent artists are being fucked over by bigcorp, AI tools aren't the battle you should pick, because it's a guaranteed loss for a lot of very logical reasons, and it's just another example of a pattern of oppression enabled by our social and political systems. Even if by magic you managed to change something there'd just be another inequality coming down the pipe shortly after.
"Even" just OpenAI alone could pocket a few of them if they need easy sources of acquiring content.
This includes the largest educational publishers. And while these publishers do not own all their content, the reality is most authors earn so little, that a "allow AI training on my work for $x extra" would give them vast amounts of content.
As for Getty, Getty has a market cap of "only" $2bn. The big players will easily afford to build or buy libraries like that.
But of course it will be the end of decent open models.
It will also guarantee that the financial means to continue making that data, that is clearly so important, would be preserved. Someone has to pay for the crafting of the data.
I don’t even have the means to start litigation, let alone see it through.
It only protects those who are already moneyed and/or famous enough to negatively impact a large corporation’s reputation - and even in those cases it’s mostly for the benefit of the lawyers and bureaucrats who make a living off it.
No, it's been years I've heard it.
Don't try to portray some people opinion as they are some AI zealot.
It's brought up on discussions about torrent, Disney, streaming platforms, music, etc...
A large chunk of the tech community was following that case and most on HN seemed to be highly sceptical of the current status quo.
Exactly like before AI you mean then? Except instead of OpenAI it was Disney, Universal and other large corporations on that same seat.
>to me personally it's crazy how many people think that we would be better off without any kind of copyright protection.
Why should I care that the old billionaire copyright corps are dying exactly? What would I benefit defending them for me as what they did was privatizing culture as far as I remember for their personal benefit and even had a large negative influence on tech.
The copyright system being so unequal and skewed towards multi billion companies dug its own grave by itself.
OpenAI is not the one who would kill a copyright. They just want their cut.
What we need is a reasonable way for people using AI to determine which parts of the text or images they have are subject to copyright.
The tool itself unambiguously is fair use.
What do you mean anyone?
Is Sony liable when you play an entire movie on their TV? Is Nuance liable when you use their Dragon screen reader to cerbalize an entire NYT article? Is Google liable when you display an entire webpage in Google Chrome? How about if you switch to Dark Mode, is that a transformative use?
Why would AI be any different? It’s just a tool at the end of the day!
(Copied from a comment of mine written over a year ago: <https://news.ycombinator.com/item?id=33582047>)
Care to share an example? I didn't hear of OpenAI or anyone else arguing or trying to sue anyone for abusing the copyright. If anything, their business decisions rely on an assumption that copyright will not help them protect their work
https://nypost.com/2023/12/18/business/openai-suspends-byted...
100% pure unadulterated hypocrisy from "OpenAI".
But then, I think that surely they would use copyright to block competition from using their model directly.
Bad things happen when you let middlemen get the upper hand, like the American health care system, or big finance disconnected from the real economy. I'll vote against the middleman every time in favor of the original value creator, because society goes down the toliet when middlemen win.
Google search at least was just a link to content we wrote. OpenAI just steals it.
This is going to end up being the music industry all over again. It's going to be impossible for any individuals or small companies to get the rights needed, and instead were going to get massive content labels selling the rights, or only giant corporations being able to hop through all these new hoops.
We don't want a repeat of that as a society, creating yet another leeching middleman and horrible industry favoring only the incumbents.
We must tell the millions of kids who doodle characters in their notebooks that this prohibited.
Specifically, the author is investigating (possible) changes in the system prompt and tools available to the model in the chat interface of ChatGPT Plus. That tells nothing about the model (GPT-4).
Sending a prompt into one vs the other, the API sends through the model and back out, the product has censorship watchers and other unknowable bolt-ons.
My use is a perfect example of many of these types of points made in the NYT Lawsuit [1].
[1] Great summary by Hoeg Law: https://www.youtube.com/watch?v=ETDBDZoRJJU
Or artists, for that matter.