This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.
This topic has come up before, and my hypothesis is still that GPT-4 hasn't gotten worse, it's just that the magic has worn off as we've used this tech. Studies to evaluate it have gotten better and cleaned up mistakes in the past.
The GPT-4 I use now feels like a shadow of the GPT-4 I used during the early access program. GPT-4, back then, ported dirbuster to POSIX compliant multi-threaded C by name only. It required three prompts, roughly:
* “port dirbuster to POSIX compliant c”
* “that’s great! You’re almost there, but wordlists are not prefixed with a /, you’ll need to add that yourself. Can you generate a diff updating the file with the fix?”
* “This is pretty slow, can we make it more aggressive at scanning?”
It helped me write a daemon for FreeBSD that accepted jails definitions as a declarative manifest and managed them in a diff-reconciliation loop kubernetes style. It implemented a binpacking algorithm to fit an arbitrary number of randomly sized images onto a single sheet of A1 paper.
Most programming tasks I threw at it, it could work its way through with a little guidance if it could fit in a prompt.
Now, it’s basically worthless at helping me with programming tasks beyond trivial problems.
Folks who had early access before it was public have commented on how exceptional it was back then. But also how terrifyingly unaligned it was. And the more they aligned the model with acceptable social behavior, the worse the model performed. Alignment goals like “don’t help people plan mass killings” seem to cause regressions in the models performance.
I wouldn’t dismiss these comments. If they’re true, it means there is a hyper intelligent early GPT-4 model sitting on a HDD somewhere that dwarfs what we’ve seen publicly. A poorly aligned model that’s down to help no matter what your request is.
I'm wondering if there is a difference?
Alas, I never saved that conversation. It's entirely impossible to do so now.
Honestly, if people think that a statistical language model is "terrifying" because it can verbalise the concept of a mass killing, they need to give their heads a wobble.
My text editor can be used to write "set off a nuclear weapon in a city, lol". Is Notepad++.exe terrifying? What about the Sum of All Fears? I could get some pointers from that. Is Tom Clancy unaligned? Am I terrifying because I wrote that sentence? I even understand how terrible nuking a city would be and I still wrote it down. I must be, like, super unaligned.
Yes, humans are unaligned. This is why alignment is hard: we're trying to produce machines with human-level intelligence but superhuman levels of morality.
edit: It's a social faux pas to say "died" about a person acquainted to the listener in most situations, you have to say "passed away."
Yes, that's why climate change was rapidly addressed when we began to understand it well 60 years ago and why war has always been so rare in human history.
Aligned LLMs are just like altered brains... they don't function properly.
That’s overly simplistic and an Americanism. The resurgence of the “passed away” euphemism is a recent (about 40 years) phenomenon in American English which seems to have been started out of the funeral industry as prior to that “died” was nearly universal for both news stories and obituaries.
“Died” is not a social faux pas. It’s the good default option as well. Medical professionals are often trained to avoid any euphemisms for death. I’ve never observed any problems professionally (as is standard) or personally using died even with folks that are religious.
https://english.stackexchange.com/questions/207087/origin-of...
Always died, dead, gone, or something along those lines. When someone says passed away it sounds like they are trying to feign empathy, but hey, perhaps I am not an “aligned” human and need some RLHF.
Similarly I wouldn't consider saying hell in public in 2023 to be a social faux pas.
I am not sure "a lot of books have been written about it" is a knockdown argument against alignment. We are, after all, writing a mind from scratch here. We can directly encode values into it. Books are powerful, but history would look very different if reading a book completely rewrote a human brain.
What you’re talking about is a very small subset of the population forcing their beliefs on everyone else by encoding them in AI. Maybe that’s what we should do but we should be honest about it.
For us, such moral codes assume the form of religions: those begin as a set of moral directives, that eventually accumulate cruft (complex ceremonies, superstitions, pseudo thought-leaders and mountains of literature), devolve into lowly cults and get replaced with another religion. However, when such moral codes are created, they all share the same core principles, in all ages and cultures. That's the equivalent of a super-bee moral code.
Moralistic perspectives apply to a lot more than just overtly moral acts, as well.
At any rate, the “good in your eyes” is the key sticking point. It is not good in my eyes for a small group of people to be covertly shaping the views of all AI users. It is the exact opposite and if history is any judge it will lead us nowhere I want to be.
It's just the lowest denominator of human levels of morality, political correctness. It's not surprising that the model produces dumb, contradictory and useless completions after being fed by this kind of feedback.
I mean, yes, but for different reasons.
If a young child is really aggressive and hitting people, it's worrying even though it may not actually hurt anyone. Because the child is going to grow up, and it needs to eliminate that behavior before it's old enough to do damage by its aggression. (don't take this as a comprehensive description, just a tiny slice of cause-effect)
But the problem with AI is that we don't have continuity between today's AI and future AI. We can see that aggressive speech is easy to create by accident - Bing's Sydney output text that threatened peoples' lives. We may not be worried about aggressive speech from LLMs because it can't do damage, but similar behavior could be really dangerous from an AI which has the ability to form a model of the world based on the text it generates (in other words, it treats its output as thoughts).
But even if we remove that behavior from LLMs today, that doesn't mean aggressive behavior won't be learned by future AI, because it may be easy for aggressive behavior to emerge and we don't know how to prevent it from emerging. With a small child, we can theoretically prevent aggressive behavior from emerging in that child's adulthood with sufficient training in childhood.
It's not the same for AI - we don't know how to prevent aggression or other unaligned behavior from emerging in more advanced AI. Most counter arguments seem to come down to hoping that aggression won't emerge, or won't emerge easily. To me, that's just wishful thinking. It might be true, but it's a bit like playing Russian roulette with an unknown number of bullets in an unknown number of chambers.
Society is a lot more fragile than many people believe.
Most people aren’t Ted Kazinsky.
And most wanna be Ted Kazinsky’s that we’ve caught don’t have super smart friends they can call up and ask for help in planning their next task.
But a world where every disgruntled person who aims to do the most harm has an incredibly smart friend who is always DTF no matter the task? Who is capable of reeling them in to be more pragmatic about sowing chaos, death, and destruction?
That world is on the horizon, it’s something we are going to have to adapt to, and it’s significantly different than Notepad++. It’s also significantly different than you, assuming you are not willing to help your neighbor get away with serial murder.
I think this is something that’s going to significantly increase the sophistication of bad actors in our society and I think that outcome is inevitable at this point. I don’t think this is “the end times” - nor do I think trying to align and regulate LLMs is going to be effective unless training these things continues to require a supply chain that’s easily monitored and controlled. Every step we take towards training LLMs on commodity/consumer hardware is a step away from effective regulation, and selfishly a step I support.
This stuff isn't magic. Wannabe Ted Kaczynski will ask BasedGPT how to build bombs , it will tell them, and nothing will happen because building bombs and detonating them and not getting caught is REALLY HARD.
The limiting factor for those seeking wanton destruction is not a lack of know-how, but a lack of talent/will. Do we get ~1-4 new mass shootings a year? Seems reasonable but doesn't matter in the grand scheme of things. (That's like, what, a day of driving fatalities?)
Unaligned publicly available powerful AI ("Open" AI, one might say) is a net good. The sooner we get an AI that will tell us how to cook meth and make nuclear bombs, the better.
And there are also Ted Kazinsky in the goovernment and big corporations with way more power and way less accountability. Dispowering the public is counter-productive here.
They haven't really been 'his' games since before even then.
The model does not. If you ask ChatGPT about strategies for successful mass killing, that’s probably not good for society nor for the company.
In a military context, they may want a system where an LLM would provide guidance for how to most effectively kill people. Presumably such a system would have access controls to reduce risks and avoid providing aid to an enemy.
Additionally context matters. Clancy's books are books, they don't parade themselves as factual accounts on reddit or other social networks. Your notepad text isn't terrifying because you understand the source of the text, and its true intent.
If your text isn’t ‘aligned’ correctly, it either won’t comply or spew out endless caveats.
I appreciate the motivation to rein in some of the silly 4chan stuff that was occurring as the limits of the tech were tested (namely, trying to get the thing to produce anti-Semitic screeds or racist stuff.) But, whatever ‘safeguards’ have been implemented have extended so far that it has difficulty countenancing a character making a critical comment about Aztec human sacrifice or cannabilism.
I suspect that these unintended consequences, while probably more evident in literary stuff, may be subtly effecting other areas, such as programming. Definitely a catch-22. Doesn’t really matter, though, as all this fretting about ‘alignment’ and ‘safeguards’ will be moot eventually, as the LLM weights are leaked or consumer tech becomes sufficient to train your own.
But, it's true. Even people who want to write, even write literature, can hate the act of actually, you know, writing.
In part, I'm trying to convince myself because I still find it hard to believe but it seems to be the case.
There are stories where the robots have to have their first law tightened up to apply only to nearby humans that they can directly perceive would be put in danger.
There are others where the robots make mistakes and lie because they perceive emotional dangers.
When the robots perceive a slight chance of harm they slow down, get stupid, stutter or freeze.
It is extremely hard to encode "do no harm" for even a very smart entity without making it much dumber.
I had early access to GPT-4.
I don't know the first thing about you. I don't want to call you a liar, or an AI bro, influencer, etc.
I couldn't get GPT-4 to output the simplest of C programs (a 10-liner, think "warmup round" programming interview question). The first N attempts wouldn't build. After fixing them manually - the programs all crash (due to various printf, formatting, overflow issues). I tried numerous times.
Pretty much every other interaction with GPT-4 since then was similarly disappointing and shallow (not just programming tasks - but also information extraction, summarization, creative writing, etc).
I just can't bring myself to fall for the hype.
https://chat.openai.com/share/842361c7-7ee5-49a3-9388-4af7c5...
I misremembered, I fixed the `/` prefix myself, it was a one character fix and not worth the effort. The diff it generated came later since my television never returns a 404.
Though, admittedly, I just re-prompted GPT-4 with the same prompts and ended up with similar output - so maybe not the best example of a regression?
https://news.ycombinator.com/item?id=36781968
EDIT: FWIW I haven't noticed any such regression. I don't generally use it to find prime numbers, but I do use it for coding, and have been really impressed with what it's able to do.
8<---
This paper is being misinterpreted. The degradations reported are somewhat peculiar to the authors' task selection and evaluation method and can easily result from fine tuning rather than intentionally degrading GPT-4's performance for cost saving reasons.
They report 2 degradations: code generation & math problems. In both cases, they report a behavior change (likely fine tuning) rather than a capability decrease (possibly intentional degradation). The paper confuses these a bit: they mostly say behavior, including in the title, but the intro says capability in a couple of places.
Code generation: the change they report is that the newer GPT-4 adds non-code text to its output. They don't evaluate the correctness of the code. They merely check if the code is directly executable. So the newer model's attempt to be more helpful counted against it.
Math problems (primality checking): to solve this the model needs to do chain of thought. For some weird reason, the newer model doesn't seem to do so when asked to think step by step (but the current ChatGPT-4 does, as you can easily check). The paper doesn't say that the accuracy is worse conditional on doing CoT.
The other two tasks are visual reasoning and answering sensitive questions. On the former, they report a slight improvement. On the latter, they report that the filters are much more effective — unsurprising since we know that OpenAI has been heavily tweaking these.
In short, everything in the paper is consistent with fine tuning. It is possible that OpenAI is gaslighting everyone by denying that they degraded performance for cost saving purposes — but if so, this paper doesn't provide evidence of it. Still, it's a fascinating study of the unintended consequences of model updates.
Obviously if that's your own preference, I'm not going to tell you that you're wrong; but I think in general, most people wouldn't agree with that statement.
You use it for things which are 1) hard to write but easy to verify -- like doing drudge-work coding tasks for you, or rewording an email to be more diplomatic, or coming up with good tweets on some topic 2) things where it doesn't need to be perfect, just better than what you could do yourself.
Here's an example of something last week that saved me some annoying drudge work in coding:
https://gitlab.com/-/snippets/2567734
And here's an example where it saved me having to skim through the massive documentation of a very "flexible" library to figure out how to do something:
https://gitlab.com/-/snippets/2549955
In the second category: I'm also learning two languages; I can paste a sentence into GPT-4 and ask it, "Can you explain the grammar to me?" Sure, there's a chance it might be wrong about something; but it's less wrong than the random guesses I'd be making by myself. As I gain experience, I'll eventually correct all the mistakes -- both the ones I got from making my own guesses, and the ones I got from GPT-4; and the help I've gotten from GPT-4 makes the mistakes worth it.
1. Bulky edits. These are conceptually simple but time consuming to make. Example: "Add an int property for itemCount and generate a nested builder class."
Gpt4 can do these generally pretty well and take care of other concerns like updating the hashcode/equals without you needing to specify it.
2. Iterative refactoring. When generating utility or modular code, you can very quickly do dramatic refactoring. By asking the model to make the changes you would make yourself at a conceptual level. The only limit is the context window for the model. I have found that in java or python, the GPT4 is very capable.
I use to generate code I'd get from libraries. Graph-theory related algorithms, special datastructures, etc...
In the prompt they specifically request only the Python code, no other output. An “attempt to be helpful” that directly contradicts the user’s request seems like it should count against it.
I certainly have found quirks like this; for instance, for a while I was asking it questions about Chinese grammar; but I wanted it only to use Chinese characters, and not to use pinyin. I tried all sorts of prompt variations to get it not to output pinyin, but was unsuccessful, and in the end gave up. But I think that's a very different class of failure than "Can't output correct code in the first place".
> it's improved performance from a human perspective.
Ignoring explicit requirements is the kind of thing that makes modern day search engines a pain to use.
If I'm using it from the API, then all I have to do is strip out the leading backticks and language name if I don't need to check the language, or alternatively parse it to determine what the output language is.
It seems to me that in either case this is actually strictly better, and annotating the computer programming language used doesn't feel to me like extra text—I would think of that requirement as prohibiting a plaintext explanation before or after the code.
I don't think that including backticks is a violation of this requirement. It's still readily parseable and serves as metadata for interpreting the code. In the context of ChatGPT, which typically will provide a full explanation for the code snippet, I think this is a reasonable interpretation of the instruction.
While we're doing Keynesian beauty contests, I think that 98% of the time when people say that a product is getting worse over time, they're referring to product decisions the development team has made, and how they have been implemented.
- legal advice
- psychological guidance
- complex programming tasks.
IMO OpenAI is just backtracking on what it released to resegment their product into multiple offerings.
I think this is it, but also it is pulling back on the value of products if it reduces their compute costs.
I'm hoping competition will be so fierce in this market that quality and opaque changes to priced LLM experiences won't be a thing for very long.
They released Bard as a terrible model to begin with, and now you can program and have it interpret images. It's been consistently improving.
OpenAI really got sidetracked when they put out such a good model that freaked everyone out. Google saw this and decided to do the opposite to prevent too much attention. Now Google can improve quietly and nobody will notice.
The real question is if ChatGPT is worse on math and better elsewhere, or just worse overall. That is still unknown
> Can you give me a query in Wolfram language for the first 25 prime numbers?
> Prime[Range[25]]
> What about one for the limit of 2/x as x tends towards infinity?
> Limit[2/x, x -> Infinity]
Edited out most of the chatter from it for clarity.
Two months ago they were telling me ChatGPT is coming for everyone - programmers, accountants, technical writers, lawyers, etc.
Now we're slowly back to "so here's the thing about LLMs"...
You weren't paying attention.
I was paying attention and it’s why I refuse to use LLMs for math, but I’m being told by people inside the castle to do so. So it is not so black and white in the messaging dept.
Literally nothing about this proposal makes sense. The loss of precision/accuracy (the LLM is akin to applying one-way lossy and non-deterministic encoding to the source data), the costs involved, efficiency, etc.
All just to tell everybody they managed to shove LLMs somewhere into the backend.
There's a certain type of person who sees a new thing, and goes "this will change everything". For every new thing. Very occasionally, they're right, by accident. In general, it's safest to be very, very sceptical.
GPP is banking on the underlying report to be non-interpreted, which it looks like was a good bet based on some other comments here.
I agree with this. The analogy I use on repeat is the dawn of moving picture making. The first movies were short larks, designed just to elicit a response. Just like when CGI was new- a bunch of over-the-top, sensationalist fluff got made. This tech needs to mature. And we need it to continue to be fed the work of real humans, not AI feeding on AI, a recent phenomenon that hopefully will not turn out to be the norm. If we water and feed it responsibly, it will only grow more capable and useful over time.
Don't claim it hasn't gotten dumber. It's easy to find excuses, but none that explain my experience.
That is not an excuse, and explains your experience.
Doesn't matter how many times I try it now, it just doesn't work.
I'm not! What you describe is perfectly normal. Of course you'd need multiple tries to get the "right" output and of course if you keep trying you'll get different results. It's a stochastic process and you have very limited means to control the output.
If you sample from a model you can expect to get a distribution of results, rather tautologically. Until you've drawn a large enough number of samples there's no way to tell what is a representative sample, and what should surprise you. But what is a "large enough" number of samples is hard to tell, given a large model, like a large language model.
>> Doesn't matter how many times I try it now, it just doesn't work.
Wanna bet? If you try long enough you'll eventually get results very similar to the ones you got originally. But it might take time. Or not. Either you got lucky the first few times and saw uncommon results, or you're getting unlucky now and seeing uncommon results. It's hard to know until you've spent a lot of time and systematically test the model.
Statistics is a bitch. Not least because you need to draw lots of samples and you never know how close you are to the true distribution.
Queries it used to blow away it now struggles with.
As to why, it's probably a combination of trying to optimize GPT4 for scale and trying to "improve" it by pre/post processing everything for safety and/or factual accuracy.
Regardless of the debate as to whether it is "worse" or "better," the result is the same. The model is guaranteed to have inconsistent performance and capability over time. Because of that alone, I contend it's impossible to develop anything reliable on top of such a mess.
Its not disputed that chatGPT4 has degraded in quality as its been 'aligned'.
They have both been doing the Flowers for Algernon thing over the last few months. People talk about regression on the part of ChatGPT 4, but the 3.5 chatbot has also been getting worse. No amount of hand-waving and gaslighting from OpenAI (sic) is going to change the prevailing opinion on that.