Microsoft's paper on OpenAI's GPT-4 had hidden information
twitter.com
twitter.com
"we asked DV3 for an explanation. DV3 replied that it detected sarcasm in the review, which it interpreted as a sign of negative sentiment. This was a surprising and reasonable explanation, since sarcasm is a subtle and subjective form of expression that can often elude human comprehension as well. However, it also revealed that DV3 had a more sensitive threshold for sarcasm detection than the human annotator, or than we expected -- thereby leading to the misspecification.
To verify this explanation, we needed to rewrite the review to eliminate any sarcasm and see if DV3 would revise its prediction. We asked DV3 to rewrite the review to remove sarcasm based on its explanation. When we presented this new review to DV3 in a new prompt, it correctly classified it as positive sentiment, confirming that sarcasm was the cause of the specification error."
The published paper instead says "we did not test for the ability to understand sarcasm, irony, humor, or deception, which are also related to theory of mind" .
The main conclusion I took away from this is "the remarkable emergence of what seems to be increasing functional explainability with increasing model scale". I can see the reasoning for why OpenAI decided not to publish any more details about the size or steps to reproduce their model. I assumed we would need a much bigger model to see these level of "human" understanding from LLMs. I can respect Meta, Google, and OpenAI's decision, but I hope this accelerates the research into truly open source models. Interacting with these models shouldn't be locked behind corporate doors.
I find it hard to see how detecting and eliminating sarcasm requires a theory of mind. It requires some association between various stylistic elements and the concept of sarcasm.
The same is true of irony.
I still wonder how many of these people have read Dennett's "The Intentional Stance", which holds that the best way to think about "intention" is as an explanatory model, not a mechanism. That is, we can say that the dog "behaves as if it has the intention to get inside" without making any claim about the internal state of the dog.
Dennett further speculates that our own self-experience of intention is a matter of turning the same explanatory model upon our own behavior, but that's an extension that isn't directly relevant to this speculation about language models.
That's also a way to frame evolution and literally all of biology.
That sounds very mechanic-ist, including Dennett's theory (I haven't read Dennett, to be exact, I'm going by your explanation). Which I guess it's par for the course when talking about a scaled up Mechanical Turk concoction like this GPT thing is. It won't "create" any Radio Erevan jokes anytime soon, that's for sure, though.
I am quite certain that detecting sarcasm can't be done based on stylistic elements" and requires some (even if implicit) estimation of what the author is thinking that contrasts with what is being said.
E.g. a relevant example I have actually seen in doing sentiment analysis on tweets to evaluate how customers perceive a company is "#CompanyName Got my order delivered in just under three hours. Thank you for great service! thumbsup-emoji" - now is this a positive review or sarcasm? And the thing is, you can't tell from the message by itself, you need an understanding of the customers' expectation (guess you might call it "a theory of mind") that for a pizza chain this is obviously sarcasm, but for a web store that sells some electronics the same thing would actually mean a fast delivery and great service. IMHO sarcasm detection is mostly about 'world knowledge' about what the implied expectations are, and not about stylistic elements at all.
If you have determined that a common element in reviews is a reference to delivery of goods or services within a given timeframe, and can identify that most of the quantitative descriptions of the timeframe are relatively short ... then coming across one that uses a much larger timeframe but is still positive will be quite noticeable.
This is actually typical of the problem with far, far too many people's interpretation of what language models are doing. The process I've described requires only a representation of language behavior. The LM can be said to understand how people talk (write) about a thing, but there is no knowledge of anything beyond language behavior.
Though, with LLMs, you might need quite a few (often more than 5) exchanges to find out that it can't actually quack.
So I guess maybe it’s saying the human was wrong after all?
Maybe more like appreciation with some humour thrown in?
Are we all a bit confused and second guessing our own use and interpretation of language now ?
"he must have never in his life seen a flick about any small towns"
That's probably the least sarcastic sentence for me, because it's just a reinforcement of the opening statement:
"I am so happy not to live in an American small town. Because whenever I'm shown some small town in the States it is populated with all kinds of monsters among whom flesh hungry zombies, evil aliens and sinister ghosts are most harmless."
I don't think it's a "positive review" I think it's neither, it's fairly neutral, and the author kind of suggests the reader watches the movie.
Out of interest I asked someone who has no idea about ChatGPT-4 and it's apparent sarcasm detection abilities and they didn't think it was sarcastic, albeit a bit 'weird' and poorly written. Confirmation bias?
We could say more, however, there are more important things to do...
He gave it a score of 7 which IMO is pretty high?
And yes it absolutely is subjective. That’s exactly the point and the power of these LLMs, to be able to handle the vagueness of human communication.
The first notable point is that the model caught that (on its first read, FWIW), even though the original human doing the labeling didn't.
And the second is that the researchers discovered this, and presumably discussed it. And yet when they wrote up the paper they not only dropped the content but denied that the analysis had been done.
> I am so happy not to live in an American small town. Because whenever I'm shown some small town in the States it is populated with all kinds of monsters among whom flesh hungry zombies, evil aliens and sinister ghosts are most harmless.
Mocking irony, in the context of a negative review.
What I find fascinating here is how generative AI has inverted all our sci-fi tropes. I mean, sure: you're right! That's the way "sarcasm" is defined in most dictionaries. But you and I both know that as the language is actually used, the term means a whole host of techniques used to convey negative emotional content in language that is not directly negative. Your (correct!) dictionary pedantry isn't interesting to me. We've been here before.
But GPT-4 wasn't trained on dictionary rules. It was trained on actual language. And it's actually better at inferring this stuff than the pedants are. Our introvert brains have trouble teasing meaning like this and have to hide behind rules and structure. The computer doesn't.
To wit: the robot is us.
There is another message that the reviewer doesn't like the genre, but that isn't a comment on the movie.
There is no way a human reading this review would think its positive.
"I'll just tell you that it's slightly reminiscent of U Turn by Oliver Stone but is a way down in all artistic properties."
Oh and UTurn only scores 6.7 so its a bit of a crap movie to start with.
Never. Ever. Reveal your sources.
They found comments that were commented out for a reason. These commented out sections aren’t a good look for the author. Usually commented out sections are either funny, notes, or provide some extra context.
He could have attempted to share this with someone in the industry of sharing information (like a reporter) who could validate it, and ask Microsoft for a comment — who will now be on the defensive instead of (potentially) forthcoming with more context about why those sections were commented out. This is a pretty fucked way of doing this.
Doesn't this sort of invalidate your point? Microsoft is going to know they accidentally leaked their comments and start stripping them.
It's not possible to perform journalism here without revealing where you got the information you're reporting. Seriously, how do you write this story without making it obvious that your big scoop is the comments from the paper?
I searched to try to decipher your comment but couldn't.
Missouri Governor Mike Parson and Cole County Prosecutor Locke Thompson (an elected prosecutor, in a reelection year [0]) trying to label a bona-fide journalist as a "hacker" were bringing ridiculous charges to try to deflect from the obvious embarrassment, instead of dealing with whatever MI state agency/ies or contractor was responsible, and had never QA'ed their webpages.
Coming back to your comment, the issue was not about reading the HTML(/JS/CSS). Can you provide me a single citation where that was the issue? (Obviously, there's a separate issue about "How do you responsibly make a disclosure when you find a leak of private information in a webpage?")
[0] https://www.newstribune.com/news/2022/feb/23/thompson-files-...
Who has been hurt by this in court, where and when?
What is the "hard line" that makes something AGI or not AGI? Because IMO it looks like GPT4 is somewhat AGI, but also the older models possibly all the way down to even Markov chains: it's just that this AGI is nowhere near human-level.
However, I think we can certainly prove when it doesn't. For instance, the fact we need an external plugin to get the model itself to return text claiming 1+1=2, tells me that GPT4 cannot reason about numbers in the abstract, and therefore lacks abstraction ability.
I think we're so strongly biased against a deeply uncomfortable reality that there may not be a hard line, we don't even want to consider the alternative.
But to make matters worse, or more muddied, 1+1=1 can be a valid mathematical statement. It simply depends upon the set and group you have, or if you're doing modular arithmetic, etc. Sometimes you're given a unital magma. So, there's still a heavy dependency on context for the problem setup, but the underlying discrete and deterministic rules that are applied to the context is less malleable than other context switches in NLP LLMs do well in (such as language styling).
Humans certainly have context windows. Try asking your CEO about some lines of code in your work. Humans have a fairly large one, I give you that and it is fuzzy.
An example about ping pong players. The pace of their movements is too fast for conscious thinking, so it's all trained reflexes with some overall strategic planning trying to keep up with events. There is no time to think about anything. Is intelligence suspended there, at least the general one? Then the same person stops playing and gives an interview about the game and full general intelligence turns on again.
Human level performance on general tasks (you can't be just good in 90%, you should avoid being terrible in the remaining 10%).
Though LLMs don't need to be AGI, to have a catastrophic economic impact.
It's pretty intelligent and rather general too, so at least by the definition that doesn't include mandatory consciousness it would mostly fit. And consciousness is pointless for a robotic system because it doesn't add anything practically useful. Just because agency has to be provided by the user doesn't make it any less of an AGI I'd say.
Content for anyone else in the same boat: https://www.techtimes.com/articles/289318/20230321/windows-1...
* https://twitter.com/DV2559106965076/status/16387694383539404...
https://www.youtube.com/watch?v=kpeODvGJE1Q
update: found a better one.
We have renewables, but those can't take up the majority capacity of a grid unless we start adding massive batteries.
Then there's grid rebalancing where we incentivize people to use and store renewable energy locally, this lessening the strain on the grid but that still results in fossil fuels or nuclear.
Hydro has been found to be environmentally destructive. I'm not sure if that was just my state or if that's ubiquitous.
What were you thinking?
I'm hopeful for nuclear, but they have a ways to go in proving themselves to the wider public. With that said, I think you could've positioned what you said a bit more fairly.
Everything we do has climate change impact. Power consumption is among the ones that's easiest to get "green", and there is significant progress specifically from cloud operators (at least Google, I assume others are similar).
https://www.reddit.com/r/MachineLearning/comments/hwfjej/d_t... suggests hundreds of GPU-years for training such a model under optimal assumptions, so let's assume thousands for the new models under real assumptions.
https://images.nvidia.com/content/technologies/volta/pdf/tes... says 300 W.
3000 GPU-years at 300 W = 7.9 million kWh. This would assume that they used those GPUs and not more efficient accelerators.
https://www.eia.gov/tools/faqs/faq.php?id=74&t=11 says 0.855 pounds of CO2 emissions per kWh. That's 388 g in normal units, and in line with what other countries are reporting (Germany was 420 g/kWh). g/kWh = metric tons per million kWh. This assumes the data centers are not using "greener" than average power.
So training such a model produces ~3000 metric tons of CO2 emissions under the assumption of 3000 GPU-years (how convenient). A round trip from SFO to London in Economy is roughly 0.9 metric tons. (https://www.icao.int/environmental-protection/Carbonoffset/P... Building a house (of unfortunately unspecified size) is 15 to 100 tons. (https://climate.mit.edu/ask-mit/how-much-co2-emitted-buildin...)
So while building one of these models does have an impact, it's on the order of other common activities that benefit far fewer people. Because once trained, those models provide values to millions of users.
Going back to the flights example, training one such massive model is likely about as bad as one larger research conference, once you consider the impact of hotels etc.
IMO the demands to justify the impact are ridiculous, coming from people who are just looking for any excuse to criticize, on par with demanding that any researcher doing any research justifies the carbon footprint of their commute as part of their research paper. Thus, it's a good thing that they didn't waste time and space in their paper addressing those claims, and we only see an early, commented out section that's equivalent to "TODO: Should we address these claims that keep getting thrown?"
I don't disagree with you either, I live in a foreign country and fly home to see family each year, so I'm far from perfect in this regard.
I guess this proves my point though, it's all justifiable to some degree: "I have to go see my family.", "this benefits millions." and on it goes.
We have to realize that it this stage, basically nothing, including the current LLMs will survive climate change if we keep this up.
Do they? Those of us in tech often take the positive value of technological progress as a given, but when looking at, e.g., the example in OpenAI's paper of GPT4 tricking a human into solving a CAPTCHA, I think it's quite clear there's possible negative value as well, and claims about "value to millions of users" probably need to be substantiated.
They should have asked it "Why do you say there is sarcasm?" A human can answer that but I don't think a bot can, can it?
Is the user referring to the ai alignment? the elitists who wants to nerf all the ai research?
I know from personal experience that I've had draft documents that were WILDLY wrong before I published to anyone but myself. Whole sections I just went back and completely deleted. In fact my senior project paper (LaTeX) in college had a whole section with big ASCII bull taking a shit on a paragraph because it was some work I'd done that didn't pan out at all. I left it in the source because I found it funny. lol, I found it: https://i.imgur.com/6Oj64AV.png
This was before I'd ever heard of a VCS system. Subversion 1.0 was released 6 months after I graduated, it turns out. So commented out code and multiple copies was all I had.
And I realise how agonisingly painful twitter threads are to consume.
It's just as bad as those YOU WONT BELIEVE WHERE THOSE CELEBS ARE NOW. Where you had to click next one by one.
https://addons.mozilla.org/en-GB/firefox/addon/nitter-redire...
I could see only the first message, the second message says This Tweet is unavailable, the third is ok, and then a bunch of random/unrelated messages.
As an early draft, them putting in a placeholder of ~'this model uses a lot of compute {TODO: put in cost estimates here?}' does not at all equal 'the authors didn't even know how much it cost to train the model!' Additionally, of course the toxicity went down. There's a world of RLHF between here and that original draft, and they've shown how the non-RLHF model has lowered the toxicity of the untrained base model significantly. If the author of the tweets had done their due diligence, they might have noticed that.
Rather obviously around the time when the model was originally being developed, text-only was sorta really the only way that LLMs were done. Them pivoting to multi-modal is just a natural part of following what works and what doesn't. This is really straightforward, I am mind-boggled that this is getting attention over discourse that is meaningful to the tidal wave of change coming with these models.
One final note that this is a bit of shoveltext is at the bottom the author offers up vague concerns followed by mass-tagging accounts with high follow counts, to include Elon Musk.
I'd encourage you not even to give the tweet the benefit of your view count and to just move on to more valuable discussions that are taking place. Why not take a look at a fun little thread like https://news.ycombinator.com/item?id=35283721 ? (not affiliated other the fact that I made the first comment on it, I just pulled from the rising threads on HN's frontpage)
Nor even how much impact it would have on capabilities.
Of course, to actually function this would also need to e.g. filter out soap operas, murder mysteries, and action films, lest it overestimate the frequency and underestimate the impact of homicide.
You: "What is grblf?"
As parents, my wife and I go through this on a daily basis. We have to explain what the behavior is, and why it is unacceptable or harmful.
The reason LLM models have such trouble with this is because LLMs have no theory of mind. They cannot project that text they generate will be read, conceptualized, and understood by a living being in a way that will harm them, or cause them to harm others.
Either way, censorship is definitely not the answer.
Theory of Mind May Have Spontaneously Emerged in Large Language Models - https://arxiv.org/abs/2302.02083
Previously discussed - https://news.ycombinator.com/item?id=34730365
(This doesn't use tokens. You can have a conversation in OpenAI Playground with text-davinci-003 and provide all the text yourself.)
Behaviours can be reinforced or dissuaded in non-verbal subjects, such as wild animals.
There's also the size of the possible behaviour space to consider: a discussion seldom has exactly two possible outcomes, the good one and the bad one, because even if you want yes-or-no answers it's still valid to respond "I don't know".
For an example of the former, I'm not sure how good the language model in DALL•E 2 is, but asking it for "Umfana nentombazane badlala ngebhola epaki elihle elinelanga elinesihlahla, umthwebuli wezithombe, uchwepheshe, 4k" didn't produce anything close to the English that I asked Google Translate to turn into Zulu: https://github.com/BenWheatley/Studies-of-AI/blob/main/DALL•...
(And for the latter, that might be why it did what it did with the Somali).
from "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? " https://dl.acm.org/doi/10.1145/3442188.3445922
That list of words is https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and...
1. medical pages/docs using the medical terms anus, rectum, nipple, and semen (note that other medical terms are not on that list).
2. pages/docs using "sex" to refer to males and females.
3. pages/docs talking about rapeseed oil or the plant it comes from (https://en.wikipedia.org/wiki/Rapeseed_oil).
The big problem with these lists is that they exclude valid contexts, and only include a small set of possible terms, so the model would get a distorted view of the world (like it learning that people can have penises, vaginas, breasts, but not nipples or anuses, and breasts cannot be big [1]). It would be better to train the models on these, teach it the contexts, and teach it where various usages are archaic, out dated, old fashioned, etc.
[1] but this is excluding the cases where "as big as", etc. are used to join the noun from the adjective, so just excluding the term "big breasts" is ineffective.
Apart from that list missing non-English words, leet, and emoji, there are also plenty of words which can be innocent or dirty depending entirely on context: That list doesn't have "prick", presumably because someone read about why you're allowed to "prick your finger" but not vice versa.
Regarding Scunthorpe, looking at that word list:
> taste my
It's probably going to block cooking blogs and recipe collections.
I've never heard the term before and would love any pointers (including enough keywords to Google for it :) )
"In the field of artificial intelligence (AI), AI alignment research aims to steer AI systems towards their designers’ intended goals and interests."
I also suggest the YouTube channel: "Robert Miles"
haha, that's very naive. There's already heaps (veritable mountains, even) of information that isn't given to the public on the public-facing instances of ChatGPT, because some info is deemed too incendiary. Filtering out "unwanted" sources of information is already a goal of information labelling, on which these entire LLMs exist. If you were to really make a LLM out of what people really thought and put on the internet instead of the current practice of castration, you wouldn't have techbros wondering about jobs, you'd have a veritable revolution on your hands.