No, GPT4 Can’t Ace MIT
flower-nutria-41d.notion.site
flower-nutria-41d.notion.site
In a way it makes perfect sense that gpt4 can score 100% on a test gpt4 also grades. To be clear the grading gpt4 has the answers so it does have more information but it still might overlook important subtleties in how the real answer differs from the generated answer due to it's own failure to understand the material.
Even this is overstating it, because for each question, GPT-4 is considered to get it "correct" if, across the (18?) trials with various prompts, it ever produces one single answer that GPT-4 then, for whatever reason, accepts. That's not getting "100%" on a test.
If as per the linked critique, some of the questions in the test set were basically nonsense, then clearly they couldn't have manually verified all the answers or they would have noticed that.
Section 2.1
Then the github repo also has wording around this:
> We double-verify manually that the grading of the test set is correct. https://github.com/idrori/MITQ/blob/main/index.html#L552
I agree it looks like this may not have actually been done given some of the questions and answers in the dataset.
Of course this was back in April when you could still get the pure unadulterated GPT4 and they hadn't cut it down with baby laxative for the noobs.
The modified cabbage-goat-lion problem [1] that GPT4 always failed to solve, it now gets it right. I’ve seen enough people run it in enough variations [2] before to know that it absolutely did change.
Maybe they didn’t “change” as in train anything, but it’s definitely been RHLFed and it’s impacting the results.
[1] https://news.ycombinator.com/item?id=35155467
[2] anecdata: dozens of people, hundreds of times total
1. People have become more accustomed to the limits of GPT-4, similar to the Google effect. At first they were astounded, now they're starting to see it's limits
2. Enabling Plugins (or even small tweaks to the ChatGPT context like adding today's date) pollute the prompt, giving more directed/deterministic responses
The API, as far as I can tell, is exactly the same as it was when I first had access (which has been confirmed by OpenAI folks on Twitter [0])
[0] https://twitter.com/jeffintime/status/1663759913678700544
How do you know?
Even if the base model didn't change, that doesn't mean they didn't fine tune it in some way over time. They also might be passing its answers through some other AI or using some other techniques to filter, censor, and/or modify the answers in some way before returning them to the user.
I don't know how anyone could confidently say what they're doing unless they work at OpenAI.
No, I don't think there's much change at all to GPT-4 (at the API level) and probably not that much at the pre/post language detection and sanitation for apparently psychotic responses.
Couple this with the reliance on crowd-sourcing to create evaluation datasets and heavy use of GPT3.5 and GPT4 by MTurk workers, you have a big fat feed-forward process benefiting only one party: OpenAI.
The Internet we know is dead - this is a fact. I think OpenAI exactly knew how this would play out. Reddit, Twitter and the like are awakening just now - to find that they're basically powerless against this wave of distorted future standards.
When sufficiently proven to pass every existing test on Earth, every institution would be so reliant on producing work with GPT that we won't have a "%100 handmade exam" anymore. No problem will be left for GPT to be tackled with.
Why? Because machine learning is not a scientific field. That means anyone can say and do whatever they like and there's no way to tell them that what they're doing is wrong. At this point, machine learning research is like the social sciences: a house of cards, unfalsifiable and unreproducible research built on top of other unfalsifiable and unreproducible research. People simply choose whatever approach they like, cite whatever result they like, because they like the result, not because there's any reason to trust it.
Let me not bitch again about the complete lack of anything like objective measures of success in language modelling, in particular. There have been no good metrics, no meaningful benchmarks, for many decades now, in NLP as a whole, but in language generation even more so. This is taught at students in NLP courses (our tutors discussed it in my MSc course) there is scholarship on it, there is a constant chorus of "we have no idea what we're doing" but nothing changes. It's too much hard work to try and find good metrics, build good benchmarks. It's much easier to put a paper on arxiv that shows SOTA results (0.01 more than the best system compared to!). And so the house of cards rises ever towards the sky.
Here's a recent paper that points out the sorry state of Natural Language Understanding (NLU) benchmarking:
What Will it Take to Fix Benchmarking in Natural Language Understanding?
https://aclanthology.org/2021.naacl-main.385/
There are many more, going back years. There are studies of how top-notch performance on NLU benchmarks is reduced to dust when the statistical regularities that models learn to overfit to in test datasets are removed. Nobody. fucking. cares. You can take your science and go home, we're making billion$$$ here!
As you increase the number of bits you are trying to comprehend, you move from quantum physics to chemistry to material science to biology to social science.
At certain points, the methods and reproducibility become somewhat of a dark art. I have experience that in my field of materials science.
Because these models are using billions or trillions of random number generators in their probability chains, it starts looking more like the harder hard sciences, it gets very difficult to track and understand what is important.
I think machine learning will be easier to comprehend than social sciences, so I wouldn't put it that high. It will be something between materials science and biology levels of difficulty in understanding.
If OpenAI ceased to be – probably for some legislative reason –, would the problems go away?
If there's one thing we can be certain of, it's that LLMs often overlooks important subtleties.
Can't believe they used GPT4 to also evaluate the results. I mean, we wouldn't trust a student to grade their own exam even when given the right answers to grade with.
We've already run the compute to run the zero-shot GPT model on all of the datapoints in the provided test set. We're going through the process now of grading them manually (our whole fraternity is chipping in!) and should have the results out relatively soon.
I can say that, so far, it's not looking good for that 90% correct zero-shot claim either.
1. Using GPT-4, generate a text explanation of a neuron's activations on sample input.
2. Using GPT-4 again, use the text explanation to simulate the neuron on some new text input.
3. Compare the result to the actual neuron's activations on the new text input.
They justify this by saying human contractors do equally poorly at coming up with text descriptions. However, the procedure is such a black box that it is difficult to make scientific conclusions from the results.
[1] https://openai.com/research/language-models-can-explain-neur...
The way OpenAI used GPT-4 is fundamentally different than how GPT-4 was used to score the answers to the MIT exam. In OpenAI's case, they had GPT-4 generate an explanation of when a neuron in GPT-3 would fire. They then gave that explanation back to GPT-4 and had GPT-4 predict when the specific neuron in GPT-3 would fire. The scoring was done by computing the correlation between when GPT-4 predicted the neuron would fire and when it actually fired. The scoring was not done by GPT-4 as was done for the MIT exam
In addition OpenAI did have human evaluators score the explanations as well to make sure they were human interpretable[0]
[0] https://openaipublic.blob.core.windows.net/neuron-explainer/...
However, both papers rely on black boxes instead of well-understood procedures. This places the papers on weaker scientific footing. For example, a poor explanation/simulation of a neuron's behavior may simply be a consequence of GPT-4. Instead, a scientist would want to "prove" some form of unexplainability.
To do this, a researcher would not use a human at all to explain neuronal behavior. Instead, a simple repeatable algorithm such as topic modeling would be applied. This would lead to significantly stronger scientific conclusions about the neurons. It also proves it is not possible to explain the neuron in some specific sense.
An interesting follow-up to the OpenAI paper might be to quantify how much "more powerful" its textual descriptions are than simpler, well-understood techniques such as topic modeling. That could at least reinforce its use.
> Several of the authors listed on the discussed paper are undergraduate researchers. Consequently, we believe it's inappropriate to hold these individuals accountable for any lapses present in the work.
Instead, we believe the responsibility should lie with the supervising authors. They are the ones who are expected to ensure that the work meets the rigorous standards of public scholarship within their field.
Recently I found that GPT4 can't even reliably create a list of german nouns with a given article (der / die / das).
It will mess up a simple list - if you ask it to analyse it, it'll be able to tell you that it's wrong.
Then you get it to correct the list and it may still be wrong.
It can take several iterations to make the list correct. I would have thought this would be a super easy task for it, but apparently not.
Seems they're enshittifying already, I hope a competitive model is released soon.
Oh, let me guess... because OpenAI told you so. OpenAI, the one istitution with the most strong incentives to tell people the model is not getting worse.
People are incredibly silly, myself included, you get old enough and see an _insane_ influx of new people, you figure pretty much exactly this is going to happen. 95% of people don't know what temperature is. Of the 5% remaining, 4.9% think its something you just tell ChatGPT to adjust.
If my Tesla went 80 mph and started going 15, we wouldn’t attribute that to the-nature-of-Teslas, or a software update: there are vast incentives for anyone who knew that was done to share that publicly.
Instead, we know something is wrong with the individual car and we take it to the dealer.
Here, the missing part is objectivity via a speedometer: I have one, I know outputs are consistent on 0315 models at 0 temperature.
It makes total sense that they would throttle due to the incredible global demand. They can't build super computers fast enough now.
> Erstelle eine Liste mit 10 Nomen, die den Artikel "der" haben.
Maybe "reliably" is doing a lot of the heavy lifting?
It's still extremely helpful, I just found it strange that it seemed like a simple task - for something that has been fed millions of documents, for it to still give some incorrect results - especially AFTER it had analysed its own results and found some noun artikels to be incorrect.
But as you say, they are still extremely helpful anyway.
I actually didn't realise it was giving me incorrect info until my gf started looking at it!
(I was trying to use it to help me learn german)
d) GPT4 is significantly worse in german and most other languages, than in english anyways.
GPT answer:
It's difficult to determine the "200 most common" nouns for a particular gender in German since the frequency of word use can vary greatly depending on the context (e.g., conversational, academic, technical). However, here's a list of 100 common masculine nouns, which use the definite article "der".
der Mann (man), der Tag (day), der Freund (friend), der Vater (father), der Bruder (brother), der Sohn (son), der Monat (month), der Kurs (course), der Tisch (table), der Stuhl (chair), der Garten (garden), der Zug (train), der Platz (place), der Himmel (sky), der Wagen (car), der Brief (letter), der König (king), der Schuh (shoe), der Baum (tree), der Hund (dog), der Bär (bear), der Fisch (fish), der Vogel (bird), der Schüler (pupil), der Regen (rain), der Wind (wind), der Wald (forest), der Berg (mountain), der Fluss (river), der Junge (boy), der Kaffee (coffee), der Lehrer (teacher), der Arzt (doctor), der Student (student), der Polizist (policeman), der Künstler (artist), der Teller (plate), der Löffel (spoon), der Geldbeutel (wallet), der Laptop (laptop), der Schalter (switch), der Fernseher (television), der Schrank (cupboard), der Helm (helmet), der Hut (hat), der Schmuck (jewelry), der Ring (ring), der Sport (sport), der Schaden (damage), der Boden (floor), der Sand (sand), der Sturm (storm), der Preis (price), der Feind (enemy), der Held (hero), der Raum (room), der Morgen (morning), der Abend (evening), der Unfall (accident), der Vortrag (lecture), der Urlaub (vacation), der Ausflug (trip), der Hafen (harbor), der Bericht (report), der Name (name), der Bauer (farmer), der Rechner (calculator), der Traum (dream), der Anzug (suit), der Geschmack (taste), der Eindruck (impression), der Zweck (purpose), der Vertrag (contract), der Krieg (war), der Kunde (customer), der Arbeitgeber (employer), der Mitarbeiter (employee), der Kollege (colleague), der Bewohner (resident), der Fahrer (driver), der Gast (guest), der Kritiker (critic), der Profi (professional), der Sieger (winner), der Kandidat (candidate), der Beamte (official), der Insasse (inmate), der Zeuge (witness), der Beweis (proof), der Schatten (shadow), der Zweifel (doubt), der Trauer (grief), der Frieden (peace), der Nerv (nerve), der Horizont (horizon), der Gedanke (thought), der Lohn (wage), der Antrag (application), der Verlust (loss), der Betrag (amount),
There's a common pattern in GPT discourse (here on HN, but elsewhere too): somebody describes a limitation, somebody else goes "no such limitation exists: look, here's GPT output", and a third person goes "no, here's why your example demonstrates the limitation".
It's interesting. People find it very hard – or are disinclined – to distrust a charismatic robot, even when warned.
> Trauer is a feminine noun. Remember that, in German, both the spelling of the word and the article preceding the word can change depending on whether it is in the nominative, accusative, genitive, or dative case.[1]
[1]: https://www.collinsdictionary.com/dictionary/german-english/...
A good AI should be smart enough to know that if you ask for "German words with the article 'der'", you're most likely to want to be given a list of masculine nouns.
The list was mostly correct, but yes it added nouns that were not "der" nouns into the list.
It then attempted to correct the list and failed at correcting it.
In terms of output, it also didn't want to give a list of 200, but I did manage to get a list of around 100 back.
It was unwilling to write "Der Jahr" when "Das Jahr" is correct.
(I use normal machine translation API for a lot of this, but you can also ask it in another context window to translate the text to other languages. I use this approach for e.g. sindarian)
Otherwise my current strategy is to put it in an analysis loop until it deems the list to be correct.
Why did you think that? This isn't meant to be critical, but I'm honestly curious, what led you to believe that technology underlying GPT-4 made it a good fit for this or any particular task?
It has probably seen the correct nouns used millions of times in the training data - and asking it to produce the correct nouns for a bunch of words is really just "tell me which case you saw most during training", which is something LLM's are really good at.
It is a purely statistical model. It does not know any "rules" about the language (it doesn't know any language at all), but it is fed data and derives from that sophisticated probabalistical relationships between words.
It shouldn't have much of a problem to generate the correct grammatical formulations, as it has been extensively trained on them. Moreso than any other technology neural networks are suited for this kind of tasks where hard rules do not exist (as a German I couldn't tell you why "rain" is masculine, but "machine" is feminine), but lots of data correctly implements the rule, does.
(https://github.com/openai/gpt-3/blob/master/dataset_statisti...)
Also tbh, with all the hype of LLMs one would think that such a task would not be such a challenge.
The strange/sad thing is that despite being "large language models" they're often hypermyopic on English..
I've done some measurements comparing generation between various languages in the prompt and no matter what I do half the time i cannot get them to not include english text or comments in code unless the request is made in japanese, chinese, or a similar very-different-language.
If you train and evaluate something mostly on a language loke English, you're going to end up with a model that thinks everything works like English, which means among other things very little morphological complexity.
Suppose you had an oracle that when asked a question gives you 5 answers, 1 of which is true. Or even 1 of which is true only 50% of the time. You would still generally be a fool to throw away that oracle.
I remember the very late nights at Burton Connor, maybe students will get more sleep now :)
Initially we started looking into it more for curiosity's sake, but as we started digging we kept finding more and more ridiculous stuff to the point where we decided to start working on blog post. Then David joined in on the action and helped a ton with the research and writeup.
No professor was involved. The paper was released yesterday, so we just documented as we went along in the investigation process. It only took like 8 hours of work to compile the doc. We finished it last night and posted to Twitter in the morning.
Most of the damning claims in the conclusion section (Obligatory: I haven't read the paper entirely, just skimmed it.) usually get ironed out in the final deadline run by the advisors anyway. I'm assuming this is a draft paper for the EMNLP deadline this coming Friday published on arxiv. So this paper hasn't even gone through the peer review process yet.
The authors could probably have carefully review all ~300 of their questions. If they couldn't they could have just reduced their sample size to say 50.
It also seems like 100% accuracy should have raised red flags, especially if you know your dataset isn't perfectly cleaned.
Is it just me, or if for the majority of papers, the effort required to understand and get value out of the paper is so much higher than the effort put in by the authors and reviewers to publish it?
Either in the browser or in node JavaScript does not have processes operating on shared memory. While I get what they're trying to say I think it's counterproductive to learning to call a task on the JavaScript event loop a process. Answering the question requires knowing that the runtime will only preempt at await statements, and calling it a process confuses that.
Also somewhat upsetting that something so low quality actually is getting published. Seems entirely driven by hype and not intellectual rigor.
"Not peer reviewed" is no excuse to publish non-information for the sake of headlines.
Again, none of this excuses this. It isn't an innocent mistake, which could have been caught later. The dataset is flawed and the methodology is questionable, still the authors published it on arxiv, with spectacular claims.
If you don't know, there has been a significant shift in how scientific papers (STEM for the most part) are distributed. Instead of Journals (which have lost almost all use in a digital world) papers are published freely accessible online without any formal quality control, before potentially later being published in some journal. Arxiv, where these papers are published has control over who gets to publish (not open to the public), but doesn't require a lengthy formal process. In mathematics this has worked remarkably well, notably one of the millenium problems was solved when the solution was uploaded to arxiv.
Poluting arxiv with low quality clickbait is destructive, "not being peer reviewed" is no excuse for bad science.
On arxiv, it is possible to "retract" an article, in the sense that you ask that it's hidden or deleted etc. THat's not the same as retracting a published paper, where you usually get some justification, and a note from the editor explaining the decision to publish, and the decision to retract, and so on. More to the point, nobody cares if you "retract" a preprint, since it's expected that it may have errors not caught by peer-review, yet. No peer review means anyone can put anything they want online, and then take it back offline as they wish.
Note that arxiv also gives you an option to publish different versions of a paper. So you can leave your paper with errors as a v1 and upload your post-review paper, with corrections, as v2. Again, no "retraction" needed.