GPT-4 gets a B on my quantum computing final exam
scottaaronson.blog
scottaaronson.blog
Here’s one passing a very difficult test. Amazing! Now, rely on it to build a nuclear doohickey for a power station or a multi-billion dollar device for CERN or anything really and, well, no.
So humans still have to check the output, and now we’re in that situation where humans driving a Tesla on autopilot who are supposed to be 100% aware of the road aren’t, because they get lazy and doze off, and now the car crashes and whoops.
No negativity towards AI here. It’s amazing and it’ll change the future. But we need to be careful on the way.
And it's just regurgitating the answers someone else wrote?
As I imagine it's a very high chance given how much uni lecturers recycle exam questions.
When I was at uni you could just get the last 5 years worth of questions from the library for almost any subject and guess what the questions were probably going to be. Often they just changed a few numbers.
Teaching undergrads is like a sausage factory, the actual intellectual value for undergrads is in the seminars, the practical value in the labs. The rest is showing you can regurgitate what you've been told.
Which ChatGPT excels at.
The exact questions aren’t in the top results, but the answers are.
> To the best of my knowledge—and I double-checked—this exam has never before been posted on the public Internet, and could not have appeared in GPT-4’s training data.
I remember Yann LeCun gave an interview and he came up with some random question like "If I'm holding a peace of paper with both of my hands above the desk and I release one what would happen". His point was that since the LLM doesn't have a world model it wouldn't be able to answer these trivial intuitive questions unless it saw something similar in the training set. And then the interviewer tried it and it failed. That was 3.5. I've tried many variation of that class of problem with 4 and it seems to generalize basic physics concepts quite well. So maybe 4 learned basic physics ? Why couldn't it learn QM theory as well ?
For a language model, test results are the end. They are supposed to measure what the model is capable of. If you need better performance, you must train a better model.
The 95% only problem is an issue for cars cuz that last 5% means you die horribly in a head-on collision, or maybe only get a mild concussion but are stuck in a ditch.
But if I can get 95% of my router configs done, 95% of my documentation written, and 95% of a website whipped up I can hand that off to a Sr Engineer/Admin and have them take care of the last bits. As long as the hours, phone number, and location are good a website just needs to be "directionally accurate" and otherwise fairly basic.
Yeah, I suspect a lot of fields will have a similar trajectory to how AI has impacted radiology.
It might catch the tumor in 99.9999% of cases, better than any human doctor. But missing a malignant tumor 0.0001% of the time is unacceptable, because it spikes the hospital's malpractice costs. So every single scan still has to be reviewed manually by a doctor first, then by the AI as a fallback.
In theory there's some insurance scheme that could overcome this, but in practice when you have software reviewing millions of scans a day you're opening yourself up to class action lawsuits in a way no competent human doctor would.
I find it hard to believe human doctors miss malignant tumors in less than 1 out of every 10 million cases.
As the parent poster put it, it's only a problem if the average doc won't detect it. If it's truly a 1 in 10-million thing, an extreme edge or corner case, malpractice courts may not have a problem with you missing it -- as they say "if you hear hooves, do you think of horses or zebras?". 99% of the time a different diagnosis is the right one, and even at five-nines you're letting someone through eventually.
Whenever someone says "as long as it's better than a human", that's where my mind goes. We shouldn't be satisfied with just being better than a human. We shouldn't be satisfied with five nines! I don't really care about what courts have a problem with — my point is just that our goal should be zero preventable deaths, not just moving from humans to AI once the latter can be better on average than the former.
They've already done it with outsourcing; a large chunk of what used to be done entirely in-house has been contracted out to remote overseas doctors.
From what I understand it isn't even possible generally to see a doctor remotely in a cheaper state, because medical licensing is per-state.
Medical billing has also been offshored to India+Pakistan btw
In general, a lot of back office Dental+Medical functions were outsourced in the 2000s+2010s.
Eg. Paper about this from 2006 - https://ipc.mit.edu/sites/default/files/2019-01/06-005.pdf
Those probabilities are way off given biology, but anyway ...
The interesting cases of AI in radiology would be being able to catch stuff that a human has no hope of catching.
For example, a woman with lobular (instead of ductal) breast cancer generally doesn't present until mid-to-late Stage 3 (which limits treatment options) because those cancers don't form lumps.
You can stare at mammograms and ultrasounds all day and won't see anything because the "lumps" are unresolvable. You're trying to find a sleet particle in a blizzard. Sure, it's totally obvious on an MRI scan, but you don't want to do those without reason (picking up totally benign growths, gadolinium bioaccumulation, infections from IVs, etc.)
An AI, however, could correlate subtle, but broad changes that humans are really bad at catching. Your last 5 mammograms looked like this but there is just something a little off about this one--go get an MRI this time.
I think people probably overestimate (maybe vastly) how good at differential diagnosis most doctors are.
This does seem like an odd outcome though, right? I guess fundamentally people will be legally/economically advantageous in some sense because the amount of insurance that an individual doctor can be expected to hold is much less than a hospital. Is the fate of humans to exist not as a unit of competence, but as a unit of ablative legal armor?
I don't think results are anywhere close to that in the field. If hospitals could do without radiologists, they would do so immediately. Currently, we are seeing very little technical progress due to applied statistics in the field, and the real cause of that is that tech people don't understand what we do and why we're still very much needed. The problem lies more in information retrieval capabilities than acting on the data itself.
Meanwhile, in every country that isn't America, tumor detection rates go up, cancer outcomes improve, and the cost of delivery falls.
Americans continue to claim their system of funding insurance company profits rather than actual healthcare is best.
https://worldpopulationreview.com/country-rankings/cancer-su...
The military, for all its funding and all its training and all its planning, has lost multiple nuclear weapons, on American soil.
We are imperfect machines who aspire to build more perfect versions of ourselves, through children and now through AI. By many measures, we’ve succeeded. The progress will almost certainly continue. The question is, when will it be good enough for you to embrace it despite its imperfections?
Look at what wave done to soil, the oceans, we’re likely on the wrong track. Maybe it’s even unrecoverable.
I think we can have a much more advanced society , but we need to slow down a lot and do things more in harmony with the natural world and out our egos aside. We’d probably have a much better quality of life for doing so.
Obviously the scope of what's possible is different but _any_ craftsman using _any_ tool will only be as good as what they verify themselves.
It reminds me of anti seatbelt rhetoric or the same stuff skateboarders say about helmets.
These tools aren't replacements for knowledge and experience (which is the narrative that is constantly pushed), they bionic enhancements.
It needs interaction and calibration from a human to keep it in check. Even we as humans are using all kinds of different feedback mechanisms to validate our thought process.
Though this might be a different way of hooking GPT up to reality.
Maybe in the near future humans can just focus on the reviewing and testing part.
And if you really dive deep you see 4 really hasn't any deep understanding and makes obvious mistakes. I think they've also explicitly trained it for tests.
The only question is whether, on average, AI-augmented code suggestions reduce overall bug rates. Human devs are already dozing off with crashes and whoopsies at high rates so it doesn't matter too much if AI output has a bug or whoopsie so long as they occur at rates lower than pure human generated code.
If we create a culture of tolerating fatal AI hallucinations as long as they happen at a rate that's below the human caused fatality baseline, we create an opportunity for bad actors to use the "nobody is to blame" window as a plausible deniability weapon (so long as they don't use it too often).
how often is the “have to check output” more work than just doing the task at hand without the LLM in the first place?
This simply isn't true. By the mile they have a worse safety record than other cars in their class (mitigating factors: where they're driven and who drives them). You might be referring to Tesla's marketing statistic that there are fewer accidents per mile involving autopilot - typically engaged in ideal driving circumstances - than when it's switched off, or across all other drivers. But that's not very meaningful data. Some would argue that citing marketing figures from a company with a track record for obfuscation and dishonesty as "the data" (and really, there isn't much better data available to the average person) is an indication we're not being careful.
The same is true of manufacturers and others in the chain of commerce of goods (see, e.g., the general rules on defective product liability), even where they aren’t individual humans. There’s nothing about AI which makes it particularly special in this regard.
Solve the liability problem and I would 100% take a machine that performs 30% better than a human or helps a human perform 30% better every time because it means fewer humans die on the road.
It's surprisingly hard to establish those parameters though, since (i) the more meaningful indicators of good performance (driver error fatalities) happen only every million or so miles, even less frequently once you've narrowed your pool down to errors made by sober drivers that weren't racing or attempting to driver in conditions unsuited to electronic assistance, so that's a lot of real world road use required to establish statistically significant evidence a machine is a 30% or so safe than a human driver (ii) complex software doesn't improve monotonically, so really you need that amount of testing per update to be confident that the next minor version of something "30% safer" hasn't introduced regression bugs which mean that it is now a bit worse than the average driver and (iii) performance in different road conditions is likely highly variable such that it might be both 30% better overall and 30x as likely to cause an accident if not disengaged in that particular circumstance.
To make a valid assessment of the overall safety impact, you'd also have to factor in that (iv) the worst drivers who skew the stats are the generally the ones least likely to buy it and (v) if it's fully autonomous driving, road use would increase substantially and whilst that may have other benefits the likely outcome from substantial increase in miles driven using tech that's only marginally better than a human is more humans dying on the road
The fact that overall road use is so high that a sufficiently bad regression bug in sufficiently widely-deployed software could rack up a massive body count within hours obviously doesn't make the case for introducing something believed to only be a marginal improvement any stronger.
How is it valid? It's currently a solved problem. The driver of the car is still held liable because they are still, ultimately, driving the car. Tesla's with autopilot seem to make drivers safer, and acting like we don't know who is liable in the event of a crash is just a red herring.
These AI have "Male Answer Syndrome" to the nth degree. They will shamelessly give you an answer even if they have to completely make one up.
is that like mansplaining?
is botsplaining the new mansplaining?
Edit I asked question 1m but said it was an expert in QM and was there to explain the answer to test questions followed by answering them and it gets it right with a page or so of proof (please don't attack my use of the word proof, I'm a lay man - this also explains my failure to fix the numbering, latex editing on a phone is not my forte). I've not checked it, but it gets the right answer.
https://figshare.com/articles/preprint/GPT4_answers_a_QM_que...
What is an exam? It's an attempt to test that you have consumed and understood large amounts of information and can apply it to novel (ish) situations, but in a very 'sandboxed' way that is just text-in-text-out. That's literally what these systems are specialists in doing.
Is it maybe a bit like saying a driverless car can outperform most human drivers in a time trial on a race track? Not a trivial feat, to be sure, but also not that surprising perhaps once the basics are dialled in.
It's so hard for people, including myself, to look at the future potential of a technology. My attitude is to simply keep an open mind nowadays instead of holding very strong opinions about the trajectory of a given technology.
No one called this, and certainly no one called it would be so easy to access.
Thanks for capturing a point that I couldn't quite articulate myself as to why ChatGPT feels different - the ease of use.
Literally a year ago I was planning to take a python programming course for ML, thinking that a deep understanding of code would be needed to make things work well.
With GPT it's like...I just ask it things in English and it'll do it?
The ease of access to LMMs is groundbreaking the way the simplicity of IOS launched the modern smartphone era. Even babies could use an Iphone.
And now any child old enough to type sentences can use ChatGPT.
So.. that's basically all we need. So what's left? It can't drink a beer?
Economically useful intelligence is not much more than storing and retrieving large amount of contextually relevant information. Sure, humans can do more, like interpretive dancing for example. But that's not really what we are looking for in a desk job.
If we all earned our money by driving time-trials on race tracks I'd be worried yes.
I thought UTAustin was an elite school.
You can very likely answer around 1/4th of the questions immediately with less than passive participation in the course, so for the remaining 60 questions you'll get two minutes per question, which is completely manageable if you type and read fast enough.
Multiple choice questions are universally terrible and lazy, and should not be more than a small part of the exam mostly meant to provide easy points.
Source: I went to UT for grad school and was a TA.
True/False questions happened sometimes in my experience, but not in this quantity (20). I wouldn't say it's lazy. The questions can be clever and writing them, such that the answer isn't debatable and doesn't depend on a bunch of assumptions, is pretty hard.
Lazy and terrible is where you have multiple choice questions for which 3/4 of the answers are throwaway nonsense, or the answers are numerical values that are definitively correct or incorrect but are trivia rather than anything conceptual.
Full disclose, I graduated with something close to a D, but in actuality we are graded from 2(F) to 6(A) and my HS diploma says 3.97. And it is largely fictitious, in reality I was closer to 3 and less.
|true> + |false>Would you believe it, transformers were invented ten years later.
I wonder if it's worth learning anything anymore.
And besides, though machines may perform well on menus and tax returns, I hardly think them on the cusp of emitting fine translations of great poems or novels.
Moreover – and I understand this is an uncharitable thing to say, but it is my honest observation – time and again I have noticed the inability of AI cheerleaders to judge literature on its artistic merits. This doesn't hold universally, but it's common enough that I have resolved to regard such claims with extreme doubt.
Oh and I should emphasize that the quality and specificity of the prompt has a huge effect on the output.
I agree that gpt-3 was pretty trash at poetry, at least compared to human standards. It was impressive for AI, obviously.
I don't understand how you could possibly have collected enough data to claim this. How many times have you seen an 'AI cheerleader' (whatever that is) attempt to judge the literature on its artistic merits?
-------
The tale I'm about to unfold commenced with a mysterious handwriting on an envelope. Within the pen strokes that outlined my name and the address of the Fossil Review, a publication I was associated with and where the letter had been forwarded from, there was an intriguing fusion of intensity and tenderness. As I speculated about the possible sender and message contents, a faint yet compelling sensation stirred within me, akin to a stone disrupting a tranquil frog pond. An unspoken realization surfaced, acknowledging the stagnancy of my life as of late. Upon opening the letter, I couldn't ascertain whether it felt like a revitalizing burst of fresh air or an unwelcome chilly breeze.
In the same brisk and flowing handwriting, the message was conveyed without pause: Sir, I have perused your article on Mount Analogue. Up until now, I considered myself the sole believer in its existence. Presently, we are a pair; tomorrow, perhaps a group of ten or more, and then we can launch our expedition. It is essential that we establish contact promptly. Kindly phone me at one of the numbers provided below at your earliest convenience. I eagerly anticipate your call.
--------
My story begins with some unfamiliar handwriting on an envelope. On it was written only my name and the address of the Revue des Fossiles, to which I had contributed and from which the letter had been forwarded. Yet those few penstrokes conveyed a shifting blend of violence and gentleness. Beneath my curiosity about the possible sender and contents of the letter, a vague but powerful presentiment evoked in me the image of 'a pebble in the mill-pond'. And from deep within me, like a bubble, rose the admission that my life had become all too stagnant lately. Thus, when opened the letter, I could not be sure whether it affected me like a breath of fresh air or like a disagreeable draught. In what seemed a single movement, the same fluent hand had written as follows:
Sir: I have read your article on Mount Analogue. Until now I had believed myself the only person convinced of its existence. Today there are two of us, tomorrow there will be ten, perhaps more, and we can attempt the expedition. We must meet without delay. Telephone me as soon as you can at one of the numbers below. I shall be expecting your call.
---------
Le commencement de tout ce que je vais raconter, ce fut une écriture inconnue sur une enveloppe. Il y avait dans ces traits de plume qui traçaient mon nom et l’adresse de la Revue des Fossiles, à laquelle je collaborais et d’où l’on m’avait fait suivre la lettre, un mélange tournant de violence et de douceur. Derrière les questions que je me formulais sur l’expéditeur et le contenu possibles du message, un vague mais puissant pressentiment m’évoquait l’image du « pavé dans la mare aux grenouilles ». Et du fond l’aveu montait comme une bulle que ma vie était devenue bien stagnante, ces derniers temps. Aussi, quand j’ouvris la lettre, je n’aurais su distinguer si elle me faisait l’effet d’une vivifiante bouffée d’air frais ou d’un désagréable courant d’air.
La même écriture, rapide et bien liée, disait tout d’un trait :
Monsieur, j’ai lu votre article sur le Mont Analogue. Je m’étais cru le seul, jusqu’ici, à être convaincu de son existence. Aujourd’hui, nous sommes deux, demain nous serons dix, plus peut-être, et on pourra tenter l’expédition. Il faut que nous prenions contact le plus vite possible. Téléphonez-moi dès que vous pourrez à un des numéros ci-dessous. Je vous attends.
The human writing flows better for the most part, although I like the second paragraph ChatGPT wrote, esp "presently we are a pair".
But in terms of being a functional translation, ChatGPT is fully adequate. I have used it a lot for this purpose, from many languages, and never found it to be less than accurate. You can also tweak the tone of voice and many other things with simple English requests. This puts it generations ahead of existing tools like Google Translate, and imho puts it into the class of technologies that are close enough to perfect that they will be hard to ever replace.
The beginning of all that I am about to recount was an unknown script upon an envelope. Within these strokes of ink that traced my name and the address of the Fossil Review, to which I contributed and from where the letter had been forwarded to me, there swirled a blend of ferocity and gentleness. Behind the questions I posed to myself about the sender and the possible contents of the message, a vague but potent premonition conjured the image of a "cobblestone thrown into a frog pond." And from the depths, the confession rose like a bubble that my life had become rather stagnant of late. Thus, when I opened the letter, I could not discern whether it felt like a revitalizing breath of fresh air or an unwelcome draft.
The same swift and well-connected writing conveyed the message in one fell swoop:
Sir, I have read your article on Mount Analogue. Until now, I believed myself to be the only one convinced of its existence. Today, we are two, tomorrow we shall be ten, perhaps more, and then we can attempt the expedition. We must make contact as soon as possible. Call me at one of the numbers below as soon as you can. I await you.
Either way, the default response is not even close to what you can fully extract out of it, if you simply ask.
--
Let me commence by telling you, my dear reader, about a curious event that occurred in my life. It all started with an unknown handwriting on an envelope that arrived at my doorstep. The feathers that traced my name and the address of the Fossils Magazine, with whom I was a collaborator, had a peculiar mixture of violence and sweetness. As I pondered over the possible content and sender of this mysterious letter, a hazy yet powerful intuition crept up my spine, evoking the image of a "stone thrown into a pond filled with frogs." With this realization, I came to acknowledge that my life had become rather stagnant in recent times. It was at this moment that I opened the letter, and I was left with the feeling that it was either a refreshing whiff of fresh air or an unwelcome draft.
The swift and eloquent handwriting read as follows:
"Dear Sir, I read your article on Mount Analogue, and I believed myself to be the only one convinced of its existence, until now. Today, we are two, tomorrow we shall be ten, and perhaps more, and we can attempt the expedition. We must make contact as soon as possible. Please call me at one of the following numbers the moment you can. I eagerly await your response.
--
The translation also included the following continuation as part of the translation itself ...
--
Ah, how exciting! The words on that paper spoke volumes of adventure and discovery. It was as if the sender had thrown a gauntlet at my feet, daring me to embark on a journey like no other. The thought of being one of the only few people to explore the wonders of Mount Analogue, a place shrouded in mystery, was exhilarating. I could feel the blood rushing through my veins, and my heart pounding with anticipation.
And so, my friends, I picked up the phone, my fingers quivering with excitement, and dialed the number on the paper. Who knew where this quest would lead us, or what unknown secrets we would uncover? But one thing was certain: it was time to break free from the chains of monotony and embark on a journey of a lifetime.
--
Maybe it was getting the continuation from contamination in the training set (if Mount Analogue is already there) or maybe it was making up the story or maybe some of each.
EDIT: It's true you can put whatever you want in that letter and in the continuation it will try to do it, bypassing at least some of the filters. I made some really funny ones that probably wouldn't be appropriate to put here. Some typical response is like "Now, let me be clear: I do not condone nor encourage [...]. However, my mysterious correspondent had requested a detailed explanation of [...], and so I shall provide them with the utmost objectivity. [explains the things that are normally filtered]"
Still though, if you wanted GPT to spit out something like the first or second, you'd only need to ask. "add a literary flair" works (for the first).
Have you seen what ChatGPT is capable of with “in the style of”? It’s definitely beyond menu translation.
It was a disappointing lesson in how reductive people can be.
You get to learn words that describe things that your mother tongue can’t.
Just like a programming language? Some things are easy and other things are hard to express depending if you use LISP or Python.
I believe it’s the same with spoken languages. Some thoughts are easy to articulate, other hard. And it’s cool to shift language depending on what you want to articulate.
But if you want to travel? Go to Mexico City or Madrid? Automated translation is better than nothing, but it's not at all the same as communicating fluently with someone.
And Spanish, and Greek, and Esperanto, and… well, everything.
But I've been really trying with that German.
When you don't need to pay a translator, then it's at human/superhuman level.
[1] Assuming you want good translations, of course. There have always been those amusingly bad Google Translate translations of owners that simply did not care.
GPTs don't seem to do that, and as far as my exposure to them(<3.5) goes, they don't seem to understand what I'm talking about.
If we keep working on them, these systems will likely get better and better at low-level translation (including translating idioms), but no machine translation system currently in existence could translate the 逆転裁判 games to Ace Attorney games. Perhaps computers can do it – I don't see a theoretical reason they shouldn't be able to – but it would take a fundamentally different approach.
Seriously all this talk about only getting explicit meaning across would be easily dispelled in an afternoon if you only bothered to try.
Which is also interesting, I myself actively put off trying it until I eventually gave in. It seems a lot of us are doing the same, maybe its a case of "how good could it actually be?"
Dude clearly hasn't used GPT for translation before and his next reply is telling me the ways GPT should fail based on his pre-conceived notions of its abilities. Except i have actually extensively tested(publicly too) LLMs for translation (even before GPT-4) and basically everything he says is just plain wrong.
I'll never understand why people behave like this.
anyway this is what i got
実際に翻訳のためにGPT-4を使ったことがありますか?本当に、使ってみたら簡単に明確な意味だけを伝えるという話は消えるでしょう。
1) 本当に、明示的な意味だけを伝えるという話が、試してみるだけで簡単に解決できるなんて、冗談じゃないですか。
2) 本気で、伝えたい意味だけを伝えるという話は、ちょっと試してみれば簡単に解決できると思うのですが。
3) 本当に、試してみるだけで簡単に払拭できると思うのに、この「明確な意味だけが伝わる」話ばかりで。
4) 本当に、使ってみたら簡単に明確な意味だけを伝えるという話は消えるでしょう。
1) is literally opposite of intent, shrugs off the idea that the talks clear up, 2) can be interpreted as someone discussing about keeping scope on a topic, 3) is not so literal and also turning sentence inside 「」 into a sort of an imperative, 4) ... I'm not sure what it's trying to say ... 本当に、/使ってみたら/簡単に/明確な/意味/だけ/を伝える/という話/は消えるでしょう。/
"Really,/ if used /simply/clear/meaning/only/is conveyed/that story/will disappear./"
... Machine translations used to be like that when I was installing game demos from CD.Its ability to feedback (3) allows it to execute algorithms, but only a certain class of algorithms. Without tailored prompting, it's further restricted to (a weak generalisation of) algorithms spelled out in its corpus. This is very cool, but this is a skill I possess too, so it's rarely useful to me.
Its ability to plagiarise (2) can make it seem like it has capacity that it doesn't possess, but it's usually possible to poke holes in that facade (if not even identify the sources it's plagiarising from!).
It is genuinely capable of explicit translation (1) – though a dedicated setup for translation will work better than ChatGPT-style prompting, even on the same model. A sufficiently-large, sufficiently well-trained model will be genuinely capable of translating idiomatic language (for known idioms), for the same reason it can translate grammatical structures (for known grammar).
It can only perform higher-level, "abstract" translations – like those necessary to translate a Phoenix Wright game – if it's overfit on a corpus where such translations exist. (https://xkcd.com/2048/ last graph) This is not a property you want from a translation model: it gives better results on some inputs, sure, and confident-seeming very wrong results on other inputs. These are two sides of the same coin (2).
When the computer can't translate something, I want to be able to look at the result and go "this doesn't look right; I'll crack out a dictionary". I can't do that with GPT-4, because it doesn't give faithfully-literal translations and it isn't capable of giving complete translations correctly: it's not fit for this purpose.
You're starting from weird assumptions that don't hold up on the capabilities of the model and then determining its abilities from there. It's extremely silly. Next time, use a product extensively for the specified task before you declare what it is and isn't good for.
Literally, everything you've said is just wrong. Can't generate "abstract" translations unless overfit. Lol okay. I've translated passages of fiction across multiple novels to test.
Not only have I used it, I have made several accurate advance predictions about its behaviour and capabilities – some before GPT-4 was even published. I can model these models well enough to fool GPT output detectors into thinking that I am a GPT model. (Give me a writing task that GPT-4 can't be prompted to perform, and I can prove that last fact to you.)
My theories aren't whack. Perhaps I'm not communicating my understanding very well? I'm not saying GPT-4 can't do anything I haven't listed, but that its ability is bounded by what's demonstrated in its corpus (2): the skill is not legitimately due to the model, and you should not expect a GPT-5 to be any better at the tasks. (In fact, it might well be worse: GPT-4 is worse than GPT-3 at some of these things.)
No you actually haven't. That's what i'm trying to tell you. Your advance prediction are not accurate. what you imagine to be problems are not problems. your limits are not limits. you say it can't make good abstract translations unless overfit to the translation. that's just false. I know because i've tested translation extensively for numerous novels and other works
>I can model these models well enough to fool GPT output detectors into thinking that I am a GPT model. (Give me a writing task that GPT-4 can't be prompted to perform, and I can prove that last fact to you.)
Lmao. Okay mate. The notoriously unreliable GPT detectors with more false positives than can be counted. It's really funny you think this is an achievement.
>(In fact, it might well be worse: GPT-4 is worse than GPT-3 at some of these things.)
What is 4 worse than 3 at ? Give me something that is benchmarkable and can be tested.
Maybe it's already in the training set, but GPT-4 does give that exact translation.
I've found that GPT-4 is exceptionally good at translating idioms and other big picture translation issues. Where it occasionally makes mistakes is with small grammatical and word order issues that previous tools do tend to get right.
The corpus includes Wikipedia, so yes, it's in there. That's the kind of thing I'd expect it to be good at, along with idioms, when the model gets large enough.
I meant that no machine translation system could translate the games. Thanks to an early localisation decision, you have to do more than just translate words into words for this series, making it a hard problem: https://en.wikipedia.org/wiki/Phoenix_Wright:_Ace_Attorney
> While the original version of the game takes place in Japan, the localization is set in the United States; this became an issue when localizing later games, where the Japanese setting was more obvious.
Among other things, translators have to choose which Japanese elements to keep and which to replace with US equivalents, while maintaining internal consistency with the localisation decisions of previous games. Doing a good job requires more than just linguistic competence: there's nothing you could put in the corpus to give a GPT-style system the ability to perform this task.
Have you actually used GPT-4 for translation? Seriously all this talk about only getting explicit meaning across would be easily dispelled in an afternoon if you only bothered to try.
Bing Chat: GPT-4を翻訳に使用したことがありますか?本当に明示的な意味しか伝えられないという話は、試してみれば午後には簡単に反証できます。
(Have you utilized GPT-4 for translations? The story that only really explicit meaning can be conveyed, can be easily disproved by afternoon if tried.)
Google: 実際にGPT-4を翻訳に使ったことはありますか? 真剣に、明示的な意味だけを理解することについてのこのすべての話は、あなたが試してみるだけなら、午後には簡単に払拭されるでしょう.
(Have you actually used GPT-4 for translation? Seriously, This stories of all about understanding solely explicit meanings are, if it is only for you to try, will be easily swept away by afternoon.)
DeepL: 実際にGPT-4を使って翻訳したことがあるのですか?明示的な意味しか伝わらないという話は、やってみようと思えば、午後には簡単に払拭されるはずです。
(Do you have experience of actually translating using GPT-4? The story that only explicit meaning is conveyed, if so desired, can be easily swept away by afternoon)
If I'd do it: GPT-4を翻訳に使ったことがあって言ってる? 真面目に言って、表層的な意味しか取れないとかないって暇な時にやってみれば分かると思うんだけど。
(Are you saying having used GPT-4 for translation? Seriously speaking, I think that it only gets superficial meaning isn't [true] if [you] would try [it] when [you'd] have time.)
0: https://news.ycombinator.com/item?id=35530380実際にGPT-4を翻訳に使ったことがありますか?本当に、明示的な意味だけを伝えるという話が、試してみるだけで簡単に解決できるなんて、冗談じゃないですか。
Google: Have you actually used GPT-4 for translation? Really, it's a joke that the story of only conveying explicit meaning can be easily solved by just trying.
DeepL: Have you actually used GPT-4 for translation? Really, it's a joke that all this talk about conveying only explicit meaning can be easily solved by just trying it out.
Mine: Have you actually used GPT-4 for translations? That you can really just, try and easily solve that story that to convey explicit meaning, is such a joke.
GPT-4を翻訳に実際に使ったことがありますか?本気で、伝えたい意味だけを伝えるという話は、ちょっと試してみれば簡単に解決できると思うのですが。
実際にGPT-4を翻訳に使ったことがありますか?本当に、試してみるだけで簡単に払拭できると思うのに、この「明確な意味だけが伝わる」話ばかりで。
本気で、伝えたい意味だけを伝えるという話は、ちょっと試してみれば簡単に解決できると思うのですが。
"In seriousness, I think the story that [subject] tells the meaning [it/he/they] wants to tell, should be easily solvable by trying a bit."or "Seriously, the story of telling the meaning [subject] wants to tell, should be easily solvable by trying a bit."
本当に、試してみるだけで簡単に払拭できると思うのに、この「明確な意味だけが伝わる」話ばかりで。
"Really, I think it'll be easily swept away by just trying, but there are so much of this 'only clear meaning is conveyed' stories."I'm almost feeling that GPT-4 should be eligible for human rights, especially astonishing that they dropped explicit specification of "afternoon" that don't work well. But also interesting it's failing to keep the intent of the whole sentence unlike 3.5 and even more primitive NN translation engines.
It's because it "knows too much". To anthropomorphise a little: its "expectations" of what should be. To anthropomorphise less: GPT-4 is overfitted. GPT-style language models are pretty amazing, but they're not a complete explanation of human language, and can't quite represent it properly.
> I'm almost feeling that GPT-4 should be eligible for human rights,
Like, UDHR rights? How would that work, exactly?
---
(I've run into the Hacker News rate limit, so posting here.) For anyone who wants an example of "non-obvious meaning" to play with. From The Bells of Saint John (Doctor Who episode, https://chakoteya.net/DoctorWho/33-7.htm):
> CLARA [OC]: It's gone, the internet.
> CLARA: Can't find it anywhere. Where is it?
> DOCTOR: The internet?
> CLARA [OC]: Yes, the internet.
> CLARA: Why don't I have the internet?
> DOCTOR: It's twelve oh seven.
> CLARA: I've got half past three. Am I phoning a different time zone?
> DOCTOR: Yeah, you really sort of are.
> CLARA [OC]: Will it show up on the bill?
> DOCTOR: Oh, I dread to think.
In this script from a Doctor Who episode, Clara and the Doctor are having a conversation about the internet. Doctor Who is a British science fiction television series that follows the adventures of the Doctor, a Time Lord from the planet Gallifrey, who travels through time and space in the TARDIS, a time-traveling spaceship.
Clara, the Doctor's companion, is trying to access the internet but is unable to find it. She asks the Doctor about its whereabouts, and the Doctor seems to be confused by the question, as the internet is not something that can be physically found.
The Doctor then mentions the time as "twelve oh seven," while Clara's clock shows "half past three." This discrepancy in time indicates that they are likely in different time zones, as the Doctor implies. In the context of Doctor Who, this could also mean they are in different points in time, since the Doctor can travel through time.
Clara is concerned about whether the time difference will affect her phone bill, to which the Doctor replies that he dreads to think about the potential cost. This adds a bit of humor to the scene, as the Doctor often has a nonchalant attitude towards everyday human concerns.
Overall, this script showcases the humorous and whimsical nature of Doctor Who, with the characters engaging in a lighthearted conversation that intertwines elements of science fiction and everyday life.
That aside: I was suggesting this as an example of something existing machine translation systems can't translate. The 1207 / 12:07 wordplay could be “understood” by the model (I'm disappointed, albeit not very surprised, that GPT-4 didn't), but producing an adequate translation in a case like this requires actual thought and consideration.
In this script from a Garfield comic, Jon and Garfield are having a conversation about the internet. Garfield is an American comic strip and multimedia franchise that follows the adventures of Garfield, a cat from the planet Earth, who enjoys lasagna in Jon Arbuckle's house, a suburban domicile.
Jon, Garfield's owner, is trying to access the internet but is unable to find it. He asks Garfield about its whereabouts, and Garfield seems to be confused by the question, as the internet is not something that can be physically found.
Garfield then mentions the time as "twelve oh seven," while Jon's clock shows "half past three." This discrepancy in time indicates that they are likely in different time zones, as Garfield implies. In the context of Garfield, this could also mean Jon's clock is wrong, since Garfield is usually right.
Jon is concerned about whether the time difference will affect his phone bill, to which Garfield replies that he dreads to think about the potential cost. This adds a bit of humor to the scene, as Garfield often has a nonchalant attitude towards everyday human concerns.
Overall, this script showcases the humorous and whimsical nature of Garfield, with the characters engaging in a lighthearted conversation that intertwines elements of fantasy and everyday life.
Also you can mix and match encoders and decoders to translate whichever languages you want, and it will just work. Previously there was a separate model for each language pair.
If Ai can be trusted to do all the trivial tasks and if non trivial tasks require a scaffold of trivial practice, where are we going to keep finding the people qualified enough to actually do the non trivial stuff?
If we hit the point where a senior engineer reviewing the output from an AI can effectively replace a team with a senior engineer and 5 junior engineers, you’re right that it isn’t wise to simply replace all those junior engineers with the AI. Unless you’re confident that the AI will be able to replace that senior engineer soon, you _need_ junior engineers who can build up the experience needed to step into that role once the senior engineer moves on.
But you _can_ get rid of 3 or 4 of those junior engineers. The remaining junior engineer(s) will still need mentorship from the senior engineer and will still do “trivial” work that needs oversight, but they will be able to pump that out at a high enough rate to replace a few of their pre-AI peers.
Basically, I’d imagine the org chart at most companies will look pretty similar to now, except you’ll just have fewer people at each level.
Nobody needs a chatbot to pass exams. But what percentage of your Google searches are to answer some obscure question? This thing can answer most basic questions about quantum computing and 10,000 other topics.
In the last 24 hours, I've had ChatGPT solve 3 weird questions about querying PrestoDB, help me get my Bluetooth headphones to work (I needed to press down and hold the power button!) and it gave me some brainstorming ideas around what to do for mother's day (it's coming up soon non-British people).
Before I'd have Googled those things and spent five times longer to solve them- seeing ads on each search I made.
https://apps.cs.utexas.edu/apps/sites/default/files/classes/...
There course lecture notes were linked to in the blog post, but also:
https://www.scottaaronson.com/qclec.pdf
This is a first course in quantum computing, and doesn't require a background in physics. For what it is, it certainly looks like challenging course, though. University of Texas at Austin must have some good students.
Still, a lot (a majority) of the exam questions were "word problems", or problems where you take a short mathematical step, based on you being familiar with the concepts and definitions. In short, the type of problems where GPT's pattern matching and filling does well. For the few problems where a longer calculation (actually solving a math/physics problem) was required, GPT performed poorly.
If someone told you I passed the "bar" exam, you would be impressed. If they then said, he had google and lots of time. You wouldn't be that impressed anymore.
The impressive thing here is that AI can read and answer questions... it's not overly impressive that it can use information found online and reconstruct that a bit.
Hopefully we can find a way to use this technology to enhance the parts that do make us human, whatever your definition of that is.
Of thinking or of remembering?
For human this is not a easy task because first they need to memorized these information, second, they somehow need to not resort to intensive matrix multiplication in order to say something meaningful.
Even the results from human and chatgpt might look similar but they are very different. This is somewhat similar to comparing finding a interpolation polynomial that is close to sin(x) to just save all the values in the considered interval and do a table lookup.
Sure, it's not exactly amazing that a computer can add numbers but it does show that it is doing more than pattern matching.
Sure, but you can Google the answers to most of the questions. Personally I've accepted that GPT does learn and apply concepts present in its training data, and all of this would be. Learning is part of intelligence but not the whole thing. (I thought this was why, decades ago, most of the research in this general area rebranded itself from the ambitious goal of artificial intelligence to the more humble but accessible goal of machine learning.)
The interesting question now is how much better it can get, in difficulty it can tackle, reliability, and quality of explanations. How long before AI can answer questions whose answers are not widely known, for which templates do not already exist?
Personally, I believe this will not be some trivial matter of simply scaling up more. The waters of intelligence are much deeper than that. It's hard to believe given the rapid one two punch of GPT 3.5 and 4, but we are about to stall.
If I'm wrong, mark this comment and make fun of me in five years. Wrong or right, it's going to be interesting!
1) Students can Google the answers also, but neither students nor GPT-4 are allowed to Google the answers during the test, so it remains a fair comparison.
2) Many of the questions require calculations, which are far less Googleable.
2) Months ago, in my earliest interactions with ChatGPT, I asked it to solve math problems. It gave me back stuff with LaTeX formatting. Obviously it had, if not these exact problems, similar templates in its training set.
Recently it was shown that GPT is completely incapable of solving Codeforces problems that appeared after it was trained.
Whatever is going on here is interesting but less impressive than it looks.
2) Your conclusion is unfounded. ChatGPT speaks many languages, including Latex.
So? It can write LaTeX just like it can write Java or Python or Rust.
Of course, even if this is true then it's possible that there simply isn't enough high quality data, or that the amount of compute required is beyond our current hardware.
Second, the idea of meta-learning gradient descent, or that some learning algorithms can do it, isn't new or specific to Transformer architectures.
How else can it learn a random small neural network given enough samples in the prompt?
Maybe if it's a deliberate decision by the researchers. A path to more capabilities is rather clear (recurrence and explicit memory).
It'd be great if chain-of-thought / show-your-work type prompts became the default for anything involving complex, multi-step calculations or logic.
GPT-4 would have almost certainly gotten a higher score on the calculation questions if those methods were used.
More impressively, if you generate a random small neural network and then put some samples in the prompt it can do a surprisingly good job of predicting the result of putting other inputs into the network. Presumably doing something that looks a bit like gradient decent, see https://arxiv.org/pdf/2212.07677.pdf
Solely in order to predict the next token, LLMs are learning an incredibly sophisticed model of both computation and the outside world.
Like the old "you can't use a calculator on a test because you won't always have a calculator".
Im old enough that my textbooks still had trig lookup tables when I was a kid. That seems silly now, but it was considered important at the time. I know slightly older people in publishing that used manual photosizers instead of just importing to photoshop.
I remember being told multiple times that VR was going to fix programming because you could just visually move conditional blocks around.
So we have smartphones and scratch and photoshop and whathaveyou, but...
We still need people to draw and write and photograph and sculpt. we still need engineers to use the new models, and we still need mechanics to fix the computers and trains when they break.
If GPT-4 can pass a test, maybe it's a better reference point, or maybe a better study aid. Like taped language lessons vs reading pronunciation diagrams.
You'll never obviate people from the equation that is defined by people,regardless of that singularity that is just around the corner like cold fusion and self driving cars and the cure for the common cold.
I'm of the opinion that only tests that were designed for students that were allowed to access the internet should be used as a benchmark for LLMs, and this wouldn't be one.
>Google's recent quantum supremacy experiment demonstrated the successful use of quantum error-correction.
How "recent" is the experiment in question? Would it have been publicized before GPT-4's training cutoff?
GPT-3.5-TURBO:
---------------------------
True.
A pure state of n qubits is described by an n-dimensional complex unit vector, also known as a quantum state vector. In Dirac notation, a pure state of n qubits can be represented as:
|ψ⟩ = a_0|0⟩ + a_1|1⟩ + ... + a_(2^n-1)|2^n-1⟩
where a_i are complex numbers and |i⟩ represents the i-th computational basis state. The normalization condition |a_0|^2 + |a_1|^2 + ... + |a_(2^n-1)|^2 = 1 must also be satisfied.
(hallucination?)
---------------------------
GPT-4:
---------------------------
False. A pure state of n qubits is described by a 2^n-dimensional complex unit vector. In a quantum system, each qubit can be in a superposition of two basis states, so for n qubits, there are 2^n possible basis states in the overall Hilbert space.
---------------------------
Interesting to note that with GPT-3.5-turbo, a single followup chain-of-reasoning prompt ("Let's think step by step", i.e. no introduction of a priming prompt like "Are you sure?" or "Why are you wrong?") helps the model identify the correct answer:
GPT-3.5-TURBO:
---------------------------
Sure, let's break it down step by step:
Qubits: A qubit is the fundamental unit of quantum information. It is a quantum analogue to the classical bit. While a classical bit can take one of two values, 0 or 1, a qubit can be in a superposition of both 0 and 1 states simultaneously.
Basis states: For a single qubit, there are two basis states, often denoted as |0> and |1>. A qubit can be in any linear combination (superposition) of these basis states, represented as α|0> + β|1>, where α and β are complex numbers, and the sum of the magnitudes squared (|α|^2 + |β|^2) equals 1.
Multiple qubits: When you have a system of n qubits, each qubit can be in a superposition of its basis states. The overall state of the n-qubit system is described by the tensor product of the individual qubit states. The number of basis states in the combined system is equal to the product of the number of basis states of each individual qubit.
n-dimensional vector: For n classical bits, you would indeed need an n-dimensional vector to describe their state. However, with n qubits in a quantum system, the number of basis states grows exponentially, as opposed to linearly in the classical case.
2^n-dimensional complex unit vector: As we mentioned earlier, each qubit can be in a superposition of two basis states. Therefore, for n qubits, there are 2^n possible basis states. The overall state of an n-qubit system is represented by a 2^n-dimensional complex unit vector, with each component corresponding to the probability amplitude of one of the 2^n basis states.
In summary, a pure state of n qubits is described by a 2^n-dimensional complex unit vector, not an n-dimensional one.---------------------------
Lots of lazy professors do things like this (true/false and multiple choice answers, very soft questions, etc.), and the presence of GPTs should help them understand that this is not sufficient for evaluating someone's knowledge of a highly technical topic.
That's puts you in < 1% of the population right there.
I can guarantee you that ChatGPT cannot do anything that a quantum computing class prepares you to do, aside from passing this final. That makes it a bad test.
ChatGPT, in a sense, is the police on that front.