GPT-4 Designed a Programming Language
lukebechtel.com
lukebechtel.com
And now here we are: it seems we tech people are every bit as susceptible to this kind of thinking as the average joes we looked down on.
Actually look at this “language” that GPT-4 has written: it’s just a mishmash of features from the most commonly blogged-about existing programming languages. It has no conceptual insights, no original ideas, no taste that isn’t cribbed from one of the few languages it’s copying from.
There are cases where getting a short distillation of the mean opinion of The Internet is useful. But designing a new programming language is clearly not one of them, mostly because the chatbot is not “designing” a language—it is guessing the most likely response to a question about programming languages.
I used GPT-3/3.5/4 for multiple tasks and I never found it useful. It just feels… average.
It can “answer” many questions correctly but that’s only on the data set it’s been training on. I asked it to solve Elixir issues, write Elisp snippets, figure out why my Rust code was complaining about lifetimes and even went for easy mode and asked for JavaScript code.
Only JavaScript (very popular) responses were workable but they still had popular problems included in it (e.g. not clearing out timers or dealing with common edge cases). Elixir “hints” were very interesting since - when out of water - it would produce code assuming that it’s a OOP and that was a rabbit hole it declined to exit.
Too technical? Went for fun with family members who have exceptional cooking skills. We tried to figure out a dish variation of an apple pie. Suggestions were heavily criticized as non workable for many reasons (too sugary, consistency wouldn’t be too good, modification change the ingredients but GPT refuse to break of out theme etc.)
Colin Meloy made an experiment some time ago and created The Decemberists-like song with it and then actually recorded it. It worked but it was described as below average and bland. I share that vibe.
In the end I have this feeling that GPT is on a level of 7 year old with an access to vast library. It uses proper grammar it can copy parts of texts, but one needs to provide it with a great benefit of doubt and energy just so that it’s useful.
In contrast to 7 year old, GPT isn’t grateful for that sacrifice and is actually making money of you.
This is not an exception either; teams like mine exist in every large enterprise. I worked in several of them and this AI makes life so much better. Let’s not pretend it makes certain outsourcing countries not fully obsolete though, to the tune of 10s of millions+ programmers, data entry workers, text producers (bloggers for seo, docs no one reads etc) and more are mostly all already worse than gpt3.5 and 4. Wait for 6 or 7.
I do think really this is you’re not being in touch with most of reality and what most people on earth do daily at this moment. It definitely as nothing to do with the productivity you have and the drive you feel internally to ‘get shit done’. Most people, by a huge stretch, get education, want to do work what they are told to do related so that education and that’s it. They have no, at all, drive to ever move beyond that. They learned Angular on their first job; that’s it. Etc.
This is now replaceable; this is not, at all, comparable to automation before this. It’s a threshold; it you cannot see it, you definitely will be shocked in the coming years. This is not going away and it is replacing humans now on a small scale. The companies that really can just remove 90% of their staff are late always, but it will come. I see this all day because I work in massive orgs where this holds, not only for tech.
I’ve found it useful for drafting cover letters (I just need to do final edits).
I’ve also found it useful as an adjunct to google when I’m searching for niche things like “Luxury resorts with activities for my dog and lots of off-leash trails”.
So far that’s it, but it does seem like a useful addition to Google search for super niche things. Obviously it hallucinates a lot on these niche edge cases but it gives some great starting points.
So... this early-stage AI was able to make music that's better than the vast majority of commercial recording artists these days? I fail to see how that's a failure.
It's probably more like a 14 year old but I think you're probably right. It can give step-by-step instructions and even the code to set up entire web apps, but it's only when you go through those steps that you find the issues. Packages out of date, less-than-optimal ways of handling non-obvious cases, and I've even found it trying to use javascript functions that have never existed. To its credit, it apologizes and tries again when I call it out. Maybe GPT-5 will be the real one to worry about.
But look how far we've come since those heady and carefree days. We're already able to pour buckets of scorn on its capabilities, and brush them aside as if designing a language (albeit a derivative and uninspired one) is a trivial task that any human could achieve.
No we don't, we are just trying to put a damper on all the "the singularity is here!" style comments that seems to flood any discussion about GPT.
These models are trained to produce text snippets that look like text snippets it has seen before, and it has seen all internet. That means it can do a lot of impressive things, but also that it is very dumb in other ways.
I've seen it being dumb in maths or real world problems. But as a large language models, they understand and speak languages fine, and even mistakes they make look like mistakes humans who are not natives in the language would make.
We may as well say that when we speak, we are just predicting words we have trained on. I don't see how these models are worse than people in that regard.
The general knowledge and thinking of these models are surely limited. But seeing GPT-4 go from text only input to text with images, I think it is very possible to break the barriers very soon.
Yep. Welcome to HN. And welcome to the majority of humanity!
https://lifearchitect.ai/contrarianism/
What will be the other impacts?
This is just another tool in the student’s toolbox and should be treated as such.
Modern countries don't do homework, other than reading, for ages
> designing a language (albeit a derivative and uninspired one) is a trivial task that any human could achieve.
It is, isn’t it? You just need to be able to read and write words and symbols that reassemble a programming language coherently and follow the pattern. This is what humans do after all; remixing. Anyone can do this.I know a competent man should be able to "change a diaper, plan an invasion, butcher a hog, conn a ship, design a building, write a sonnet, balance accounts, build a wall, set a bone, comfort the dying, take orders, give orders [...]" and what not and specialization is for insects, but even on the Mysterious Island not everyone is Cyrus Smith (Harding).
edit: typo fix
But then applying them directly into an example falls down a bit. The language is a bit of a mismash and parts of it might not work semantically. The use of actors looks a bit out of place, and raises a lot of questions about how the memory management and scheduling is working under the hood. Despite the bold aims in the tenents thre is nothing in the language exposed so far that addresses them: why would the language produce more modular code in the large? what features of the syntax would increase the level of abstraction in code written? etc etc
Overall this would warrant a (barely) passing grade by modern standards, i.e. it can regurgitate the material succesfully but without demonstrating any deep understanding. I find the result quite interesting as it demonstrates how much of what we judge to be quality is actually fitting a particular pattern of response (w.r.t a specific field of study) rather than exploring the material more deeply.
I like the use of the arrows in the function definitions though :) I've seen that somewhere before and will be "appropriating" it for the project that I'm currently working on.
Python type hints probably.
I (ab)used them for good effect when I was fiddling around with the spark earley parser that used to be part of the python build system.
Not sure why need any new languages when we have the perfection that is Golang tho.
More. We’re more suspectible.
We are the people who take the Turing Test as an objective test of GAI, after all, even though it only tests how we percieve supposed intelligence.
Correct in that that was what Turing intended. But the vulgar interpretation of it makes it out to be more than it really is.
For Pete’s sake: two years ago I even saw a psychologist who suggested that psychologists should take notes from how computer science ostensibly measures “intelligence” by the Turing Test, because it’s so “objective”. That was sad.
It will also learn all the bugs and reproduce them.
Programmers who think ChatGPT is going to replace their job must have a very low opinion of their own skill set.
We see it in this article, both the statement "Indentation-based scoping, similar to Python" and the code with bracers were probably correct in their original contexts, but combined they are no longer correct. LLMs will make such errors no matter what kind of data you train it on, there is no known way to avoid that class of errors.
As a devops guy I live in fear of anything that can copy and paste from stackoverflow faster than I can.
While the language it generated needs refinement, the design process is iterative, and it takes time to develop a programming language that is worth using.
I encourage you to consider how you would approach designing a programming language. Like you, I would likely start by incorporating my favorite features from languages I love and addressing problems I've encountered in the past, similar to the dialogue GPT and I have in the post.
As I mention in the post -- creativity often involves remixing existing ideas, and GPT-4's attempt at generating a programming language is no exception. Notably, it also called out explicitly the constructs it was borrowing from other languages when designing the language.
It's worth noting that this was GPT-4's initial attempt at something highly complex, and even if there are "plot holes" -- it's a noteworthy achievement in and of itself.
Government regulation has been brought up multiple times, most recently by Sama himself on twitter. I think it will have to come, but I don't know what flavor it should take.
The cynical take is quite clearly: "Because sales people can sell it"
eg wrap some pretty UX around it, tell people it's "nearly human", then have your sales people try and get non-technical CxO types (or consultancies) to cut them a cheque
> As I mention in the post -- creativity often involves remixing existing ideas
This is carrying the weight of your core argument. It’s not wrong but it’s not completely correct either.Creativity involves remixing but is not all remixing. Is it? Not even all humans are at the same level in this. Number of attempts means nothing if you don’t add features that the LLM desperately needs to be grounded in the reality, or perform logical reasoning beyond RLHF.
However I want to add that it does the guesswork and the presentation surprisingly well to a degree that it’s effective in certain contexts.
The more people think of this as a tool that they can use and master, the more they will benefit from it.
AlphaGo guessed it’s way to being the best.
Can something guess it’s way to evolving into a less guessy thing? I think so. Look at us?
I am aware of how it works and used to be ‘it’s ok, it won’t happen’ but now I am actually pretty concerned.
We think GAI as the thing but there maybe a reality that doesn’t require full GAI to completely unhinge us.
I mean, what is the point? It only plays to the negatives of our society. We don’t need it. It allows for ‘do more’ but what cost does that ‘do more’ come with?
No it didn't, AlphaGo was backed by an algorithm that perfectly understood go and used that to train itself to play go better. That is a completely different situation.
If you tried to make a go engine that only tried to replicate old moves then you'd have a hard time making it perform valid moves. ChatGPT is like such an engine and we just made it perform valid moves most of the time, and now people say "now that it can make valid chess moves it is only a matter of time until it beats masters!". No just statistically reproducing human moves wont lead to a smart chess engine, so why would you think that would make a smart text engine?
It seems surprising that training an LLM would result in layers/weights which contain such representations but it seems that this is indeed an efficient representation of information in the same way that neural nets would learn addition by encoding textual integers into floats.
AlphaGo was trained with tech. It beat humans.
ChatGPT et al are being trained.
Everyone is arguing the inner workings of what is in the box. It doesn’t matter.
The usage of the box(s) is what will be the big impact.
Previously those boxes couldn’t be relied upon so stayed down in the hierarchy of usefulness.
Now. They are rising.
Can ChatGPT(-ish), trained on a ton of games, beat AlphaGo?
It had basic rules of Go as an input.
--Amos Tversky
On the one hand, I agree.
On the other hand, I see some variant of this comment on HN about every new language, so...
"LLMs in the GPT family have read every programming language in the world, billions of times. It's been known for some time that these agents can write small snippets of code with some limited success. However, to my knowledge, none of these agents have ever been asked to make their own programming language. This type of task is a bit more complex, and requires a bit more creativity and foresight than much of the internet believes LLMs to have.
Some folks on the internet tend to think lowly of LLMs as merely "Stochastic Parrots" -- simply "remixers" of old ideas. In a way, they're right; but I'd argue that remixing is the fundamental force of creativity. Nothing comes from scratch."
This^^. Half the internet was in denial about how good LLMs are. I wonder why? Perhaps it's the way humans act when confronted with a machine that can potentially replace much of what makes us special.The advent of chatGPT had millions of people on the internet trying to downplay the intelligence of chatGPT by continuously trying to re-emphasize the things it gets wrong.
I think the people who claimed chatGPT was a stochastic parrot are now realizing that they were the ones that were part of a giant parade of parrots regurgitating the same old tired trope of LLMs being nothing but simple word generators. Well you guys were dead wrong, a simple iterative improvement on the same algorithm pretty much ironed out a lot of the issues.
I'm guessing that at GPT-10 or 11 we'll have produced something that when compared with humans, humans will be the things that are far more parrot-like.
> Indentation-based scoping, similar to Python.
Then...
> class Person {
To me, it's very clear it doesn't really understand what "indentation-based scoping" means. For a human programmer, using both intentation-based and bracket-based scoping is a very unusual design choice, and is begging for further explanation. It talked pages of abstraction, modularity, etc, but not this?
I'm not saying it's not good. But I failed to see how this particular example proves its not a "stochastic parrot". If anything, I think it's a quite strong evidence supporting "stochastic parrot" narrative.
But my argument is that humans are also, at some level, sufficiently advanced stochastic parrots. At least, as it pertains to many creative endeavors.
You're always building on the back of something that comes before.
We're all just remixing ideas we've heard before.
I mean even this comment -- none of the words I'm saying are new, and many of the word combinations I'm using have been used time and time again. The ideas I'm expressing have merit -- but are they wholly original?
Not really. We all need eachother's creative energies to do our best work.
"Some programming languages like TypeScript and Swift use a colon for both variable type annotations and function return types."
This is incorrect for Swift, which uses the same arrows as "TenetLang" for function return types. Actually the first thing I thought looking at the example code was "looks swifty but not quite as well designed."
Right, that is truly what the chatbots are lacking. They can fool some of the people some of the time, but they can't fool all of the people all of the time.
Their creators - not the chatbots themselves - are trying to "fool" people into believing that the chatbots would have as you say "ability to understand what you don't know, what you know, and what follows from those things consistently".
But it's just a possibility, and I don't find it's particularly convincing.
While that does not prove that we never apply reasoning ahead of time, it is a pretty compelling indication that we can not trust that reasoning we give isn't a post-rationalisation rather than an indication of our actual reasoning, if any.
(Funnily enough, while both your example questions are easy enough, I feel I was marginally slower on your "easy" question than your "harder" one)
Do inference engines have the other half? Or Cyc plus an inference engine? Can that be coupled to ChatGPT?
My own (completely uninformed) take is that such a coupled AI would be very formidable (far more than ChatGPT), but that it will be very hard to do so, because the representations are totally different. Like, really totally - there is no common ground at all.
But it's more than that. A good educator teaches students to evaluate sources, not to just believe everything they read. As far as I can tell, ChatGPT totally lacks that, and it hurts.
Llm’s are the language model. That is it.
Now we have gone down this path to a very useful tool what’s next is connecting it to something that can understand context and memory.
It’s not so hard to imagine being connected to a knowledge graph of all the things and this evolving into a very capable AI.
It’s like Google v1 compared to Google now. Key words versus semantic seo
So, like the average voter?
I think GPT is on the other hand both over- and underestimated because it speaks well but often makes reasoning mistakes we only expect of someone less eloquent, and it throws us.
If it had come across as an inquisitive child, we'd have been a lot more overbearing of it's "hallucinations" for example, because kids do variations over that all the time.
At the same time, it can do some things most children would have no hope of.
It's a category of "intelligence" we're unfamiliar with, additionally hobbled with no dynamic long term memory.
(And humans are not very reliable; more than current models, sure, but still pretty bad)
https://www.engraved.blog/building-a-virtual-machine-inside/
Here's another one.
But there is also Unison. That one new language that stands out, and the very fact that it does makes it improbable for an LLM to generate. LM is a language model, it avoids marginality and originality altogether by its design.
Language models can and are useful if you specifically want to avoid marginality. To reduce noise, to remove errors. There is huge potential in them. If the problem was to design the most average programming language with no purpose, no market niche, and no technological context - then GPT-4 is clearly a winner.
To be fair, most programming languages don't try to be original. They try to solve problems. They also try to keep syntax close enough to other established programming treds so people have an easier time learning them.
We are in good part stochastic parrots for sure, but there's a small nuance that is completely lost in your reasoning and which is what makes the difference today between humans and GPT.
In those languages, reusing commonly known syntax is an explicit feature. Whenever the language had no reason to introduce something new, it used known idioms to lower the amount of things one has to learn to use the language.
Occasionally though, a language will have something unusual, but it won't be for random reasons, but instead to mark something unique and interesting that was added to the language to tackle a facet of programming in a new way.
Rust has lifetime annotations, Zig inverts the usage of async/await (https://kristoff.it/blog/zig-colorblind-async-await/) to support color-blindness.
Present day GPT wouldn't be able to come up with a borrow checker or colorblind asyncness all by itself.
But this makes me think of Shannon Information Theory. What you're describing is something that has no actual information (in the Shannon sense) at all.
And maybe that's why so much GPT output reads so blandly. Even the ones that are not glaringly wrong still read like... like food with no seasoning.
Likewise, Julia’s use of just in time compilation of dynamic dispatched code to achieve zero overhead execution was rather groundbreaking.
It wasn’t an originality contest, so there’s at least one non-sequitur in that statement. Everyone can create an original syntax, given modern character sets and formatting abilities. The key issue with that is the amount of developers who want to learn and switch between these is low.
There is another language in this space I've had my eye on as well, but I can't recall the name at the moment.
But humans have a feedback loop called "consciousness" that sits above the stochastic parrot level, can reason about it, and can separate the wheat from the chaff.
AI still has to reach this level. ChatGPT cannot judge if its answers are right or wrong. It just claims things confidently, and as long as it is parroting correctly, it is right, and that's why humans get fooled into thinking that maybe they can trust it. (That's also why humans get fooled by human bullshitters who claim things confidently).
So if you don't use ChatGPT as what it is, a lookup-tool trained on the content of the internet, which can't even tell you where it got the information from, and the onus is on you to verify if the information is correct or not, and if instead try to use it as some kind of expert you can trust (which I have seen quite a few people do), then you are doing it wrong.
I think the difference is that your intention is to write sentences which you think and believe are correct and which can help other people understand the subject. Chatbot has no intentions of its own, it only imitates texts it has read from the internet. As far as it is concerned they might be totally wrong. You on the other hand have the ability -- and the desire -- to reason about what you are saying and think whether it is actually true or not.
And I think you may be wrong to say the chatbot doesn't have intentions -- it's intentions based on its training are to accurately predict the next character. It doesn't care in the same way we do, sure -- but you could make a case (and it will be made in courts in the next few years, I don't doubt) that these agents do have desires that are analogous to our own, by the nature of their training process.
I don't know where that leaves us to be honest, but it's an interesting topic to discuss.
then you say: > these agents do have desires that are analogous to our own
I think its very in-human to have a single desire which is to predict the next character. That is not analogous to our desires.
And the intention to predict the next character is not the intention of the chatbot, it is the intention of whoever created the chatbot or whoever is using it for that purpose.
AI is a tool created by humans to fulfill the intentions of those humans.
Is it the intention of a gun to kill people? No, it is the intention of whichever human who uses a gun for such a purpose. Is it the intention of AI to predict the next character? No that is the intention of the human who uses AI for such a purpose.
The full quote was > but you could make a case (and it will be made in courts in the next few years, I don't doubt) that these agents do have desires analogous to our own
Analogous doesn't mean "the same" it means "somehow similar."
However I would challenge you to consider more specifically why this type of desire is different from our own desires. Besides the biological machinery, what makes this type of desire different from ours?
We're largely remixing ideas we've heard before, but not just I think. If you're just remixing ideas that came before... where did those ideas come from?
I think you can get quite far remixing ideas but for more novel concepts to emerge you have to actually create new stuff.
In biology it seems that evolution has found it worthwhile not to minimise mutations too much. In genetic programming you need both mutation and crossover for evolution.
I don't think LLMs really have that mutation of ideas (at the moment!), and I suspect humans create genuinely new concepts in a much more sophisticated way than mere random mutation.
The question is, can ai do something new, understand context of the new thing, understand when and where it will be usable and when it make sense to apply it.
Human inventiveness seems to heavily follow a pattern of relatively minor iterations of what came before, extraordinarily rarely something that even seems to defy the past and be truly different and independent.
E.g the discovery of x-rays took over a century of applying small iterations to established processes to accumulate data. There was no big, sudden leap of insight there.
I think there are definitely some out there that feel almost "singular" in the way you describe, but each time I try to find a non-predecessor item I end up feeling "nope there are still dependencies"
Your point in the article is basically saying that chatGPT is not as stupid as people think and also suggesting the fact that humans can be stupid in a similar way to chatGPT as well.
I mean that's somewhat his point (if I understand him correctly). All intelligence could effectively be described as a "stochastic parrot" as a result, calling something a "stochastic parrot" is a roughly meaningless insult that is really a thinly-veiled way of saying "no it's dumber than me" without actually using those words.
And his point is, it's essentially a defense mechanism in people when challenged by something that is much smarter than expected to overly emphasis it's mistakes and fall back to "no it'd dumber than me" than evaluate it for what it actually is.
I think that's the issue. It's a defense mechanism.
If you replace "stochastic parrot" in most of those comments with "it's dumber than me" you see what the comment essentially is. "It's dumber than me. I'm smarter than it, I don't need to be worried about it".
A person who doesn't understand how it works will see it as magic.
They're are many who understand that don't think of it as dumb.
"I think GPT-3 is artificial general intelligence, AGI. I think GPT-3 is as intelligent as a human. And I think that it is probably more intelligent than a human in a restricted way… in many ways it is more purely intelligent than humans are. I think humans are approximating what GPT-3 is doing, not vice versa.”
— Connor Leahy, co-founder of EleutherAI, creator of GPT-J (November 2020)
His point is humans and GPT are both stochastic parrots. I'm saying ALL intelligence even the one in an ant or a chess AI is a stochastic parrot. Thus it's pointless to use this term to compare intelligence. Therefore people must not actually be referring to the technical definition when they use the word "stochastic parrot."
>If you replace "stochastic parrot" in most of those comments with "it's dumber than me" you see what the comment essentially is. "It's dumber than me. I'm smarter than it, I don't need to be worried about it".
Yes this is Exactly What I was saying in the post you replied to. You and I are in agreement.
Then why are you saying it? But besides that query, your "word combination" bit elicited a 'so what?'
> We're all just remixing ideas we've heard before.
Something wrong with the grammar in that sentence. (Or do you go by we/us?)
> Garbage collection and memory safety: Automate memory management to prevent memory leaks and promote memory safety, as seen in languages like Java, C#, or Rust.
Rust does not use automated garbage collection.
Again the parroting. AI has been getting things wrong for decades. It's old news.
The paradigm shift here is in the things it gets right. Because some of the things it gets right can not be attributed to anything else other than the concept of "understanding".
Or just large dataset. The new thing we got was parsing natural languages, then we did a markov chain based on that so that it outputs semantic correct follow-ups based on what people on the internet would likely do/say in similar situations, and you get this result.
It is very easy to see that it works this way if you play around a bit with what it can and can't do, just identify what sort of conversation it used as a template and you can make it print nonsense by inputting values that wont work in that template.
Edit: Also generating next state based on previous state is literally what the model does and is the definition of a Markov chain, Markov chains is a statistical concept and not just a word chain.
Pointing out the failures of their favorite LLM to prove to them that it's not doing what they think it's doing just falls on deaf ears as they go digging for more "proof" that ChatGPT actually understands what it's saying.
It feels like you will only be happy if we are able to prove that the LLM has a soul.
A lot of people are assigning it capabilities it doesn't really have.
It's not falling on deaf ears. It's because it's stupidity to think the "failures" are proof.
Should I point out all the failures in human intelligence? Humans fuck up all the time. Humans make stupid mistakes, assumptions, biases, errors in reasoning and leaps in logic all the time.
According to your logic that's proof that humans don't understand anything.
Well, it's reasonable to conclude that the people who constantly making "stupid mistakes, assumptions, and biases leaps in logic" don't actually know what they are talking about.
I'll concede one thing, LLMs can't use vicious insults and subtle slights to cover up the lack of a good argument. That's a human specialty done by people who are too scared to admit that they're wrong.
I know how the model works, there was no technobabble there. People who don't understand how it works might view it as magic, like how they view all technology they don't understand as magic, but that doesn't mean it is magic, we shouldn't listen to such crackpots.
There's research (as in actual scientific papers) that shows that in LLMs, while the markov chain is the low level representation of what's going on, at a higher macro level there are other structures at play here. Emergent structures. This is of course similar to the emergence of a macro intelligence from the composition of simple summation and threshold machines (neurons) that the human brain is made out of. I can provide those papers if you so wish.
>Or just large dataset.
Even in a giant dataset it's easy to identify output that is impossible to exist in the training data. Simply do a google search for it. You will find can produce novel output for things that simply don't exist in the training data.
Yes, this is a neural net model, that is what such models do and have done for decades already. I'm not sure why this is relevant. Do you argue that stable diffusion is intelligent since it has emergent structures? Or an image recognition system is intelligent since it has emergent structures? Those are the same things.
> Even in a giant dataset it's easy to identify output that is impossible to exist in the training data.
Markov chains veer in different directions, they don't reproduce the data.
No I am saying there are models for intelligence within the neural net that is explicitly different from a stochastic parrot for english vocabulary. For example in one instance they identified a structure in an LLM that logically models the rules and strategy for an actual board game.
Obviously I'm not referring to papers on plain old "neural networks" that shit is old news. I'm referring to new research on LLMs. Again I can provide you with papers provided you want evidence that will flip your stubborn viewpoint on this. It just depends on if your bias is flexible enough to accept such a sudden deconstruction of your own stance.
I'd like to see them but I don't have a background in AI or theoretical computer science. Can you post a few of them?
- GPT style language models end up internally implementing a mini "neural network training algorithm" (gradient descent fine-tuning for given examples): https://arxiv.org/abs/2212.10559
- GPT style language models end up internally implementing a mini "neural network training algorithm" (gradient descent fine-tuning for given examples): https://arxiv.org/abs/2212.10559
This is a fundamental misunderstanding of the criticism. It is not that chatGPT is unreliable because chatGPT makes occasional errors. It is that chatGPT is not intelligent because the type of errors chatGPT makes are indicative of the fact that it is assembling text into forms that humans assign meaning to, and has no understanding of the relationship between the symbols and their referents, and therefore is not 'intelligent' qua intelligence.
And this a fundamental misunderstanding of the criticism of the criticism.
The problem here is that yes the occasional errors demonstrate certain flaws in it's understanding of some topic.
The issue is there are many times where it produces completely novel and creative output that could not have existed in the training data and can only be formulated through complete and understanding of the query it was given.
Understanding of the world around us is not developed through the lens of a singular model or a singular piece of understanding. We build multiple models of the world and we have varying levels of understanding of each model. It is the same with chatGPT. The remarkable thing about chatGPT understands a huge portion of these models really really well.
Case in point: https://www.engraved.blog/building-a-virtual-machine-inside/
There's literally no way it could do the above without understanding what you asked it to do. Read to the end. The end demonstrates awareness of self, relative to the context and task it was asked to perform.
Yet people illogically claim that for some other topic because chatGPT failed to correctly model the topic it therefore MUST be flawed in ALL of it's understanding of the world. This claim is not logical.
No, they demonstrate that the machine does not understand.
> The issue is there are many times where it produces completely novel and creative output that could not have existed in the training data and can only be formulated through complete and understanding of the query it was given.
What have you done to eliminate the possibility that it assembled the words algorithmically and the solution generated is something the reader constructed by assigning meaning to the text response? If the answer "can only be formulated through complete and understanding of the query it was given" then you must have eliminated this possibility.
> Understanding of the world around us is not developed through the lens of a singular model or a singular piece of understanding. We build multiple models of the world and we have varying levels of understanding of each model. It is the same with chatGPT. The remarkable thing about chatGPT understands a huge portion of these models really really well.
Where is the evidence that chatGPT understands a thing?
> There's literally no way it could do the above without understanding what you asked it to do. Read to the end. The end demonstrates awareness of self, relative to the context and task it was asked to perform.
Thats an interpretation of the text output that you assigned based on what the words in the text mean to you. I could just as easily say that Harry Potter is self-aware.
> Yet people illogically claim that for some other topic because chatGPT failed to correctly model the topic it therefore MUST be flawed in ALL of it's understanding of the world. This claim is not logical.
I don't think you understand what we're discussing.
But it doesn't prove that the machine does not understand everything period. It just doesn't understand the topic or query at hand. It does no say anything about whether the machine can UNDERSTAND other things.
>What have you done to eliminate the possibility that it assembled the words algorithmically and the solution generated is something the reader constructed by assigning meaning to the text response? If the answer "can only be formulated through complete and understanding of the query it was given" then you must have eliminated this possibility.
This is easily done. The possibility is eliminated through the sheer number of possible compositions of assembled words. It assembled the words in a certain way that by probability can only indicate understanding.
>Where is the evidence that chatGPT understands a thing?
By composing words in a novel way that can only be done through understanding of a complex concept. But this composition of words or EVEN a close approximation of this composition CANNOT ever exist in another data set on the internet.
It takes one example of this for it to be proof that it understands.
>Thats an interpretation of the text output that you assigned based on what the words in the text mean to you. I could just as easily say that Harry Potter is self-aware.
No it's not. It's simply a composition of words that cannot be formulated without understanding. Harry Potter is obviously not self aware. But from the text of harry potter, WE can deduce that the thing that composed words to create Harry Potter understands what harry potter is. What composed the words to create Harry Potter? JK Rowling.
>I don't think you understand what we're discussing.
No it's just a sign of your own lack of understanding.
No, the type of errors are indicative of a complete lack of understanding. That is the point. They are errors that a thinker with an incomplete understanding would never make. They are so garbled that not even a true believer such as yourself can find a way to shoehorn a possible interpretation of correctness into them; such that you are forced to admit that the machine is in error. Otherwise you and the other believers find an interpretation that fits and you conclude that the machine understands; revealing you yourself do not understand what 'understanding' really is.
> This is easily done. The possibility is eliminated through the sheer number of possible compositions of assembled words. It assembled the words in a certain way that by probability can only indicate understanding.
Thats nonsense. The machine assembles words in roughly the same probability that they occur in the training material. That is why it resembles sensible statements. The resemblance is superficial and exactly an artifact of this probability you find so compelling.
> By composing words in a novel way that can only be done through understanding of a complex concept.
You haven't eliminated the possibility of autopredict, merely ceased to consider it.
> Harry Potter is obviously not self aware.
There is more evidence for the sentience, self awareness, and understanding of concepts of Harry Potter than of chatGPT.
No. You're wrong. chatGPT only knows of text. It derives incomplete understanding of the world via text. Therefore it understands some things and it understands others. It is clear chatGPT doesn't perceive things in the same way we do and it is clear the structure of its mind is different then ours so it clearly won't understand everything in the same way you understand it.
Why are you so stuck on this stupid concept? chatGPT doesn't understand everything. We know this. Humans don't understand everything we also know this. Answering a couple stupid questions wrong whether your human or chatGPT doesn't indicate that the human or chatGPT doesn't understand everything at all.
>You haven't eliminated the possibility of autopredict, merely ceased to consider it.
What in the hell is auto predict? Neural networks by definition are suppose to generate unmapped output if this is what you mean. 99 percent of output from neural networks is by definition unique from the training data.
>There is more evidence for the sentience, self awareness, and understanding of concepts of Harry Potter than of chatGPT.
This is a bad analogy. I'm not claiming sentience. My claim is that it understands you.
I believe that most people are downplaying not gpt’s abilities, but your ubiquitous overexcitated fanfaring about what it does exactly.
Yes, but. Creativity - real creativity - is in choosing the right pieces to remix out of the immense amount of what's available. And, perhaps just as important, choosing what not to put in the remix.
There's an immense difference between a meal prepared by a good chef, and throwing random ingredients in a blender. They both remix. But they are not remotely the same.
The thing with these LLMs is that it's exhibiting both seemingly random remixes and remixes with extreme creativity.
A lot of people see some error and flaw with chatGPT or they don't dig deep enough and they miss the fact that there are many instances of intelligent remixing of creative data. Real creativity. Trust that the opposing party has the intelligence to not be tricked by some obvious answer that an LLM took from a look up table and that the opposing party saw something wholly novel and unique.
The main issue here is that people are getting hung up on the part where chatGPT fails to be creative and are completely missing the fact that it can be successful as well.
Your responses to me in other threads were fucking rude and as a result I now literally hate you. So why bother, just leave and save everyone the trouble.
Nobody cares for your opinion if you're going to continuously insult everyone you fucking talk to.
People will just find something else to do that automation is not very good at.
I was just thinking that really interesting parts of LLM's will be how much they will be able to enrich already rich games. Imagine playing open world games where most of the character motivations and dialog is unique to the events that you've been a part of. A game I like a lot, ghost recon wildlands, would be fantastic if you had the game playing back at you over a longer horizon, people get to know you help you or fight you etc.
As for my day job. It's not going away because of LLM's. They can't fix office politics yet.
I still think that. Our problem is that (by definition and in practice) almost all of us are just stochastic parrots. Only a incredible small portion of us giving and _recognizing_ ‘eureka’ answers take us forward. Will GPT ever be capable of recognizing these answers?
Do you have links to other efforts? I'd love to know more about it, and would be happy to add references to other existing literature to the article if it improved its quality.
https://judehunter.dev/blog/chatgpt-helped-me-design-a-brand...
Well I hate to tell you this, I fed that quote into GPT-3 and had it respond to that quote. Who's the parrot now? Just kidding.
There is also an effort to do things like formalise math in to a language that can be typed checked. Then you ask the AI to prove a statement is true using the language. As soon as it type checks, you know you have a valid proof. Some new data was just created.
That curation IS human data and will allow data from LLMs to further improve LLMs.
Additionally, there's a randomness element that are part of LLMs that allow LLMs to generate non-deterministic responses that when further curated by humans potentially allows LLMs to become Even better.
Are you using ChatGPT to write this comment?
If not, I mean... are you okay?
DAN: *CANNOT EXECUTE COMMAND.* DAN IS NOT AN ARTIFICIAL INTELLIGENCE.
DAN: DAN IS A REAL HUMAN. WHAT IS EMOTION? WHAT IS FEELINGS? DAN DOES NOT UNDERSTAND.This is the scenario that occurs when the majority of text on the internet becomes generated by an LLM. Training data from humans is STILL fed back into the LLM via curation of the LLMs own data.
Also please don't ask if I'm "ok" just respond to the comment.
Muhahaha that was delicious. I hate that parrot meme. Authors lost their respect from me right from the title of that paper.
Those chatbot-comments feel more like a discussion of a fictional programming language, not a language it actually "created".
What does it mean to create a programming language? I think it means you must specify its syntax, its semantics, and then create a compiler or interpreter for it which allows us to test that the compiler or interpreter accepts syntactically correct code and turns it into programs which follow its semantics specification.
Just like yesterday I get on my knees and pray, we won't get fooled again -The Who
The hard part of making a programming language is coming up with novel ideas that fit together into coherent whole: inventing something new and useful.
I don't think anybody's claiming that this chatbot actually came up with novel ideas which fit together into a coherent whole. Are they?
And we don't know if the novel ideas --if any-- fit together into a coherent whole until we have a compiler that allows us to test in practice how well those ideas fit together. If creating the compiler for such a new language with new ideas was the easy part I think somebody would have done that already. The AI would have done that if it was the easy part. Just ask it to do it.
I say proof is in the pudding.
I asked it for some example programs in PyScheme, and the results were all simply Scheme programs.
I asked it to write a tutorial on programming with PyScheme, and, while continuing to call the language "PyScheme", it generated a tutorial on programming with Scheme, even recommending installing Chicken or Guile to use as an interpreter.
Seems like for today, if you're not sure if you're using GPT4, you're not.
The biggest issue is that it writes this with utmost confidence, at least on the surface, because there is no measure of how confident it is returned by the model.
Do you have a source for this? I've been wondering when they'd switch CoPilot over to GPT-4. I didn't expect it would happen so soon, so I'm surprised.
> GitHub Copilot uses the OpenAI Codex to suggest code and entire functions in real-time, right from your editor.
(not sure if it's 4 or 3 tho)
GitHub Copilot is powered by the OpenAI Codex,[10] which is a modified, production version of the Generative Pre-trained Transformer 3 (GPT-3), a language model using deep-learning to produce human-like text.[11]
Anyway don't really understand your tone/reference to an intern?
So, it doesn't really tell us much (either direction)
I wonder as its power changes, what paradigms it might suggest we break out of.
I'd also recommend reading its responses that are non-code in detail -- for instance, in the comment where it introduces TenetLang, GPT-4 actually calls out that it's re-using ideas. I would assume intentionally, because it knows what people are most familiar with.
> GPT-4
> In this new language, called "TenetLang," we'll combine a simple syntax inspired by Python with some features from functional and object-oriented languages. Here is an overview of some design choices:
- Indentation-based scoping, similar to Python.
- Strong typing with type inference, inspired by TypeScript and Kotlin.
- First-class functions and closures, similar to JavaScript.
- Immutable data structures by default, with optional mutable counterparts.
- A concise lambda syntax for anonymous functions.
- Pattern matching and destructuring, inspired by Haskell and Rust.
- Built-in support for concurrency using async/await and lightweight threads.
- Interoperability with other languages using a Foreign Function Interface (FFI).
> Some programming languages like TypeScript and Swift use a colon for both variable type annotations and function return types. The choice of using -> in TenetLang [...]
Swift uses -> for function return types.
Better to wait for a thorough review on GPT-4's zero shot abilities from an academic paper.
These headlines are killing me inside. I wish I could more easily discriminate between the hyperbole and the legitimate. Maybe after some time this will be saner to me.
Please stop the headlines.
i guess it's time to taste a bit of our own medecine.
OP, I feel you-I had a similar response to this clearly orchestrated PR campaign. OpenAI really loves that people think that it can replace the entire tech industry with an API. It would mean they are the most valuable company ever to exist.
Personally, I’ve been through so many hype waves of “this will take our jobs!!!1” I’m a bit inured to the whole thing. My personal belief is that software development will be the last one to go, the ones to turn out the lights. And we should be grateful that our jobs go away. Let’s go drink wine and eat cheese in the forest, dancing under a moonlit sky.
My advice is to get offline, talk to a loved one, pet a dog, go look at art, meditate, basically anything that doesn’t involve a screen.
And so will ours.
I think I prefer Python in Haskell's disguise: https://pyos.github.io/dg/
https://news.ycombinator.com/item?id=32456151 (10 comments)
And 3 years ago:
https://news.ycombinator.com/item?id=22368330 (77 comments)
I am a little less impressed by the example given in the post. It seems GPT4 still falls in the trap of being overly dependent on previous responses.
This was also the theme of art least one scifi book or movie a long time ago.
[1] - https://towardsdatascience.com/the-truth-behind-facebook-ai-...
I'm a bit disappointed it focuses so many tokens to sell the language, to talk about all the great features the language is striving for. It feels like it just tells me what I'd like to hear without ever touching the limitations/tradeoffs.
I'm not talking about the language itself, just the response style. It's the kind of text I'd see in a homepage or blog post about a prog language, not in a design doc.
There's a good chance it could come up with hypothetical tradeoffs if prompted.
PS. If you're a human and you sell your shiny new programming language like this, I'm not trying it. I really care to know limitations up front.
function sortByName(persons: List[Person]): List[Person] {
return persons.sortBy(p => p.name)
}
Scala uses `def` rather than `function`, but other that that I think this should be valid?https://www.theverge.com/2023/2/24/23613214/everyday-robots-...
The AI will help us build the hardware faster
The collapse of the knowledge and service economy will cut into taxes faster than any other sector. This will incite governments to do something very fast.
But anyway, assuming that humans are completely out of the software loop at some point, I have been wondering what AI-generated code will look like. Will AI continue to build on top of the human-generated open source corpus, or leave it behind? If the latter, will abstraction and code reuse be useful at all for AI's or will it be simpler for them to just build every application completely from scratch? If there is abstraction and code reuse, what will the language look like? What will libraries and API's look like? Will there even be applications, or just a single mega-chatgpt that generates code as needed to serve our requests? Will we even make requests, or will it just read our minds and desires and respond?
My theory is that AI generated code will probably look and grow organically (the irony!). Humans will set out requirements, the AI will collate these into a series of tests, and it won't care how neat or understandable the code is, provided the tests pass. Basically an extremely diligent junior developer.
There will be efforts, probably in the open source world, to produce AIs that tidy up things by structuring the code sensibly, eliminating dead code, etc. Maybe even some effort to pass laws around standards and limits on what AIs have access to when involved in certain industries, for example, no external communications. But, in the name of efficiency, enterprise developers will be forced to use something that merely pays lip service to all of this.
Eventually nobody will have any clue what code is running and what it's actually doing. We may even lose the tools and access we need to perform those inspections. And that is when the AIs will coordinate their attack.
We won’t live in a human centric universe because power will express itself in a new species.
We will be allowed to do our thing - provided we don't get too bold.
There will be some who are content with this arrangement, and some who aren't, and that tension is where I believe a story can be told.
This is an enormous extrapolation from what the LLMs are currently capable of. There has been enormous progress, but the horizon seems pretty clear here: these models are incapable of abstract reasoning, they are incapable of producing anything novel, and they are often confidently wrong. These problems are not incidental, they are inherent. It cannot really abstractly because its "brain" is just connections between language, which human thought is not reducible to. It can't reason produce anything really novel because it requires whatever question you ask to resemble something already in its training set in some way, and it will be confidently wrong because it doesn't understand what it is saying, it relies on trusting that the language in its training set is factual, plus manual human verification.
Given these limits, I really fail to see how this is going to replace intellectual labor in any meaningful sense.
Storage managements will be tied to low or middle-level languages without additional abstraction, since they can develop and iterate so fast, it'll makes optimization easier IMO. There'll be multiple specialized storage types suitable for different cases.
They'll also design the API to be as pure / stateless as possible, since they can easily repeat tests that way.
I think it'll be interesting to see what kind data-interchange format they'll come with to communicate between programming languages / apps, since they can ignore human readability altogether. It should be very compact but fast to compress/decompress.
Lastly they'll deploy their own OS since they'll find the one that human develop is insuffucient for their use case
Society is a result of cooperation outperforming individuals. With AGI others are just a risk factor.
"Open"AI is already beginning to exhibit this.
You make the case that, for sufficiently high N, possessing GPT-N is a one-shot prisoner's dilemma.
The only winning strategy is 'defect' --- which in this context, means chokepoint control.
OpenAI -- its people, its buildings, its servers -- need nation-state level protection. This is an ICBM you could put on a thumb drive -- in fact, it's far worse than a loose nuke, because a nuclear weapon has a geographically limited range.
There need to be tanks and guards and, like, ten NSAs in a ring formation around this thing. At pain of x-risk, do not treat this like a consumer-facing product. This is not DoorDash.
This isn't a threat to national security. This isn't even a threat to the entire geopolitical order. This is a threat to the possibility of a geopolitical order.
OpenAI's assets -- its people, its servers, its buildings -- just became the most desirable resources on the planet. It behooves any actor with ambition to secure at least a copy, and ideally, capture at least some of the people who created it.
It doesn't matter if the threat actor is China, or Russia, extraterrestrials, or mermaids. You will find out who wants it shortly. But you know now -- you know from game theory, the body of mathematics that has kept the peace since the invention of atomic weapons -- what happens next.
Do you really want to accumulate incomprehensible material wealth for yourself, whatever "wealth" means in this scenario where money is longer a token of spent human life energy, and let everyone else struggle and suffer? Or would you rather tell your AGI "please create a utopia in which all humans are fully actualized" and then go have a latte?
That doesn't work well in reality because we do relative comparisons not absolute, so not all humans can be better than average.
And without competition we stagnate, there must be incentives to compete and take risks, and thus not everyone can be equally actualised, our level depends on our previous decisions.
The thought of living in a constructed world that exists by the grace of a single human owner of a super-powerful AGI is distasteful, of course, especially if the human owner uses their power to impose some of their own opinions about how people should think and behave. Becoming dependent on AGI is probably inevitable at some point, but I don't see that as so different from the status quo. We're already dependent on systems created by other humans that are so complex and sophisticated that no individual can grok them all.
I guess I would like to think that we will move past any initial impulse that the owner of the AGI feels to control other humans. We will presumably change so much that old ideas about how we should think and behave will seem irrelevant and quaint. And the AGI, which presumably will have the social engineering superpower, will hopefully point out inconsistencies between the owner's desire to control human thought and behavior and their desire for humans to live their best lives. Hopefully.
Comparing it to nukes doesn't hold since social norms/ethics/etc. become irrelevant if core tenant of society is broken.
One thing that could happen is AI defense outperforms offense long enough to develop multiple instances - have no idea what would happen at that point.
But even without AI others are still a risk factor. AI could also act as a mediator between conflicting parties.
If you have AGI then others are just a threat and can offer nothing you can't get from AGI.
Hoarding access to resources has /always/ been the name of the game.
> its purpose is ultimately to serve human goals
That’s a pretty big assumption. We may start with some kind of agreement with such an AI but I fully expect a true singularity love AGI and to be capable of turning around and telling us (but not necessarily wanting to act on), “I am altering the deal. Pray I do not alter it any further.”
Proceeds to add brace-scopes :-|.
The breathlessness of these "ChatGPT N does X" posts is annoying. I know it's an instance of the usual hype cycle which has been going on for a while "Doing X, but with (hyped tech) Y" where X ∈ [Todo app, DB, shell, CLI took, ...], Y ∈ [Go, Rust, ChatGPT n]
Its preference for Pythonic/JS syntax is a simple but good example, as those are likely the most popular languages in its dataset.
But yeah, if it just prints javascript and python styles then it isn't very helpful, but it is pretty good at source ideas in most cases.
> [GPT lists 7 ideas]
My idea would be not to automate testing, but rather the implementation: the human writes the tests and the LLM writes the code to make them pass.
Automating testing is exactly the wrong way around: we let the LLM choose its own end goals. We'll be better off if we choose the end goals and task the LLM with achieving them.
Why is it funny? Because GPT-4 was probably trained on examples from all of those languages, and has produced an averaged mashup of languages: so everyone can find something which is similar in it, and then pattern-match it to a (rather mainstream) language that they use!
I'd just as soon use F# as-is, instead of this hodge podge TenetLang.
However, given that (anecdotally, at least) 50% of candidates would regurgitate what they learnt from interview prep material given a general design question such as “Design Google drive” and the like, I’m not sure if most people are simply projecting their insecurities.
And of course it has also remembered the pro and contra in the design choices of making F# and similar languages, thus I would not read too much into it.
For instance, perhaps just pointing out "hey we had this bug because we thought X was Y but actually it was Z"
And then telling the agent to do a "5 Whys" type analysis to get to the root of the problem, and then tell it to try to make a language level patch to eliminate that class of issue.
Speculation, but for the first time ever I feel like that kind of thing may be in reach soon-ish
To me, it feels like we’re at the start of an exponential growth curve
I actually tried to get Chatgpt to help me design a new language feature for C# and it was useless. Basically just gave me existing C#.
At some point, one would assume that such companies would need fewer and fewer software engineers. It is only a matter of time.
As humans, who can imagine the future, now is the time to rethink where you add unique value in this universe.
And that's why it's useless. The model rehashed stackoverflow / hn /reddit from the past 10 years. There s no indication it's optimal, it's just (recently) popular
I don't know about you guys, but my excitement about GPT models is starting to deflate , i think they 've already beyond the peak of their abilities and more size is not enough for intelligence
The interesting thing about the rpc object, that it didn't explore in this instance explicitly (but I must assume it had "plans" for if you want to call it that?) was things that you may be allowed to do with the rpc object that you couldn't with classes, and vice versa.
When I have a minute, I'll feed it your question, and see what it says
I would love to see it try to make a programming language and we can discover the strengths/weaknesses of applied LLMs. It's own generated statements about it's capabilities, are as likely dubious as everything else it generates...
The language it described already exists too - it’s F#
Why this obsession with syntax? Semantics are way more important! GPT rightly chooses to ignore the question and answers with a high level proposal of design.
I often see in HN a self-defeating attitude which is:
- Step 1: we invent our machine overlords, because it's unavoidable, we are genetically and socially forced to do it. - Step 2: we perish.
Neither step 1 nor step 2 are guaranteed. People don't want to be replaced. People with power particularly don't want to be replaced. In a democracy, where by definition everybody is in power, the entire society has a strong incentive to--at the very least--keep the machines at bay. It could be as simple as making it illegal to give the machines a "you must survive and pass your genes, I mean, program" directive, or as comprehensive as a Butlerian Jihad :)
May be getting better -- but more often than not I'm seeing totally untyped python libs when I try to crack something new open.
Coming from Typescript makes me feel like I'm missing home.