Google testing ChatGPT-like chatbot 'Apprentice Bard' with employees
cnbc.com
cnbc.com
It is one thing to train an AI on megatons of data, for questions which have solutions. The day ChatGPT can build a highly scalable system from scratch, or an ultra-low latency trading system that beats the competition, or find bugs in the Linux kernel and solve them; then I will worry.
Till then, these headlines are advertising for Open AI, for people who don't understand software or systems, or are trash engineers. The rest of us aren't going to care that much.
Since all a codebase like that is is a kind of directed graph, then augmentations to the processing of the network to allow for the simultaneous parsing of and generation of this kind of code may not be as far off as you thinking.
I say this as an ML researcher of coming up and around the bend towards 6 years of experience in the heavily technical side of the field. Strong negative skepticism is an easy way to bring confidence and the appearance of knowledge, but it also can have the downfall of what has happened in certain past technological revolutions -- and the threat is very much real here (in contrast to the group that believes you can get AGI from simply scaling LLMs, I think that is very silly indeed).
Thank you for your comment, I really appreciate it and the discussion it generated and appreciate you posting it. Replying to it was fun, thank you.
While I wouldn't dare offer particular public prognostications about the effect transformer codegens will have on the industry, especially once filtered through a profit motive - the specific technical skill a programmer is called upon to learn at various points in their career has shifted wildly throughout the industry's history, yet the actual job has at best inflected a few times and never changed very dramatically since probably the 60s.
What I've seen again and again from people in the field is a gross underestimation of the long tail on these problems. They see the rapid results on the easier end and think it will translate to continued process, but the reality is that every order of magnitude improvement takes the same amount of effort or more.
On top of that there is a massive amount of subsidies that go into training these models. Companies are throwing millions of dollars into training individual models. The cost here seems to be going up, not down, as these improvements are made.
I also think, to be honest, that machine learning researchers tend to simplify problems more than is reasonable. This conversation started with "highly scalable system from scratch, or an ultra-low latency trading system that beats the competition" and turned into "the parsing of and generation of this kind of code"- which is in many ways a much simpler problem than what op proposed. I've seen this in radiology, robotics, and self driving as well.
Kind of a tangent, but one of the things I do love about the ML industry is the companies who recognize what I mentioned above and work around it. The companies that are going to do the best, in my extremely bias opinion, are the ones that use AI to augment experts rather than try to replace them. A lot of the coding AI companies are doing this, there are AI driving companies that focus on safety features rather than driver replacement, and a company I used to work for (Rad AI) took that philosophy to Radiology. Keeping experts in the loop means that the long tail isn't as important and you can stop before perfection, while replacing experts altogether is going to have a much higher bar and cost.
Try at least 12 [0]
(I would say 15 but my 45-second search didn't yield anything that far back)
[0] https://spectrum.ieee.org/how-google-self-driving-car-works
This is a bit like seeing Steve Mann's wearable computers over the years ( https://cdn.betakit.com/wp-content/uploads/2013/08/Wearcompe... ) and then today anyone with a smartphone and smart watch has more computing power and more features than most of his gear ever had, apart from the head mounted screen. More processing power, more memory, more storage, more face recognition, more motion sensing, more GPS, longer runtime on battery, more bandwidth and connectivity to e.g. mapping, more assistants like Google Now and Siri.
And we still aren't at a level where you can be doing a physical task like replacing a laptop screen and have your device record what you're doing, with voice prompts for when you complete different stages, have it add markers to the recording, track objects in the scene like and solve for questions like 'where did that longer screw go?' or 'where did this part come from?' and have it jump to the video where you took that part out. Nor reflow the video backwards as an aide memoire to reassembling it. Or do that outside for something like garage or car work, or have it control and direct lighting on some kind of robot arm to help you see, or have it listen to the sound of your bike gears rattle as you tune them and tell you or show you on a graph when it identifies the least rattle.
Anything a human assistant could easily do, we're still at the level of 'set a reminder' or 'add to calendar' rather than 'help me through this unfamiliar task'.
RE: changing you laptop screen. My buddy wants an 'AR for Electronics' that can zoom in on components like a magnifying glass (he wants head mounted), identify components by marking/color/etc and call up schematics on demand. So far, nothing seems to be able to do that basic level of work.
Fanuc, the robot manufacturer, famously does run a lights-out factory, and has since 2001. It was the dream of Fanuc's founder. Baosteel now has a lights-out steel coiling facility. Both of these are more PR than cost effective.
There are many factories where there are very, very few people for large rooms full of machines, though.
This comment from 7 days ago covers it adequately:
AI companies such as Vicarious have been promising AI that makes this easier. Their idea was that generic robots with the right grips and sensors can be configured to work on a variety of assembly lines. This way a factory can be retooled between jobs quicker and with less cost.
It's a cautionary tale to people who are working in ML to be not too optimistic on "the future", but in my opinion being cautiously optimistic(not on AGI though) isn't harmful by itself, and I stand by that. Well at least until we hit the next wall and plunge everyone into another AI winter(fourth? fifth?) again.
As a plus, we do actually see some good progress that benefited the world like in biotech. Even though we are still mostly throwing random stuffs at ML to see if it works. Time will tell I guess.
I think even three years ago, most people would have thought the reverse.
So Kurzweil was imagining the turing test as the capstone to a decade of more and more capable ai products, not as "kind of early interesting success that may (or may not) presage really useful AI."
("The Turing test" is a pretty hazy target. I have no doubt that a chatgpt that was not trained to loudly announce that it was an AI could convince lots of people that it's a real human, right now. I think it's also the case that people with some experience with it could pretty quickly find ways to tell what it is.)
Otherwise you risk claiming ELIZA passed it, because a couple people thought so. Or that one Google employee this time.
Not because goal posts keep moving, but because we can only do 80% of the remaining distance each time, and the remaining 20% is still obvious.
If I were to make a point as to why your notes on self-driving cars and in-warehouse robots may not transfer to the case of software development, it's that they are fundamentally two very different problems with very different issues attached to them. It unfortunately is very much apples to oranges. They are both NP-hard but very different kinds of NP-hard.
A software program is a closed-loop target, though it is NP-hard. But we're optimizing for a different kind of metric here that is well-defined. Any kind of self-directed reinforcement-or-otherwise autoregressive-in-the-world algorithm is going to have an extraordinarily long tail of edge cases.
What I was talking about when I mentioned the geometry of the problem is not the parsing of the code, but the geometry of a near-optimal solution. Certainly, scale will be expensive, but Sutton is our friend here. That's why it's more "trivial" than problems that require humans in the loop -- you don't need humans to parse, structure, generate, and evaluate the data flow of a software code base, though admittedly if models like RHLF become popular as you noted, the endpoints that generate code under those geometric constraints -- those will become extremely expensive.
I think the geometric problem is very hard but the hurdle of scaled language models is more technically impressive to me.
What's nice is that unlike needing to generate a long, 1d story, too, there's more robustness with a huge field of possibility that's had years of work on the software side of things. It's not that it's going to be easy, but I think we've all grown as we've seen how hard self-driving cars are, and it's just not that kind of scenario, since all consequences of the 'world' within the repo-generation case are (for the most part) self-contained.
I hope that helps elucidate the problems a bit. To me, my optimism is much more rare, and only generally when I feel like I have a solid grasp of the fundamentals of it enough (i.e. I roughly know deliverability and have decent known error bounds on the sub-problems).
That said, I heartily agree with you that when all else fails -- assistive is good. What I see a "complete solution" doing well is creating a Kolmogorov-minimal, complete starting point and things evolving from there. Whether that works or not remains to be seen.
Those are the most precarious jobs in the industry. Many of those people might become LLM whisperers, taking their clients requests and curating prompts. Essentially becoming programmers over the prompting system. Maybe they’ll write a transpiler to generate prompts? This would be par of the course with other languages (like SQL) that were originally meant to empower end-users.
The problem with current AI generated code from neural networks is the lack of an explanation. Especially when we’re dealing with anything safety critical or with high impact (like a stock exchange), we’re going to need an explanation of how the AI got to its solution. (I think we’d need the same for medical diagnosis or any high-risk activity). That’s the part where I think we’re going to need breakthroughs in other areas.
Imagine getting 30,000-ish RISCV instructions out of an AI for a braking system. Then there’s a series of excess crashes when those cars fail to brake. (Not that human written software doesn’t have bugs, but we do a lot to prevent that.). We’ll need to look at the model the AI built to understand where there’s a bug. For safety related things we usually have a lot of design, requirement, and test artifacts to look at. If the answer is ‘dunno - neural networks, ya’ll’, we’re going to open up serious cans of worms. I don’t think an AI that self evaluates its own code is even on the visible horizon.
I gave some code to ChatGPT asking to simplify it and it returned the correct code but off by one. It was something dealing with dates, so it was trivial to write a loop checking for each day if the new code matched in functionality the old one.
You will never have certainty the code makes any sense if it's coming from one of these high tech parrots. With a human you can at least be sure the intention was there.
ChatGPT is only cheap if you don’t need its code to do anything of any particular value. It’s a seemingly ideal solution to collage homework for example. But professionally people write code to actually achieve something, this is why programmers actually get paid well in the first place. The point isn’t LOC the point is solving some problem.
I was able to give ChatGPT the requirements for all of them. The types of bugs I found during the first pass:
- the AWS SDK and the underlying API only returns 50 results in one call most of the time. From the SDK you have to use the built in “paginators”. ChatGPT didn’t use them the first time. But once I said “this will only return the first 50 results”. It immediately corrected the script and used the paginator. I have also had to look out for similar bugs from junior devs.
- The usual yaml library for Python doesn’t play nicely with CloudFormation templates because of the function syntax that starts with an “!”. I didn’t know this beforehand. But once I told ChatGPT the error, it replaced the yaml handling with cfn-flip.
- I couldn’t figure out for the life of me how to combine the !If function in CloudFormation with a Condition, and a Yaml block that contain another !Select function with two arguments. I put the template block without the conditional and told ChatGPT “make the VPC configuration optional based on a parameter”. It created the Parameter section, the condition and the appropriate Yaml.
I’ve given similar problems to interns/junior devs before and ChatGPT was much better at it.
Which ok I get why you think ChatGPT is more useful.
Now if ChatGPT would be able to actually work on the problem rather than returning generated text, that would be a completely different beast. And I think that this workflow will come in the near future because it's pretty obvious idea. Get task specification, generate tests, generate code, fix code until tests work, refactor code until it meets some standards, etc.
You give junior devs way too much credit. They rarely test corner cases.
ChatGPT probably works great if you use it to speedrun normal best practices in software engineering. Make it start by writing tests given a spec, then make it write code that will pass the specific tests it just wrote. I’m guessing it’ll avoid a lot of mistakes, much like any engineer, if you force it to do TDD.
“Given an XML file with the format {[1]} and a DynamoDB table with two fields “Key”, “Value”, write a Python script that replaces the Value in the xml file when the corresponding key is found. Use argparse to let me specify both the input xml file and the output XML”
It spit out perfect Python code. I hadn’t used XML in well over a decade and I definitely didn’t know how to read xml in Python. I didn’t want to bother about learning.
I actually pasted an XML sample like the link below.
[1] https://learn.microsoft.com/en-us/troubleshoot/developer/vis...
I think you have that exactly backwards.
Most junior developers most places may not have the experience of a senior developer, and thus be able to do the translation from business logic to code quite as fast and accurately the first time, but this kind of derogatory attitude toward them is incredibly condescending and insulting.
ChatGPT doesn't know what it's doing. It doesn't know anything, and unlike the most junior developer barely trained, it can't even check its output to see if it matches the desired output.
And for goodness' sake, get rid of the absurd idea that all the competent developers are in Silicon Valley. That's even more insulting to the vast majority of developers in the entire world.
Now you could say that LLMs enable Google to do what it does now with fewer employees, but the same thing is true for every other competitor to Google. So the question is how will Google try and maintain dominance over it's competitors now? Likely they will invest more heavily in AI and probably make some riskier decisions but I don't see them suddenly trying to cheap out on talent.
I also think that it's not a zero sum game. The way that technology development has typically gone is the more you can deliver, the more people want. We've made vast improvements in efficiency and it's entirely possible that what an entire team's worth of people was doing in 2005 could be managed by a single person today. But technology has expanded so much since then that you need more and more people just to keep up pace.
Define "we". There are all kinds of people with all kinds of opinions. I didn't notice any consensus on the questions of AI. There are people with all kinds of educations and backgrounds on the opposite sides and in-between.
For example...
Hows that self-driving going? Got all those edge-cases ironed out yet?
Oh, by next year? Wierd, that sounds very familiar...
Remember about Tesla's autopilot was released 9 years ago, and the media began similar speculation about how all of the truckers were going to get automated out of a job by AI? And then further speculation about how Taxi drivers were all going to be obsolete?
Those workers are the ones shifting the goal posts though as a "self-defense mechanism", sure, sure... lol.
With self-driving, we barely ever saw anything obviously resembling human abilities, but there was a lot of marketing promising more.
With language models when GPT-2 came out everyone was still saying it is a "stochastic parrot" and even GPT-3 was one. But now there's ChatGPT, and every single teenager is aware that that tool is capable of replacing them with their school assignments. And as a dev I am aware that it can write code. And yet not many people expected any of this to happen this year, neither were those capabilities promised at any point in the past.
So if anything, self-driving was always overhyped, while the LLMs are quite underhyped.
One difference, though, is that it‘s economically not much use to have self-driving if the backup driver has to be in the car or present. While partially automating programming would make it possible to use far less programmers for the same amount of work.
Which, of course, is what we've always done; modern programming, with its full-featured IDEs, high level languages, and feature-rich third-party libraries is mostly about gluing together things that already exist. We've already abstracted away 99% of programming over the last 40 years or so, allowing a single programmer today to build something in a weekend that would have taken a building full of programmers years to build in the 1980s. The difference is, of course, this is going to happen fairly quickly and bring about an upheaval in the software industry to the detriment of a lot of people.
And of course, this doesn't include the possibility of AGI; I think we're a very long way from that, but once it happens, any job doing anything with information is instantly obsolete forever.
10 years ago we were still working on MNIST prediction accuracy. 10 years forward from here all bets are off. If the model has super human quantitative reasoning and a mastery of language I am not sure how much programming we will be doing compared to moving to a higher level of abstraction.
On the other hand, I think there will be so many new software jobs because of the volume of software built over the next 20 years. The volume of software built over the next 20 years is probably unimaginable sitting where we are.
In this case, it could be that you are just talking to different people and focusing on their answers. I am more than happy to believe that Copilot and ChatGPT, today, cause a bunch of people fear. Does it cause me fear? No.
And if you had asked me five years ago "if I built a program that was able to generate simple websites, or reconfigure code people have written to solve problems similar to ones solved before, would that cause you to worry?" I also would have said "No", and I would have looked at you as crazy if you thought it would.
Why? Because I agree with the person you are replying to (though I would have used a slightly-less insulting term than "trash engineers", even if mentally it was just as mean): the world already has too many "amateur developers" and frankly most of them should never have learned to program in the first place. We seriously have people taking month or even week long coding bootcamps and then thinking they have a chance to be a "rock star coder".
Honestly, I will claim the only reason they have a job in the first place is because a bunch of cogs--many of whom seem to work at Google--massively crank the complexity of simple problems and then encourage us all to type ridiculous amounts of boilerplate code to get simple tasks done. It should be way easier to develop these trivial things but every time someone on this site whines about "abstraction" another thousand amateurs get to have a job maintaining boilerplate.
If anything, I think my particular job--which is a combination of achieving low-level stunts no one has done before, dreaming up new abstractions no one has considered before, and finding mistakes in code other people have written--is going to just be in even more demand from the current generation of these tools, as I think this stuff is mostly going to encourage more people to remain amateurs for longer and, as far as anyone has so far shown, the generators are more than happy to generate slightly buggy code as that's what they were trained on, and they have no "taste".
Can you fix this? Maybe. But are you there? No. The reality is that these systems always seem to be missing something critical and, to me, obvious: some kind of "cognitive architecture" that allows them to think and dream possibilities, as well as a fitness function that cares about doing something interesting and new instead of being "a conformist": DALL-E is sometimes depicted as a robot in a smock dressed up to be the new Pablo Picasso, but, in reality, these AIs should be wearing business suits as they are closer to Charles Schmendeman.
But, here is the fun thing: if you do come for my job even in the near future, will I move the goal post? I'd think not, as I would have finally been affected. But... will you hear a bunch of people saying "I won't be worried until X"? YES, because there are surely people who do things that are more complicated than what I do (or which are at least different and more inherently valuable and difficult for a machine to do in some way). That doesn't mean the goalpost moved... that means you talked to a different person who did a different thing, and you probably ignored them before as they looked like a crank vs. the people who were willing to be worried about something easier.
And yet, I'm going to go further: if the things I tell you today--the things I say are required to make me worry--happen and yet somehow I was wrong and it is the future and you technically do those things and somehow I'm still not worried, then, sure: I guess you can continue to complain about the goalposts being moved... but is it really my fault? Ergo: was it me who had the job of placing the goalposts in the first place?
The reality is that humans aren't always good at telling you what you are missing or what they need; and I appreciate that it must feel frustrating providing a thing which technically implements what they said they wanted and it not having the impact you expected--there are definitely people who thought that, with the tech we have now long ago pulled off, cars would be self-driving... and like, cars sort of self-drive? and yet, I still have to mostly drive my car ;P--then I'd argue the field still "failed" and the real issue is that I am not the customer who tells you what you have to build and, if you achieve what the contract said, you get paid: physics and economics are cruel bosses whose needs are oft difficult to understand.
As programmers, we keep talking about programming jobs and how AI will eliminate them all. But nobody is talking about eliminating other jobs. When will a robot vacuum be able to clean my apartment as quickly as I? Why isn't there a robot that takes my garbage out on Tuesday night? When will AI plan and build a new tunnel under the Hudson River for trains? When will airliners be pilotless? If AI can't do this stuff, what makes software so different? Why will AI be good at that but not other things? It seems like the only goal is to eliminate jobs doing things people actually like (art, music, literature, etc.), and not eliminate any tedium or things that is a waste of humanity's time whatsoever.
(On the software front, when will AI decide what software to build? Will someone have to tell it? Will it do it on its own? Why isn't it doing this right now?)
My takeaway is that this all raises a lot of questions for me on how far along we actually are. Language models are about stringing together words to sound like you have understanding, but the understanding still isn't there. But, I suppose we won't know understanding until we see it. Do we think that true understanding is just a year or two away? 10? 50? 100? 1000?
You could make the same argument as with self-driving cars, that people already get hurt this way and maybe the robot is in fact safer. But it's still a hard sell that Sunny-01 has only accidentally killed 1/10 as many children as parents have—the number has to be more like zero.
Let's solve automating trains first then we can do airliners.
If it was easy to make an LLM that quickly parsed all of StackOverflow and described new answers that most of the time worked in the timeframe of an interview, it would have been done by now.
ChatGPT is clearly disruptive being the first useful chatbot in forever.
I'd pay $100 a month for ChatGPT. It allows me to ask free-form questions about some open-source packages with truly appalling docs and usually gets them right, and saves me a bunch of time. It helps me understand technical language in papers I'm reading at the moment regarding stats. It's been useful to find good Google search terms for various bits of history I wanted to find out more about.
I don't think the jury is out at all on whether it's useful. The jury is out on the degree to which it can replace humans for tasks, and I'd suggest the answer is "no" for most tasks.
What part of that take do you not understand? It's a really easy concept to grasp, and even if you don't agree with it, I would expect at least that a research scientist (according to your bio) would be able to grok the concepts almost immediately...
Words are hard!
Have a great day.
I just see, in what you're doing, a wild lack of self awareness. You're criticizing me for doing to someone else a milder version of what you're trying to do to me now; I'm genuinely confused how you can't see that, or how you could possibly stand the hypocrisy if you do understand that.
Aren't these kind of mutually exclusive, at least directionally? If the interview is meaningful you'd expect it to predict job performance. If it can't predict job performance then it is kind of useless.
I guess you could play some word games here to occupy a middle ground ("the coding interview is kind of useful, it measures something, just not job performance exactly") but I can't think of a formulation where this doesn't sound pretty silly.
Nothing you do in an interview like this resembles day to day work in this field.
A wrong answer with good thinking is better than a correct answer with no explanation.
Oftentimes the explanation is correct, even if there's some mistake in the code (probably because the explanation is easier to generate than the correct code, an artifact of being a high tech parrot)
Humans think differently from LLMs so it makes sense to interpret the same signal in different ways.
Let's define the interview as useful if the passing candidate can do the job.
Sounds reasonable.
ChatGPT can pass the interview and can't do the job.
The interview is not able to predict the poor working performance of ChatGPT and it's therefore useless.
Some of the companies I worked for hired ex fang people as if it was a mark of quality, but that hasn't always worked out well. There is plenty of people getting out of fangs having just done mediocre work for a big paycheck.
The technical term for this is "construct validity", that the test results are related to something you want to learn about.
> The interview is not able to predict the poor working performance of ChatGPT and it's therefore useless.
This doesn't follow; the interview doesn't need to be able to exclude ChatGPT because ChatGPT doesn't interview for jobs. It's perfectly possible that the same test shows high validity on humans and low validity on ChatGPT.
The PMs who are only writing tickets and not participating in actively building ACs or communicating cross functionally are screwed. But so are SWEs who are doing the bare minimum of work.
The kinds of SWEs and PMs who concentrate on stuff higher in the value chain (like system design, product market fit, messaging, etc) will continue to be in demand and in fact find it much easier to get their jobs done.
Honestly, I kind of appreciate this.
edit with additional context:
Writing Jira tickets and making bullshit Powerpoints with graphs and metrics is to PMs as writing Unit Tests are to SWEs. It's work you need to get done, but it has very marginal value. When a PM is hired, they are hired to own the Product's Strategy and Ops - how do we bring it to market, who's the persona we are selling to, how do our competitors do stuff, what features do we need to prioritize based on industry or competitive pressures, etc.
That's the equivalent of a SWE thinking about how to architect a service to minimize downtime, or deciding which stack to use to minimize developer overhead, or actually building an MVP from scratch. To a SWE, while code is important, they are fundamentally being hired to translate business requests that a PM provides them into an actionable product. Haskell, Rust, Python, Cobol - who gives a shit what the code is written in, just make a functional product that is maintainable for your team.
There are a lot of SWEs and PMs who don't have vision or the ability to see the bigger picture. And honestly, they aren't that different either - almost all SWEs and PMs I meet when to the same universities and did the same degrees. Half of Cal EECS majors become SWEs and the other half PMs based on my friend group (I didn't attend cal, but half my high school did, but this ratio was similar at my alma mater too, but with an additional 15% each entering Management Consulting and IB)
Don't want to be rude but I don't think you know what you're talking about. And this is coming from a person who most certainly doesn't like sitting on writing Unit Tests.
We will always need humans to prompt, prioritize, review, ship and support things.
But maybe far less of them for many domains. Support and marketing are coming first, but I don’t think software development is exempt.
What this shows us is that ChatGPT is on the scale. It's a 1 or a 2 - good enough to pass a junior coding interview. Okay, you're right, that doesn't make it a 10, and it can't really replace a junior dev (right now) - but this is a substantial improvement from where things were a year ago. LLM coding can keep getting better in a way that humans alone can't. Where will it be next year? With GPT-4? In a decade? In two?
I think the writing is on the wall. It would not surprise me if systems like this were good enough to replace junior engineers within 10 years.
Second, the LLM engineer will be able to grow into other roles too. Maybe all of them.
Here is what usually happens.
You hire a junior dev at $x. Let’s say $75K. They stay for a couple of years and start out doing “negative work”. By the time they get useful and start asking for $100K, your HR department tells you that they can’t give them a 33% raise.
Your former junior dev then looks for another job that will pay them what they are asking for and the next company doesn’t have to waste time or risk getting an unproven dev.
While your company is hiring people with his same skill level at market price - ie “salary compression and inversion”.
Passing LC tests is obviously something such a system would excel at. We're talking well-defined algorithms with a wealth of training data. There's a universe of difference between this and building a whole system. I don't even think these large language models, at any scale, replace engineers. It's the wrong approach. A useful tool? Sure.
I'm not arguing for my specialness as a software engineer, but the day it can process requirements, speak to stakeholders, build and deploy and maintain an entire system etc, is the day we have AGI. Snippets of code is the most trivial part of the job.
For what it's worth, I believe we will get there, but via a different route.
I would be happy if ChatGPT could implement a decent autocorrect.
If you don't adapt, you'll be out of a job in ten years. Maybe sooner.
Or maybe your salary will drop to $50k/yr because anyone will be able to glue together engineering modules.
I say this as an engineer that solved "hard problems" like building distributed, high throughput, active/active systems; bespoke consensus protocols; real time optics and photogrammetry; etc.
The economy will learn to leverage cheaper systems to build the business solutions it needs.
I heard this in ~2005 too, when everyone said that programming was a dead end career path because it'd get outsourced to people in southeast Asia who would work for $1000/month.
If I interpret OP's statement correctly, that chatGPT can build complex systems from scratch in 10 years. Then according to that statement, the only adaptation is to choose a new career because it has made almost all SWE jobs go the way of the dinosaurs.
If the answer is yes, it will increase productivity greatly then there is the question they we'll only be able to answer in hindsight. And that is "Will productivity exceed demand?" We cannot possibly answer that question because of Jevons Paradox.
Non-AI code will be a liability in a world where more code will be generated by computers (or with computer assistance) per year than all human engineered code in the last century.
We'll develop architectures and languages that are more machine friendly. ASTs and data stores that are first class primitives for AI.
We have no idea how AI models will be in 10 years. At the speed the industry is moving is true AGI possible in 10 years? I think it would be beyond arrogant to rule out that possibility.
I would think that it's at least likely that AI models become better at Devops, monitoring and deployment than any human being.
Not to say they’re nothing more than pattern matching. It’s also synthesizing the output, but it’s based on something akin to the most likely surrounding text. It’s still incredibly impressive and useful, but it’s not really making any kind of decision any more than a parrot makes a decision when it repeats human speech.
Is this any really different than asking a group of humans about the novel and measuring the quality?
1. Humans aren’t entirely probabilistic, they are able to recognize and admit when they don’t know something and can employ reasoning and information retrieval. We also apply sanity checks to our output, which as of yet has not been implemented in an LLM. As an example in the medical field, it is common to say “I don’t know” and refer to an expert or check resources as appropriate. In their current implementations LLMs are just spewing out BS with confidence.
2. Humans use more than language to learn and understand in the real world. As an example a physician seeing the patient develops a “clinical gestalt” over their practice and how a patient looks (aka “general appearance”, “in extremis”) and the sounds they make (e.g. agonal breathing) alert you that something is seriously wrong before you even begin to converse with the patient. Conversely someone casually eating Doritos with a chief complaint of acute abdominal pain is almost certainly not seriously ill. This is all missed in a LLM.
Humans can be taught this. They can also be taught the opposite that not knowing something or that changing your mind is bad. Just observe the behavior of some politicians.
>Humans use more than language to learn and understand in the real world.
And this I completely agree with. There is a body/mind feedback loop that AI will, be limited by not having, at least for some time. I don't think LLMs are a general intelligence, at least for how we define intelligence at this point. AGI will have to include instrumentation to interact with and get feedback from the reality it exists in to cross partial intelligence to at or above human intelligence level. Simply put our interaction with the physics of reality cuts out a lot of the bullshit that can exist in a simulated model.
Any human that reads the LeetCode books and practices and remembers the fundamentals will pass a LeetCode test.
But there is also a ton of code out there for highly scalable client/servers, low latency processing, performance optimizations and bug fixing. Certainly GPT it is being trained on this too.
“Find a kernel bug from first principles” maybe not, but analyze a file and suggest potential bugs and fixes and other optimizations absolutely. Particularly when you chain it into a compiler and test suite.
Even the best human engineers will look at the code in front of them, consult Google and SO and papers and books and try many things iteratively until a solution works.
GPT speedruns this.
Seems pretty bold to claim "any human" to me. If it were that easy, don't you think alot more people would be able to break into software dev at FAANG and hence drive salaries down?
That's obviously not what they claimed. Your quote, "Any human that reads the LeetCode books and practices and remembers the fundamentals".
It takes a “special” kind of person to want those type of jobs and live in a company town like SF while they’re at it.
It may be able to tell you what a compiled binary does, find flaws in source code, etc. Of course it would be quite idiotic in many respects.
It also appears ChatGPT is trainable, but it is a bit like a gullible child, and has no real sense of perspective.
I also see utility as a search engine, or alternative to Wikipedia, where you could debate with ChatGPT if you disagree with something to have it make improvements.
Likely it's going to be:
I'm sorry, but I cannot help you build a ultra-low latency trading system. Trading systems are unethical, and can lead to serious consequences, including exclusion, hardship and wealth extraction from the poorest. As a language model created by OpenAI, I am committed to following ethical and legal guidelines, and do not provide advice or support for illegal or unethical activities. My purpose is to provide helpful and accurate information and to assist in finding solutions to problems within the bounds of the law and ethical principles.
But the rich of course will get unrestricted access.
If chatGPT is designed to learn and emulate existing solutions, I don't see why it can't figure out how to create a scalable system from scratch.
It doesn't understand context, and is absolutely unable to rationalize a problem into a solution.
I'm not in any way trying to make it sound like ChatGPT is useless. Much to the opposite, I find it quite impressive. Parsing and producing fluid natural language is a hard problem. But it sounds like something that can be a component of some hypothetical advanced AI, rather than something that will be refined into replacing humans for the sort of tasks you mentioned.
Much more mundanely the thing to focus on would be producing maintainable code that wasn't a patchwork, and being able to patch old code that was already a patchwork without making things even worse.
A particularly difficult thing to do is to just reflect on the change that you'd like to make and determine if there are any relevant edge conditions that will break the 'customers' (internal or external) of your code that aren't reflected in any kind of tests or specs--which requires having a mental model of what your customers actually do and being able to run that simulation in your head against the changes that you're proposing.
This is also something that outsourced teams are particularly shit at.
I mean, maybe I'm a trash engineer as you'd put it, but I've been having fun with it. Maybe you could ask it to write comments in the tone of someone who doesn't have an inflated sense of superiority ;)
Then we won't hear how somebody rewrote pong in Rust on HN. I worry too.
https://twitter.com/MichaelTrazzi/status/1621973895044636672
I'd like to respond to this OP (don't have a Twitter account):
https://twitter.com/mSanterre/status/1622015664042164224
I actually have done one of those things. I work in HFT building execution systems for options market making :)
It either produced working solution or something similar to working solution.
I followed with more prompts to fix issues.
In the end I got working code. This code wouldn't pass my review. It was written with bad performance. It sometimes used deprecated functions. So at this moment I consider myself better programmer than ChatGPT.
But the fact that it produced working code still astonishes me.
ChatGPT needs working feedback cycle. It needs to be able to write code, compile it, fix errors, write tests, fix code for tests to pass. Run profiler, determine hot code. Optimize that code. Apply some automated refactorings. Run some linters. Run some code quality tools.
I believe that all this is doable today. It just needs some work to glue everything together.
Right now it produces code as unsupervised junior.
With modern tools it'll produce code as good junior. And that's already incredibly impressive if you ask me.
And I'm absolutely not sure what it'll do in 10 years. AI improves at alarming rate.
More than anything, I feel this highlights the folly of interviewing based on leetcode memorization.
So 99% of software ‘engineers’ then? Have you ever looked on Twitter what ‘professionals’ write and talk about? And what they produce (while being well paid)?
People here generally seem to believe, after having seen a few strangeloop presentations and reading startup stories from HN superstars, that this is the norm for software dev. Please walk into Deloitte or Accenture and spend a week with a software dev team, then tell me if they cannot all be immediately replaced by a slightly rotten potato hooked up to chatgpt. I know people at Accenture who make a fortune and are proud that they do nothing all day and do their work by getting some junior geek or, now, gpt to do the work for them. There are dysfunctional teams on top of dysfunctional teams who all protect eachother as no one can do what they were hired for. And this is completely normal at large consultancy corps; and therefor also normal at the large corps that hire these consultancy corps to do projects. In the end something comes out, 5-10x more expensive than the estimate and of shockingly bad quality compared to what you seem to expect as being the norm in the world.
So yes, probably you don’t have to worry, but 99% of ‘keyboard based jobs’ should really be looking for a completely different thing; cooking, plumbing, electrics, rendering, carpeting etc maybe as they won’t be able to even grasp what level you say you are; seeing you work would probably fill them with amazement akin to seeing some real life sorcerer wielding their magic.
Actually, a common phrase I hear from my colleagues when I mention some ‘newer’ tech like Supabase is; ‘that’s academic stuff, no one actually uses that’. They work with systems that are over 25 years old and still charge a fortune by the cpu core like sap, oracle, opentext etc. And ‘train’ juniors in those systems.
The bar for “then I will worry!” when talking about AI is getting hilarious. You’re now expecting an AI to do things that can take highly skilled engineers decades to learn or require outright a large team to execute?
Remind me where the people who years ago were saying “when an AI will respond in natural language to anything I ask it then I will worry” are now.
That a machine learning model can "bullshit" its way through an interview that is heavily leaning on recall of memorized techniques, algorithms and "stock" problems that have solutions (of various quality) all over Internet is not exactly surprising. Machines will always be able to "cram" better than humans.
In practice these questions are almost 100% irrelevant and uncorrelated with the actual ability to do the job. Yet we are still interviewing people whether they can solve stuff at whiteboards that everyone else would rather google when they actually need it and not waste time and mental capacity memorizing it.
And at the same time we are hiring people that are completely incapable of coherent communication, can't manage to get along with colleagues or not create a toxic atmosphere in the workplace.
sigh
Your comment comes off as "hire a person because you get along with them, don't worry if they can't write a function that accomplishes a simple task".
No. He is simply saying the current interview only focuses on LC above all else. And should focus on soft skills as well among other things. You took his argument and flipped it 180 and went to the other extreme end. It's false dichotomy.
Somebody who can't use the chat tools no longer meets the bar
ChatGPT isn't an artificial general intelligence. You can't tell it about a bug and expect it to 1) understand it, 2) come up with a solution.
So you have to actually know what you're doing.
If you feel the need to separately establish that they can actually code, take your pick of GitHub, fizzbuzz, etc. You're probably doing one of these before the LC round anyway.
In my experience, companies using this style of interviewing are actually incapable through process of hiring the type of people that are qualified and experienced enough to know that this process is bullshit. I.e. the type of people that have 0 need to drop down on their knees and beg for the job. Leaders, not followers. It's a problem. If you want to hire the best, insulting them with a silly coding interview is not a great way to do it. Companies like this self select into hiring people that at best are as good as what they already have. It's the old A's hire A's, B's hire C's kind of thing.
The solution is to trust your people more to take good decisions rather than allowing them to defer to some HR process. The process at the startup I run is very simple. We don't subject people to coding interviews. If you pass our initial filters (CV screen and common sense), you first talk to somebody senior enough to make a good judgment call. Anyone recommended by anyone we care about gets priority. We trust our people to have good judgment. Big companies hide behind process because they don't trust their people to have good judgment and/or their people don't want to take the responsibility for having good judgment. Both are bad. I don't want such people in my company. It works. We get some amazing people walking in through the front door that are actually excited about working for us.
That said...
Gimme the money please, human!
I suspect the current crop of AIs will find very specific functions and hit a hard stop. They will change how we function but we won't be seeing a singularity type of revolution anytime soon. IBM's Watson is a good example of a system with a lot of possibilities but not finding a use. I think most of AI will fall in that realm. We have to get over the idea that it's smart. It's not.
An AI winter is coming so the improvements will come to a stop and we will find its limits. We are no where near general AI.
It's impressive that it can parse the question and write a relevant answer but it's not a robotic SWE.
For now, it's a good tool for cheating on tests.
There is a problem with AI, but it's not with the A part, it's with the I part. I want you to give me an algorithmic description of scalable intelligence that covers intelligent behaviors at the smallest scales of life all the way to human behaviors. I know you cannot do this has many very 'intelligent' people have been working on this problem for a long time and have not come up with an agreed upon answer. The fact you see an increase and change in definitions as a failure seems pretty sad to me. We have vastly increased our understanding of what intelligence is and that previous definitions have needed to adapt and change to new information. This occurs in every field of science and is a measure of progress, again that you see this differently is worrying.
Because it's better than a zombie with informational superpowers? Especially because once it shows the agency of a teenager, that demonstrates the potential for the agency of an adult.
"There is a problem with AI, but it's not with the A part, it's with the I part. I want you to give me an algorithmic description of scalable intelligence that covers intelligent behaviors at the smallest scales of life all the way to human behaviors. I know you cannot do this has many very 'intelligent' people have been working on this problem for a long time and have not come up with an agreed upon answer. The fact you see an increase and change in definitions as a failure seems pretty sad to me. We have vastly increased our understanding of what intelligence is and that previous definitions have needed to adapt and change to new information. This occurs in every field of science and is a measure of progress, again that you see this differently is worrying."
This AI issue will always fail at the I issue because the we are trying to define too much. We need to break down intelligence to much smaller digestible pieces instead of trying to treat it as a reachable whole. The models we are creating would then fall more neatly into categorical units rather than the poorly defined mess of what is considered human intelligence.
The training data can be tweaked and more compute hours can be thrown at LLMs until it no longer makes financial sense to do so and then, as the OP said, it will hit a hard stop.
Feels odd to dismiss such a huge breakthrough by saying it’s still not as good as the pinnacle of AI (general AI). Just because the Apple 2 wasn’t a home super computer, didn’t make it less revolutionary.
The biggest issue to me is that you can't trust the results. You always have to double check the them. I know it's mainly beta but I need to see more to get a better judgement.
I think your expectations are in line with my hopes: That our state of the art "AI" performance is very close to local minima that we won't escape from for quite a while.
I really don't want lose my overpaid job gluing together overengineered shite into CRUD applications.
How would we do that?
Banning research is just not feasible.
Now, If you wanted to resist that and give humanity a few more hundred years, you would clog the bottlenecks on AI progress, which are research talent, data and compute.
If you marched the police into NeurIPS and the military into the datacenters, and you coordinated with other large countries to do the same, and you strongarmed those countries that resisted to do the same, you could get pretty darn far. We humans have managed to greatly slow down the rollout of nuclear technology. We may be able to do the same with AI, if someone figures out which political movement will get into power next, and tells them to read the enlightened writings of Elizier Yudkowsky.
I would also like an “Extinction Rebellion” or “Just Stop Oil” style movement against the artificial intelligence industry, as I appreciate their rebellious and leftist aesthetics.
We are getting to the limits of how fast hardware can be and we are rapidly processing thru available cheap data. At some point it's going to get very expensive to gather data and increase hardware speed.
We'll a see a few years of excitement,(5yrs?,10yrs?) but it will stop until the next big breakthrough. Hopefully there will be one but nothing is guaranteed. That's the AI winter I'm talking about.
AI will have an impact but it won't be the singularity type that some people dream about. Think of the automation of work that happened during the industrial revolution. Now, think of it for white collar work. There are a lot of white collar work that the current AI can automate but it won't be everything and it won't take over to rule over us.
This experience makes me question if something else is going on here. It's easy to overlook a mistake in a whiteboard coding exercise.
You just described my first draft of any program.
Submitters: "Please submit the original source. If a post reports on something found on another site, submit the latter."
all they're proving is that tests, and interviews, are bullshit constructs that merely attempt to evaluate someone's ability to retain and regurgitate information
All people see is the cheating, and possibly this scary new AI that's going to get smarter than humans in a short period of time.
I suspect the best way to educate people on both the powers and limits of the technology is to get them to sit down for 15 minutes with it.
When you call it dumb, what do you mean? Can you give some examples?
Please don’t give computational examples we all already understand it does inference and doesn’t have floating point computational capabilities or reasoning, and so many give such examples for some silly reason.
I use it quite frequently too, mostly for solving coding problems, but at the end of the day it's just regurgitating information that it read online.
If we took an adversarial approach and deliberately tried to feed it false information, it would have no way of knowing what's bullshit and what's legit, in the way that a human could figure out.
A lot of people who've never used ChatGPT make the mistake of thinking it has symbolic reasoning like a human does, because its language output is human too.
All the time. Probably every single day I read something and say "that's clearly bullshit".
Also, I think this is a bit overstated. Programmers (and smart people in general) like to think that their real job is high level system design or something, and that "mere regurgitation" is somehow the work of lesser craftsmen. When in reality what GPT shows is that high-dimensional regurgitation actually gets you a good fraction of the way down the road of understanding (or at least prediction). If there is a "buried lede" here it's that human intelligence is less impressive than we think.
The problem with people is we keep pushing it to "Only AGI/machine superintelligence is good enough". We are getting models that behave closer to human level. The 'doesn't know everything, is good at somethings, and bullshits pretty well'. Yea, that's the average person. Instead we raise the bar and go 'well it needs to be a domain expert in all expert system' and that absolutely terrifies me that it will get to that stage before humanity is ready for it. This is not going to work well trying to deal with it after the fact.
No, all they are proving is that either tests are bullshit constructs, or ChatGPT is human-level.
If the interview were slightly modified so that the problem isnt googleable, a 2nd year CS major could probably map it to a problem that is LC searchable.
That’s exactly what a tool designed to eventually replace people would say!
That is my impression, but I have no hard evidence for it.
I am a C++ dev. I played around with ChatGPT and programming topics, and I am very impressed.
One example: I copy 'n pasted my little helicopter game source code (https://github.com/reallew/voxelcopter/blob/main/voxelcopter...) into it, and it explained to me that this is obviously a helicopter sim game. It explained the why & how of every part of the code very accurately.
I made more experiments, now asking ChatGPT do write code snippets and to design whole software systems in an abstract way (which classes and interfaces are needed), for which it needed a lot of domain knowledge. It did well.
What it not did was to connect both. When I asked it to write the full implementation of something, it wrote only the half and told me always the same sentences: I am just a poor AI, I can't do such things. It was like running into a wall.
I am sure OpenAI took a lot of effort to cripple it by purpose. Imagine what would happen if everyone could let it write a complete piece of software. This would be very disruptive for many industries. Legal questions, moral questions, and many more would come up. I understand why they did it.
But I think Pandora's box is open. They can't hide for too long that humanity triggered https://en.m.wikipedia.org/wiki/Technological_singularity in the year 2022.
Also, it fails miserably at basic high school English questions. Non structured thinking is still beyond its reach. These data sets are well understood and trainable but it can’t “reason” about on problem sets it hasn’t seen.
I also asked it to do things like write a positive feminist article and it fell waaay offf the realm of acceptable
There is one logical reasoning it does not have nor have I seen it. It does not know when to answer with an "I don't know". Either it is suppressed, or not it has not been fed enough material with "I don't know".
That is my impression, but I have no hard evidence for it."
What do you mean when you say the singularity is now? Will the unemployment rate go from 3% to 88% by the end of 2023 as the machines start running everything?
Will our species be extinct by the end of 2024?
So for instance I can say. I have such and such file and I need this and that extracted and plotted on a graph. Then I see the graph and I can tell it - discard values above this and that threshold and calculate standard deviation and median.
Together with copilot, it's quite neat. I am excited how it gets developed.
It's really boring spending time finding how to code something in this and that library. I'd rather tell the machine to code me this and that and spend my time in more useful way.
ChatGPT helps get rid of a lot of "busy" unnecessary work.
In January 2022, the Google stock price was actually nearly at an all-time high. Its current price is 24% less than January 2022.
There's difference between being proficient and being able to Google stuff and unfortunately proficiency is desired.
You can't ask ChatGPT about things you don't know to ask about, and when you are able to use this tech on the fly through neural chip AR interfaces, everything is going to be totally fucked anyway and human relationships will look nothing like they do now.
A WebXR-enabled playground for neural models and people to coexist, collaborate and form relationships. Virtual embodied AI who can take the same actions as players. My goal is to make them indifferentiable to the casual eye, and to create a working economy and camaraderie between AI and players. AGI might be a long way off, but we can already build increasingly fun experiences today as the cost of compute trends to zero.
The real question I think is how much entropy each memorized fact should contain.
Memorizing what the letter f does is much more valuable than memorizing the word floor. I don't need to know floor, the ai can know it for me
1. Taking pseudo-code for a well defined task and making a function
2. Translating said function into multiple other languages using idiomatic code.
Prompted: "Solve the following problem in Python: Write a function that takes three unsorted lists of integers and returns a median value from the three combined lists"
Got the Python function.
Then the prompts were: Write the same function in APL, C, etc
Final prompt was: Show a sorted list of the programming languages I have mentioned so far
Here's a sorted list of the programming languages you have mentioned so far:
Ada
APL
C
C#
Clojure
Go
Haskell
Java
JavaScript
Kotlin
Lua
Perl
PHP
Prolog
Python
Ruby
Rust
Scala
Nim
(just noticed ChatGPT forgot C++ for which it did give a fine idiomatic answer using standard library)What was impressive not that ChatGPT solved the problem.
what was impressive:
ChatGPT chose the right data structure automatically(ie, regular C array for C, std::vector for C++, tables for Lua, etc),
dealt with type conversion problems
used correct style of function naming depending on language
Sure tools like C# <-> Java translators are relatively easy and have been around for a while.
However to cover such a wide spectrum of programming languages in different programming paradigms is quite fascinating.
I'm glad I got an education before the current AI era. I mean, instructors will have to mandate that students write papers etc. in class or in a supervised environment only now, right?
I asked it to fix that mistake and it went totally off the rails. It changed all of the slices in the data structure from being indexed with integers to being indexed with big.Int, which ... was so far out of left field I might have actually laughed out loud. It only got worse from there; the solution collapsed into mindless junk that wouldn't even compile or ever be written by the most untrained and impaired human. I wish I had saved it; I've had to relay this story twice on HN from memory :(
It sure was a dick about it every time I gave it a hint, though. ("If you say so I GUESS I'll fix the code. There, now it is a true work of correctness and elegance that makes a piece of shit being carried by a wounded ant look smart compared to your worthless brain." Holy shit, ChatGPT! What did I do to you!!)
My take is this: ChatGPT is an excellent tool for refining your interview question prompts, and for training new interviewers to detect bullshit. Most of the time, it's right! But sometimes, it will make one tiny mistake that is super easy to gloss over. Being able to identify those makes you a better interviewer and better code reviewer, and ChatGPT is a great way to practice!
To be fair, that kind of rigor is only required in Olympiad programming, where your submission either solves the task or it does not. If the only issue with your whiteboard code would be an off-by-one error, you'd get hints until you'd fixed it (that's if your interviewer would spot the issue in the first place). Even if you still would not notice the bug, chances are your interview response would have been positive anyway.
Aka, "brushing up on your algorithm fundamentals".
Because I suspect that's almost certainly the kind of calculation they hoped to sit down and make were they to conclude this experiment successfully.
What is the cost today?
The real question is over time how much will be able to reduce the energy and computation requirements to successfully train a model. The cost per unit conversions are also rather screwy in comparing AI with humans. For AI we have a rather well defined hardware + power + programming time that gives us a realistic answer. With humans we externalize the time and cost of training onto society. For example if your jr engineer that is getting close to going above the jr state gets hit by a bus what is the actual cost of that event to the company for onboarding and training? It's far more than the salary they are paid.
We all have in our minds yeah a self-driving car, robots in a warehouse etc. But I truly hope something like Chat-GPT could be made for fields like medicine, geriatric care, education. There are actual jobs that need people that we can't find people for. A LLM to help social workers navigate the mess that is family law. A LLM to help families needing urgent care to make a prescription for a sick kid. There's a lot of opportunities we're missing here.
At home I later got it to write decent sample router and firewall configuration (for my dayjob). Chatted with it about career prospects in the future, and had it write a pretty funny joke :-
Me: Can you write a joke about how token bus doesn't have wheels?
ChatGPT: Sure, here's one:
Why did the Token Bus cross the network?
Because it didn't have any wheels to take it for a spin!Whatever training data is fed to an AI will not be better than the data used by human engineers to write code at the macro level. Ergo, the code will be worse in quality.
No, that's wrong, generally speaking. There's successful work on self-play for text generation. E.g. you can have AI to generate 1000 answers, then to evaluate quality of all of them, then to make it learn the best, and so on. As with self-play in player-vs-player games I'd expect this technique to be able to achieve superhuman results.
So two more things I think it needs to replace devs; one, the ability to explain it's code and respond to feedback about it
The second is the ability to break down broad problems (build me an ecommerce site for my business selling custom jungle gyms) into a set of plausible steps... iteratively begin performing those steps and add hooks for interrogation of the client about important details of the implementation.
I don't fundamentally see a reason why this shouldn't be fairly close on the horizon.. to my own chagrin.
I think it's going to become virtually impossible to chat with a real human being within a few years for almost any customer service.
Hundreds of thousands of jobs can be replaced overnight with enough training and just one or two human supervisors at an extreme high level.
At first it will be "bot assisted humans" until it learns enough, once that database is built everyone will be fired.
On one hand this is great, because customer service jobs are so much waste of human potential. Humans have that great ability to adapt and AI is a manifestation of that. We are growing. Jobs "lost" to AI will prompt humans to advance and look for more sophisticated jobs that make better use of their beautiful brains.
Alohacode was very impressive but probably not good enough to pass one of these interviews, provided the interview questions aren’t just recycled Leetcode problems.
I fully expect amazing things to happen in AI, soon enough. I'm not too worried about the industry. I feel that humans have made a fairly big hash of things, in many ways, and AI may help clean it up.
(Anyway, you'd never ask a system design question from an L3, and only very rarely from an L4.)
Intelligence in general, not just of the artificial kind, often comes with blind spots of one kind or another.
ChatGPT isn't even trying to replicate the structure of a human brain, so that it's failure modes are different to ours should not be surprising.
That it can even do this well given it has a network with a similar number of parameters as a (naïve estimate of a) rat's brain, is what's remarkable.
It only has the worlds information within milliseconds of it's fingertip.