Yes, I know calculators can encourage people to not think; I should know because I wrote one. [1]
But the current "AI" tech is so much worse on that front. It's a difference of degree, and that degree does matter.
By the way, when I was learning to fly helicopters [2], I used my calculator to calculate weight and balances, but I also did it by hand!
[1]: https://git.gavinhoward.com/gavin/bc
[2]: https://gavinhoward.com/2022/09/grounded-for-life-losing-the...
You act like encouraging people not to think is a problem. Thing is, you'd be wrong. We want people not just to think, but to focus. If I'm a pilot and I have to worry about the runtime environment of the command-line calculator I use to hand-calculate my route and cockpit configuration, is that a good use of my focus? I think most people would say no. We definitely want to discourage the pilot from actively thinking about that kind of stuff. Should they have a grasp of the basics in case of emergency? Sure. Do we have a sustainable and efficient system of transportation if that's how our pilots spend their time? No.
Aviation is about redundancy. Redundancy is a good use of a pilot's focus. That's why I did both. I didn't blindly trust my calculator to not have bugs (even though I wrote it!), and I didn't blindly trust my hand calculations to be correct.
If they agree, though, it's a good sign that everything is in good order. That's what redundancy is for, to ensure that a problem in one thing does not lead to another problem, like in the Swiss cheese model of accidents.
If you're just trying to get off Gilligan's island, that's another thing entirely.
A commercial pilot friend has told me that they still check the calculations. When they don't, they get accidents like the Gimli Glider.
It's like saying that putting all of the legal checks on the lawyers is not going to give you a sustainable business. But we all know that's wrong.
When I wrote that calculator, I didn't write it for flight. I wrote it as a general calculator and just used it for flight. I would have used the GNU bc if I didn't write my own.
So it's a bit disingenuous to claim that I am claiming that pilots should write their own.
And pilots are included in those maintaining the system; they're not just using it.
Sure you can open an accountancy business and refuse to use calculators, but that's just working with a strange self-imposed limit rather than using technology to best support your business.
Yes, calculators do that. I'm arguing that "AI" does not let programmers write better code faster. It lets them write worse code faster, or better code slower.
Turns out that really doesn't matter. I think your argument is incredibly weak; the fact that some people don't use these tools effectively doesn't mean that nobody can. Whoever figures this stuff out is going to win, that's just how it works.
That is to say, there will always be a niche for people who refuse to move up the chain of abstraction: they're actually incredibly necessary. However, as low-level foundations improve, the possibilities enabled higher up the chain grow at an exponentially-higher rate, and so that's where most of the work is needed. Career-wise it might be better to avoid AI if that's what you want to do, but as a business I can't see a dogmatic stance against these tools being anything but an own goal.
Except that it does!
For every level of abstraction, you lose something, and abstractions are leaky.
The lower levels of abstraction make you lose the least, and they are also the least leaky. The higher you go, the more you lose, and the more leaky.
What I'm claiming is that these "AI" tools have definitely reached the point where the losses and the leaks are too large to justify. And I'm betting my career on that.
You're not betting anything because the cost for you to change your mind and start working with AI tools is exactly 0. This rhetoric is just marketing. I'm sure you'll find the customers that are right for you, but you can at least admit that this kind of talk is putting the aesthetic preference of what you want work to look like above what's actually the most effective. Again, I'm sure you'll find customers who share those aesthetic preferences, but to pretend like it's actually an engineering concern is marketing gone too far.
Did I ever deny that? Sure, some of those layers are worth it. That doesn't address my assertion that these "AI" tools are not.
> Telling yourself that the layer above you isn't feasible isn't going to do you any favors but it does generate buzz on social media which seems like it's the goal here.
You're halfway there.
> You're not betting anything because the cost for you to change your mind and start working with AI tools is exactly 0.
And here is where you contradict yourself.
If I'm getting loud about this bet, and making customers because of this bet, then it will cost me a lot to start working with "AI" tools. My customers will have come to be because I don't, so if I start, I could easily lose all of them!
> This rhetoric is just marketing.
Yep! But that's what makes my best actually cost something. I'm doing this on purpose.
> I'm sure you'll find the customers that are right for you, but you can at least admit that this kind of talk is putting the aesthetic preference of what you want work to look like above what's actually the most effective.
No, I will not admit that because I believe very strongly that my software will be better, including engineering-wise, than my competitors who use these "AI" tools.
The idea isn't to write better code faster, it's to build better products faster.
Although IMO in the future, AI will probably also enable programmers to write better code too (faster, less bugs, more secure, more frequently refactored etc)
All else being equal, better code means better products.
Also, to have a better product without better code, you're implying that the design of the product is better and that these "AI" tools help with that.
Until they can reason, they cannot help with design.
And I would bet that AI design would help things where the existing designers are bad, e.g. so much open source UI (that is, not cli UX) written by devs, but it is still a bit away from the top quality like Steve Jobs.
Maybe this is like the transition from hand crafted things to machined things; we go from a world some some excellent design and some meh design to a world with more uniform but less great designs.
"AI" design will not help until we have a true AI that can reason. (I don't think we ever will.)
Why is reasoning necessary? Because design is about understanding constraints and working within them while still producing a functional thing. A next-word-predictor will never be able to do that.
What is your definition of reasoning that you do not think GPT-4 would demonstrate signs of?
Heh, there have been many attempts to define reasoning. I haven't seen a good one yet.
However, I'm going to throw my hat into the ring, so be on the lookout for a blog post with that. I've got a draft and a lot of ideas. I'm spending the time to make it good.
Otherwise it’s just moving the goalposts.
To demonstrate this, ask it to prove something that most or all people believe. Say some "intuitive" math thing. Perhaps the fact that factorial grows faster than exponential functions.
And no, don't just have it explain it, have it prove it, as in a full mathematical proof. Give it a minimal set of axioms to start with.
Merriam-Webster's definition of "reasoning" [1] says that reasoning is:
> the drawing of inferences or conclusions through the use of reason
So starting GPT4 off with some axioms would give it a starting point to base its inferences on.
Then, if it does prove it, take away one axiom. Since you started with a minimal set, it should now be impossible for GPT4 to prove that fact, and it should tell you this.
Having GPT4 prove something with as few axioms as possible and also admit that it cannot prove something with too few axioms is a great test for if it is truly reasoning.
Take this problem instead which certainly requires some reasoning to answer:
“Consider a theoretical world where people who are shorter always have bigger feet. Ben is taller than Paul, and Paul is taller than Andrew. Steve is shorter than Andrew. Everyone walks the same number of steps each day. All other things being equal, who would step on the most bugs and why?”
I think it’s a logical error to say “AI can’t reason about this, so that proves that it can’t reason about anything at all” (particularly if that example is something most humans can’t do!). The LLMs reasoning is limited compared to human reasoning right now, although it is still definitely demonstrating reasoning.
Because Ben is the tallest, his feet are the biggest, and because he takes the same amount of steps as the others, the amount of area he steps on is larger than the area that the others step on.
Therefore Ben is most likely to be the one to step on the most bugs.
Easy. And I'm not brilliant.
The problem with testing these tools is that you need to ask it a question that is not in their training sets. Most things have been proven, so if a proof is in its training set, the LLM just regurgitates it.
But I also disagree: if the "AI" can't reason about that, it can't reason because that one is so simple my pre-Kindergarten nieces and nephews can do it.
But even if not, the LLM's should have "knowledge" about exponential functions and factorial because the humans who wrote the material in their training sets did. So it's not a lack of knowledge.
And I claim that most humans could rediscover theorems from basic axioms; you've just never asked them to.
Ben (tallest) Paul Andrew Steve (shortest) Since shorter people have bigger feet in this world, we can also deduce the following order for foot size:
Steve (biggest feet) Andrew Paul Ben (smallest feet) Assuming that everyone walks the same number of steps each day and all other things being equal, the person with the biggest feet would be more likely to step on the most bugs simply because their larger foot size would cover a greater surface area, increasing the likelihood of coming into contact with bugs on the ground.
Therefore, Steve, who is the shortest and has the biggest feet, would step on the most bugs.”
GPT4 solved it correctly. You didn’t.
And GPT4 didn't solve it correctly. It's a probability, not a certainty, that the shortest person will step on more bugs.
At the very least, this should be evidence that the problem wasn't a totally-trivial easy pre-kintergarden level problem though, and it did manage to correctly solve it.
It required understanding new axioms (smaller = bigger feet) and infering that people with bigger feet would crush more bugs without this being mentioned in the challenge.
Your dismissal that the AI messed up because it didn't phrase the correct answer back in the way you liked is a little harsh IMO, as the AI's explanation does make it clear it is basing it on likelihoods ("the person with the biggest feet would be more likely...").
Equally if it can help devs launch a month earlier, that’s a huge advantage in terms of working out early product/market fit.
All things being equal, I would rather have a company with better product/market fit than one with great code (even though both are important!).
That's a very big "if", and one I just don't think will exist.
Also, that only helps at the beginning. Add the product gets more complex, I believe the AI will help less and less until velocity will become slower than companies like mine.
And product/market fit is just a way for companies to cover up the fact that their founders wanted to found a company, not solve a real problem. If you solve a real problem first, founding a company is simple and you "just" have to sell your solution.
Floating point errors creeping in is why we have to use quaternions instead of matrixes for 3D games. Apparently. I'd already given up on doing my own true-3D game engine by that point.
In some sense we "know why" humans make mistakes too — and in many fields from advertising to political zeitgeist we manipulate using knowledge of common human flaws.
On this basis I think the application of pedagogical and psychological studies to AI will be increasingly important.
Lying requires knowledge that what you are saying is not the truth, and usually there's a motive for doing so.
I don't think ChatGPT is there yet... or is it?
"Confabulate" however, appears to be a good description. Confabulation is, I'm told, associated with Alzheimer's, and GPT's output does sometimes remind me of a few things my mum said while she was ill.
Outside of the code use case, what should I rely on ChatGPT for that won't have me also looking for the information somewhere else? I suppose subjective soft things, like writing communications. But I can't rely on it for information.
I wouldn’t ride in a vehicle it designed tho, based on my week of asking it to do go programming.
Yes, there's a difference between a deterministic outcome and a non-deterministic one. But throw humans into the loop, and it becomes more interesting. I can't count the number of times I've listened to someone argue their answer must be right because they got it from the calculator. And it's not just students; as a teacher I've always paid attention to how adults use math.
With calculators or GPT tools, or any other automated assistant, judgement and validation continues to matter.
Answers from calculators are always right! But the human may have asked the wrong question.
Edit: slapping a few more in here:
https://learn.microsoft.com/en-us/office/troubleshoot/excel/...
Just because they are both both decks doesn't mean they are the same.
I don't mean to be pedantic. I teach coding to elementary school students, and this is something fundamental I try to make them understand. A computer will always do what you tell it to do. A bug is when you accidentally tell a computer to do something different than what you'd intended.
Going back to the calculator example, if a student used a calculator and got the wrong answer, the problem didn't come from the calculator. This is useful to understand; it can help the student work backwards to figure out what did go wrong.
AI is different in that we've instructed the computer to develop and follow its own instructions. When ChatGPT gives the wrong answer, it is in fact giving the right answer according to the instructions it was instructed to write for itself. With this many layers of abstraction, however, the maxim that computers "always do what you tell them" is no longer useful. No human truly knows what the computer is trying to do.
It's wrong in the same way that saying 1/1 = 1.0004 is wrong. It's not a matter of chosen precision in that it doesn't make the answer correct when you increase the number of zeros between 1 and 4.
In the case of translation of floating point numbers from base-2 to base-10 we have to make approximations which will often be slightly wrong forever without regard for amount of precision.
With AI, depending on the pre-conditions, the AI could be stuck in a state of being slightly wrong forever for a specific question without regard to further refinement of the query.
These are both still useful as tools. We just need to be able to work on the amount of refinement of the answer that the AI gives, which may be able to be solved fairly well through prompt engineering, if not through the advancement of GPT itself.
I'm sorry in advance, but this reply is just to meet pedantry with pedantry.
> A computer will always do what you tell it to do.
This is the Bohr model of computers. It's the kind of thing you tell elementary school students because it's conceptually simple and mostly right, but I think we know better here on HN. Pedantically, computers don't always do what you tell them to, because the don't always hear what you tell them, and what you tell them can be corrupted even when they do hear it.
For instance, random particles from outer space can cause a computer to behave quite randomly: https://www.thegamer.com/how-ionizing-particle-outer-space-h...
why was nobody able to pull it off, even when replicating exactly the inputs that DOTA_Teabag had used? Simple: this glitch requires a phenomenon known as a single-event upset, which is very much out of any player's control.
I don't think we can reasonably say that in this instance, the computer behaved according to what the user told it to do. In fact, it responded to the user and the environment.My underlying point is that, at least in 99.999% of cases, the problem isn't the calculator, it's the human using the calculator incorrectly. And although you could draw some parallels between calculators and AIs with regard to selecting the right tool and knowing when and how to use it, I'd say the randomness involved in an LLM is fundamentally different.
It’s mostly fine until it isn’t. AI will probably operate in the same capacity. We already have so much incorrect information out there that’s part of our pop culture. Even down to things like the fact that Darth Vader never said, “Luke, I am your father,” and Mae West never said, “Why don’t you come see me sometime?”
Even basic movie quotes are beyond our ability to get right. Hilariously, I just asked ChatGPT about these quotes and it explained that these are common misquotes, told me what was actually said in these movies, and explained some relevant context.
Sherlock never said, “Elementary, my dear Watson” even once in the books. Kirk never said, “Beam me up, Scotty.” We’re much less correct than we like to think. And somehow we’ve survived.
ChatGPT is fallible just like we are. We’ll manage, just like we always have.
I actually agree with you, but, in the same vein, does it not mean that user did not ask correct prompt?
And in this case you don't get to put in the formula.
We know this wording is just show but people still get swayed by it and believe it must be true.
Also another point while I'm here: Many many humans I've met are often confident when incorrect as well and can and will bullshit an answer when it suits their comfort.
That is, for me, the output of ChatGPT or other AI tools is the starting point of my investigation, not the end output. Yes, if you just blindly paste the output from an AI tool you're going to have a bad time, but we also standardize code reviews into the human code-writing process - this isn't that different.
Just giving one specific example, I find ChatGPT to my an incredibly efficient "documentation lookup tool". E.g. it's great if I'm working with a new technology or API and I want to know "what my options are", but I don't know what keywords to search for, it can help give me a really good "lay of the land", and from there I can read on my own to get more specifics.
I can't buy any of this hype for a "word-putting-together" algorithm. It's not real intelligence.
For example, I commented last week that I've found ChatGPT to be a great tool for managing my task list, and for whatever reason the "verbal" back-and-forth works much better for my brain than a simple checklist-based todo app: https://news.ycombinator.com/item?id=35390644 . But, I also pointed out how it will get the sums for my "task estimate totals by group" wrong. But it's so easy to see this mistake, and after using it for a while I have a good understanding for when it's likely to occur, that it doesn't lessen the value I get from using the tool.
The code in the post is wrong. For this "trivial" example, if you just blindly copied it into your code, it would not do what you want it to do. I love this example not just because it's ironic, but because it's a perfect illustration of how you need to know the answer before you ask for the solution. If you don't know what you're doing, you're gonna have a bad time.
I'm not at all concerned about the value of programmers falling to zero. I'm concerned that a lot of bad programmers are going to get their pants pulled down.
[1] https://skventures.substack.com/p/societys-technical-debt-an...
(Edit: and as a totally hot take, while I'm not worried about good programmers, I think the marginal value of multi-thousand word, think-piece blogposts is rapidly falling to zero. Who needs to pay Paul Kedrosky and Eric Norlon to write silly, incorrect articles, when ChatGPT will do it for free?)
Also, I think the coding example in that substack highlights that one of the most important characteristics of good programmers has always been clarifying requirements. I had to read the phrase "remove all ASCII emojis except the one for shrugs" a couple times because it wasn't immediately clear to me what was meant by "ASCII emojis". I think this example also highlights what happens when you have 2 "VC bros" who don't know what they're talking about highlighting the "clever" nature of what ChatGPT did, because it is totally wrong. Still, I'd easily bet that I could create a much clearer prompt and give it to ChatGPT and get better results, and still have it save me time in writing the boiler plate structure for my code.
In short, I don't know if we "agree", but I think OP is/was correct that GPT generates lots of subtle mistakes. I'd go so far as to say that the folks filling this thread with "I don't see any problems!" comments are probably revealing that they're not very critical readers of the output.
Now for a wild prediction of my own: maybe the rise of GPT will finally mean the end of these absurd leetcode interview problems. The marginal value of remembering leetcode soltutions is falling to zero. The marginal value of detecting an error in code is shooting up. Completely different skills.
I'd possibly argue that knowing how to ask the right question is tantamount to knowing the answer.
Whatever, time will tell. I still haven’t quite figured out how to make good use of GPT-4 in my daily work flow, tho it seems it might be possible.
Has anyone asked it to make an entry for the IOCC?
I was also using it to generate some Java code for a bit. That is, until it started giving me maven dependencies that didn't exist, and classes that didn't exist, but definitely looked like they would at first glance.
OK, wow - that example kind of perfectly proves my point. If I were to ask ChatGPT an extremely specific, low-level question about an extremely niche topic, then I would absolutely be on "high alert" that it wouldn't know the answer. And while I agree the "confidence" with which ChatGPT asserts its answers (though I'd argue the GPT-4 version does a much better job at not being over-confident than 3.5) is off-putting, I think it's pretty easy to detect where it's wrong.
I'd also be curious about your Java example. There was a good YouTube video of a guy that got ChatGPT to write a "population" game for him. In some cases on first try it would output code that had compile errors, e.g. because it had wrong versions of Python dependencies. He would just paste the errors back in to ChatGPT and ChatGPT would correct itself. Again, though, this highlights my point that I use ChatGPT as the start of my processes, a 1st draft if you will. I don't just ask it to write some code, then when I get an error throw my hands up and say "see how dumb ChatGPT is." To each their own, though.
I don't consider a popular video game from 2009 to be "extremely niche", and I also shouldn't have to know what ChatGPT knows. And no, I don't think it's easy to detect where it's wrong if you don't know the right answer, and it's actually pretty useless when you have to spend time confirming answers.
LLMs are great for “creative” work: images, poems, games - entertainment based on imaginary things.
Because of how they're used.
If you think of AI as a source of truth, obviously you're going to run into trouble: it "lies"! But if instead of thinking of it in isolation, you think of the person+AI producing results, then you should trust that person exactly as much as you would whether or not they use AI.
A tool is a tool. it has good uses and not good uses. as the human you figure out where it works and where it doesn't.
As a software engineer, I'm not concerned about people using AI to write simple functions. That's not where my value is - it absolutely incidental to me.