An "oh fuck" moment in time
ghuntley.com
ghuntley.com
It’s really easy to learn how to use software assistants. Those engineers can hop on or off any time. It’s really hard to build up fundamental engineering skills.
If there is little to no professional work in a scope, people don't get smarter, they get dumber, and it happens through simple atrophy.
Good engineering is the constant mental build up, natural tear down (atrophy) involved, where reality favors us with resiliency.
AI replacing large swathes of work will cause those neurons to not fire as often, eventually you'll lose the connections.
I don't think there is any safe amount of use that doesn't already create a persistent, self-fulfilling cycles of demand, where overuse would naturally occur just to compete in the labor market.
Its effectively a n-body problem with no visibility since the measurements would ultimately be based on value which is subjective. It collapses the value of inherent skilled labor down to 0 paving the way for slavery or death when natural dynamics collapse the economic production system.
Anyway, like any analogy, it has its limits, and trying to base the whole discussion in that analogy is not productive
I somewhat disagree with this statement. Just look in this thread and you will see people who have tried using AI assistants and got no utility out of them. Prompting is a skill. Here are some areas that a skilled user of AI needs to develop:
* What sort of questions are good for AI to answer
* What context/information is needed for an AI to give a good solution.
* Identify when the AI is hallucinating and going in loops.
* Which of the many AI tools to use in a given situation.
However I agree that engineers still need to build fundamental engineering skills. I am not sure what the next generation will do, but in general I think the kids will be alright.I jest, of course, but I have very specialized skills in an industry that requires esoteric bespoke things, so I feel less at immediate risk. That said, I should start building those skills, and I do see a lot of older engineers being left in the dust if they don’t adapt. Especially web / front end dev and dev ops, any field that has a bunch of boilerplate and repeated conventions will suffer first.
Imagine if it was driving a truck instead. Anyone, realistically, could've learnt to do it. But if you don't already know how to drive one, you're never going to get hired as one. And by the time the truck driving jobs are the only jobs there is left, it's a tad late - you're already behind everyone else who's got some years of experience.
Some of those cases are:
* Sometimes the Accept changes button doesn't work
* Sometimes it DELETES a lot of your code, without you noticing
* Sometimes it duplicates code
* Sometimes it gets stuck in a loop trying to fix an error, or it can't fix trivial errors
* Sometimes it suddenly can not access your files.
* Sometimes it adds "// ...rest of code goes here" comments instead of writing the code.
In my experience, the best use-cases for AI assistants/LLMs are:
* Generate boilerplate/scaffold for a new project
* Refactor code (rename variables, functions, or slightly alter functionality like use "while loop instead of for")
* Quick way to find the name/signature of an API/function without having to google the docs of that library (e.g. "add a new React MaterialUI list")
I have used it for about 2 months, but in my experience it actually gets worse with time, not better. The problem now doesn't seem to be with the actual writing of code, but with the UI and how it applies the changes.
that seems bad, IDEs are already quite good at these three points without AI.
Sometimes I have multiple variables, like:
var firstJob;
var lastJob;
var totalJobs;
And then, if I want to rename Job -> Task, with LLMs, I just edit first variable, and then it auto-detects my intention and suggests all those changes with a single Tab press: var firstTask;
var lastTask;
var totalTasks;
With the IDE you can only F2 to rename individual variables, or replace all occurences, which might overwrite other variable names too (also maintaining casing is tricky sometimes).LLMs are also good at predicting what you're going to type, especially with adding comments or hard-coded content.
Like, when you write
var firstTask;
var l...
Now you will get autocomplete for lastTask; in one tab, faster than typing it.Also amazing for conversion, let's say you have a YAML file and you want it converted to JSON, you just select it and ask "convert to JSON", done.
yq -o json .For example, you can just copy-paste a json snippet from the web, pasted it into the chat window and say "pass this as inline YML".
I set a far lower bar for paying for windsurf - I just wanted a smarter grep. I was blown away when I pointed it at some code and asked it questions I knew the answer to. But then I asked it questions I didn't know the answer to. I can't recall it helping in the slightest on any code base I was familiar with.
Maybe it will help writing unit tests. It sure as hell didn't help for a Python 2 to 3 conversion. 50% of it was just plain wrong, and most of the remaining 50% needed fixing. So far, for this person who has been programming for 40 years, it a net drag on productivity.
May some day need no pilots? Yes. Not today, no tomorrow.
LLMs are great at translations, like "convert this SQL to JooQ" or "convert this Chinese text to English". If you were going from Haskell -> Rust it might have trouble fighting with the borrow checker, or it might not if your memory usage patterns were straightforward or you can "RC all the things". Haskell's a particularly good target because less state means less complexity thinking about state.
It doesn't happen very often—maybe 1 out of 50 times?—but it's definitely a somewhat regular occurrence. And ofc, it's not just a result of the fact that every language is different and so the same meaning will need to be expressed differently—it's actually hallucinating bona fide content.
With a tricky copyright declaration and licensing.
A mechanical translation of Rust code under the copyright of ___?
Maybe also involving stealing bits of Haskell code without attribution from material copyrights of ___, ___, and ___?
Cleanroom engineering is a thing when it comes to "make me one like that, except I have copyright of it". The LLM laundering approach isn't that.
Ripping off open source code isn't new. Various companies and individuals would do it. Occasionally they got caught.
Microsoft/GitHub had to reassure companies that cared about LLM copyright violations, by indemnifying them. I guess so that open source projects ripped off would have to face the legal resources of Microsoft.
I can't just take something of yours, and make it public domain.
I guess I know why not: all of the major copyright holders would probably rather prefer to settle such issues in the courts and/or contracts, to get their share of the rent extraction, and AI companies also don't have much incentive to have AI-generated output be public domain either, so I'd imagine the lawmakers would try to reconcile theirs opinions somehow instead of adopting this point of view of mine (I am, after all, a random person on the Internet and have no lobbying power).
Eventually, AI (as it is today) will run out of commoditizable knowledge.
It is impressive, but there is a man inside the machine writing wrappers and talking about them to allow this magic to happen.
I've seen people dismiss Musk, saying that it's not intelligence, it was luck.
I've seen people dismiss intelligence in the general domain of "being successful", saying "grit" (or sometimes "common sense") is more important.
Before Deep Blue, Chess was seen as something that needed human intuition; after, it was "just brute force search, but Go, now that's a real sign of intelligence"… until AlphaGo, after which that too was no longer claimed.
"Kilo-girl-hours" used to be a measure of compute, back when "computer" was a profession rather than an appliance.
> It is impressive, but there is a man inside the machine writing wrappers and talking about them to allow this magic to happen.
https://en.wikipedia.org/wiki/Standing_on_the_shoulders_of_g...
And, to borrow a quote: If you immediately know the candlelight is fire, then the meal was cooked a long time ago.
By this I mean: If an AI only counts as "intelligent" when it can reproduce everything in the entire path-history leading up to modern civilisation without dependence on training data (the way AlphaZero can specifically in the domain of Go), then point at which the AI bested us is comprehensively behind us.
Consider the following thought experiment[1]: MuZero plays Donut Go, a variant of Go where a hole has been punched out of the middle of the board, but has not been retrained. Lee Sedol and a competent amateur would both be able to play competent amateur Donut Go by qualitatively adjusting their strategies to the new board in a "zero-shot" fashion: perhaps Lee would do much better, and the amateur makes a few boneheaded mistakes, but both would be competent.
MuZero would not be competent, it would make severe blunders over and over again because it's not capable of adjusting its strategy to the donut hole. It would need to be retrained. There's a categorical difference in flexibility, which is why the competent amateur is intelligent and MuZero is not.
[1] I am not sure if they've done a Donut Go experiment, but DeepMind did show that the same algorithm trained on Breakout could solve Level 1 with superhuman skill yet completely fail level 2. I am fairly confident my thought experiment holds water, but it would be interest to run it IRL. There is an issue in that I'm not sure how you would get MuZero to "see" the hole... if it has a deterministic check for move legality then perhaps that could be updated.
To be equivalent to a frozen/no retraining allowed AI, the human would need to have their brain chemistry prevented from altering synaptic weights.
The option to freeze weights in AI models is the only way they can be safely productised. Or training runs fairly compared — didn't Kasparov allege that IBM cheated by changing the algorithm during the game?
When it comes to "what even is this 'intelligence' thing?" there's definitely an argument about how much advantage these models take from the fact that silicon is around a million times faster than synaptic chemistry; but in practical terms, when it comes to "are we nearly obsolete?", it's the wall-clock time not the subjective time that matters — and AlphaZero took only 8 hours to beat its predecessor AlphaGo Zero, which itself beat AlphaGo Master, which in turn won 60:0 against professional players. So, how long would it take a human to get good at the new rules, and would it improve faster or slower than the wall-clock time of an AI learning the new constraints?
> I am not sure if they've done a Donut Go experiment, but DeepMind did show that the same algorithm trained on Breakout could solve Level 1 with superhuman skill yet completely fail level 2. I am fairly confident my thought experiment holds water, but it would be interest to run it IRL. There is an issue in that I'm not sure how you would get MuZero to "see" the hole... if it has a deterministic check for move legality then perhaps that could be updated.
Should be easy, given that MuZero did go, chess, shogi, and a standard suite of Atari games…
But yeah, 2019 is an eternity in AI, should be easy to do what you suggest today.
Instead of a small dude inside the machine, there is training data.
Instead of pulls and levers connected to chess pieces, we have vectors and weights.
If we scaled the turk to have millions of people inside, it would appear to be intelligent as well.
I don't need it to reproduce everything. Just show me one convincing non-anectodal evidence that this is not another mechanical turk.
And your brain does it with chemistry. The mechanism isn't important.
> I don't need it to reproduce everything. Just show me one convincing non-anectodal evidence that this is not another mechanical turk.
Define your test.
Every single test that existed when I was a kid has been bested.
All the new tests from all the domain experts to measure AI, are getting rapidly saturated.
https://ourworldindata.org/grapher/test-scores-ai-capabiliti...
Coming up with a test is itself a challenge, these days.
Alledgedly.
> Define your test.
A pass is a model with no human training data input that is able to reason on human content.
> Every single test that existed when I was a kid has been bested.
The mechanical turk fooled humans for 80 years and passed all the tests at its time. As it turns out, the tests were not good enough.
> Coming up with a test is itself a challenge, these days.
My test is simple to understand and conceive. We just don't have any valid candidate to put through it. My requirements disqualify LLMs.
If you can't agree on our brains being chemical in nature, you're engaged in magical thinking and we have no more to discuss.
> A pass is a model with no human training data input that is able to reason on human content.
What does it mean to "reason on human content"?
Both "reason" and separately "human content"?
The former because some argue that "reason" is the mechanical application of logic, where a single transistor alone would be sufficient to pass; others require it to be a conscious application, at which point you've got to deal with there being about 40 different definitions of "consciousness".
The latter, because a model with no human training data will not be able to read any human language for the same reason that no human alive can read Linear A… but also, our content includes the games we play, and AI trained on the rules of Go with zero examples of our games beat human Go players at the game of Go, which definitely seems to me to be doing things on "human content".
> My test is simple to understand and conceive. We just don't have any valid candidate to put through it. My requirements disqualify LLMs.
Where did you write this "simple to understand and conceive" test? Because what you've given here is not even a test.
If you can't come up with a test, there's nothing to distinguish you personally from the AI that you're dismissing.
I agree that our brains have a lot of chemistry. I can't deny that. Science shows we can mess up brains by messing up their chemistry.
So far, though, we still don't understand how it works. We can't use chemistry to improve it, just to break it or fix it. That alone should be a huge red flag for the reductionist point of view.
It's our best hypothesis, but is not good enough to travel disciplines to IT. We don't know enough to compare the two, hence the skepticism.
> What does it mean to "reason on human content"?
Pass the turing test should be enough. I'm just changing the requirements (no tiny humans inside, complete or otherwise).
> no human alive can read Linear A
If we had a living human fluent in Linear A, I believe we could decipher it. Machines have us, living humans speaking a living language, to learn.
> If you can't come up with a test, there's nothing to distinguish you personally from the AI that you're dismissing.
I can. Open the machine and prove me that there are no tiny humans or human pieces inside.
Caffeine and theobromine are both commonplace cognitive enhancers, AKA nootropics. There's plenty that aren't available over the counter, too, like amphetamines.
> Pass the turing test should be enough.
LLMs already do that, and it's a problem because it's a fraud opportunity.
> If we had a living human fluent in Linear A, I believe we could decipher it. Machines have us, living humans speaking a living language, to learn.
How is this supposed to be different from what LLMs already do?
> I can. Open the machine and prove me that there are no tiny humans or human pieces inside.
That reads like an LLM — and not even a good one, I mean like the GPT-2 based AIDungeon — wrote it, given there's obviously no tiny human inside a computer.
What did you actually mean? Be specific, be testable.
We still don't fully understand how they work. Of course, we didn't make them, nature did.
> amphetamines
Those break more than enhance.
> LLMs already do that
As I said, disqualified for being a cauldron of souls.
> How is this supposed to be different from what LLMs already do?
My requirements enforce a no human content during training. It's exactly that requirement that makes the test easy. If you allow human content to be part of it, then tests become ellusive.
> What did you actually mean? Be specific, be testable.
I'm saying that LLMs are the spiritual successor to the mechanical turk. The entire concept falls apart if you take the human part of it out.
There are living humans that had their works put inside and are publicly saying "I don't want to be inside this machine". What more do you want?
Yes, they are impressive. I already said that. It's just not intelligence, it's an almost mythological trick we should have been more guarded against.
Would you like help moving those goal-posts for your supposedly easy to understand test?
This claim:
> My requirements enforce a no human content during training. It's exactly that requirement that makes the test easy. If you allow human content to be part of it, then tests become ellusive.
Is directly incompatible with this one:
> If we had a living human fluent in Linear A, I believe we could decipher it. Machines have us, living humans speaking a living language, to learn
If there's no human content allowed, then we here today can't use a living human fluent in Linear A to decipher it. Nor, for that matter, even so much as the mere historical record of Linear A without any human.
You are demanding AI pass a standard that literally no human would ever be capable of, because if I simply switch the nouns you get: "A pass is a human with no Linear A training data input that is able to reason on Linear A content".
> There are living humans that had their works put inside and are publicly saying "I don't want to be inside this machine". What more do you want?
I want you to be specific and to be testable. You still have a test which is… have you ever had a client who wanted you to build "Uber for aircraft" or something else equally vague? They didn't even realise they didn't have a plan. That's what I see from you. Your position is pre-paradigmatic.
Consider for example: a human who grows up with no human works put inside their head is what is called a "feral child" — they do not magically understand anything of the rest of civilisation, this is because that understanding has to be taught by… observing humans as they talk to each other, mimicking them, trying to talk with them to get feedback, etc., which is what the LLMs are all doing.
It looks like a double standard to me, that you demanding that an AI, who is raised by actual wolves does better than a human who is raised by actual wolves.
What you have actually said so far, would refuse to recognise as intelligent an AI that is given control of an android body and had to read books by physically holding them and turning the pages with robot hands, which it operates in human timescales, and which gets there in the same timeframe as an actual human.
It would also rule out an actual human, because we learn from the works of other humans.
No. I'm estabilishing a previous important goal-post: no machine turks. It's the current industry that laxed the requirements, not me that got more strict.
You don't seem to understand the difference between training data and context. Read about it, and re-read what I said with that concept in mind. All your answers are already there.
What you demand for AI, is equivalent to denying humans with the requirement: "A pass is a human with no Linear A training data input that is able to reason on Linear A content".
YOU came up with the idea of tests, and I tried to entertain you with some possibilities so you can understand better my point of view.
YOU asked an impossible thing. You asked me to devise a test that proves that a fraudulent claim (LLMs have intelligence) is truthful. The best I can do is show you how silly that claim is.
I actually don't need to prove anything. I just disproved something on a fundamental, ontological level. You're trying to make me go back on that ontodefinition or make a contradiction, which is not going to work.
OK, then humans fail hard. A baby raised with "no human training data" (no conversation, no active teaching, no social interaction, obviously no books or videos) would be a total imbecile... or dead.
LLMs have training data and context. I'm talking about training data.
Conversation, interaction... those happen within an LLM context window. LLMs learn nothing from those experiences. What they can learn within a context window is super limited and disposable.
You can't really compare the two. Humans are much, much more.
How that input is used isn't relevant. Humans don't have separate training and inference training phases, but that does not mean that a human doesn't have to learn from other humans. Including just plain ingesting streams of stuff without interaction, although that's not the dominant mode for humans.
By the way, any LLM you have ever interacted with has undergone reinforcement post-training in which it did in fact converse. The results of that are baked into the weights, not in the context. Not that it matters for the unreasonableness of your "test". And I predict that within a couple of years the training-versus-inference distinction in LLMs will start to break down. Which still doesn't matter when you're evaluating their current performance.
False.
Why do you think e.g. ChatGPT has buttons for indicating if an answer is good or bad?
Even when a specific version is frozen, this is a choice rather than a fundamental requirement, and the data from human feedback is used to continue training the system as a whole.
Can we at least agree that this cannot be compared with a human baby?
You are indeed born being fed a firehose of experience. Is it all the content available in the world? No, it's much less ordered than that.
Also much less content. If your point were to have been "current AI are thick because they need a lot of examples to learn anything", you'd have a much easier time with your claims. (Some people still argue that point because vision is high bandwidth, but nevertheless it would be easier to argue than setting the bar so high that no human could possibly get there).
> Then, when you need to learn something, you just firehose it all again from scratch with some added notes you made on a napkin.
I have no idea what this is supposed to mean in this analogy.
> Can we at least agree that this cannot be compared with a human baby?
https://pmc.ncbi.nlm.nih.gov/articles/PMC27565/
You can compare almost anything, if you want to.
But why would it matter, that it's not like a baby? A baby treated the way you propose treating an AI, would still be somewhere between "feral" and "dead". That, even in isolation, is enough to show that your test is not well-conceived.
So, you're a self-declared sophist. We can't get anywhere from here, I am not.
And then ask yourself what point you're even trying to make, because without clarity on that, my point is that your point is unclear.
AI are like brains in certain ways, by design even, and different in others. You have yet to provide any suggestion for which specific ways to test for that simultaneously pass humans and fail AI.
For anything more complex or esoteric, LLMs shit the bed pretty quickly. A recent example: I needed to implement a particular setup for an Argo CD ApplicationSet (default values plus a per-stage override, at the generator level). Every single suggestion by the LLM was wrong. I decided to dig into the documentation because I assumed there was no way this wasn't possible, and lo and behold - matrix generators are a thing, and my exact use case was an example in the doc. I asked the LLM about it and it knew what they were, so it clearly had indexed it before.
It's far from the only issue - the most common one is the below-average quality of code in some cases, with plenty of wrong implementations strictly due to hallucinations. Then Cody has its own issues on top of that - not applying changes correctly is the most common one. Put that together with engineers that are in a rush or are using the LLM as a crutch... and I can see a future where devs with little experience can put together huge codebases that will be as brittle as a house of cards.
…
> This wasn't something that existed, it wasn't regurgitating knowledge from Stackoverflow. It was inventing/creating something new.
But it was something that existed… it took a rust library and converted it to haskell. That’s not invention, that’s conversion
We may very well go the way of the professions of "computer", burlak, lamplighter, postilion, samurai, seneschal and reeve: https://en.wikipedia.org/wiki/List_of_obsolete_occupations
I won't bet on how quickly or slowly that will change, in part because I want it to be a slow change so I don't have to worry about learning a replacement trade.
This can definitely help with some of the tech debts. But I wonder how do we deal with changing requirements? Use Cases? Experiments & new directions from Product? Software has a life after building.
Why leave out an explanation of why you needed, or how you’ll use, this Haskell audio api wrapper?
In the absence of these this just feels like resume slop.
If it is, do recruiters see “I asked a thing to do the work for me” as desirable in a potential hire?
The original library’s license is Apache 2.0. Does that have any bearing on the authors opinion on sharing this generated code?
I've started to think of them merely as smarter auto-corrects, like for when you forget the syntax or which order a functions arguments go in, you can just use copilot's inline chat to write "Convert this array of tuples to a dict" and it saves you the 20 seconds it takes to look up the syntax. Or "See above line where I deserialized a uint32 from a binary file. Do the same thing for all the other fields in the foobar struct".
I just never even thought to ask it to simply build my whole app for me.
But I do agree that the code is pretty crap in most cases. I feel like if I was seriously doing this then I'd be mostly refactoring the generated code and writing tests.
I'm looking at coders coming out of University who chat-GPT'd every coursework and they can't program their way out of a paper bag. I think what is being lost is the ability to solve non-chat-GPT-able problems. I'm having my 'oh fuck' moment, but I'm looking at the new coders.
Maybe I am just super behind the progress train, and that programmers will become a thing of the past. Maybe we can setup some adversarial LLM training scenario and have LLMs that reliably code better than humans. My concern remains valid though, that there is now a generation being created that turns to an LLM rather than reasoning about a problem [1].
[1] https://www.reddit.com/r/ChatGPT/comments/1hun3e4/my_little_...
I've been using Cursor since the summer and wrote up a few tips here https://betaacid.co/blog/cursor-dethrones-copilot
In my experiments with LLM's one of the most obvious issues is project structure and organization, particularly in the context of future objectives. It is way too easy to have an LLM go in circles when analyzing ideas for project structures. I have seen them propose layouts that, while workable, are far from ideal.
The same applies to code. It can write usable code faster than a human. However, and, again, in my experience, you are constantly refactoring the results.
In all, this is a new tool and, yes, it is useful. It's amazing to me considering the idea that the whole thing is based on guessing the next token. It doesn't really understand anything, yet, you'd be hard-pressed to prove this to those who don't know how we got here. I've had a few of those conversations with non-coders using ChatGPT for a range of tasks. It's magic.
If I want to invest time in fixing some junior developer's probably broken code, I'd much rather do it with a real junior developer who has the potential to learn over time and become a not junior developer
I'll give you an example of something that I moved to LLM code generation.
I have used Excel for probably more than two decades to generate certain types of code structures that are simply tedious to both write by hand and also maintain. One such example are lookup table definitions for state machines. Using Excel for this, once setup, makes it fast and infinitely more maintainable.
Today, I can do this with LLM's in a heartbeat. The code is good and it is easy to maintain.
Over the years I have written custom code generation tools for the same reason. Sticking with the state machine example, if you have 100 states and each state requires entry, during and exit handler functions, well, you have to take that lookup table and write 300 functions. In a language like C this means both the declaration and implementation skeleton. LLMs are good for this as well.
Like I said, it's a tool in the toolbox. It can be used a million different ways.
As I said, this topic is likely to remain controversial for some time. I have a feeling that a lot of people are reacting to it rather than actively using and understanding where we are today.
I have been using it actively and across a wide range of application domains (not just coding) for probably somewhere around a year. This use has been under a "trust but verify" modality, being careful and skeptical. And, yes, I have come across issues here and there. And yet, in the aggregate, it would be impossible to envision not using this tool. It is most-definitely a potentially huge productivity and time-to-market booster. And it will only get better.
I believe in this to the extent that we decided to pivot a hardware-centric startup we've been pursuing to make it software-centric with AI agents at the core of the offering. These tools will change the way we conduct business, write software, design hardware, conduct research, administer healthcare and more. They might not be ideal today. These are just babies learning to walk.
My impression so far is that they're already in diminishing returns territory. They might just be puppies learning to walk, in the sense of having a fundamental limit to their potential
Well, the demand from every potential customer we have spoken to is very real. And the pain point we have identified cannot be addressed without using AI. Now it is just a matter of getting to a reasonable MVP. We already have customers who want to be in the beta program.
Here is a logic puzzle that I need some help solving: Samantha is a girl and has two brothers and four sisters. Alex is a man and also one of Samantha's brothers. How many brothers and sisters does Alex have? Assume that Samantha and Alex share all siblings.
And I get back a very well written, multi-step response that leaves no doubt in anyones mind that: To solve this logic puzzle:
Samantha has 2 brothers and 4 sisters.
This means there are 7 children in total (Samantha, her 2 brothers, and her 4 sisters).
Alex is one of Samantha's brothers. Since Samantha and Alex share all siblings, Alex has:
1 brother (the other brother besides himself).
4 sisters.
Final Answer:
Alex has 1 brother and 4 sisters.
Maybe it's like with Apple and I am using it wrong.To get back to the "intern"-comparison. I could usually tell when an intern was struggling, there just were human telltale signs. When AI is wrong, it still presents its results with the confidence of someone who is extremely deep in the Dunning-Kruger hole but can still write like a year-long expect on the topic.
What I've learned is that it's good primarily for tasks like the following:
- Tasks which take time to do, but are then easy to verify.
- Tasks which effectively boil down to translating something from one format to another. Which might e.g. be "read this technical document and implement it in code, as for style, look at these sample code files as a reference”.
- Tasks which are about exploring unknown unknowns. E.g. I write down a design, and then I ask the AI to roast it. The point is not that all the points it'll make are good and I need to please the AI, it's that out of 20 points it will list, 2-3 might both make sense, and haven't been thought of by myself.
Finally, AI requires good writing skills, and asking questions in an unbiased way, otherwise the AI will gladly hallucinate to reinforce your bias.
Logic exercises which are easy to verify are a moderately good fit for "reasoning models" which will go through many iterations of an LLM and basically write out the whole reasoning process. In practice though, this can be very expensive to get good results with.
Your takeaways are good and fit that model
> I read all of these amazing things that they supposedly all can do.
You seem to be implying people are confused (or lying?) about the things they are able to get LLMs to do.
If you give it an honest effort to solve some real problems you are facing then you may be able to speak with more authority. Often it comes down to prompting skill. Try to read about different prompting approaches as that may help you.
In general, you need to be specific about what you need, and you need to give all relevant details. Like the post author said, treat it like a junior programmer or an intern.
What you call "gotcha word problem", I'd compare to typical math problems where you need to understand a text, extract the required information, solve the issue, and then present your results. Maybe this is a toy-example, but compared to reading the specs of some Microprocessors, this is rather easy. These AIs seem apparently be able to solve school or even college level math problems. Shouldn't my example be a walk in the park, then? Especially since it's a large LANGUAGE model?
> You seem to be implying people are confused (or lying?) about the things they are able to get LLMs to do.
I am merely stating observations and was hoping for an explanation. What good does it me if I accuse people of lying?
> Often it comes down to prompting skill. Try to read about different prompting approaches as that may help you.
"You are using it wrong" it is, then. So how do I differentiate between a good sounding but wrong answer, whether that came to be due to my apparently lack of prompting skills or else? They all sound equally well, it just starts "being wrong" at some point.
> In general, you need to be specific about what you need, and you need to give all relevant details.
What details should I have added in the given example? The prompt was probably more comprehensive and detailed than if this task was given in primary school.
> Like the post author said, treat it like a junior programmer or an intern.
I would, if it acted like a junior programmer or like an intern. For them, you can usually see if they are unsure or making things up (if they do these things). For an AI I've yet to see something like "hey, I might be wrong about this, but this is my best effort, maybe we can have a look together."
Let me help you solve this step by step.
First, let's identify what we know about Samantha:
Samantha has 2 brothers
Samantha has 4 sisters
We also know that:
Alex is one of Samantha's brothers
Alex and Samantha share all siblings
Now, from Alex's perspective:
Alex is one of Samantha's brothers, so he has 1 other brother (since Samantha has 2 brothers total)
Alex has Samantha as a sister, plus her other 4 sisters
So Alex has 5 sisters total (Samantha + her 4 sisters)
And Alex has 1 brother (the other brother besides himself)
Therefore, Alex has:
1 brother
5 sisterswhy not ask it to produce said program, and then evaluate it on that output, rather than proxy it via asking a logic puzzle?
It's like doing a coding interview when hiring an employee with these brain teaser puzzles, and when they fail you disqualify them, rather than asking them to do a real task that would be something they'd encounter on the job.
The back and forth wasn’t fun and it flat out refused to use seaborn for some reason, but it worked and was fine overall.
I then used aider+claude to help me work with yjs. Led me down a rabbit hole based on an incorrect description of the yjs sync protocol. Took 2 days to untangle everything. Yjs is fairly new though, so I didn’t fault it too much.
I thin tried using it for work to deal with some surprisingly intricate back button logic. Again, incorrect understanding (on both our parts) of the underlying API caused a few days of headache. I would’ve been better off just reading the docs than trying to use an AI assistant.
Using AI actually frustrated me to the point where it convinced me to suck it up and just read the Specs and sources of the tools I’m using. I’ve been doing that for a few months now and just RTFM is better for me than AI assistants have been.
By the way, o1 has no trouble with this problem: https://chatgpt.com/share/6785429c-d1d4-800c-8209-02c542468a...
I do this every time I'm in a different city. Most of the data that I use is from the last three years.
ChatGPT o1 got the answer correct with no tweaks to the prompt.
o1 does fine with that btw. https://chatgpt.com/share/6785626a-6ea8-8009-83ab-673433465c...
Do you search Google or Reddit and wish you could just 'get the answer' instead of wading into pages/posts?
Do you compare two long documents together and not want to invest a few hours into a close reading of them?
Do you write code that consists of trivial functions or trivial text manipulation?
Do you want a 3-hour podcast summarized into a few bullet points for a particular audience?
Do you want to send a saucy limerick to your friend on their birthday?
Do you want to compare Kant's view on <topic> with <new_metaphysical_school_of_thought>?
Do you want to analyze 250k rows in an Excel file of user support tickets and summarize the top issues?
etc etc etc.
Totally fine if you don't do any of these things, but these are the things most people are using LLMs for.
Well, yeah. You are. It's built to answer questions people actually ask, not solve new logic puzzles.
Despite what the marketing says, it's not a perfect-infinite-knowledge oracle. You should think of it more as a really, really big database with all of the Internet's "knowledge". When you ask it "2 + 2 = ?", it isn't parsing those into numbers and math operations, it's searching its database for occurrences on the Internet where someone answered the question "2 + 2 = ?" and filling in the closest answer it found. If you ask it what "120938120938120931 + 1209389120381208390" is, it'll probably get it wrong, because no one has asked that before. But you should probably be using a calculator instead.
If you ask it something it hasn't seen before such as your logic puzzle, it's not going to parse it like a person would and synthesize an answer. It's going to try to find something similar to what it's seen before and return that. Odds are good this will be a wrong answer, since it's not addressing what you actually asked.
However, if you ask it something it has seen before, like a programming problem, it will return something appropriate. It turns out the Internet is pretty big, so it has seen a lot of stuff, and so often works pretty well. Hence the success you're seeing from others who are using it as-intended, i.e., asking it real questions, not logic puzzles.
It's not very hard to come up with a scenario that has never been put on the Internet, so it's pretty easy to make it dig up the "wrong" answer and do something stupid, as you've found.
The real trouble is that it can't tell you whether it's guessing, or found an actual match. Hence the "confidently wrong" thing, which absolutely destroys user trust. If it's confidently wrong about this thing I know a lot about, how can I trust it to be accurate for something I know little about?
Isn't the causality inverted here? It's trained on questions people have asked before, so that's what it's better at. New logic puzzles illustrate this flaw
> ”First, list out Samantha’s siblings explicitly: • Brothers (2 total): Alex + one other brother • Sisters (4 total): Samantha + three other sisters
Since Alex shares all siblings with Samantha, let’s see it from Alex’s perspective. Alex himself is one of the 2 brothers. Therefore, from Alex’s point of view: • He has 1 other brother (the second brother besides himself). • He has 4 sisters (including Samantha).
Thus, Alex has 1 brother and 4 sisters.”
This just moves the goalpost. I’m sure someone can give the next example where it fails. I find it useless as well, but at the same time it really feels like criticizing a talking dog about their lack of understanding.
It's non-deterministic, everyone gets different responses
Recently though, I've had a different "oh fuck" moment: what if the period of general global prosperity is coming to an end, that Fukuyama's end of history is coming to an end, that climate change won't come under control. Maybe I'll look back and wish that I could be a redundant SWE in a 2010s economy. With such a backdrop, how scary could AI really be?
All signs point to this, or it's already reached a tipping point - certainly nothing we are already doing is reducing total carbon output, just slowing its acceleration. There was that one weird year in 2020 where emissions went down as the entire world stopped commuting - regardless of what you think about the realistic ability of new/unproven tech and regulations will be at preventing 3+ celsius change from industrial average by 2100 (catastrophe of biblical proportions) - if AI is indeed as critical to the economy/humanity's future as its proponents (marketers) say, the enormous amount of power and other resources that often don't get factored in, like the massive amounts of increasingly scarce water to cool the massive datacenters - will certainly make such efforts at reducing emissions extraordinarily unlikely, if not impossible, short of converting our entire power grid to sustainable energy, which there seems to be little to no political will for. It's grim, I'm extremely skeptical of optimism at this point on this particular topic. I am bracing for catastrophe at some point in my life (looking at the LA fires very near me, I would argue it's already here) and think the world governments would be wise to do so as well rather than pretend we're just going to magically figure it out at the last possible second. Not even to mention there are a lot of people, including people in power, who don't even believe climate change is happening at all.
We're living in Don't Look Up.
On a personal level, sure.
On a society-wide level…
Even if the AI doesn't improve, consider that there are people who think the answers to every question can be found in the actual literal Bible, and then ask what such a personality would do when there's a small box that can opine confidently about basically everything and is right often enough to be competitive with people who stay in full time education until age 19 or 20.
Or, if the Bible analogy doesn't appeal, consider how it took until the Renaissance for European society to seriously attempt to go past that of classical Rome and Greece.
But speed doesn’t matter. When code is going to live for several years, no one will care if it took minutes to create vs a week.
If you are an incompetent human, you might never be able to generate the same code some LLM generates. This will cause you to sing it’s praises and worship the LLM as your savior.
In general, there is a correlation between weak developers and how hyped they are by LLM coding assistants.
No, but... if it took minutes instead of a week, that leaves you the rest of the week to create N more things. That matters, and may still matter in several years.
That said, forgetting all of the AI doom or hype, this seems frankly incredibly cool. I haven't managed to find a great use case for AI code generation where it hasn't failed miserably (except for minor boilerplate uses, like with IntelliJ's LLM auto complete for Java), but actually, generating FFI libraries seems like a pretty kickass use case for automation. I dig this. I will definitely try it.
Maybe not writing the code itself, but creating a product/business still requires deciding how things will work, what technologies to use, what's the public API, etc.
So instead of writing the code to do it, you still need someone to ask the AI to do it: "create a to-do list with login system, that has e2e encrpytion and stores the data in a NoSQL database", etc. So you still need some programming knowledge to use it, otherwise you don't know what to ask for.
there's a perfectly working LLM for which you can query the answer! Ask it to tell you what are the important aspects of product XYZ.
Also, this is an incredible achievement, but I don't understand this statement:
> This wasn't something that existed, it wasn't regurgitating knowledge from Stackoverflow. It was inventing/creating something new.
I mean you literally told it to copy an existing library. I understand the Haskell code is novel, but translation is not the same as inventing.
If it was real, why aren't all the unicorns in this space worth $100B already?
It could be. It also depends on whether there is a moat around the tech or if we'll have a dozen fungible clones.
But I'd certainly expect more signal if these models could do as advertised.
This simply is not the case as far as we know.. I'm yet to be convinced that anything under the hood is more than a clever predictive generator for now.
It can appear to do really impressive things, but that's only while it's behaving correctly. All it takes is one bad line and the snowball effect of the problem is in motion.
Lastly, it is true :
> "the better the type system of the programming language and verbosity of compiler errors the harder LLM's go"
The LLM has the best opportunity to succeed in this case. Nice to have but inventing, I doubt.
Mine was InstructGPT, a few months before the public release of ChatGPT. Everything that followed was an unsurprising extrapolation of what I saw in that moment.
Question is, how far will it continue? That I cannot say. At least a factor of 3 from foreseeable compute improvements… unless a war is started over the manufacture of compute. But also, Koomey's law (energy) is slower than Moore's (unit cost + feature size), and if one presupposes that manufacture of compute is unbounded, then we rapidly hit limits of manufacture of power sources for the compute, and that in turn causes huge economic issues well before there's enough supply to cause mass unemployment… unless there's a lot of huge algorithmic improvements waiting to happen, because that might be enough to skip electricity shortages and go straight to "minimum and cheapest calories to keep human alive are more expensive than the electricity to run the servers that produce economic value equal to the most productive human".
I have no confidence in any forecast past the end of 2028.
https://gwern.net/creative-benchmark
I can't wait to see if AI can get past its centripetal plateau.
> convert this rust library at https://github.com/RustAudio/cpal/tree/master to a haskell library using ghc2024.
The end goal is something more like
> Build me a website where users can make and share a link to a text document like a shopping list and edit it collaboratively
This guy thinks LLMs are so great, but if he grew up with LLMs, not only would he not have the skill set to competently ask the LLM to produce that output, he wouldn't have any idea the output he's trying to produce.
OpenAI, Google, Microsoft are not trying to sell the idea of helping out your senior engineers. They are trying to sell the idea that any of your interns could be senior engineers, so that you no longer have to pay for expensive senior engineers.
Between a human and a current-day LLM that would take a day, tops, to implement.
Assistants are so easy to use yet we should drop everything to "learn"/"adopt" them? Why not just wait until they're not complete dumpster fires then start using them?
Maybe it is far enough along that we can at least see what the final form would look like and start getting comfortable with working inside those constraints? I don’t know.
Except it was. It wasn't regurgitating the exact answer it gave you, but it was using the collective knowledge of StackOverflow/other code to generate the solution.
Give your head a shake.
Fundamentally I only know about a language's functions/keywords/types/... because I've read them in some documentation or code snippet somewhere, but if I combine them in a new way to solve some previously-unseen task, I'd consider that to be creating something new.
Anything I've ever done at work could have been pieced together from StackOverflow. Anything you have ever done can most probably be pieced together from StackOverflow as well.
Also, I never said it was a "downside". I'm simply stating that the author is exaggerating what the LLM is doing. It isn't creating novel code.
I used LLMs all the time to spit out boilerplate when it makes sense.
Review: great fun story, I agree with the first part of the take and I like the post and I like some ai tools, I disagree with the last bit. There are really good SWE’s who are in spaces where ai is not that useful. SWE is too big now, it’s too vague, ai doesn’t apply to all. You don’t need to use ai to survive, but I agree you should try it out and see before dismissing it. And retry it every so often cuz it’s developing and changing.