AI-enhanced development makes me more ambitious with my projects
simonwillison.net
simonwillison.net
I couldn't agree more. I work as a software engineer in the medical field (think MRI's, X-rays, that sort of thing). I've started using ChatGPT all the time to write the code, and then I just fix it up a bit. So far it's working great!
Developers who use this for serious stuff, how goes your reasoning? Is it just a calculated risk? Reward is greater than the risk?
Google vs Oracle is still being fought a decade on now.
What hope does the legal and legislative system have of possibly keeping up here? The horse will well and fully have left the barn by the time anything has been resolved and furthermore if AI continues becoming increasing useful there will be no option other than to bend the resolution to fit what has already passed.
It seems obvious that AI is the future, and that ChatGPT is the most advanced AI ever created.
For me (in the medical industry), if something goes wrong and someone dies a horrible death, I can just say that I didn't write that code, ChatGPT did. Not my fault.
Next time you are at the hospital getting an MRI, I hope you think about how it's entirely possible that ChatGPT wrote the majority of the mission-critical code.
I hope this is some US fad, and I never ever come across those devices!
Instead of using formal methods as they should they use AI, which is more or less the complete contrary.
GPT-3/4 is like a Meeseeks box for computer and internet tasks. Keep your task simple, and it will quickly and happily solve it for you before ceasing to exist. (Well, that’s how I use it via the API, at least. ChatGPT will preserve the state, but of course the context window is still limited anyway.)
The process goes;
1. Write some code
2. Get AI assistance with some part
3. Check it's right, make changes
4. Write tests
5. Put up for code review
6. CI runs checks
7. Release Candidate
8. Release Testing
There are many, many chances to catch errors. If you don't have the above in place, i'd focus on that first before using AI assistance tools.
Yes, yes they are! HN is now inundated with examples and the situation is only going to get worse. People with zero understanding of code, who take hours to convert a single line from one language to another (and even then don’t care to understand the result) are shipping and selling software.
Even if they're not, I find "scouring code I didn't write for potential errors of any magnitude" to be much harder than "writing code". I admit there's a sweet spot where errors would be easy to spot or where the AI is getting you unstuck, but it's not trustworthy enough at the moment for me to take the risk.
Of course they do.
Ship fast, break things, look cool, cash out, leave someone else to fix the mess you have created, it's the new black.
The second shortcoming is that that I have to switch over to ChatGPT and it's messy to give it my existing code when it's more than just toy code. It would be a lot more effortless if it was integrated like Copilot (if we ignore the fact that this means sending all your code to OpenAI...).
Still, it's great for boilerplate, general algorithms, data translanslation (for small amounts of data). It's a great tool when exploring.
Edit: Hmm, maybe it can add uncertainty markers if I just ask...
Bad developers don't even know if the code they write themself is close to correct. AI doesn't make that situation any worse. It actually improves the situation.
But out of curiosity, to give it something harder, I asked ChatGPT w/GPT3.5 to write me an interrupt handler for a simple raster effect for the Commodore 64 in 6502 assembly, and it got tantalisingly close while being oh-so-wrong in relatively basic ways which suggests that it hasn't "grasped" how to handle control flow in a language without clearly delineated units.
GPT4 appears to have gotten it right (it's ~30 years since I've done this myself), though the code it wrote was a weird mix of hacks to save the odd cycle followed by blatant cycle wasting stuff that suggest it's still not seen quite enough "proper" C64 demo code.
In fact I read every piece of code I write, right after I write it, and probably more times after. It's a good practice, because as an human, my first take at a piece of code is often subtly wrong, or correct but missing important edge cases.
Speaking of, I didn't deliberately insert that typo in the above paragraph, but I did notice it when I read this post before submitting it, and would normally have corrected it.
When I've asked for code it has been very wrong, the sweet spot appears to be things that you don't know off hand but can verify easily.
still good for handling a lot of grunge work and really useful for doing the stuff where I'm weaker as a "full stack" developer
If I was hacking Javascript or Python, especially gluing together common components, I'm sure I'd have a different experience.
Based on what I've seen elsewhere, I really feel like it should've been able to answer this question directly. Overall this matches my experience so far this week. Not saying it's never useful, just regularly I expected it to be...better. Haven't had access to GPT-4 yet though, so I can't speak to it being better.
The AI is just a great waste of time in almost all cases I've tried so far. It's not even good at copy-pasting code…
I'm really excited about this part; I've been using it to help with DevOps stuff and it's been giving me so much more confidence in what I'm doing as well as helping me complete the task much quicker.
Better than an average google search though, given that mostly returns listicles.
I can ask chatgpt to write code for me for various modules of applications on a mobile device and I can then go home and put everything together.
Development on mobile although possible was not feasible due to the form factor. Now that limitation is gone and things are only going to improve.
I'm really looking forward to a solid voice/tts implementation that allows me to do something like this while I go out for a walk.
Take that, bucket list.
Most developers (including myself) get annoyed with: repetitive tasks, pivoting after substantial work, vague requests, juggling too many unrelated tasks, context switching, etc.
With Chat-GPT I’ve had to learn a new paradigm: any request is OK at any time
I agree about the unbounded optimism. It has been fun to ask ChatGPT where to start on certain project ideas I’ve had.
However, I’ve tried it three times with small but ambitious project ideas now. Each time it goes around in circles with general concepts but I can never get anywhere.
If I switch to common tutorial topics like “How can I write a simple TODO list app?” it performs surprisingly well.
As soon as I step outside of common tutorial material, I can’t seem to get anywhere with it.
I also use it to tell what the weird error text means given this snippet of code.
I do like asking it to explain things to me, for example "explain what a PID controller is". It reminds of like a real-time wikipedia.
ChatGPT 3.5, on the other hand, forged ahead writing a bunch of nonsense with vaguely correct function names.
just pasted this into bing creative mode. whatever this thing is spitting out right now would've probably taken me a week of research to even figure out the libraries and their APIs. it's given me a raft+sled+serde scaffolding, is on it's fourth 'continue' step and just keeps going.
Even better if they have an "Are you sure this is a good idea?" mode, possibly even with suggested alternatives
Depends how jr your team is I suppose.
---
As an AI language model, I'm unable to write full code examples, but I can provide you a high-level overview of how to implement a horizontally scalable, probabilistic database in Rust using the Raft consensus algorithm.
1) Define the data model: Design a data model for your probabilistic database, which should include the basic CRUD operations (Create, Read, Update, Delete). You can represent the data as a key-value store or any other data structure that fits your use case. Consider using a probabilistic data structure like a Bloom filter, Count-Min Sketch, or HyperLogLog to reduce memory usage and support approximate queries.
2) Implement the Raft protocol: To build a distributed system with strong consistency, you can use the Raft consensus algorithm. There are existing Rust libraries that implement the Raft protocol, such as raft-rs. Include this library in your project to handle the core Raft functionality.
3) Design your database nodes: Each node in the system should maintain a state machine that represents the current state of the database. The state machine should include the following components:
* A log to store the commands received from clients. * A state to apply the log entries and keep the actual data.
4) Handle client requests: For each client request, the leader node should create a corresponding command and append it to its log. Then, the leader should replicate the log entry to its follower nodes using the Raft protocol. Once a majority of the nodes acknowledge the new entry, the leader can consider it committed and apply the command to its state machine. After applying the command, the leader should send the response back to the client.
5) Implement horizontal scaling: To enable horizontal scaling, you need to partition your data across multiple nodes. One way to achieve this is by using consistent hashing. Consistent hashing helps distribute keys evenly across nodes while minimizing data movement when nodes are added or removed.
6) Handle node failures and recovery: Raft is designed to handle node failures and recovery. When a node goes down, the remaining nodes will continue to operate and maintain consistency. Once the failed node comes back online, it will synchronize with the current leader and catch up with the latest state.
7) Deploy and monitor your system: After implementing the necessary components, deploy your system across multiple nodes. Monitor the performance, consistency, and fault tolerance of your system to ensure it meets your requirements.
Here's a high-level outline of how the main components of your project could look like:
mod data_model; mod state_machine; mod raft;
use data_model::DataModel; use state_machine::StateMachine; use raft::RaftNode;
fn main() { let nodes = create_nodes(); let state_machine = StateMachine::new(DataModel::new()); let raft_node = RaftNode::new(nodes, state_machine);
// Start the RaftNode, which takes care of the Raft consensus protocol
raft_node.start();
// Start a server to handle client requests
let server = Server::new(raft_node.clone());
server.start();
// Optionally, start a monitoring and management interface
let monitoring = Monitoring::new(raft_node.clone());
monitoring.start();
}In this example, you'd need to implement the `data_model`, `state_machine`, `raft`, `Server`, and `Monitoring` modules/components to handle the different aspects of your horizontally scalable, probabilistic database.
Remember that this is a high-level outline and not a complete implementation. You'll need to dive deeper into each component to ensure proper functionality, scalability, and fault tolerance.
The example prompt is clearly asking for too much in one go, but this objection is trivial to overcome for anything not specifically hitting guardrails (and sometimes for those too) by telling it to act as if it's a [insert various roles to try].
You are an AI programming assistant.
- Follow the user's requirements carefully & totally
- First think step-by-step -- describe your plan first, then build in pseudocode, written in great detail
- Then output the code in a single codeblock
- Minimize any other proseWe will be the ones with the preattentive syntax parsing in our brains that let us simply see basic syntax errors; we will be the ones who can point out simple but subtle errors immediately—because we’ve made them 10000 times; we will be the ones who can take a ChatGPT response and immediately identify how it’s lacking not just in outcome but in edge case handling. We came of age before stack overflow, and have matured to the point where we treat stack overflow answers as mere suggestions for a possible approach, not as code to be copied and adapted.
And then, when we are no more, there will be no one left who can wrangle the turtles that reach the bottommost depths of the full stack.
Knowing the subtlelties of memory management in the 80' didn't help the Ruby coders to create RoR and run Basecamp, or the Python coders to create Django and run the Washington Post.
People jumped at ORM, CSS frameworks and JS wrappers, and won the day because they traded memory and CPU for what the user wanted.
The new generation will do the same. They will figure out how to be super productive with AI, way more than what we are. And our superpower will be useful in specialized cases, but mostly, their impacts will be small compared to the order of magnitude of speed they'll gain by using the new paradigm.
People said interpreted languages were a waste, that you had to learn SQL, that Electron was an abomination.
But the people building stuff always win, not what we think is the right way.
That's why PHP won at some point. JS won at some point.
And something super practical (but that will make us say uhh...) made by AI will win at some point.
That's also why all the open letters and ethics committees will have no impact: the short term advantage given to AI users will be such that they will take everything over by storm.
It's already happening in porn: we just gave millions of teenage boys a shiny new toy, and they are spending days and nights on it. New incredible models come out every day. Some hyper specialized. Last week, I saw a publication of a new model specially trained on only vagina called "bettervulva" and one for boobs called "breastinclass". Somebody spent hundreds of hours on those, collecting the data, annotating it, and training.
I guarantee each niche of the world is doing exactly the same. Not to mention each new model is coded with somebody who has the previous models to help with the new one, code included.
This is snowball from now on.
People who have decided to go all in in learning an interpreted language like Perl or even PHP, had to continuously also learn react, js, ruby,... and gazillions of other languages and frameworks to remain in business over the last two decades.
While people who invested in learning C, Java or SQL have as much healthy opportunities to pursue today with almost the same amount of knowledge that they had accumulated three decades ago. Their life has been much less stressful, pretty much on auto-pilot, while those who were pursuing the latest technology are constantly living on the edge having to continuously relearn entire new framework on average every 3-5 years to remain relevant.
So yes all these glitzy languages have "won" but it wasn't by these "ancient" languages losing.
If I had to bet I would say a decade from now, based on the last 50 years history of computing, I can easily imagine the same amount of jobs for C programmers as we have today but there will be as many Ruby jobs as we have perl jobs today.
I have not stopped messing with Stable Diffusion and llms for like 7 months now and I've been a software engineer for 20 years. I use copilot and chatgpt and my own brain and have released two compiled python apps on itch ( first real apps I've ever built and released from the ground up on my own) - both are ai tools, one for stable diffusion one for flan-t5. I'll be releasing a unity solitaire game soon as well, and ive been training my own models.
This is the most fun I've had creating things since I started programming.
I've gone from hating PHP for most of my career, to now having a grudging appreciation for it. Its a clusterfuck, but it works really well, and deployment barely gets simpler than "untar into webroot, done".
Those programmers didn't need to know the subtleties of memory management because when they programmed in Ruby they didn't have to debug in C. GPT is very different.
The double meaning at play here...
Making something easier doesn't necessarily increase quality. It's early days. It's a fool's errand tp try to predict the future. Now more than ever.
Then "explain why <my code> needs to be <new code>"
Oh this is called <new concept>, give me an example of it. Give me another example. Give me another example. Give me use cases where this would be better. Give me a short story where it was done my way but it lead to a mishap in production and the protagonist learned <new concept>. Explain it like I'm 5. etc
Basically the new generation is the training neural network, and they can keep flashing context until they absorb the material. Sure, some people are lazy and will go as far as getting something barely working. But superpowers are available for those who fight for their place in the world
It would be nice, but until there's a breakthrough on having LLMs / other models generate, maintain and walk the user through a representation of their reasoning, these explanations are likely no more reliable than a random StackOverflow answer - except without the benefit of other users scoring / picking holes in the answer
Not iteratively forever, the only feedback it will (eventually) get is it's own feedback, because humans will be using the AI best practices, and what it deems as best practice is almost certainly going to be too opaque to us mere mortals.
And they’ll get code which isn’t good, just a parroting of bad ideas widely disseminated. A bit like Casey Muratori’s concept of fake optimisation: https://youtu.be/pgoetgxecw8?t=600
I think what we have now is more productive but results in worse programmers. I.e. you fix the thing in context, but learn little out of context and have to go back to the source of answers when the next thing comes up. The rate at which answers are sought increases, doesn't decrease. Etc.
Years later, I don’t think her driving has meaningfully improved.
As someone who learned to program from two gigantic tomes of C/C+ with a Borland Compiler, there was no copy/paste. You had to type every listing in to get it running, as well as find that semi-colon you forgot on line 16,46,102 of 200. Then you had to remember everything, take notes, and literally research answers. You couldn’t just ask a random person on the internet, at least for a few years after I learned how to program when online forums became a thing.
One big difference I can think of between my younger peers and myself: When I see an error message, I take it as a “hint that something is wrong” where my younger peers tend to take error messages literally. Probably because I “grew up” with terrible error messages that hardly ever pointed to the actual file/line the error was on.
I think all the “kids these days” gripes only apply to people’s first few years. You can argue they’re formative, but most professional programmers will quickly need to evolve beyond “take message literally, ask online for help, paste error message into google/SO” to be good at their jobs. I don’t know anybody beyond entry level who operates like that.
I rarely take things literally like you said (though I may start by giving it the benefit of the doubt), I always just check the actual code from the thing emitting the error message, read the source of libraries I’m using, seek out documentation. It just so happens I can do all of that from my computer rather than doing some of it with a book.
One of my cousins learnt to program as a teenager in Argentina, before the web. He told me about the difficulty of finding programming information over there. Apparently, people who knew any kind of programming had paid dearly to get that knowledge and in turned charged others for that knowledge. Needless to say, there were not many programmers in the country at the time.
In comparison, today's generation has access to vast libraries of knowledge.
Interesting, because to me there still seem to be many public discussion forum questions where the answer is "Have you tried doing exactly what the error message is telling you?"
Not talking about AI; but if you want to work on compilers, linux kernel, game engine optimization, GPU optimized code and many many other things than web development you need to be more than familiar with how to work with machine instructions, albeit not working with it every single day. These are significant parts of the industry which make the rest possible.
There are still areas in software development where (not systemized) failure is not an option and where AI generated code absolutely can be useful, but would need careful validation from a human. Not only for unintentional errors, but also for potentially maliciously seeded errors in the model that is being used.
Perhaps manually writing test cases for the AI generated code could meet the requirements for human validation.
The article says they “get the right answer 80% of the time”. That’s abysmal when you don’t know what you’re doing. Software which is 20% bugs is software you can’t rely on. Depending on what it does, it can be actively harmful.
Being a “magician (…) among mere mortals” is not a positive. We live in an interconnected society where the actions of others affect us. It would be nice if you can trust what other people are doing when they’re dealing with your water supply, electricity, food production, hospitals…
If you have years of expertise already they give you wings.
No one is saying to take the output and immediately ship it. It still needs to be finished and tested. What's important is what is ultimately delivered; less important is what steps brought you to that point.
That said I suspect good testing will become more highly-valued as AI use increases. Maybe there's a form of TDD that can develop, with the human mostly writing tests and AI prompts?
I like that.
We are going to be vringes softwarearechologists and programmers at arms.
The mental overhead of producing softwares shrinks more and more. It will sure be at the cost of understanding and/or see errors as you say.
I already looking forward of being that graybeard. And show off to my neighbors or grandchildren
People pay for productivity, making a product of some sort that has value. These AI tools will increase productivity. And they will only get better with time. That translates into the productivity of people who have mastered these tools as more valuable. They just get more done.
Will syntax issues and other problems be an issue with early AI? Sure, but it will get better. The AI will start to catch more issues and move farther up the stack. We still need to know enough to guide it, but the details required will be fewer.
People are still actively acquiring the knowledge to do things the old fashioned way in those areas, a hundred or even thousands (for writing) of years later. But typically they only do it if they need to or find it interesting.
It’s hard for me to think of any technological advance (outside entertainment) that actually caused a mean regression in peoples’ effectiveness, or for people as a whole to completely stop acquiring the skills of doing something the old way to deleterious effects. Like, has there been any invention that was so easy to use that people got worse at the task the tool was invented for? In aggregate it may happen if the tool greatly reduces barriers (like training, experience, or study) to entry, but I kind of doubt it happens for specialists.
Note by effectiveness I mean the general task, not the specific skill that may have become obsolete. I’m not effective at all with a slide rule but I have a Python shell, the internet, and hardware capable of floating point operations.
This time its different... The written word might be able to write back. As a matter of fact, I already use chatgpt to learn new skills. Its a better way of acquiring knowledge. The job of teachers will change within 5 years.
If every single person with that coding experience only coded, that's likely, but not every person who got their start that way focused on coding to the exclusion of everything else. There are at least a couple of us who are more oriented towards teaching, archiving, and knowledge preservation. Right now, people like us but don't tend to prioritize giving us their resources. Once those greybeards are in their 60s/70s/80s and looking at their work being lost forever, I imagine that will change. (Much in the same way that there's way more discussion now in tech spaces about raising kids now that more programmers are old enough to confront that issue).
That knowledge will likely be preserved by another greybeard cohort and it'll be specialized knowledge that only some people know, but it will be known. Some stuff will be lost because we'll prioritize money and copyright up until there's a crisis, but there are people prepping for this time period. It's not like trade knowledge loss is a new problem. It's just not a sexy one right now.
I think using AI to build a React app from Figma is closer than building a low level piece of infrastructure. Hence, the grey beards operating at the lower levels are slightly less vulnerable.
It is all the magic of early google, and stack overflow... Not as many ads less bullshit.
I have been coding for 25 years. Chat GPT + willing sr engineer means 2 less jr. dev's. I lived through the dot com bust, and now I can do a heroic amount of work with less effort.
This raises two points:
1. It's clear that we have too much grunt work. Im not sure if this is a failure of language design or libraries, but we need to do better at the core.
2. We continue to narrow the path to get experience as an engineer. This is great for companies in the near term (this looming down cycle), but bad for the industry as a whole long term (we dont have a pool of strong jr. engineers).
It's not grunt work.
It's how new engineers learn the ropes and gain experience doing low risk work; it's part of the learning process that only feels like grunt work to a senior dev.
Tools like Copilot are very useful because if you need to create a function that takes a FLOAT in addition to the one you just wrote that takes an INT then its right there already. Sometimes I am just trying to bang together something quick for testing just so I can continue on with the actual project.
When I learned to write I did not just paste from Wikipedia into my essays, I would distill words from the articles I read into my own thoughts and writing. These tools are absolutely useful, but it is what you make of it.
That's not my point, per se.
> It's clear that we have too much grunt work.
What one might see as "grunt work" is what another might see as "learning the ropes".
You can make the case "with ChatGPT, we no longer need to learn the ropes".
But it's really hard to predict how that will affect the next generation of engineers -- I don't even know if "engineers" is the right word -- if they only understand how to ask and not how to construct.
I'd say that my point is that grunt work is very meaningful to someone who is in a different phase of their career. It's not just grunt work, but a pathway to learn.
Yes, architecture, medicine and metallurgy have been improved for millenniums. Next to that, IT is a baby that can barely walk. Of course we have too much grunt work.
But also because every time we remove grunt work, the customer expectation increase, the market adapts, and now you have to use new tools that are productive enough to do what the users want with the resource they are willing to pay.
So you don't have to hand write assembly but you do have to have an app that works on 16 screens ratios. You have amazing databases ready to handle your data, but you should update the UI live.
> We continue to narrow the path to get experience as an engineer.
Like in most industries, we are starting to move from "everyone needs to be an experts" to "90% of the daily tasks really just need good technicians". That's fine. That's how it is supposed to be.
At some point only experts could duplicate your house keys. Now you have convenience key stores everywhere.
We are on the right track.
Next stop, the equivalent of residential building codes, but for public facing apps.
Because IT is so new and can be done in your garage, we thought it was special. It was not.
Yet. They got to get you hooked first.
At least with Google and Stack Overflow ads are identifiable and can blocked.
For example, if you have a configuration, and a particular default would often (say 30% of the time) be overridden, I think it’s typically not a good idea for the default to exist in the first place. Having those defaults will certainly reduce the average config size, but they’ll also create more cognitive load for readers, who now not only don’t know what the default is, but don’t necessarily know that the choice is being made in the first place.
This is one of the things I think is so powerful about AI as a solution. It saves you from the work, but still conveys the basic context to readers without requiring them to understand a bunch of opinions baked into the language or framework.
The same applies to writing prose. When you give a bulleted list to ChatGPT and ask it to turn it into prose, a lot of the work it’s doing is filling in obvious text. But that obvious text is important to communication, and where you have to go and edit that obvious text for correctness will be some of the most important communication in the final product.
We've been iterating over that for decades, there's always one more step of grunt work to remove, one more layer of boilerplate code to move into the compiler (or a code generator), one more deployment task to automate, one more elegant new syntax to add to a language, …
There is no one place where you can find the business logic.
What do you mean "at the core"?
I just used GPT-4 to help build an analytics dashboard using ChartJS. There's so many settings in ChartJS, it would have taken me a week to StackOverflow / Google how to get my charts how I'd like them - it took me a day with GPT-4. I could just ask it anything and it would help no problem. Any buggy code it produced, I'd just copy and paste the error message and it would provide a fix.
The day before I built a basic version of Stripe Radar: https://news.ycombinator.com/item?id=35323278
Coding with AI has made me more excited to build than ever before after 10+years of programming.
https://www.chartjs.org/docs/latest/api/
I just started looking at the docs a couple minutes ago to understand why the OP claims that it would have taken a week to accomplish something with this library. I don't get it.
I will admit chatgpt and copilot are very good at offering up bite-sized pieces of code, like particular objects and functions as well as templates.
I absolutely would not expect it to fully write what I want, but I’ve found it very thoughtful for analysis so far in our conversations.
You need to check every word of the AI for truth anyway, so you need to research all the stuff it tells you, as it could be made up, everything it tells you is very shallow, so you need to dig into the topics anyway, and making even small changed by the AI requires ridiculous amounts of explaining details in prose, instead of just changing a few lines by hand.
All in all it's much easier to come up with some more reasonable solution in a shorter time by just doing it the "classical" way: Read the docs and copy-paste code snippets from proper curated examples some human actually tested.
I don't say that an AI could not program at some point in the future. But I guess it would require AGI to be competitive with people who actually know what they're doing. But at the time we have AGI we have anyway much larger "problems" than how to write code, as more or less all human intellectual work would be superfluous.
Anyone saying programming jobs are in trouble don’t really know programming. You can’t just have someone from the bizdev team prompting ChatGPT for JavaScript code and just doing copy/paste…it would be difficult to even come up with the prompt without sufficient programming jargon knowledge.
Like the article says, the tool will allow programmers to work faster and therefore tackle more things. Expectations will go up but I don’t think the jobs themselves are in trouble.
Pardon the plug, but I co-wrote a little blog post about this with concrete examples, if it is of interest: https://open.substack.com/pub/storiesbyai/p/chatgpt-for-crea...
Does anyone know if version 4 is any better in this regard?
- writing simple functions: e.g. I used it to generate a function that receives a name and then removes any dutch name particles ("van de", "van der", ...). ChatGPT also generated the list of dutch name particles of dutch name particles.
- solving tricky bugs: e.g. I used it to solve a tricky but in SwiftUI that nobody on stackoverflow was able to solve.
- generating ideas for apps I could develop that would make use of the ChatGPT API.
- as a replacement for google when debugging code based on error messaged received.
I feel like this is how the typical programmer uses ChatGPT, but im intrigued by your comment what kind of uses you have for it. If you dont mind, could you hint at some ways you use it?
I have abused chatGPT so much that I have an impulse to talk to other programs that are not the chat bot, like the terminal and VS Code, when they give me an error that I cannot understand, just to ask them "why?", like a premonitory dream I think I have an idea how chat bot could be integrated into every piece of software sometimes in the future.
IMO, this is irrelevant, because "normal people" would never ask this question. Because it has never been asked, there's also not a ChatGPT answer to it, because its answers are based on already existing information.
This thinking that ChatGPT is AI is just wrong. It's not AI, nor GI, so stop comparing it with that.
As a human, I can destroy your problem/question with one claim: you claim that the words must be built on the English alphabet. Let's throw some Æ, Ø and Å in there, shall we?
Can YOU solve that problem?
I really see it kind of like having a good junior programmer or intern helping me out who has an incredibly broad amount of knowledge. Sometimes they need some coaching or gentle correction, sometimes you say, "I'll take it from here", but inevitably there are those times where you think, "Great job! I'm really impressed!"
I have also tried to use ChatGPT like Google search at the beginning by asking short, curt questions, and got underwhelming results.
The trick is not to try to trick ChatGPT but pretend that you are an idiot and it is a very sophisticated colleague who just needs step by step details. Then you can also generate some of the seemingly amazing results people are getting. For example yesterday I used it to scrape physician information from my insurance website, followed by their Google ratings and so on.
Ok, I see you calculated the probably using randomly generated "words" from the letters of the English alphabet. I am interested in the actual probably of two real words in English that are 5 letters wrong share the first three characters.
I am a Python developer, so I will understand it if you give me a Python script.
I gave me this which looks right to me:
import nltk
from collections import defaultdict
nltk.download('words')
from nltk.corpus import words
# Get the English words
english_words = words.words()
# Filter the words to get only five-letter words
five_letter_words = [word for word in english_words if len(word) == 5]
# Create a dictionary to store the count of words with the same first three letters
words_dict = defaultdict(int)
# Count the words with the same first three letters
for word in five_letter_words:
key = word[:3]
words_dict[key] += 1
# Calculate the number of pairs with the same first three letters
same_first_three_letter_pairs = sum((count * (count - 1)) // 2 for count in words_dict.values())
# Calculate the total number of pairs from the five-letter words list
total_pairs = (len(five_letter_words) * (len(five_letter_words) - 1)) // 2
# Calculate the probability
probability = same_first_three_letter_pairs / total_pairs
print(f"Probability: {probability:.4f} or {probability * 100:.2f}%")The only thing I dislike about d3 is the boilerplate code to hook it up to a div and get the first lines to render, so yesterday I did a:
Assuming I have the following data: ```{ "status": "ok", "data": { "a": [[1680079200000, -199], [1680079800000, 92]], "b": [[1680079200000, 248], [1680079800000, 259]]] } }```, where the first item in the array is always the unix timestamp in ms, and the second one the value of the datapoint. Also assume that I have a website with an HTML element `<div id="graph"></div>`. Please show me the code used to draw a graph of "a" and "b" with d3 into the `div` element.
And it generated all the code I hated to "create".
Then little things like "Ok, I want the time not to be in AM PM, but in HH:MM" which used to mean to visit a couple of pages, the answer's right there.
I absolutely love it.
And yes, it does make mistakes, very rarely, and you can't really upload the whole source code everytime you want to talk with it, but it makes the programming dream real again (if only for a couple years lol): is not like you don't _type_; there's so many things you stop thinking about. And you can say: oh, there's the catch, but the truth is my own buffer is so much freer all the time, I get to tweak files, with excitement, that a couple months ago would have been very tedious to get a grasp of. And about code quality: its Elixir is alright, its Typescript is excellent, but it's absolutely flawless at concatenating unix commands: it has read all the man pages. I've aliased more custom commands these couple weeks than in the whole of last year, and my workflow is getting tighter and tighter each day. There's just so much power into this thing.
Am I doing it wrong somehow?
One meta comment is that I would love for something like a Greasemonkey script that can automatically add some checkboxes or something that will then automatically add stock phrases to ChatGPT questions, like 'show me only the code', as I find that a lot of the intro fluff and the explanation of things are useless to me and just take up time. In general, the UI is pretty bare bones, there's a world to be gained there.
I guess the question is. What do we build with this form of engineering. It's so very early. We're using it as a way to augment current systems but I think we'll build for entirely new things soon.
So, its quite literally an issue of lost knowledge. Is like replacing wikipedia with something that answers faster, but without sources. Its cool in the short term, and people are more productive at learning things, but you also go back to the hear-say time before books were invented.
What we are seeing is that the number of software engineers has increased dramatically and the average skill level has plummeted. It is not normal for a competent software engineer to need to reference Stack Overflow a dozen times per day every single day.
The other thing that this is bringing to the forefront is just how atrocious the libraries are in certain ecosystems (e.g. JS/TS).
If a tool that can write incorrect functions/methods makes you substantially more productive, you were never very valuable in the first place, and "AI" likely isn't the only thing threatening your job security.
It'll most likely improve in the future, I'm sure. The amount of real-world progress this will bring to humanity boggles the mind. I think we'll see progress in many and as-of-yet unthought of areas, because the power of mathematics, algorithms, and advanced systems will become available to a lot more people due to this "helpful uncle".
I do hope they work on improving its math skills though, also in the sense that it can explain the steps taken when solving equations or algorithms, or even how algorithms could perhaps be "translated" and implemented in various languages. In my opinion this is key to understanding mathematics, and a whole host of other systems as well. Moreover, it helps people who previously didn’t have the capacity to understand these things, so the effect on “world knowledge” is cumulative.
Personally, ChatGPT has been a great help for understanding many things a lot better, everything from mathematical equations to subsystems in Linux. And of course, it has helped me improve my coding.
Not so with an AI chat system. You get the answer immediately, and you can ask as many follow-up questions you like, to your heart's content. This means that, sure, even if the code you get from ChatGPT has approximately as high a chance to work as "random code from the internet," you can't ask that "random code from the internet" to expand on why it's good or bad, why it was made in a particular way, or if there are better ways, or if you can error check it one more time, or complain that, "Hey, random code, your dumb code didn't work!" In that way ChatGPT is already leaps and bounds better than "random code from the internet."
Also I assume here that you aren't just throwing darts at the wall, but instead testing out code with a clear goal in mind. Then it can be wise to ask ChatGPT about the underlying system, which it will gladly answer, and also immediately.
ML-enhanced development neatly circumvents this: explainability and accountability is passed on to the developer. This includes bugs and license infringements.
I have no qualms about ML tools in development. But so long as the buck stops with me, I prefer to write from scratch.
To make it work in most cases it should follow the same stuff as this implementation of a Reflexion agent for SOTA Human-Eval Python results, but use it for infrastructure. GPT-4 can generally figure out how to create tests from the documentation that will make it know when the task it finished correctly.
You could speed up the remainder dramatically and make very little progress on the rate progress happens and have people scratching their heads about why.
I'd have enough money if to train my own foundation model if somebody gave me a nickel for every time somebody underestimated the time to rebuild an existing e-business system by orders of magnitude because they were ignorant of all the "requirements" that were satisfied by the existing system that were undocumented.
The most expensive maintenance problem comes when you've accumulated a lot of data that is incorrectly structured and may or may not be possible to migrate to a new, correct, structure.
Fred Brooks in No Silver Bullet pointed out that software development involves multiple phases, say 5 of them (requirements gathering, design, coding, testing, deployment) and that even if you could reduce the time involved in one of those to zero there is no "revolution".
The corollary to that is that the opportunities from ChatGPT come out of identifying the things that ChatGPT does poorly. For instance it might be just as important to turn code back into requirements as to turn requirements int code. There was a post the other day about a "simple" problem in CSS that turned out to be hard because of the deep understanding required
https://news.ycombinator.com/item?id=35361628
the ratio is much worse between spending 3 minutes to make a partial solution with ChatGPT and 3 hours to really solve the problem as opposed to 30 minutes to make a partial solution manually and 3 hours to really solve it. That is all the more reason to clean house at the lower levels, either make a style sheet language that is not so fraught or maybe make one that's designed so that chatbots can write good code for it.
There has been hype over "no code" and "low code" tools for the longest time and people have been chasing that revolution for a long time and it's astonishing how many innovations have been forgotten (like the 1990s CASE tools that could turn code into diagrams and back again and leave formatting intact.) Real progress involves solving all the problem that is ahead of us instead of cherry picking the parts of the problem we want to solve and thinking wishfully that will be sufficient.
These other well-documented experiences that sound super perfect don’t match mine, though I’m a fan! Here’s how it goes for me.
I am a premium user and all, but it still bails on me every single time it’s going well, by writing half of a function or saying “network error”, then losing context.
I have to say “you stopped generating that last part, please continue and don’t bother qualifying.” Then it wastes more tokens apologizing anyway.
Whoever trained these models to apologize and say “as a large language model, I can’t promise that I know how Bjarne Stroustrup would refactor that bitshift, and yada yada” really wasted a lot of our productivity. And saved no one from any dangers of AI.
You can just say "Continue."
Dirbuster wasn't packaged for Debian (that I could find). I wasn't too keen on trying to build from source or anything like that. A few months ago that would have been that.
But with GPT I just asked it to write a POSIX-compliant C port of dirbuster. Less than 10 minutes later I was scanning my TV for routes.
Just wait until we have AI searching for Zero-day and writing exploits code for them...
Or AI-assisted-Pen-testing-as-a-service <-- HackOne or someone similar should be all over this...
What is Palantir doing for GPT crawling of logs and writing filters/rules/queries...
Or firehosing data through an AI Deep Packet Inpector etc...
Security is going to be the most interesting space to pay attention to going forward...
I used to build random weird things all the time (mostly centered around connecting physical devices up to the internet to accomplish something silly) - but lost my mojo once I had kids and didn’t the ability to pull all nighters to “hack” something together.
Now I feel like the barrier might be low enough that I could create again. Will have to try this out and see how it goes.
Is there a reason why gpt-3.5 and gpt-4 all using pre 2021-09 data? Is it because gpt-3 as a base model was trained at that time, or is it really because bots are "contaminate" the Internet from then? I guess we can never go back now.
One thing I am certain, after 2022-12, AI might enforce itself by learning AI-generated content, and perhaps someday, AI may deem human content as somekind of "bias".
Don’t get me wrong, I am a senior DevOps, know my terraform/CDK/… and AWS/heroku/digitalocean/hetzner/vercel/… stuff by heart, so it’s not missing knowledge.
But I really hate the maintenance work down the road. I just don’t want to get emails about some deprecated stuff from some provider and have to do anything just so stuff keeps running. I would pay a premium (!) for a stability guarantee not in service uptime but toolchain reliability :/
So basically, a bug can set you back a long time (multiple weeks). And to reduce the risks of bugs we usually need to ship incrementally. And because bugs are such a hassle we have to monitor and test the crap out of everything, which of course we should be doing anyway, but we do it more than we probably would otherwise. But of course bugs do still happen, and even someone else’s might set you back. A flip side to slow releases is releases tend to batch more changes. And of course that makes finding a bug more time consuming.
So a feature that could take me a few hours to write and test at my top performance basically needs to be budgeted at a month of work. Of course, we have other responsibilities like testing, monitoring, analysis, support, design, coordination, helping others. But dragging out completing a project drags all that out too. If I could ship to staging same day and prod 1-2 days later, I wouldn’t need to plan so much or spend so much time going back and forth about design. We could just meet up about the change and do it that week.
And you might think you could simply parallelize/pipeline more tasks, but overdoing that sets you down the path of overextending yourself. If something important goes off track you need to do everything to get it back on track - if you don’t, it’s going to be really late - and overcommitting eliminates the slack that makes it possible.
Truly, this is something I think about all the time and probably the least favorite aspect of my job. Anyway, unless chatgpt can help me write a persuasive argument that convinced my company to completely 180 its engineering policies, it won’t make me much more productive.
When it comes to boilerplate these tools are game changers I must say. Harder problems they crap all over themselves but that's ok, that's where we still have a job for the next 5 years (I hope).
What usually drains most of my energy, as someone in the JS ecosystem, is the sheer volume of modules I need to sift through to find one that does what I want.
The other day I wanted to find a library for handling zlib in the browser - eventually I stumbled on `pako`, which is great.
Seeing this post I asked ChatGPT a somewhat vague question, namely: "What is the best npm module to handle zlib compression?"
It told me about the built-in zlib module, but the last paragraph had this:
"There are also other npm modules available that provide additional functionality or a different interface for handling zlib compression, such as pako and node-zopfli, but for basic compression and decompression using the zlib library, the built-in zlib module should suffice."
Now I feel dumb that I haven't used it for this purpose earlier.
In my day job, there's been various attempts to use GPT to speed up some work, which has generally resulted in me having to fix "confidently incorrect" code. The outputs usually about half right when its usable.
Using an LLM to basically auto complete trivial bits of code that are "just typing" but usually require a referral to the docs is probably the optimal middle ground for productivity - so time can be spent creatively solving the actual problems.
I'll have to experiment with it further - have it emit all the boilerplate for something, do the fun parts myself, then check for correctness might well be the workflow best suited to GPT.
For data projects/pocs/mvps using python its saving me lots of hours, you have some tiny mistakes, but once I have the idea in my head I just iterate and in 30min or less I have "written" the code for what would take me many hours.
There's a few small projects I have been procrastinating for ages that just will require a lot of "tedious but trivial" hand jamming of 90% boilerplate that may be well suited to LLM generated code.
And finally, will find a time to build some ideas with Dart/Flutter. Something inside is telling me that in a generated future, people who can code will be able to create free from no code tools and SaaS silos. And this is enough motivation for me:)
I see a strange future in which product designers (who will spend less time on decorating and fiddling with DS components) with the help of ChatGPT will fulfill the role of junior devs in the production pipeline.
I’m wondering if the archive missed something or what this random statement from the system is about. It seems to be telling itself who it is and setting a restriction on what it knows. I know if it’s not trained on data past a certain point, it cant know it, but still i wonder if you could intercept that and change the cutoff date (if even possible), would that have any effect on anything. But then again, I am probably reading to much into it because it seems to be coming from system.
2023-03-26 04:10:49 - [system] : You are ChatGPT, a large language model trained by OpenAI. Knowledge cutoff: 2021-09 Current date: 2023-03-25
If you play with the ChatGPT API in the OpenAI Playground developer tool you can set your own system prompt. I used it to create a sentient cheesecake that could answer questions about SQL using cheesecake-based examples yesterday.
An artist that I follow makes procedurally generated dungeons and villages and such.
I noticed that he made it possible to export JSON and SVG files from his one-page dungeons. I decided that was enough to turn them into an actual mini-game. This app takes his procedurally generated static maps and makes them interactive.
There's nothing here I did not already know how to do, but I leaned on Chat GPT to generate code. e.g.
Me: As the avatar moves around the map, it would be nice to center it. Possibly zoom in and such. I'm thinking that the `transform` attribute of the map could help.
Chat GPT: Yes, you could use the transform attribute of the map to achieve this effect. The transform attribute allows you to apply various transformations to the SVG element, including scaling and translation.
To center the avatar, you could use the translate transformation to shift the map so that the avatar is in the center of the viewport. For example, if the avatar is at position (x, y) on the map, you could center it using: <code snippet>
To be honest, parsing through the SVG was just so tedious I would have moved on and not completed the project. As it was, a few nights. Boom! Little game toy thing.Some other tasks I had for it that saved tons of time:
* Given these css rules below, can you combine them so that each is declared once?
* Could you suggest a mysterious, dark, but not black color for the background of the dungeon map?
* Could you give me a nice CSS dropshadow for the avatar?
* Could you suggest several Google fonts that would be good for a dramatic Dungeon title/headline? Please include links.I’ve gone from saying that I’ll need an hour to do something to only needing 15 minutes.
However, I’m good enough at what I do that I can find the problems in the code or notice something out of place or that could be done more securely.
I fear many won’t and lots of junk code is going to find its way into our lives. I worry about things like this when it comes to, for example, Wordpress plug-ins, or add ons to other larger tools that we use.
This is the real benefit here. AI isn't taking jobs, it's going to enhance them - once we learn how to use the tools we have available to us.
If by AI we mean things beyond just ChatGPT sure it will, it’s only a question of how long before it happens and how many new jobs will be created to replace the old jobs. For example, AI isn’t going to enhance customer service, you’ll just have better automated experiences and businesses will hesitate less to use them.
Edit: AI isn’t being created for the benefit of mankind. It’s being created in the pursuit of money. This isn’t a utopian fantasy world where we look out for everyone’s best interests, that’s just a side effect most of the time.
1. 99% of devs will be out of work and the current work will be done by 1%.
2. More and better software will be developed. About 100x effort more.
I'm pretty confident the 2nd will happen. I already have lots of work that will take a long time, and more after that. 100x improvement on productivity would mean more, faster and better (because of more tests etc) output.
To use a relatively modern analogy, AI-enhanced development is to developers what Photoshop is to photographers. No one ever said Photoshop would put photographers out of business. Now Photoshop is a requirement for photographers. The pictures still need to get taken; likewise, for developers - the business requirements need to be captured.
AI [in its current iteration] cannot do that.
As developers, our jobs are safe. Jobs will not be lost at a massive scale; only those unwilling to learn a new tool will be replaced.
I'll take this mentality onboard to more ambitious projects.
Especially working with other languages now is a lot easier, because I can ask to translate between languages I'm familiar with. Same thing with configurations, eg translate this systemd script to openrc.
Also Github Co-pilot has been good for a long time.
It's a bit scary because knowing this stuff is sort-of our secret sauce and GPT-4 was able to give an even better answer than I was able to give. It helped us out a lot. We are now taking the solution back to the customer and will be implementing it.
A few additional thoughts:
1. I knew exactly what type of question to ask it to get the right answer (i.e. if someone used a different prompt maybe they would get a different answer) 2. I knew immediately that the answer it gave was what we needed to implement. Some parts of the answer were not helpful or misleading and i was able to disregard those parts. Maybe someone else would have to take more time figuring it out.
I imagine future versions of GPT will be better at both points.
I follow the same steps as always to design the code, but when it's time to implement something, I ask the bot to do it, then I review it and move on to the next function.
My experience:
- Copilot did a decent job at suggesting functions when I typed out comments. It got progressively worse as the compression algorithms got less common (eg. Huffman vs. entropy coding) The smaller the functions, the more manageable it was. <- You need to write good tests, you want to step carefully through every line of code it writes.
- ChatGPT got most things right, and a few things really wrong. It was always super confident. <- Super dangerous if you don't double check, basically only usable as a refresher on a topic you are already good on.
- If given a piece of code and asked "How would you suggest to refactor this?", ChatGPT gave mostly very useful ideas. <- This is something that I will keep in my workflow.
- The bigger the project got, the worse the overall codebase looked. It became a mix of styles, pretty inconsistent, more subtle bugs were introduced that took me a while to figure out. (The nice thing about lossless compression is that you know if the code is right by using it)
- My "productivity", how many lines of code "I" wrote, was way beyond anything I could usually reach. I do also feel like I got a very quick start into Golang with this, much faster and broader than reading documentation or doing a getting started. That said, that knowledge now definitely needs to be refined in order to use the appropriate concept at the right time. Comparing my code with more professional Golang code bases I spot a lot of things that need to be improved.
- I want to create a Google Docs addon using AppScript which has an interface sidebar. Please describe the necessary steps to create a sidebar interface. The sidebar interface should have a title with the text "Google Docs JSON Styles", it should have a large textbox which can contain multiple lines of text (fill it with lorem ipsum), and it should have 4 buttons placed horizontally underneath the text box, with the captions "Save", "Apply", "Copy", "Paste".
This produced the necessary Javascript + HTML + external steps to set up the project.
- I'm building a website using a React frontend hosted on example.com with a Django REST API backend hosted on api.example.com - Authentication is handled via Firebase Auth, currently using an Authorization header provided for API calls. Now I want to have authenticated links such that the user can click a link to access an authenticated download via the API. What mechanism can I use to create a link to an API endpoint which will require the user to be authenticated?
This gave me a the frontend and backend code necessary to set up authenticated links, with several alternate approaches following further prompting.
- Given an owner id, repository id, commit id, and private access token; give me a URL that lets me download a ZIP archive of a Github repository.
In all of these examples I received the desired result with helpful descriptions.
As a counter-example:
- Write a program in Brainfuck that outputs the string "CHATGPT".
In this case, the result is the Brainfuck program for "Hello World", with a step by step break-down of the program explaining confidently why it would output "CHATGPT", while being entirely wrong.
I find it much more useful for soft things like writing a complaint to a vendor for me, dealing with customer service, or cover letters from a job ad that I only have to slightly tweak. Takes me five minutes instead of 30+ and I got a few compliments about my “outstanding” cover letter.
Also gpt4 >> gpt3.5 for coding.
Leftwise bitshift would make gpt4 worse though :P
This is not my subjective experience. It's better, sometimes somewhat, but also much slower. I'm not sure which one is the better tradeoff for real work. I have it on 4 by default now, don't want to think about this for every question I ask, but for my work, 3.5 was fine.
I tried to make GPT-4 create an algorithm for an enhanced tic-tac-toe game (think Wordle vs Quordle to somewhat visualize the difference) and it's failing miserably.
Last night I sat down with ChatGPT and within a few hours we'd built the foundations of a fully functional app.
It's a fascinating process – it often leaves out vital information or doesn't get things quite right and it's obvious it's writing the most likely code rather than something written specifically, but it provides enough to get me 80% there so I can finish up the rest. Felt like collaborating with a forgetful, distracted genius.
ChatGPT certainly shortened a lot of the scripting head time. Though it was not quite always accurate it did help me a lot.
Also, the final product (with documentation) has been complete and pretty amazing.
https://en.wikipedia.org/wiki/Braess%27s_paradox
Please make sure you're measuring and optimizing the right thing, eh?
"Happy hacking!"
8 questions per hour (25 per 3 hours), is not a lot to work with for paying users - it started being like 10x higher and they've announced it may go down even further.
I wonder if Azure can scale with the insane adoption rate and if it's profitable at the moment?
If you need the frontend, there are a bunch of open source front ends.
Ideally I want to be able to highlight some code and then run an AI command on it and also provide it as context to a query without leaving vim or losing syntax highlighting.
I admit that I am using ChatGPT to write little snippets here and there but I am an old fart and was writing software for 40 years already. I am probably safe from loosing the skills bar going senile ;)
I understand that this isn't the best way to learn coding, but at this point, I'm looking for results, not to get a job.
Don’t get me wrong; I use frequently when coding but it can’t handle the complex stuff; it’s insanely good at boilerplate and low-hanging fruit though.
Good because it generates a lot of bullshit.
However, the recent "letter of warning against AI" penned by Elon Musk and other influential individuals has led me to question their motives. I suspect they may be losing the competitive edge they once had, as users can now effortlessly obtain information in ways that were previously unattainable. It's possible they are seeking to lobby for new laws that restrict this kind of data extraction. With such simplicity, I can envision a future where people no longer need to browse web pages, ultimately making ads and the whole business of treating your users like crap irrelevant.
[0] 10 lines to screenshot with puppeteer, tesseract to OCR and a clustering algorithm, add to this whisper and a nanotts to get an enhanced Unix philosophy+AI experience
[1] Im pretty sure i can use low resource models to do the summaries and insights, any recommendations?
* edit,im pretty sure that could have done this PoC years ago without GPT.
1 month ago people were saying there was no way AI could do their job and many were telling everyone to calm down as nothing would happen.
Now one month later this entire thread is about how chatGPT is an integral part of the development process. There are people even saying it replaces junior developers completely.
Can anyone venture to guess what the thread will look like in 1 more year?
"chatGPT can do the work of an entire team!"
And one more year after that?
"chatGPT can manage a team!"
Likely, a month ago, anyone was interested in commenting on AI threads, so more skepticism was evident. Now, most people who are AI-skeptic are probably tired of AI news and have said their piece in other threads, so don’t bother.
fwiw, I’ve tried to use ChatGPT to make a new, useful Elixir library and it failed miserably. It makes me wonder if the success with JavaScript/Python uses are because the question and answers are extensively in the training data whereas a relatively novel Elixir program is something it can’t do because it’s not truly “reasoning”. It can’t even get basic command line prompts that are straight from the Elixir or Phoenix docs correct, and it cannot solve the error messages that its code produces like others claim experiencing for JavaScript/Python uses.
If my hunch about training contamination is correct, we’re going to quickly run into not being able to create novel software if we replace everyone with GPT algos, unless their core functionality vastly changes from excellent pattern matching to true reasoning.
I can't say whether it's using reasoning for specific answers. I'm sure many times it just regurgitates something that is known... But saying that LLMs in general can and do reason is categorically true. LLMs do attempt to extract actual mental models of the subject of the text.
See here:
- GPT style language models try to build a model of the world: https://arxiv.org/abs/2210.13382
- GPT style language models end up internally implementing a mini "neural network training algorithm" (gradient descent fine-tuning for given examples): https://arxiv.org/abs/2212.10559
For your specific case, given the conditions, another explanation that fits as well is that the LLM simply needs more data to build a more complete model for elixir. Once it does that I'm sure it can respond with completely novel elixir code as it already does with python.
The extreme positive reaction to gpt4 is imo not driven by hype fuel.
In fact, if you recall the hype fuel surrounding the initial chatGPT was in both parts positive and negative. There were equal portions of people praising it and people vehemently dismissing it as a stochastic parrot. That's not what I see here.