Everything I built with Claude Artifacts this week
simonwillison.net
simonwillison.net
All of the PRs I ever submitted touched a handful of files in my project’s subdirectory.
Or what "yes" looks like to you? It can do all the work itself, for a 50m-file monorepo, without a human guiding it which files to look at?
If it were true then human programmers would have been considered obsoleted today. There would be exactly zero human programmers who make any money in 2025.
Out of curiosity what does your IDE do when you do a global symbol rename in a repository with fifty million files?
I'm absolutely a real human, and I think this just might be too much context for me! Perhaps I am not general enough.
Since it’s git based, it makes it very easy to keep track of the LLMs output. The agents is really well done too. I like to skip auto commit so I can “git reset —hard HEAD^1” if needed but aider has built in “undo” command too.
These tools aren’t magic. But they do certain tasks remarkably well.
People do work on monorepos with 50 million+ files, though…
Based on my personal experience it works well as long as each file is not too long.
It works good until it doesn't.
It's definitely a useful tool and I'll continue to learn to use it. However it is absolutely stupid at times. I feel there's very high bar to use it, much higher than traditional IDEs.
I do find it very useful, but I agree that one of the main issues is preventing it from making unnecessary changes. For example, this morning I asked it to help me fix a single specific type error, and it did so (on the third attempt, but to be fair it was a tricky error). However, it persistently deleted all of the comments, including the standard licensing info and explanation at the top of the file, even when I end my instructions with "DO NOT DELETE MY COMMENTS!!".
https://github.com/Aider-AI/aider/blob/main/aider/coders/edi...
excerpt: """ Act as an expert software developer. Always use best practices when coding. Respect and use existing conventions, libraries, etc that are already present in the code base. {lazy_prompt} Take requests for changes to the supplied code. If the request is ambiguous, ask questions.
Always reply to the user in the same language they are using.
Once you understand the request you MUST: """ ... etc...
I was blown away when I realized some Haskell functions have only one possible definition, for example. I think most people haven't worked with type systems like this, and there are type systems far more powerful than Haskell's, such as dependant types.
There's not much reason to worry about low level quality standards so long as you know it's correct from a high level. I don't think we've seen what a deep integration between a LLM and a programming language can do, where the type system helps validate the LLM output, and the LLM has a type checker integrated into its training process.
We're not quite there yet, but while regular programming is quite tough for AI due to how fuzzy it is, formal proofs are something AI is already very good at.
Compiling does not differentiate between True and False, so no safety for that escape pod door.
But I definitely want as much as possible to be automated and formally correct, which is why I wrote what I wrote.
Even with, the aforementioned "have an LSP agent work through type errors", it may be faster to just do it yourself than wait for an LLM to spit out what may be correct.
- Has obvious bugs, many at runtime - Has subtle bugs. - Is inefficient.
All of which it will generally fix when asked. But how useful is it if I need to know all the problems with its code beforehand? Then it responds with the same-ish wrong answer the next time.
Still a long way to go IMO.
Is this a brother or a cousin of the "sufficiently advanced compiler"? :-)
it merely remains to build a debugger for your Turing-complete type system, and the toolchain will be ready for production
I believe the claim was that a sufficiently advanced compiler could do a lot of optimization and greatly improve performance. Maybe my claim here will turn out the same.
Even if you could, you probably wouldn't want to make any change a breaking change by exposing implementation details.
* build a codegen for Idris2 and a rust RT (a parallel stack "typed" VM)
* a full application in Elm, while asking it to borrow from DT to have it "correct-by-construction", use zippers for some data structures… etc. And it worked!
* Whilst at it, I built Elm but in Idris2, while improving on the rendering part (this is WIP)
* data collators and iterators to handle some ML trainings with pausing features so that I can just Ctrl-C and continue if needed/possible/makes sense.
* etc.
At the end I had to rewrite completely some parts, but I would say 90% of the boring work was correctly done and I only had to focus on the interesting bits.
However it didn’t deliver the kind of thorough prep work a painter would do before painting a house when asked for. It simply did exactly what I asked, meaning, it did the paint and no more.
(Using 4o and o1-preview)
That's for this: https://tools.simonwillison.net/ocr
Again, please don’t be offended; what you’re doing is great and I dearly appreciate you sharing your experience! Just be aware that the stuff you’re demonstrating isn’t (hasn’t been, for me at least) capable of producing the kind of complexity I need while using the languages and tooling required in my environment. In other words, while everything of yours I’ve seen has intellectual and perhaps even monetary value, that doesn’t mean your examples or strategies work for all use-cases.
As such, for one-shot apps like these there's a strict limit to how much you can get done purely though prompting in a single session.
I work on plenty of larger projects with lots of LLM assistance, but for those I'm using LLMs to write individual functions or classes or templates - not for larger chunks of functionality.
That’s an important detail that is (intentionally?) overlooked by the marketing of these tools. With a human collaborator, I don’t have to worry much about keeping collab sessions short—and humans are dramatically better at remembering the context of our previous sessions.
> I work on plenty of larger projects with lots of LLM assistance, but for those I'm using LLMs to write individual functions or classes or templates - not for larger chunks of functionality.
Good to know. For the larger projects where you use the models as an assistant only, do the models “know” about the rest of the project’s code/design through some sort of RAG or do you just ask a model to write a given function and then manually (or through continued prompting in a given session) modify the resulting code to fit correctly within the project?
In my experience most of effective LLM usage comes down to carefully designing the contents of the context.
My experience with Copilot (which is admittedly a few months outdated; I never tried Cursor but will soon) shows that it’s really good at inline completion and producing boilerplate for me but pretty bad at understanding or even recognizing the existence of scaffolding and business logic already present in my projects.
> but I mostly just paste exactly what I want the model to know into a prompt.
Does this include the work you do on your larger projects? Do those larger projects fit entirely within the context window? If not, without RAG, how do you effectively prompt a model to recognize or know about the relevant context of larger projects?
For example, say I have a class file that includes dozens of imports from other parts of the project. If I ask the model to add a method that should rely upon other components of the project, how does the model know what’s important without RAG? Do I just enumerate every possible relevant import and include a summary of their purpose? That seems excessively burdensome given the purported capabilities of these models. It also seems unlikely to result in reasonable code unless I explicitly document each callable method’s signature and purpose.
For what it’s worth, I know I’ve been pretty skeptical during our conversations but I really appreciate your feedback and the work you’ve been doing; it’s helping me recognize both the limitations of my own knowledge and the limitations of what I should reasonably expect from the models. Thank you, again.
I'm very selective about what I give them. For example, if I'm working on a Django project I'll paste in just the Django ORM models for the part of the codebase I'm working on - that's enough for it to spit out forms and views and templates, it doesn't need to know about other parts of the codebase.
Another trick I sometimes use is Claude Projects, which allow you to paste up to 200,000 tokens into persistent context for a model. That's enough to fit a LOT of code, so I've occasionally dumped my entire codebase (using my https://github.com/simonw/files-to-prompt/ tool) in there, or selected pieces that are important like the model and URL definitions.
A formal proof is simply a type that allows only the correct implementation at the term level. And specifying this takes ages, as in hundreds of times longer than just writing it out. So you want to use an LLM to save this 1% of time after you've done all the hard work of specifying what the only correct implementation is?
I know it's not what you meant, but tbh I don't think going deeper than Haskell on type-system power is the route to achieve mass adoption.
It’s an approach similar to how I’ve dealt with junior devs in the past. You specify an interface for a class, provide examples as a spec, and you get what you want without colliding with the main project.
For sanity’s sake, I keep these AI generated modules in single files just so it’s an easy copy and paste into ChatGPT.
Your experience with and approach to juniors is different than my own. I don’t ask juniors to write a single class file; I give them a design document, an API spec, and documentation for standard practices, then I work as closely with them as needed to get the results I need and the experience they need. This approach works well for me and the vast majority of juniors with whom I’ve worked because we can pretty quickly identify gaps in their knowledge so we can then provide experience and education that benefits both parties. The same approach has failed miserably for me when pairing with an LLM for anything other than trivial code-generation tasks. The models don’t learn (sure, some providers offer “memory” but the limits of those features are pretty obvious once you try to use ‘em in practice for anything other than “don’t forget that I like tacos, the color blue, and sci-fi”)
> For sanity’s sake, I keep these AI generated modules in single files just so it’s an easy copy and paste into ChatGPT.
That’s not acceptable for production-quality code—at least not in my environment.
I know the organizational style doesn't fit a typical "production" set up, but the reality is the code produced is very good. I only set it up this way so I guarantee I can continue iterating on a module without too much pain.
Also, who cares if I have a way more files if I'm still building features for my customers?
Using GitHub copilot, I just tell it to style its code like an example and it gets pretty close.
The easier way to integrate into an existing code base is just to refactor the code yourself. AI gives a working version, you refactor and move on. For me this has been a huge productivity boost from writing everything from scratch
Tell claude or your favorite LLM to write a full plan to implement what you need in such a way that your coworker can implement it.
Copy the result into aider, and check the results!
That's a feature, not a bug. Complexity is something to avoid.
Being able to tackle complex tasks is still a real challenge for the current models and approaches and not all problems can be solved with elegant solutions.
Complexity is inherent in many problem spaces.
"It's even crazier to me that we've just... Accepted it, and are in the process of taking it for granted. This type of technology was a moonshot 2 years ago, and many experts didn't expect it in the lifetimes of ANYONE here - and who knew the answer was increasing transformers and iterating attention?
And golly, there are a LOT of nay-sayers of the industry. I've even heard some folks on podcasts and forums saying this will be as short-lived and as meaningless as NFTs. NFTs couldn't re-write my entire Python codebase into Go, NFTs weren't ever close to passing the bar or MCAT. This stuff is crazy!"
Neither can LLMs. They can produce output that looks like a plausible re-write of your codebase, but on closer inspection turns out to have many minor and major errors everywhere.
The problem is that the closer inspection part is very often more work than writing the code by hand in the first place.
There hasn't been enough evidence for me that this will be possible to fix.
You're posting on a thread that hyperlinks to a list of code and Claude Artifacts for pet-projects that can make thousands a month with some low-effort PPC and an AdWords embed, and some mid-size projects that can be anything from grounds to a promotion at a programming role - to the MVP for a PMF-stage startup.
What, specifically, would pivot your pre-conceived notions?
You'd be surprised how big the "(simple task) online" search query market is, and how much they are usually multi-visit monthly customers, and how much their ad space is worth.
I cannot stress this enough, just because it's simple does not mean it's not lucrative.
Besides all of this is completely besides the point. This isnt useful for a programmer. These examples are barely useful for a layperson. And said layperson is paying money and time for this.
Not to attack you but from your profile it sounds more like your the typical marketing grifter talking big. Why is none of those projects in the list you mention there?
Looking deeper you got lots of projects with parts of your websites just broken and seem to be peddling what looks like life insurance scams.
Sure, there's probably more projects of mine, over the years, that are more broken than not. I've cast several wide nets for product creations and iterations over the years, and kept maintaining the more "fittest" of the bunch. Billit's probably the only one that's broken AND I have no control over it; I sold it. I don't know what else to tell you here, perhaps you value a lesser repertoire with higher rigidity?
I'm not sure how to address your pre-conceived notions that a single industry I've worked in, at large, is a scam. Also, the one company mentioned in life insurance doesn't have a backlink on Lead EnGen - so I especially don't know what you're talking about when you say "peddling".
Not surprised at all; my inability to find examples of /how/ someone might get an LLM to produce—or even intelligently collaborate on—something useful, well… it says a lot about how much junk is out there contributing to the noise.
Is this an exaggeration? Because this is absolutely not true. I'm a beginner in JavaScript and other web stuff and I absolutely can't build it in many days.
Of course if you show me the answers I will think I can do it easy, because answers in programming are always easy (good answers anyways). It's the process of finding the answer that is hard. And I'm not a bad programmer either, I'm at least mediocre, I'm just unfamiliar with web technology.
I've seen a complete no-code person install whisper x with a virtual Python environment and use it for realtime speech to text in their Japanese lessons, in less than 3 hours. You can do a simple library call in JavaScript.
Why don't you give that a go? See if you can knock out a QR code reading UI in JavaScript in less than 3 minutes, complete with drag-and-drop file opening support.
(I literally built this one in a separate browser tab while I was actively taking notes in a meeting)
I say three minutes because my first message in https://gist.github.com/simonw/c2b0c42cd1541d6ed6bfe5c17d638... was at 2:45pm and the final reply from Claude was at 2:47pm.
Those prompts might be sufficient enough to result in deployable HTML/JS code comprised of a couple hundred lines of code but that’s fairly trivial in my definition. I’m not trying to be rude or disrespectful to you; within my environment, non-trivial projects typically involve an entire microservice doing even mildly interesting business logic and offering some kind of API or integration with another, similarly non-trivial API—usually both. And they’re typically built on languages that are compiled either to libraries/executables or they’re compiled to bytecode for the JVM/CLR.
Again, I’m not trying to be disrespectful. You’ve built some really great stuff and I appreciate you sharing your experiences; I wish I knew some of the things you do—you keep writing about your experiences and I’ll keep reading ‘em, we can learn together. The problem is that I’m beginning to recognize that these models are perhaps not nearly ready for the kinds of work I want or need to do, and I’m feeling a bit bummed that the capabilities the industry currently touts are significantly more overhyped than I’d imagined.
I have a bunch more larger projects on my blog: https://simonwillison.net/tags/ai-assisted-programming/
I do a whole lot of API integration work with Claude, generally by pasting in curl examples to illustrate the API. Here's an example from this morning: https://til.simonwillison.net/llms/prompt-gemini
But it's more than that, isn't it? It has a whole interface, drag and drop functionality etc. Front end code is real code mate.
The project i see people build in Java classes on the other hand is a CLI version of Battleships. And honestly that is more complex than the presented projects solved by Claude.
Your personal experience is one point of many. That these projects seem hard to you doesn't make it so for the average person. When i say "a beginner can do it", there's bound to be some who can't. I'm sorry, if these projects take you weeks that is a problem.
And it is strange that Claude picked the AudioRecorder when the MediaRecorder exists. I'd wager a beginner would have used the latter(i don't use javascript and am not better than a beginner in any way, but i found that) since it outputs a straight wav file and doesn't need the encoding step. And since the data isn't streamed to OpenAI there's no need for the audio chunks that AudioRecorder provides. So Claude did it in an unnecessarily complex way, that doesn't make the problem complex.
If you want details of more complex projects I've written using Claude here are a few - in each case I provide the full chat transcript:
- https://simonwillison.net/2024/Aug/8/django-http-debug/
- https://simonwillison.net/2024/Aug/27/gemini-chat-app/
- https://simonwillison.net/2024/Aug/26/gemini-bounding-box-vi...
That said, after looking at a couple of your sessions, I don’t see anything you’re doing that I’m not—at least in terms of prompting. Your prompts are a bit more terse than mine (I can be long-winded so I’ll give brevity a try with my next project) but the structure and design descriptions are still there. That would suggest the differences in our experience boils down to the languages with which we choose or are required to work; maybe there’s a stylistic or cultural difference in how one should prompt a model in order to generate a Python project and how one should prompt for a Haskel or Scala/Java project; surely not though, right?
I’m not giving up and I’ll keep playing with these models but for now, given my use-case at least, they still seem to be far more capable at rubber-ducking with me than they are as a pair programming partner.
20 years later that’s still not the case, because it turns out NN/ML can do some very impressive things at the 99% correct level. The other 1% ranges in severity from “weird lane change” to “a person riding a bicycle gets killed”.
GPT-3.5 was the DARPA grand challenge moment, we’re still years away from LLM being reliable - and they may never be fully trustworthy.
So? Neither are humans. Neither is google search. Chatgpt doesn't write bug free code, but neither do I.
The question isn't "when will it be perfect". The question is "when will it be useful?". Or, "When is it useful enough that you're not employable?"
I don't think its so far away. Everyone I know with a spark in their eye has found weird and wonderful ways to make use of chatgpt & claude. I've used it to do system design, help with cooking, practice improv, write project proposals, teach me history, translate code, ... all sorts of things.
Yeah, the quality is lower than that of an expert human. But I don't need a 5 star chef to tell me how long to put potatoes in the oven, make suggestions for characters to play, or listen to me talk about encryption systems and make suggestions.
Its wildly useful today. Seriously, anyone who says otherwise hasn't tried it or doesn't understand how to make proper use of it. Between my GF and I, we average about 1-2 conversations with chatgpt per day. That number will only go up.
That’s not remotely true. I am an expert, and it’s incredibly clear to me how bad LLM are. I still use them heavily, but I don’t trust any output that doesn’t conform to my prior expert knowledge and they are constantly wrong.
I think what is likely happening is many people aren’t an expert in anything, but the LLM makes them feel like they are and they don’t want that feeling to go away and get irrationally defensive at cogent criticism of the technology.
And that’s all it is, a new technology with a lot of hype and a lot of promise, but it’s not proven, it’s not reliable, and I do think it is messing with people’s heads in a way that worries me greatly.
For context, I'm an expert too. And I had the same experience as you. When I asked it questions about my area of expertise, it gave me a lot of vague, mutually contradictory, nonsensical and wrong answers.
The way I see it, ChatGPT is currently a B+ student at basically everything. It has broad knowledge of everything, but its missing deep knowledge.
There are two aspects to that to think about: First, its only a B+ student. Its not an expert. It doesn't know as much about family law as a family lawyer. It doesn't know as much about cardiology as a cardiologist. It doesn't know as much about the rust borrow checker as I do.
So LLMs can't (yet) replace senior engineers, specialist doctors, lawyers or 5 star chefs. When I get sick, I go to the doctor.
But its also a B+ student at everything. It doesn't have depth, but it has more breadth of knowledge than any human who has ever lived. It knows more about cooking than I do. I asked it how to make crepes and the recipe it gave me was fantastic. It knows more about australian tax law than I do. It knows more about the american civil war than I do. It knows better than I do what kind of motor oil to buy for my car. Or the norms and taboos in posh british society.
For this kind of thing, I don't need an expert. And lots of questions I have in life - maybe most questions - are like that!
I brainstormed some software design with chatgpt voice mode the other day. I didn't need it to be an expert. I needed it to understand what I was saying and offer alternatives and make suggestions. It did great at that. The expert (me) was already in the room. But I don't have encyclopedic knowledge of every single popular library in cargo. ChatGPT can provide that. After talking for awhile, I asked it to write example code using some popular rust crates to solve the problem we'd been talking about. I didn't use any of its code directly, but that saved me a massive amount of time getting started with my project.
You're right in a way. If you're thinking of chatgpt as an all knowing expert, it certainly won't deliver that (at least not today). But the mistake is thinking its useless as a result of its lack of expertise. There's thousands and thousands of tasks where "broad knowledge, available in your pocket" is valuable already.
If you can't think of ways to take advantage of what it already delivers, well, pity for you.
But just now had a fairly frequent failure mode: I asked it a question and it gave me a super detailed and complicated solution that a) didn’t work, and b) required serious refactoring and rewriting.
Went to Google, found a stack overflow answer and turns out I needed to change a single line of code, which was my suspicion all along.
Claude was the same, confidentially telling me to rewrite a huge chunk of code when a single line was all that was needed.
In general Claude wants you to write a ton of unnecessary code, ChatGPT isn’t as bad, but neither writes great code.
The moral of the story is I knew the gpt/claude solutions didn’t smell right which is why I tried Google. If I didn’t have a nose for bad code smells I’d have done a lot of utterly stupid things, screwed up my code base, and still not have solved my oroblwm.
At the end of the day I do use LLM, but I’m experienced so it’s a lot safer than a non-experienced person. That’s the underlying problem.
My point is that even now, you're only talking about using chatgpt / claude to help you do the thing you already know how to do (programming). You're right of course. Its not currently as good at programming as you are.
But so what? The benefit these chat bots provide is that they can lend expertise for "easy", common things that we happen to be untrained at. And inevitably, thats most things!
Like, ChatGPT is a better chef than I am. And a better diplomat. A better science fiction writer. A better vet. And so on. Its better at almost every field you could name.
Instead of taking advantage of the fields where it knows more than you, you're criticising it for being worse than you at your one special area (programming). No duh. Thats not how it provides the most value.
It’s like false memories of events that never occurred, but false knowledge - you think you have learned something, but a non-trivial percent of it, that you have no way of knowing, is flat out wrong.
It’s not a “helpful B+ student” for most people , it’s a teacher, and people are learning from it. But they are learning subtly wrong things, all day, every day.
Over time, the mind becomes polluted with plausible fictions across all types of subjects.
The internet is best when it spreads knowledge, but I think something else is happening here, and I think it’s quite dangerous.
The news has an equivalent: The Gell-Mann amnesia effect, where people read a newspaper article on a topic they're an expert on and realise the journalists are idiots. Then suddenly forget they're idiots when they read the next article outside their expertise!
So yes, I agree that its important to bear in mind that chatgpt will sometimes be confidently wrong.
But I counter with: usually, remarkably, it doesn't matter. The crepe recipe it gave produced delicious crepes. If it was a bad recipe I would have figured that out with my mouth pretty quickly. I asked it to brainstorm weird quirks for D&D characters to have, some of the ideas it came up with were fabulous. For a question like that, there isn't really such a thing as right and wrong anyway. I was writing rust code, and it clearly doesn't really understand borrowing. Some code it gives just doesn't compile.
I'll let you in on a secret: I couldn't remember the name of the gell-mann amnesia effect when I went to write this comment. A few minutes ago I asked chatgpt what it was called. But I googled it after chatgpt told me what it was called to make sure it got it right so I wouldn't look like an idiot.
I claim most questions I have in life are like that.
But there are certainly times when (1) its difficult to know if an answer is correct or not and (2) believing an incorrect answer has large, negative consequences. For example, Computer security. Building rocket ships. Research papers. Civil engineering. Law. Medicine. I really hope people aren't taking chatgpt's answers in those fields too seriously.
But for almost everything else, it simply doesn't matter that chatgpt is occasionally confidently wrong.
For example, if I ask it to write an email for me, I can proofread the email before sending it. The other day asked it for scene suggestions in improv, and the suggestions were cheesy and bad. So I asked it again for better ones (less chessy this time). I ask for CSS and the CSS doesn't quite work? I complain at it and it tries again. And so on. This is what chatgpt is good for today. It is insanely useful.
It’s never the product or its marketing that’s at fault; only my own.
In my experience, the value proposition for ChatGPT lies in its ability to generate human language at a B+ level for the purposes of a an interactive conversation; its ability to generate non-trivial code has proven to be terribly disappointing.
This is just not true. My reaction to the second challenge race (not the first) in 2005 was, it was a 0-to-1 kind of moment and robocars were now coming, but the timescale was not at all clear. Yes you could find hype and blithe overoptimism, and it's convenient to round that off to "everybody" when that's the picture you want to paint.
> 20 years later that’s still not the case
Also false. Waymo in public operation and expanding.
And despite everything, Waymo is not quite there yet. It's able to handle certain areas at a limited scale. Amazing, yes, but it has not changed the reality of driving for 99.9% of the population. Soon it will, I'm sure, but not yet.
Fact is Google will never break even on the investment and it’s more or less a white elephant. I don’t think it’s even accurate to call it a Beta product, at best it’s Alpha.
... followed by speculation about the future.
> [not everywhere]
The standard you proposed was "on the road". In their service areas (more than "one", they've been in Phoenix for some time) anyone can install their app and get a ride.
I shouldn't have poked my nose in here, I was just kind of croggled to see someone answer "ideological battle" by bringing up another argument where they don't seem to care about facts.
It's the long, long, long tail of edge cases - not just porting them, but even identifying them to test - that slow or doom most real-world human rewrites, after all.
You have to review the code these thing write for you, just like code from any other collaborator.
oh ok. this is quite different than what I was picturing. So far this is my favorite use case of LLMs, they seem very good at this.
I mistakenly thought you were using it almost as a black box compiler. "look it ported it to Rust, I can't make sense of it, but it seems to work and no segfaults!".
What you say sounds pretty sensible, and it is a very nice practical example of the power of LLMs.
I will tell you this, the second most-used language in my day-to-day (TypeScript) is one that I've seldom sat down and learned, rely on AI for me to create and streamline, and has not given me any issues for 16 months running (since the project has started).
AI won't replace jobs; but someone who knows how to use it better will.
So where do I go to learn how to use ‘em better? Or, at least, examples of what works so I can understand what I’m doing wrong?
1.) https://platform.openai.com/docs/guides/prompt-engineering 2.) https://www.promptingguide.ai/
“Since I’m dealing with models rather than other engineers should I expect the process of breaking down the problem to be dramatically different from that of writing design documents or API specs? I rarely have difficulty prompting (or creating useful system prompts for) models when chatting or doing RAG work with plain English docs but once I try to get coherent code from a model things fall apart pretty quickly.”
Said another way, I long ago learned as an engineer how to do the things you’re suggesting (they are skills I’ve used and evolved over more than twenty years as a professional software engineer) but, in my experience, those same skills do not seem to apply when trying to do non-trivial code-generation tasks for Java/Scala/Python projects with an LLM.
I’ve tried prompting ’em with my design documentation and API specs. I’ve tried prompting ’em with a pared-down version of my docs/specs in order to be more succinct. I’ve tried expanding my docs/specs to be more concrete and detailed. I’ve tried very short prompts. I’ve tried very detailed and lengthy prompts. I’ve tried tweaking system prompts. I’ve tried starting with prompts that limit the scope of the project then expanding from there. I’ve tried uploading the docs/specs so that the models can reference them later. I’ve tried giving ‘em access to entire repositories. I’ve tried so many things all to no avail. The best solution I’ve thus far found in these threads is to just try to fit the entirety of a project within the limits of the context window and/or to just keep my whole project in a few short files; that may be sufficient for small projects but it’s not possible nor even reasonable given the size and complexity of projects with which I work.
As I’ve said elsewhere, I dearly /want/ these things to work in my environment and for my use-cases as capably as they do in/for yours—this stuff is really interesting and I enjoy learning how to do new things—but after reading all the comments in this thread and others I don’t think the needs of my environment are supported by these models. Maybe they will be someday. I’ll keep playing with things but as of right now I see a significant impedance mismatch between the confidence others have in the models’ ability to do complex coding tasks compared to the kinds of tasks I’ve seen demonstrated here and elsewhere.
Truth is, this has been a learning process for us all with these tools, but it needs to be understood -- especially going in -- that these models excel at translation tasks and constrained problem spaces but can struggle with generating cohesive, large-scale code without specific hand-holding.
This is generally what I do:
1. Start with the "Whole" Picture: Models often work best when they know the final goal and the prompter has worked backwards from them. Think ontologically: define the problem as if you’re describing it to a junior dev colleague who only understands outcomes, not methods. Instead of just prompting with specs, explain the end-state you want (even simple features like error handling or specific libraries). If you have ANY achieved method or inclusion for what the end-state should include, you write it out clearly.
2. Break Down the Process: Models handle complexity better if it's broken down into micro-tasks. Instead of expecting it to design an entire feature, ask for components step-by-step, integrating each output with the rest manually. There is a very decent chance that you have to do this across multiple new chats, after 3-5 iterations in, the AI will most likely crash and burn. At that point, you open a new chat, paste in the whole working codebase, and picked up from where you left off in the last chat. You have to do this A LOT.
3. Iterative Refinement: When the model generates code, go over it closely. Check for errors, then use targeted prompts to fix specific issues rather than requesting whole rewrites. Point out exact issues and ask for specific fixes; this prevents the model from “looping” through similar incorrect solutions.
Some Hacks I Use As Well:
1. Contextual Repetition: Reinforce key components (e.g., function structure, file organization) to avoid losing them in longer prompts.
2. Use “As if” Phrasing: Prompt the model to act “as if” it’s coding for a hypothetical person (e.g., a junior dev). It’s surprisingly effective at generating more thoughtful code with this type of frame.
3. Ask for Questions: Have the model ask you clarifying questions if it’s “unsure.” This can uncover key details you may not have thought to include.
4. Remind It What It Is Doing: Sounds counter-productive, but almost all of my code chats end with a description of what exactly I expect from the AI, iterated over the various stunts and "shortcuts" that it has taken over the years I've used it. I generally say "Write the code in full with no omissions or code comment-blocks or 'GO-HERE' substitutions" (this is directly because AI has generally pulled "/rest of code goes here/ on me several times), "write the code in multiple answers if you must, pausing at the generic character limit and resuming when I say 'continue' in the next message" (because I've had "errors" from code generation in the past because the chat reply processing time had timed out).
It's a labor of love and it's things you learn over time, and it won't happen if you don't put the work in.
-------------------------
I wrote all of this haphazardly in a Google Doc. GPT-4 organized it for me cleanly.
Beware of trying to get the LLM to output exactly the code you want. You get points for checking code in git and sending PRs, not tokens the LLM outputs. If it's being stupid and going in circles, or you know from experience that the particular LLM used will (they vary greatly in quality), you can just copy the code out (if you're not using some sort of AI IDE), fix it, then paste that in and/or commit it.
Some may ask, if you have to do that, then why use an LLM in the first place. It's good at taking small/medium conceptual tasks and breaking them down, and it's also a faster typer than me. Even though I have to polish its output, I find it easier to get things done because I can focus more on the higher level (customer) issues while the LLM gets started with lower level details on implementing/fixing things.
That said, l feel like there’s a mutual-exclusivity problem between ‘Start with the "Whole" Picture’ and ‘Break Down the Process’.
For example, how does this from your first suggestion:
> explain the end-state you want (even simple features like error handling or specific libraries). If you have ANY achieved method or inclusion for what the end-state should include, you write it out clearly.
not contradict this from your second suggestion:
> Instead of expecting it to design an entire feature, ask for components step-by-step
Additionally, you said:
> There is a very decent chance that you have to do this across multiple new chats, after 3-5 iterations in, the AI will most likely crash and burn. At that point, you open a new chat, paste in the whole working codebase, and picked up from where you left off in the last chat. You have to do this A LOT.
But IME, by the time the model chokes on one chat, the codebase is already large enough that pasting the whole thing into another chat typically results in my hitting context-window limits. Perhaps, in the kinds of projects I typically work, a good RAG tool would offer better results?
To be clear, right now I’m only discussing my difficulties with the chatbots offered by the model providers—which, for me, is mostly Claude but also a bit of ChatGPT; my experience with Copilot is outdated so it probably deserves another look, and I’ve not yet tried some of the third-party, code-centric apps like aider or cursor that have previously been suggested, though I will soon.
As for your recommended hacks, these look to be helpful; thank you! The only part I find odd is your inclusion of “Write the code in full with no omissions or code comment-blocks or 'GO-HERE' substitutions”; I myself feel like I get far better results when I ask the model to 1) write full code for the methods that are likely to be the kinds of generic CS logic that a junior would know, 2) write stubs for the business logic, then 3) implementing the more complex business logic myself manually. IOW—and IME—they’re really good at writing boilerplate and generating or reasoning about junior-level CS logic. That’s indeed helpful to me, but it’s a far cry from the kinds of “ChatGPT can write entire apps with minimal effort” hype I keep seeing, and it’s only marginally better, IME at least, than what I’ve been able to do with the inline-completion and automatic boilerplate features that have been included in the IDEs I’ve used for over a decade.
> It's a labor of love and it's things you learn over time, and it won't happen if you don't put the work in.
Indeed. I do love playing with this stuff and learning more. Thank you again for sharing your knowledge!
> I wrote all of this haphazardly in a Google Doc. GPT-4 organized it for me cleanly.
I am regularly impressed at how well these models behave when asked to summarize a document or even when asked to expand a set of my notes into something more coherent; it’s truly remarkable!
I don't think either of you are wrong; it just heavily depends on the complexity of the app and how familiar LLMs are with it.
E.g. rewriting a web scraper, CRUD backend or a build script? Sure, maybe. Rewriting a bootloader, compiler or GUI app? No chance.
"Yes, AI can make human sounding sentences, but can it play chess?"
"Well yes, it can play chess. But no computer can beat a human grandmaster at chess."
"Well it beat Kasperov - but it has no hope of beating a human at Go."
"Its funny - it can beat humans at go but still can't speak as well as a toddler."
"Alright it can write simple problems, but it introduces bugs in anything nontrivial, and it can't fix those bugs!"
I write bugs in anything nontrivial too! My human advantages are currently that I'm better at handling a large context, and I can iterate better than the computer can.
But - seriously, do you think innovation will stop here? Did the improvements ever stop? It seems like a pretty trivial engineering problem to hook an AI up to a compiler / runtime so it can iterate just like we can. Anthropic is clearly already starting to try that.
I agree with you, today. I used claude to help translate some rust code into typescript. I needed to go through the output with a fine toothed comb to fix a lot of obvious bugs and clean up the output. But the improvement over what was possible with GPT3.5 is totally insane.
At the current rate of change, I give it 5-10 years before we can ask chatgpt to make a working compiler from scratch for a novel language.
Speaking as a software engineer, I feel seen.
I think you're going to be very disappointed.
"There is superstition about creativity, and for that matter, about thinking in every sense, and it's part of the history of the field of artificial intelligence that every time somebody figured out how to make a computer do something - play good checkers, solve simple but relatively informal problems - there was a chorus of critics to say, but that's not thinking."
That's from 1979! https://simonwillison.net/2024/Sep/13/pamela-mccorduck-in-19...
Is there something particularly unique about biological circuits that allow thought, as opposed to electronic ones?
> I'm still not convinced birds can fly any more than a rock shaped like a bird would convince me that it's flying.
I don't propose harder tests myself, because it doesn't make sense within my philosophy about this. When those tests are passed, to me it doesn't prove that the AI proponents are right about their systems being intelligent; it proves that the test-setters were wrong about what intelligence entails.
Nobody made any claim in this thread that modern AIs have thoughts.
What these (increasingly complicated) tests do is demonstrate the capacity to act intelligently. Ie, make choices which are aligned with some goal or reward function. Win at chess. Produce outputs indistinguishable from the training data. Whatever.
But you're right - I'm smuggling in a certain idea of what intelligence is. Something like: Intelligence is the capacity to select actions (outputs) which maximise an externally defined given reward function over time. (See also AIXI: https://en.wikipedia.org/wiki/AIXI ).
> When those tests are passed, [..] to me it proves that the test-setters were wrong about what intelligence entails.
It might be helpful for you to define your terms if you're going to make claims like that. What does intelligence mean to you then? My best guess from your comment is something like "intelligence is whatever makes humans special". Which sounds like a useless definition to me.
Why does it matter if an AI has thoughts? AI based systems, from MNIST solvers to deep blue to chatgpt have clearly gotten better at something. Whatever that something is, is very very interesting.
Yes, you understand me. I simply come in with a different idea.
>AI based systems, from MNIST solvers to deep blue to chatgpt have clearly gotten better at something. Whatever that something is, is very very interesting.
Certainly the fact that the outputs look the way they do, is interesting. It strongly suggests that our models of how neurons work are not only accurate, but creating simulations according to those models has surprisingly useful applications (until something goes wrong. Of course, humans also have an error rate, but human errors still seem fundamentally different in kind.)
The idea of ChatGPT being asked to "think" just reminds me of Pozzo from Waiting for Godot.
It just can't translate the kinds of programs I write between languages on its own. Today.
Another way to look at it is that we're refining our understanding of the capabilities of machine learning in real time. Otherwise one could make basically the same argument about any field that progresses - take our theories of gravity for example. Was Einstein moving the goalposts? Or was he building on previous work to ask deeper questions?
Set against the backdrop of extraordinary claims about the abilities of LLMs, I don't think it's unreasonable to continue pushing for evidence.
I mean, we first put up a ladder and we could reach the peaches! Next, we put a ladder next to the apple tree and we could pluck those. Now, in their incessant goal post moving people said, great, now setup a ladder to the moon. There is no reason to assume this won’t work. None at all. People are just complaining and being angry at losing their fancy jobs.
More specific: it cannot learn, because it has no concept of learning from first principles. There is no way out, not even a theoretical one.
We've seen that the rate of change went up hugely when LLMs came around. But the rate of change was much lower before that. It could also be much slower for the foreseeable future.
LLMs are only as good as their training materials. But a lot of what programmers do is not documented anywhere, it happens in their head, and it is in response to what they see around them, not in what they scrape from the web or books.
Maybe what is needed is for organizations to start producing materials for AI to learn from, rather than assuming that all they need is what they find on the web? How much of the effort to "train" AI is just letting them consume the web, and how much is concsiously trying create new learning materials for AI?
If it gives you a solution that is wrong, you have to point it at, then it will give you a second version , if that is also wrong, it will then slightly modify the same solutions over and over again instead of actually fixing the issue.
It gets stuck in a loop of giving you 2-3 versions of the same solution with the slightly different outputs.
It's only useful for boilerplate code and even then, you have to clean it up..
LLMs are insidious, it feeds into "everything is simple" concept a lot of us have of the world. We ask an LLM for a project plan and it looks so good we're willing to fire our TPM, or a TPM asks the LLM for code and it gives them code that looks so good they question the value of an engineer. In reality, the LLM cannot do either role's job well.
While I appreciate the suggestion that I might be an expert, I am decidedly not. That said, I’ve been writing what the companies I’ve worked for would consider “mission critical” code (mostly Java/Scala, Python, and SQL) for about twenty years, I’ve been a Unix/Linux sysadmin for over thirty years, and I’ve been in IT for almost forty years.
Perhaps the modernity and/or popularity of the languages are my problem? Are the models going to produce better code if I target “modern” languages like Go/Rust, and the various HTML/JS/FE frameworks instead of “legacy” languages like Java or SQL?
Or maybe my experience is too close to bare metal and need to focus on more trivial projects with higher-level or more modern languages? (fwiw, I don’t actually consider Go/Rust/JS/etc to be higher-level or more “modern” languages than the JVM languages with which I’m experienced; I’m open to arguments though)
> LLMs are insidious, it feeds into "everything is simple" concept a lot of us have of the world.
Yah, that’s what I mean when I say I feel gaslit.
> In reality, the LLM cannot do either role's job well.
I am aware of this. I’m not looking for an agent. That said, am I being too simplistic or unreasonable in expecting that I too could leverage these models (albeit perhaps after acquiring some missing piece of knowledge) as assistants capable of reasoning about my code or even the code they generate? If so, how are others able to get LLMs to generate what they claim are “deployable” non-trivial projects or refactorings of entire “critical” projects from the Python language to Go? Is someone lying or do I just need (seemingly dramatically) deeper knowledge of how to “correctly” prompt the models? Have I simply been victim of (again, seemingly dramatically) overly optimistic marketing hype?
Claude does benefit from some architectural direction. I think it's better at extending than creating from whole-cloth. My workflow looks like:
1) Rough out some code, say a smart contract with the key features
2) Tell claude to finish it and write extensive testing.
3) Run abigen on the solidity to get a go library
4) Tell claude to stub out golang server event handlers for every event in the go library
5) Create a react typescript site myself with a basic page
6) Tell claude to create an admin endpoint on the react site that pulls relevant data from the smart contracts into the react site.
6.5) Tell claude to redesign the site in a preferred style.
7) Go through and inspect the code for bugs. There will be a bunch.
8) For bugs that are simple, prompt Claude to fix: "You forgot x,y,z in these files. fix it."
9) For bugs that are a misunderstanding of my intent, either code up the core loop directly that's needed, or negotiate and explain. Coding is generally faster. Then say "I've fixed the code to work how it should, update X, Y, Z interfaces / etc."
10) for really difficult bugs or places I'm stumped, tar the codebase up, go to the chat interface of claude and gpto1-preview, paste the codebase in (claude can take a longer paste, but preview is better at holistic bugfixing), and explain the problem. Wait a minute or two and read the comments. 95% of the time one of the two LLMS is correct.
This all pretty much works. For these definitions of works:
1) It needs handholding to maintain a codebase's style and naming.
2) It can be overeager: "While I was in that file, I ..."
3) If it's more familiar with an old version of a library you will be constantly fighting it to use a new API.
How I would describe my experience: a year ago; it was like working with a junior dev that didn't know much and would constantly get things wrong. It is currently like working with a B+ senior-ish dev. It will still get things wrong, but things mostly compile, it can follow along, and it can generate new things to spec if those requests are reasonable.
All that to say, my coding projects went from "code with pair coder / puppy occasionally inserting helpful things" to "most of my time is spent at the architect level of the project, occasionally up to CTO, occasionally down to dev."
Is it worth it? If I had a day job writing mission critical code, I think I'd be verrry cautious right now, but if that job involved a lot of repetition and boiler plate / API integration, I would use it in a HEARTBEAT. It's so good at that stuff. For someone like me who is like "please extend my capacity and speed me up" it's amazing. I'd say I'm roughly 5-8x more productive. I love it.
Claude has the added benefit that you can yell at it, and it won’t hold it against you. You know, speaking of pairing with a junior dev.
Yet.
I don’t look forward to the day these models are trained on all the context we’ve fed their predecessors; if AGI is possible, it’s gonna hate us. :)
It can also index docs pages of newer APIs and/or search the web to find latest info of newer libraries, so you won't struggle with issue #3
Learn how to think ontologically and break down your requests first by what you're TRULY looking for, and then understand what parts would need to be defined in order to build that system -- that "whole". Here's some guides:
1.) https://platform.openai.com/docs/guides/prompt-engineering 2.) https://www.promptingguide.ai/
> Learn how to think ontologically and break down your requests first by what you're TRULY looking for, and then understand what parts would need to be defined in order to build that system -- that "whole".
Since I’m dealing with models rather than other engineers should I expect the process of breaking down the problem to be dramatically different from that of writing design documents or API specs? I rarely have difficulty prompting (or creating useful system prompts for) models when chatting or doing RAG work with plain English docs but once I try to get coherent code from a model things fall apart pretty quickly.
Oh, wow, it honestly just occurred to me that examples of how to prompt a model to produce a certain kind of content might be considered, more or less, some kind of trade secret vaguely akin to a secret recipe. That would be a bit depressing but I get it.
- https://gist.github.com/simonw/0a4b826d6d32e4640d67c6319c7ec... - most of the work
- https://gist.github.com/simonw/a04b844a5e8b01cecd28787ed375e... - some tweaks
Lots more details in my full post about it here: https://simonwillison.net/2024/Oct/18/openai-audio/
I’ve found this to be the difference between writing 50+ prompts / back and for the to get something useful, and when I can get something useful in 1-3 prompts. If you look at Simon’s post, you’ll see that these are all self-contained tools, whose entire scope has been constrained from the outset of the project.
When you go into a large codebase and have to change some behavior, 1) you usually don’t have the detailed solution articulated in your mind before looking at the codebase. 2) That “solution” likely consists of a large number of small decisions / judgements. It’s fundamentally difficult to encode a large number of nuanced details in a concise prompt, making it not worth it to use LLMs.
On the other hand, I built this tool: https://github.com/gr-b/jsonltui that I now use every day almost entirely using Claude. “CLI tool to visualize JSONL with textual interface, localizing parsing errors” almost fully qualifies this. In contrast, my last 8 line PR at my company, while it would appear much simpler on the surface level, contains many more decisions, not just of my own, but reflecting team conversations and expectations that are not written down anywhere. To communicate this shared implicit context with Claude would be so much more difficult than to perform the change myself.
This caught my eye and I’m genuinely curious about what you mean by it. Part of our success with Claude is that we don’t do abstractions, “perfect architecture”, DRY, SOLID and other religions that were written by people who sell consulting in their principles. If we ask LLMs to do any form of “Clean Code” or give them input on how we want the structure, they tend to be bad it.
Hell, if you want to “build from the bottom” you’re going to have to do it over several prompts. I had Claude build a blood bowl game for me, for the fun of it. It took maybe 50 prompts. Each focusing on different aspects. Like, I wanted it to draw the field and add mouse clickable and movable objects with SDL2, and that was one prompt. Then you feed it your code in a new prompt and let it do the next step based on what you have. If the code it outputs is bad, you’ll need to abandon the prompt again.
It’s nothing like getting an actual developer to do things. They can think for themselves and the probability engine won’t do any of that even if it pretends to. Their history for building things from scratch also seems to be quickly “tarnished” within the prompt context. Once they’ve done the original tasks I find it hard to get them to continue on it.
Within my environment, some of those “religions” are more than a requirement; they’re also critical to the long-term maintenance of a large collection of active repositories.
I think one of the problems folks tend to have with following or implementing a “religion” (by which I mean specific structural and/or stylistic patterns within a codebase) comes down to a fear of being stuck forever with a given pattern that may not fit future needs. There’s nothing wrong with iterating on your religion’s patterns as long as you have good documentation with thorough change logs; granted, that can be difficult or even out of reach for smaller shops.
I don’t think any of them are inherently bad, but they lead to software engineering where people over complicate things. Building abstractions they might never need. I’ve specialised in the field of taking startups into enterprise, and 90% of the work is removing the complexity which has made their software development teams incapable of delivering value in a timely manner. Some of this is because they build infrastructures as though they were Netflix or Google, but a lot of times it’s because they’ve followed Clean Code principles religiously. Abstractions aren’t always bad, but you should never abstract until you can’t avoid it. Because two years down into your development you’ll end up with code bases that are so complex that it makes them hard to work with.
Especially when you get the principles wrong. Which many people do. Over all though, we’ve had 20 years of Clean Code, SOLID, DRY and so on, and if you look at our industry today, there is no less of a mess in software engineering than there were before. In fact some systems still run on completely crazy Fortran or COBOL because nobody using “modern” software engineering have been capable of replacing them. At least that’s the story in Denmark, and it hasn’t been for a lock of trying.
I think the main reason many of these principles have become religions is because they’ve created an entire industry of pseudo-jobbers who manage them, work as consultants and what not. All people who are very good at marketing their bullshit, but also people who have almost no experience actually working with code.
Like I said, nothing about them are inherently bad. If you know when to use which parts, but almost nobody does. So to me the only relevant principle is YAGNI. If you’re going to end up with a mess of a code base anyway, you might as well keep it simple and easy to change. I say this as someone who works as an external examiner for CS students, where we still teach all these things that so often never work. In fact a lot of these principles were things I was thought when I took my degree, and many haven’t really undergone any meaningful changes with the lessons learned since their initial creation.
This is VERY different from my own experience. The bugs introduced by the code I’ve tried to generate via LLMs (Mostly Claude, some GPT-4o and o1-preview, and lots of one-off fiddling with local models to see if they’re any better/worse than commercial products) are considerably more numerous (and often more subtle) than what my fellow engineers—juniors included—tend to introduce.
I /want/ these tools to be useful; they haven’t been so far though and I’m kinda stuck on understanding if I’m just not using ‘em right or if they’re even capable of what I want to do. Like I said in a previous comment; I don’t know if I’m being gaslit or if I’m being naive but it feels a lot more like gaslighting.
I mostly agree with what the others are saying. It can generate boilerplate and it can generate simple API calls when there are lots of examples in the training set.
Generating Go is probably easier because at least you get compiler feedback.
Right now the only place it saves me time are with languages I don't know at all and with languages like Bash and SQL where I just can't bring myself to care enough to remember the long tail of more esoteric points that I don't use every day.
Yes, mistakes may happen. However, I’ve used it to translate a fairly complex MIP definition export into a complete CP-SAT implementation.
I use these models all the time for complex tasks.
One major thing that is perhaps not immediately obvious is that the models are only good at translation. If I give it a really good explanation of what I want in code or even English, and ask it to do it another way or implement it with specific tools, I get pretty good output.
Using these to actually solve problems is not possible. Give it a complex problem description with no instructions on how to solve it, and they fail immediately.
For me they save a lot of time on research or general guidance. But when it comes to actual code - not really useful.
I don’t get good output, not when trying to get code that matches a detailed spec (which includes the languages I wish to use, the structure of the APIs, and the libraries I think might be useful) so your suggestion that we’re “just wrong” and your claim that the tools can be used for coding “complex tasks” is difficult for me to swallow.
I’ll admit that perhaps I’m not using ‘em right but that’s why I’m here—to get advice on /how/ to use ‘em correctly; to date, the implications found in the advice I have received is to: - limit my scope to strict HTML/JS (not web/ui frameworks), or - limit the size of the project to a handful (less than ten) of very short files, or - limit my scope to code translation only, or - limit the size of my chat sessions.
Unfortunately, those limitations don’t fit the needs of my environment.
What you are looking at is mangled other people’s work. Great. Thanks AI, for digging it up, but let’s not get too excited here.
I’ll be getting excited when we give it some first principles and it can actually learn on its own.
I completely disagree with this viewpoint. I've created terminal games with my own rules, and that shows me the tool can take what it knows about Rust and assemble code to complete a task. It's essentially doing the same thing a human would.
While I understand the criticism, I sometimes feel that the cynical perspective we bring into these discussions prevents us from offering more meaningful critique.
If your Rust tools are truly non-trivial, I’d love to know /how/ you prompted ChatGPT to accomplish what you want—ideally with chat session transcripts that include the generated artifacts.
I recognize that you may not want or be able to share such things; if that’s the case, can you share or discuss the resources you used in order to learn how to do it? I’ve not yet been able to find any resources that demonstrate what I’d consider to be non-trivial and I’m hoping that’s only a failure of my Google-fu rather than an indictment of the purported capabilities of these models.
We follow a YAGNI approach to our code architecture and abstractions, meaning it’s very straight forward with things happening where they are written and not in 9 million places like Clean Code lovers try to do. Our C services and Libraries are also fairly small and “one purpose”. I’m not sure you would be wrong on larger code bases, at least not right now.
With what we see Claude do now though, I don’t think we’re far from a world where Software Developers are going to do significantly different work. I also think quite a lot of the stuff we do today will no longer exist.
Since mid 2023, I've yet to have an issue
My experience with Python and Cursor could've been better though. For example when making ORM classes (boilerplate code by definition) for sqlalchemy, the assistant proposed a change that included a new instantiation of a declarative base, practically dividing the metadata in two and therefore causing dependency problems between tables/classes. I had to stop for at least 20 minutes to find out where the problem was as the (one n a half LoC) change was hidden in one of the files. Those are the kind of weird bugs I've seen LLMs commit in non-trivial applications, stupid 'n small but hard to find.
But what do I know really. I consider myself a skeptic, but LLMs continue to surprise me everyday.
Since the most contested one in this thread is the "tipping" prompt hack: https://arxiv.org/pdf/2401.03729
If we can’t explain why or how a thing works, we’re going to continue to create things we don’t understand; relying upon our lucky charms when asking models to produce something new is undoubtedly going to result in reinforcement of the importance of those lucky charms. Feedback loops can be difficult to escape.
What matters is it meets functional and non functional requirements.
One of my juniors wrote his first app two years ago fully with chatgpt, could figure out by iteratively asking it how to improve it and solve the bugs.
Then he learned to code properly fascinated by the experience. But the fact remains, he shipped an application that did something for someone while many never did even though they had a degree and a black belt in pointless leet code quizzes.
I'm fully convinced that very soon big tech or a startup will come up with a programming language meant to sit at the intersection between humans and LLMs, and it will be quickly better, faster and cheaper at 90% of the mundane programming tasks than your 200k/year dev writing forms, tables and apis in SF.
Good luck expressing novel requirements in complex operating environments in plain English.
> Then he learned to code properly fascinated by the experience. But the fact remains, he shipped an application that did something for someone while many never did even though they had a degree and a black belt in pointless leet code quizzes.
It's good in the sense that it raises the floor, but it doesn't really make a meaningful impact on the things that are actually challenging in software engineering.
> Then he learned to code properly fascinated by the experience. But the fact remains, he shipped an application that did something for someone while many never did even though they had a degree and a black belt in pointless leet code quizzes.
This is cool!
> I'm fully convinced that very soon big tech or a startup will come up with a programming language meant to sit at the intersection between humans and LLMs, and it will be quickly better, faster and cheaper at 90% of the mundane programming tasks than your 200k/year dev writing forms, tables and apis in SF.
I am sure there will be attempts, but if you know anything about how these systems work you would know why there's 0% chance it will work out: programming languages are necessarily not fuzzy, they express precise logic and GPTs necessarily require tons of data points to train on to produce useful output. There's a reason they do noticeably better on Python vs less common languages like, I dunno, Clojure.
That's the hard engineering part that gets skipped and resisted in favour of iterative trial and error approaches.
Yes, we do. Good code (which, by my definition, includes style/formatting choices as well as the code’s functionality, completeness, correctness, and, finally, optimized or performant algorithms/logic) is critical for the long-term maintenance of large projects—especially when a given project needs integration with/to other projects.
> It it cant write good code consistently,
You moved the goal post within this post.
It starts with the first world and is very perceivable.
How did we perceive cars replacing horses? Well for one they were replaced in the first world... now imagine how fast a piece of software can change reality.
It's not there yet, and you can't perceive it because so.
What a weird reply.
It's literally everywhere around me.
Coworkers, friends in other companies, business owner friends writing their first code, NGO friends using it to write grants.
I'm not sure where you are, but you appear to be isolated from the real world.
Did you not read the context of the comments you're replying to or something?
People using the tool isn't the same as those people being replaced by the tool. Why would anyone think those are the same?
Our jobs aren't replaced yet because they can't be.
const jsonObj = jsyaml.load(yamlText);
const jsonText = JSON.stringify(jsonObj, null, 2);
Don't get me wrong, if you want to convert YAML to JSON then using a battle tested library is the way to do it, but Claude doesn't deserve a round of applause for stating the blaringly obvious.There is an awesome power and innovation to the entire of edifice of targeted advertising. The first time, perhaps, we were all "suggested" something that was in fact quite relevant, was in its own way a giddy-inspiring moment. But we have learned to hate it, not even considering the externalities it brings.
Just always remember: if you are paying for it, its not your friend!
> please write a rust library implementing a variant of simple8b integer compression augmented to use run-length encoding whenever it's beneficial to do so.
Initially I was sort of impressed, it quickly generated a program which looked like rust code, and provided an explanation that, while not as technically detailed as I'd hoped, seemed to be at least related to the topic.
Then I tried to compile the program. Turns out the bot didn't quite actually write rust, it had written something closely resembling rust though, and the compiler errors helped me fix it.
Then I tried to run the tests--yes! the bot even wrote tests, although it did so in a totally bone-headed way by writing multiple distinct tests in one test function--not good. Panic on integer overflow trying to left shift a value. There were also multiple pages of compiler warnings complaining about dead code, unused functions, enum variants, etc. I always fail on warnings.
This is not a lot of code. 190 lines including tests. At this point, given that I already have concerns about its correctness, I don't think there's anything I can really use here. I'm worried the deeper I dig the worse it'll get, so better to cut my losses now, sit down and read the simple8b paper, and implement this from first principles.
Every time I try to use one of these things it's the same story. I cannot understand the hype. I'm genuinely trying but I just can't understand it.
EDIT: Oh my. After digging into the code I found this gem:
fn encode_rle(&self, value: u64, count: usize) -> u64 {
let selector = Simple8bSelector::RLE as u64;
(selector << 60) | ((count as u64) << 30) | (value & 0x3FFFFFFF)
}
And this one: fn try_rle(&self, input: &[u64]) -> Option<(u64, usize)> {
if input.is_empty() {
return None;
}
let value = input[0];
let mut count = 1;
for &x in input.iter().skip(1) {
if x != value || count >= 0x3FFFFFFF { // Max 30-bit run length
break;
}
count += 1;
}
Some((value, count))
}
What even is going on here? Compare to an actually sane implementation like[1] or[2].[1]https://github.com/lemire/FastPFor/blob/master/headers/simpl... [2]https://github.com/timescale/timescaledb/blob/403782a5899c75...
[1] https://play.rust-lang.org/?version=stable&mode=debug&editio...
4o had been happy to attempt to help me fix my function, o1 just went "well that's interesting, meatbag, but have you considered reading the manual?"
EDIT: I don't mean to suggest programming languages aren't tools of human communication--they absolutely are. In fact, that's their primary purpose--to communicate ideas about the structure of a computation to other programmers. But starting with structural ideas about the implementation rather than conceptual ones about the nature of the problem and the shape the solution should therefore take is putting the cart before the horse.
I guess I'm holding it wrong? Is there a better way I could phrase my query?
> The program you wrote doesn't compile. Please fix it such that it compiles.
Then, maybe, if we're lucky, we progress to the second step:
> Ok, now the program compiles but there are tons of warnings about dead code, unreachable code, blanket trait implementations which aren't actually used, etc. Could you please fix those?
Then assuming we clear that hurdle,
> Great! The program compiles without warnings, but when I run the tests it panics due to an integer overflow. I see in your encode_rle function you're inexplicably left-shifting a small unsigned integer by 60, which will absolutely for certain cause it to overflow and panic. Would you mind explaining why in the actual fuck you did this and please fix it? Kthx.
And on, and on... You know what? No. Fuck that shit. I refuse. I have absolutely no confidence this process will come up with a working, trustworthy implementation of the algorithm.
- write me a file with the function definitions for this problem. - compile that - write a test that test x outcome - compile that - then have it start writing functionality
If it's trying to one shot a complex problem that you would typically break up, your prompt is probably too vague.
I think this only works for things where it just doesn't matter whether it actually works correctly, which to me seems synonymous with "problems that aren't worth working on".
[1] just look at this https://play.rust-lang.org/?version=stable&mode=debug&editio...
first attempt: https://onecompiler.com/rust/42w2duuqh
final result: https://onecompiler.com/rust/42w2e3jr4
PS: it "thought" about it for 2 x 60 seconds
EDIT: To be clear, I do understand that I'm making an unreasonable demand. I know the process that came up with this program has no ability to justify this or that "decision" (deliberately scare quoted because it doesn't actually have agency and can't decide at all). And that's the problem. That's why I find it very difficult to trust it.
[1] https://github.com/lemire/FastPFor/blob/52e45deeab9c3a481daa...
Once these are at a point where we can automatically interpret them into usable workflows, it's going to be incredible how quickly you can develop your ideas. I'm really excited for it.
Some examples of outputs:
https://image.non.io/cd90cc33-4a6a-41d8-abd2-045d3a272010.we...
https://image.non.io/5a0c3fc7-37f8-4e72-aba9-cd61f3c18517.we...
https://image.non.io/920adf7c-a554-41bd-a29c-77bebed1cdad.we...
Other things that are important are img2img and inpainting flows to give the model more context for generation.
Why that is there may be several reasons one being that when you try to improve existing code you would need to know all the untold dependencies in it to not break anything. Whereas when you write new code, there are no dependencies you don't know about.
But so, if you take an AI provided implementation, and try to fix it, you are basically doing just that, trying fix old code you don't know much about. You are working on a (AI-provided) legacy code-base in essence.
It's writing code that can be read, changed, understood. Today, next month, after ten years of random freelance engineers abusing it. By caffeinated me, tired me, bored me, that way too smart junior and that boneheaded senior.
What you are trying to propose is that computer generated code should somehow make a programmer feel better about one self because they are given the opportunity to improve something that no one knows where is coming from and without a context.
I think you are forgetting that computers can not generate context aware systems or programs. Look at the list by the author - useless in a different context than the one the author lives in. It does not improve anything, it simply adds more things someone else potentially has to worry about. Furthermore it adds to the same cognitive load than any other code - one needs to read it before one can change it, and changing it is really the first step to fully understanding it.
You are not fixing anything old with AI generated code. Ask yourself also, where is the upstream you are trying "to fix".
I wonder though if I have an intern and I tell them what is wrong with their code, they do learn from my feedback. One purpose for having interns write code is for them to learn. But is it the same with AI? Does it really learn from my feedback, or will it then just try something else, until I'm happy with its output?
When (and if) AI learns from my feedback, does it then apply it's learning when other people ask it to do similar tasks?
The difference between using AI and working with an intern is that the intern learns from you while the AI doesn't... which means that YOU need to learn from your interactions with the AI so you can prompt it more effectively next time.
All of my Claude interactions include "(no react)" because I learned from past experience that without that it writes a React component which is much harder to export out and use separately.
I've also learned to remind it to use 16px text sizes on input boxes (to avoid a zoom effect when selecting an input field on Mobile Safari), and I habitually say "Add a copy to clipboard button that changes its text to Copied! for 1.5s after you click it" because then I get a better UX for the copy feature.
The first stab was basically correct, but then I needed to prompt him several times to correct an issue where the disk that's currently picked up was being rendered as a vertical strip instead of as a disk, and then I just told him to convert the React into pure JS and IIRC it worked first try.
I see some arguments about high-level languages, eg “I’d rather program in Python than assembly - this is just another step-up”. But I feel natural language is altogether different and obliterates any skill/knowledge you’ve built-up in programming.
I can think other things, for example music - I like playing guitar even though I’ll never create something totally original or do better than a machine could do. But for me, programming combines the fun of creation with the satisfaction of the end result - something you wanted to exist now exists.
To clarify, I’m not talking about “usefulness” or accomplishing some business objective, I’m talking about the joy and satisfaction of programming.
Looking up how to accept a drag and dropped file (and then implementing it) in a JavaScript application isn't really that fun to me - certainly not for the tenth time.
I thrive on variety and building interesting things. LLMs let me incorporate WAY more tools into my work - I can build with Go and AppleScript and Bash and ffmpeg and jq, all things I have never climbed the learning curve enough to feel confident using in the past.
For your example, "How do I accept a drag and dropped file in JS?" - An LLM can spit out 100 lines that does what you want, OR someone writes a nice library that does what you want in a single function call (say, for the common case. More complicated usage requires more arguments, etc.)
(Of course, another option is LLMs are the ones writing these library functions).
I guess I am one of those programmers who is allergic to boilerplate (for better or worse), so having LLMs split out lots of code bothers me.
At this point their limited output limit is far behind o1/o1-mini. I really hope they significantly improve that next.
https://github.com/cline/cline
check out the gif to get a look.
Not sure why no discussion at all - maybe the design is underwhelming.
And while tool use is being worked on, the results I've seen are are the "that's an interesting tech demo" level rather than the mind-blowing change when InstructGPT demonstrated the ability for a language model to generate any meaningful code at all from natural language instruction.
that's what rag is supposed to solve: they chunk your 60M loc, and then retrieve and process only relevant depending on your inquery.
I use copilot every day, but it's only so good.
The LLM hype feels like it's been driven by FOMO.
Ever saw someone really bad at googling? It's the exact same thing with LLMs (for now). They're not magic crystal balls, and they certainly can't read everyone's minds at the same time. But give them a bunch of context, and they'll surprise you.
Other than the automation aspect, it is a pretty good alternative to in-depth googling.
BTW, why the disparaging reference to "little toy apps"?
As a programmer (which is the requisite to build such tools even with LLMs), I have a plethora of tools to do the tasks, what I choose and how much time I invested in in that depends on something similar to this chart, but with an added dimension: interest.
Take for example the URL extraction. For one single occasion, I'd probably use VIM and macros to quickly do it. If it were many pages, I'd write a script. If it were infrequent, but recurrent, I'd take the time to write a better script and would only write a web page if the use case was shared with other people or if I wanted a cross platform solution.
I believe the first question one should ask before building is why. That leads you to find a better UX than shoehorning everything inside a web app.
In that aspect, I am hopeful. Maybe if "waste of time" activities are commoditized, "professionals" can instead focus on "what is important," whatever that might be.
What LLMs promise is endless drag. I try to structure my work to ensure that the final velocity is high.
This app for example - which runs OCR against PDF files entirely in the browser - was assembled by pasting in an example of PDF.js usage and an example of Tesseract.js usage and having it figure out the rest: https://simonwillison.net/2024/Mar/30/ocr-pdfs-images/
The kind of project I work on is more like this: Build an Android app for a quiz game. The quiz takes a list of random question from a set. Each set is a package that can be installed and upgraded when online. While the app is free, there is an activation code to be able to download the main packages. The app should work offline except for the activation and downloading packages. It also should notify when a new version is ready for a package. etc...
I don't know if LLMs could have helped me at the time (pre 2020), but I doubt it. Not because the code was complex, but mostly how cohesive the whole thing should be while taking care they're not tightly coupled and be maintainable by a single person. The IDE was a great helper once I got the design and the architecture outlined, mostly because it was deterministic and I already know what the end result should be.
The best current LLMs GPT-4o, Gemini 1.5 Pro, Claude 3.5 Sonnet - are just about at the point now where I'd expect them to be able to get a useful chunk of your Android spec there done. Which is pretty wild!
It's an unmaintainable, single use piece of software (that doesn't even implement the features, it just glues together already existing code) that any CS student could write in a week. Congrats on getting a really fast CS student I guess ? Not to mention the fact that perfectly viable, better alternatives are available in many places.
It's like me nailing two 2x4s together to make a shelf. Yeah, sure, I made it myself and I didn't need any woodworking knowledge, but let's just hope I don't put grandma's heavy china on it.
The upside is that I can produce a shit-ton of one-shot code in record time, so I've got time to face the downside.
search_documents "search term" --bm25 --table-prefix some_project | fulltext -hl -v | less
search_documents "search term" --bm25 --table-prefix some_project | metadata
It inserts documents by piping paths into another script, eg, find /some/path/*.pdf | insert_documents --table-prefix some_project
The documents end up in a Postgres database with pg_search bm25, tsvector, and semantic embeddings (from a local model).I would estimate that I only wrote 5% of the code in the project with the rest coming from the LLM.
Sure, it's just a few hundred lines of code but it's been stable and helpful to get through some very large tranches of discovery material.
Perhaps how quickly we become jaded should be taken as evidence of how quickly the world is changing right now.
When I looked at the examples they seemed like the kind of one off scripts, of limited complexity, that we’ve seen many times in the last year or so.
Just saw someone recommend:
Thought there might be more worth exploring from this community. ;-)
It's not magic, but it is very useful. I find that it works best if you have a commit-sized piece of work in mind; something that might take you half an hour but can be described in a few sentences.
I need to update my post to emphasize this, but that's kind of the point.
Every one of these 14 tools (with the possible exception of the OpenAI Audio debugger one, that one's quite hard) is something any web programmer could build relatively quickly.
... but not as quickly as I did with an LLM, because they almost all took less than 5 minutes from idea to finished implementation.
The key point is that if I didn't have Claude to help build these, I wouldn't have built them at all. None of them would justify even an hour of work - they weren't essential tools that I needed to get stuff done, they were just things I built because building them is now so cheap (in terms of time) that there was no reason not to.
That's the real magic here. The cost of knocking out a single page app that does something simple is often now lower than even the cost of spending a few minutes on Google trying to find an existing tool that solves the same problem.
I'm not a programmer.
It's life-changing.
If all someone does is write code based on specifications handed over by someone else then yes, they have cause to be worried - but in my career as a software engineer the "typing code into a computer" bit has only ever been 10-20% of the work that I do.
The big challenge of software development has always been turning human needs into working software. That requires a great depth of experience in terms of what's possible, what isn't possible, how software works and how to architect and design software to deliver value today while still staying flexible for future development.
LLMs can accelerate that process a bit, but I don't think they can replace it. Someone still has to drive the LLMs. I think people with software development skills are best placed to do that.
I think LLMs mean developers can build stuff faster, which reduces the cost of developing software.
My optimistic scenario is that this expands the market for custom software, a lot. Companies that would never have considered developing their own software - because they'd need six developers working for twelve months - can now afford to do so, because they need two developers for three months instead.
The result is more jobs for engineers, and engineers become more valuable because they'd can get more done.
I'm not an economist so I won't pretend I'm confident this will happen, but it's my optimistic scenario.
I believe that we can already observe that modern tools/languages have made programming a lot more accessible, and that the average quality of software has decreased dramatically (not that all software is bad: just that this new accessibility brought a lot more bad software than good software).
Your example is interesting: it says "it's good because people will be able to produce more", not "developers will have more time to focus on fixing bugs and optimizing their code".
Anyone taking bets on how many years till the operating system doesn't install any software anymore, and just dynamically generates whatever software you need on the fly? "Give me a calculator app" is doable today "give me an internet browser" isn't but it should be a matter of time.
Based on what? Magical thinking?
I just don't see that as being a sound solution. If the user requesting to solve a solved task then using an existing tool will always be a more efficient path.
What I do see as a desired option is where AI could take an existing tool and personalize it to your specific use case. In this example it takes web browser as a GUI lib and wraps it around generic solutions like library that quotes HTML entities which is kinda this just very poorly done so far.
I'd imagine that future programs will be much more component driven with AI connecting components to produce personalized solutions as this is really the only viable option until AI can reason at least in some reasonable capacity to fix its own mistakes.
Since Elon is so interested in that model, if xAI had Claude's capabilities, they would surely go with that angle
Currently also on the front page https://news.ycombinator.com/item?id=41926067
Chat-based AI is a remarkable set of functionalities to build with. It isn't the only improvement in Tech, let alone AI/ML, but it is massive.
I quite enjoy learning more about it from dedicated folks like simonw.
curl -sL $URL | htmlq 'a' -a hrefYou don't memorize them. You learn the foundational knowledge (in this case how http works and the html format, and a bit of shell scripting), then read the manuals and compose the commands. And as days pass, you save interesting snippets somewhere. Then it becomes easier each time you interact with the tools.
Anyone would find ffmpeg or imagemagick daunting if they don't know anything about audio or graphics.
I understand the fundamentals of how git works under the hood. But the cli commands are in no way shape or form, intuitive (just one example). But LLMs nail them every time when I forget.
But more to the point ... in Phaedrus he's not talking about "who will memorize the Iliad now that we have the written word", he's talking about "can the written word _teach_". And the answer (as always) is "no and yes".
> and now you, who are the father of letters, have been led by your affection to ascribe to them a power the opposite of that which they really possess. For this invention will produce forgetfulness in the minds of those who learn to use it, because they will not practice their memory. Their trust in writing, produced by external characters which are no part of themselves, will discourage the use of their own memory within them. You have invented an elixir not of memory, but of reminding; and you offer your pupils the appearance of wisdom, not true wisdom, for they will read many things without instruction and will therefore seem [275b] to know many things, when they are for the most part ignorant and hard to get along with, since they are not wise, but only appear wise.
https://www.perseus.tufts.edu/hopper/text?doc=Perseus%3Atext... and https://www.gutenberg.org/files/1636/1636-h/1636-h.htm#link2....
As is it just spits out migraine-inducing "it-works-doesn't-it" solutions from someone starting to learn to program.
For three, add one more subtle requirement to the task, and now you're reading awk manpages and trial-and-erroring perl oneliners.
And I would have had to use GPT to give me the syntax for that command line anyway :)
Also it doesn't look like htmlq can handle pages that render their content with JavaScript. If you want to do that you might find my shot-scraper CLI utility useful: https://shot-scraper.datasette.io/en/stable/javascript.html
shot-scraper javascript https://simonwillison.net/ 'Array.from(document.links).map(a => a.href)'These aren't even the "hard" problems that are beyond the reach of LLMs today; they seem like things they should be able to do. It's just that, today, they just aren't achieving the spectacular results that many are claiming; it's mostly pretty crappy.
shot-scraper looks nice.
gifsicle --unoptimize input.gif '#-1-0' > reversed.gif
AI makes docker compose app. Cloud providers that cannot deploy a docker compose app simply and without errors will miss out.
https://tools.simonwillison.net/jina-reader?
{"data":null,"code":451,"name":"SecurityCompromiseError","status":45102,"message":"Your request is categorized as abuse. Please don't abuse our service. If you are sure you are not abusing, please authenticate yourself with an API key.","readableMessage":"SecurityCompromiseError: Your request is categorized as abuse. Please don't abuse our service. If you are sure you are not abusing, please authenticate yourself with an API key."}
Lots of TUI interfaces try to approximate this, but I think I really just need to build out something a bit like https://anvil.works/
Contact me if you need some help getting going…
I think Claude offers me 10x productivity, especially for all these helper apps and technical POCs that I typically create during the week.
And that's without even mentioning mail chain replies, analysis of legal or financial documents, helping my kids with their math assignments,...
It's a huge enabler for me, and it's getting better every month.
We are getting up the abstraction ladder faster and faster, and I cannot even imagine where we will end up within a few months, or a few years.
it is basically a tool to add some additional text to json text files and interpret it as comments for each line. I did it with ChatGPT.
"generate an index.html for {idea}".
It's so much faster to just work within a single file. Of course, you have to be limited in scope, but for quick tools such as these it's excellent.
The parsing was terrible, but it worked at all, which is impressive. It suggested multiple next steps, one of which was to work harder on parsing, so I told it to do that. The resulting app showed a similar UI, but was non-responsive. I noticed it had output a message in small text that said "Claude’s response was limited as it hit the maximum length allowed at this time."
So I told it that it had gone over its own limits and to try again, making an effort to stay within its limits. It tried again, and this time the UI didn't even render beyond a text outline of the UI elements, and I got the same message.
So: pluses and minuses.
Second easiest is to take the code and paste it in a Gist - like this: https://gist.github.com/simonw/14a2c3ef508839f26377707dbf5dd...
And then take the Gist ID and add it to this URL:
https://gistpreview.github.io/?14a2c3ef508839f26377707dbf5dd...
That gives you a URL you can load in your browser.
My preferred route is to host the generated HTML directly myself. I mainly use that via GitHub Pages - I can drop an extract-urls.html file into https://github.com/simonw/tools and about 20 seconds later it becomes available at https://tools.simonwillison.net/extract-urls
Those last two options only work for Artifacts that didn't use React (that's why I use "no react" in most of my artifact prompts). If you DID use React you can turn that into a standalone HTML and JavaScript app that you can deploy using https://github.com/claudio-silva/claude-artifact-runner - I wrote some notes on using that here: https://simonwillison.net/2024/Oct/23/claude-artifact-runner...
Do you do anything special to make the tools directory work like that?
- https://til.simonwillison.net/github/custom-subdomain-github...
I was just trying to be helpful, since it was relevant to content in the post…
Works in every single web browser. No calls to OpenAI needed, and I'm rusty on Javascript. Make it a bookmarklet, and you don't even need to run a dedicated webpage on your machine for that.
We're past the POC stage. LLMs can generate code for simple programs. It's when you try to tweak the requirements and point how a program introduces a bug that you eventually realize they still fail to take you through the last mile just as they did year and a half ago.
That's what's so wild about this: that's true, and yet in most of those cases it's still faster and more productive for me to ask Claude to build me a brand new tool _from scratch_ than it is for me to try and find an existing one via Google.
The problem with trying to Google for these kinds of things is that you have to evaluate the results that come back and figure out which one of them correctly solves your problem. That's a few extra steps.
It's genuinely faster to prompt something like this instead:
> Build an artifact (no react) where I can paste text into a textarea and it will return that text with all HTML entities - single and double quotes and less than greater than ampersand - correctly escaped. The output should be in a textarea accompanied by a "Copy to clipboard" button which changes text to "Copied!" for 1.5s after you click it. Make it mobile friendly
Done: https://claude.site/artifacts/46897436-e06e-4ccc-b8f4-3df90c...
In this case I knew exactly what I wanted: it had to do less than, greater than, ampersand, double quotes AND single quotes. I know from past experience that many tools like this forget about single quotes, so I'd have to evaluate any tools I found to check that they do that. And I was on my phone so I wanted a "copy to clipboard" button.
I have hundreds of projects which are small single page HTML apps, single purpose command line, utilities or plugins for my larger projects.
All of these are short enough to fit into the context window of an LLM.
If all I worked on were 1-2 hundred thousand or million line applications I would get significantly less value out of LLMs.