A commenter on another thread mentioned it but it’s very similar to how search felt in the early 2000s. I ask it a question and get my answer.
Sometimes it’s a little (or a lot) wrong or outdated, but at least I get something to tinker with.
A commenter on another thread mentioned it but it’s very similar to how search felt in the early 2000s. I ask it a question and get my answer.
Sometimes it’s a little (or a lot) wrong or outdated, but at least I get something to tinker with.
If this is the state of the art of coding LLMs, I really don't see why I should waste my time evaluating their confident sounding, but wrong, answers. It doesn't seem like much has improved in the past year or so, and at this point this seems like an inherent limitation of the architecture.
I ask it questions mostly about libraries I’m using (usually that have poor documentation) and how to integrate it with other libraries.
I found out about Yjs by asking about different operational transform patterns.
Got some context on the prosemirror plugin by pasting the entire provider class into Claude and asking questions.
It wasn’t always exactly correct, but it was correct enough that it made the process of learning prosemirror, yjs, and how they interact pretty nice.
The “complete” examples it kept spitting out were totally wrong, but the information it gave me was not.
I had a suspicion that what I was trying to do was simply not possible with that library, but since LLMs are incapable of saying "that's not possible" or "I don't know", they will rephrase your prompt and hallucinate whatever might plausibly make sense. They have no way to gauge whether what they're outputting is actually correct.
So I can imagine that you sometimes might get something useful from this, but if you want a specific answer about something, you will always have to double-check their work. In the specific case of programming, this could be improved with a simple engineering task: integrate the output with a real programming environment, and evaluate the result of actually running the code. I think there are coding assistant services that do this already, but frankly, I was expecting more from simple chat services.
Specific is the specific thing that statistical models are not good at :(
> how do I do X with library Y?
Recent research and anecdotal experience has shown that LLMs perform quite poorly with short prompts. Attention just has more data to work with when there are more tokens. Try extending that question like “I am using this programming language and am trying to do this task with this library. How do I do this thing with this other library”
I realize prompt engineering like this is fuzzy and “magic,” but short prompts have a consistent lower performance.
> In the specific case of programming, this could be improved with a simple engineering task: integrate the output with a real programming environment, and evaluate the result of actually running the code.
Not as simple as you’d think. You’re letting something run arbitrary code.
Tho you should give aider.chat a try if you want to test out that workflow. I found it very very slow.
I'm aware of that. The actual prompt was more elaborate. I was just mentioning the gist of it here.
Besides, you would think that after 30 minutes of prompting and corrections it would arrive at the correct answer. I'm aware that subsequent output is based on the session history, but I would also expect this to be less of an issue if the human response was negative. It just seems like sloppy engineering otherwise.
> Specific is the specific thing that statistical models are not good at
Some models are good at needle-in-a-haystack problems. If the information exists, they're able to find it. What I don't need is for it to hallucinate wrong answers if the information doesn't exist.
This is a core problem of this tech, but I also expected it to improve over time.
> Tho you should give aider.chat a try
Thanks, I'll do that eventually. If it's slow, it can get faster. I'd rather the tool be slow but give correct answers, than it slowing me down by wasting my time error correcting it.
Thankfully, these approaches can work for programming tasks. There is not much that can be done to verify the output of any other subject.
Codebases are exploding in size. Feature development has slowed down.
What might have been a carefully designed 100kloc codebase in 2018 is now a 500kloc ball of mud in 2024.
Companies need many more developers to complete a decent sized feature than they needed in 2018.
Once you start shifting it to micro services the business logic gets spread out and duplicated.
At the same time each micro-service now has its own code to handle rest, graphql, grpc endpoints.
And each downstream call needs error handling and retry logic.
And of course now you need distributed tracing.
And of course now your auth becomes much more complex.
And of course now each service might be called multiple times for the one request - better make them idempotent.
And each service will drift in terms of underlying libraries.
And so on.
Now we have been adding in LLM solutions so there is no consistency in any of the above services.
Each dev rather than look at the existing approaches instead asks Claude and it provides a slightly different way each time - often pulling in additional libraries we have to support.
These days I see so much bad code like a single microservice with 3 different approaches to making a http request.
Maybe I'll have better luck next time, or maybe I need to improve my prompting skills, or use a different model, etc. I was just expecting more from state of the art LLMs in 2024.
A lot of the devs in my office using Claude/gpt are convinced they are so much more productive but they aren't actually producing features or bug fixes any faster.
I think they are just excited about a novel new way to write code.
Right now, I find each tool better at different things.
If I can only describe what I want but don't know key words, LLM are the only solution.
If I need citations, LLMs suck.
Google used to prioritise big comprehensive articles on subjects for desktop users but mobile users just wanted quick answers, so that's what google prioritised as they became the biggest users.
But also, per your point, I think those smaller simpler less comprehensive posts are easier to fake/spam than the larger more compreshensible posts that came before.