I think it's a good reality check for the claims of impending AGI. The models still depend heavily on being able to transform other people's work.
I think it's a good reality check for the claims of impending AGI. The models still depend heavily on being able to transform other people's work.
It's my understanding that LLMs change the code to meet a goal, and if you prompt them with vague instructions such as "make tests pass" or "fix tests", LLMs in general apply the minimum necessary and sufficient changes to any code that allows their goal to be met. If you don't explicitly instruct them, they can't and won't tell apart project code from test code. So they will change your project code to make tests work.
This is not a bug. Changing project code to make tests pass is a fundamental approach to refactoring projects, and the whole basis of TDD. If that's not what you want, you need to prompt them accordingly.
The whole point is having the LLM figure out what you want from vague hand-wavy descriptions instead of precise specification.
You don't need an LLM to parse a precise specification, you have a compiler for that.
It's not a problem. It's in fact the core trait of vibe-codig. The primary work a developer does in vibe coding tasks is providing the necessary and sufficient context. Hence the inception of the term "context engineering". A vibe coder basically lays out requirements and constraints that drives LLMs to write code. That's the bulk of their task: they shift away from writing the low-level "how" to instead write down the high-level "what".
> The whole point is having the LLM figure out what you want from vague hand-wavy descriptions instead of precise specification.
No. The prompts are as elaborate as you want it to be. I, for example, use prompt files with the project's ubiquitous language and requirements, not to mention test suites used for acceptance tests. You can half-ass your code as much as you can half-ass your prompts.
I assume in this case you mean a broader conventional application, of which an LLM algorithm is a smaller-but-notable piece?
LLMs themselves have no goals beyond predicting new words for a document that "fit" the older words. It may turn 2+2 into 2+2=4, but it's not actually doing math with the goal of making both sides equal.
Not necessarily. If you prompt a LLM to limit changes to some projects or components, it complies with the request.
It's not a bug if we're talking about a mischievous jinn granting wishes instead of a productivity tool.
"I'm Mr. Meeseeks! Look at meeee!"
It's a big mess.
0. https://github.com/isaacs/semicolons/blob/main/semicolons.js
As a test recently I instructed an agent using Claude to create a new MCP server in Elixir based on some code I provided that was written in Python. I know that, relatively speaking, Python is over-represented in training data and Elixir is under-represented. So, when I asked the agent to begin by creating its plan, I told it to reference current Elixir/Phoenix/etc documentation using context7 and to search the web using Kagi Search MCP for best practices on implementing MCP servers in Elixir.
It was very interesting to watch how the initially generated plan evolved after using these tools and how after using the tools the model identified an SDK I wasn't even aware of that perfectly fit the purpose (Hermes-mcp).
Last night I tried to build a super basic “barely above hello world” project in Zig (a language where IDK the syntax), and it took me trying a few different LLMs to find one that could actually write anything that would compile (Gemini w/ search enabled). I really wasn’t expecting it considering how good my experience has been on mainstream languages.
Also, I think OP did rather well considering BASIC is hardly used anymore.
The models don’t have a model of the world. Hence they cannot reason about the world.
They don't need a formal model, they need examples from which they can pilfer.
I have had it writing LambdaMOO code, with my own custom extensions (https://github.com/rdaum/moor) and it's ... not bad considering.