But if you want it to generate chunks of usable and eloquent Python from scratch, it’s pretty decent.
And, FWIW, I’m not fluent in Python.
But if you want it to generate chunks of usable and eloquent Python from scratch, it’s pretty decent.
And, FWIW, I’m not fluent in Python.
I think that's the point of the article.
In a dynamic language or a compiled language, its going to be hallucinating either way. If you vibe coding the errors are caught earlier so you can vibe code them away before it blows up at run time.
You can say that again.
I was looking into the many comments for this particular comment and you did hit the nail on the head.
The irony is that it took the entire GenAI -> LLM -> vibe coding cycle to settle the argument that typed language is better for human coding and software engineering.
My repos all have pre-commit hooks which run the linters/formatters/type-checkers. Both Claude and Gemini will sometimes write code that won't get past mypy and they'll then struggle to get it typed correct before eventually by passing the pre-commit check with `git commit -n`.
I've had to add some fairly specific instructions to CLAUDE.md/GEMINI.md to get them to cut this out.
Claude is better about following the rules. Gemini just flat out ignores instructions. I've also found Gemini is more likely to get stuck in a loop and give up.
That said, I'm saying this after about 100 hours of experience with these LLMs. I'm sure they'll get better with their output and I'll get better with my input.
A few things beyond your question, for anyone curious:
I've also poked around with a custom MCP server that attempts to teach the LLM how to use ast-grep, but that didn't really work as hoped. It helps sometimes but my next shot on that project will be to rely on GritQL. Smaller LLMs stumble over the YAML indentation. GritQL is more like a template language for AST aware code transformations.
Lastly, there are probably a lot of little things in my long term context that help get into a successful flow. I wouldn't be surprised if a key difference between getting good results and getting bad results with these agentic LLM tools is how people are reacting to failures. If a failure makes you immediately throw up your hands and give up, you're not doing it right. If instead you press the little '#' (in claude code) and enter some instructions to the long term context memory, you'll get results. It's about persistence and really learning to understand these things as tools.
Also interesting note on the docs, though, Claude does try to use cargo doc by itself sometimes.
I was actually wondering why GritQL did not have an MCP, this seems like a natural fit. Would be interested to know if this works for you.
I'm always a bit hesitant to add things to the long term context as it feels very finicky to not have it be ignored and having more seems to make it more likely to be ignored. Instead I usually just repeat myself.
Thank you for the answer, seems there is still lots of things to try.
Why not have static analysis tools on the other side of those generations that constrain how the LLM can write the code?
What I'd like to see is the CLI's interaction with VSCode etc extending to understand things which the IDE has given us for free for years.
We do have it, we call those programmers, without such tools you don't get much useful output at all. But other than that static analysis tools aren't powerful enough to detect the kind of problems and issues these language models creates.
Largely I think LLMs struggle with Rust because it is one of very few languages that actually does something new. The semantics are just way more different than the difference between, say, Go and TypeScript. I imagine they would struggle just as much with Haskell, Ocaml, Prolog, and other interesting languages.
But that is all independent of how the LLMs are used, especially in an agentic coding environment. Strong/static typed languages with good compiler messages have a very fast feedback loop via parsing and typechecking, and agentic coding systems that are properly guided (with rulesets like Claude.md files) can iterate much quicker because of it.
I find that even with relatively obscure languages (like OCaml and Scala), the time and effort it takes to get good outcomes is dramatically reduced, albeit with a higher cost due to the fact that they don't usually get it right on the first try.
'I have a database table Foo, here is the DDL: <sql>, create CRUD end points at /v0/foo; and use the same coding conventions used for Bar.'
I find it copies existing code style pretty well.
At the end of the day this is a trivial problem. When Claude Code finishes a commit, just spin up another Claude Code instance and say "run a git diff, find and fix inefficient and ugly code, and make sure it still compiles."