I used to think of it as a decent sr dev working alongside me. Not it feels like an untrained intern that takes 4-5 shots to get things right. Hallucinated tables, columns, and HTML templates are its new favorite thing. And calling things "done" that aren't even half done and don't work in the slightest.
Yes, I know. That’s what the test was for.
My fear when using Claude is that it will change a test and I won't notice.
Splitting tests into different files works but it's often not feasible, e.g. if I want to write unit tests for a symbol that is not exported.
(I couldn't find that documentation when I went looking just now.)
Step 2: Type 'Allowed Tools'
Step 3: Click: https://docs.anthropic.com/en/docs/claude-code/sdk/sdk-headl...
Step 4: Read
Step 5: Example --allowedTools "Read,Grep,WebSearch"
Step 6: Profit?
> allow zoned access enforcement within files. I want to be able to say "this section of the file is for testing", delineated by comments, and forbid Claude from editing it without permission.
Maybe rtft ? Read the fucking thread.
At least with local LLM, it's crap, but it's consistent crap!
Likely the common young startup issues: a mix of scaling issues and poorly implemented changes. Improve one thing, make other stuff worse etc
So it could be a matter of serving more highly quantized model because giving bad results has higher user retention than "try again later"
Also yesterday tried to use it to debug some AWS issue and it tried to send me down so many wrong paths, and suggested changes that were either plain wrong or had unintended consequences, that if I didn't actually know my stuff and had followed blindly, the results would have been pretty bad or at least a huge time waster. When I called it out it would quickly reverse course ("You're right of course!") and it did provide some helpful snippets but I was unimpressed.
What I find it excellent at is for throw-away scripts to do small jobs or automate little things--stuff I could do but would take me a lot longer (especially in bash).