69 karma · joined May 21, 2011
For evaluation, there is command to record session, repository commit, and observed problems that stored in special database, so each session can be reproduced. Developer commit reports, I do analysis, refine and evaluate system.
<system-reminder> IMPORTANT: this context may or may not be relevant to your tasks. You should not respond to this context unless it is highly relevant to your task. </system-reminder>
[1] https://www.capitalone.com/tech/open-source/announcing-vulnh...
This is that I do not see. My journey, just couple weeks ago, Claude Code + Opus 4.8. The task was not too complicated, 4 new API endpoint plus events streamed from client by websocket.
1. Multiply iterations on API definitions, refine request/response models, database schema, whole flow. A lot of corrections, removing contradictions, manual changes in document. Opus went of rails all the time. 500+ lines final document
2. API Integration tests. Once again, back and forth. AI was unable to create tests directly from document, so 2 iterations: Create placeholders with Given-When-Than comments, review an correct by hand. Second iteration was to implement tests. A lot of mistakes corrected after review.
3. Implementation. CC got api document, working tests ( modifications blocked by hook ), 6+ "best practices" skills ( most promptly ignored ), "rubber duck" and "code simplifier" agents, pre cooked scipts to run tests, linter, and check for compilation errors. Plan + execution + review, multiply corrections on the way. Feature implemented, all tests passed.
4. Code review. At average, found one issue per 20 lines of code. Not count code style, things like: Use in memory semaphore in kubernetes service (deployment described in CLAUDE.md ), 8 database calls to update the same record during a single request. One column at a time! Read-modify-save without transaction. Mistakes in business logic, failure recovery, authorization.
The result: almost one workweek, $100+ in tokens, and one thought: did it worth the effort ? P.S. I have a team of 2 developers. Just got PR to review from one of them. 80% slop.
The average numbers have a little sense. By our measurements, it was't an "even" distribution, but some "hot spots" that had 10x time radiation level than surrounding territory.
The article focuses on cancer, but for me and my buddies, the worst was an impact on immune system. In the six months after disaster, the healthy boy before, I got: pneumonia, chicken pox ( and 2 more with me ), furunculosis ( the whole company got it as well ), and endless flu/fevers. The weak immune continued for about 10 years.
WWI era poison gas, tear gas, potassium cyanide, and bunch of explosives like acetone peroxide.
LLMs have all of that knowledge in training data
I did create my own MCP with custom agents that combine several tools into a single one. For example, all WebSearch, WebFetch, Context7 exposed as a single "web research" tool, backed by the cheapest model that passes evaluation. The same for a codebase research
Use it with both Claude and Opencode saves a lot of time and tokens.
And while security rules created enormous roadblocks for work, whey also left enough holes to be exploited. Before getting required permissions, I managed to create dual boot with linux and share files between 'approved' and 'illegal' systems
I see it's necessary ALL the time. The AI generated code can be used as scaffolding, but it's newer get close to real production quality. The expierence from small startup with team of 5 developers. I do review and approve all PRs, and none ever able to pass AI code review from the first iteration.
1. create several hundreds github repos with projects that use your product ( may be clones or AI generated )
2. create website with similar instructions, connect to hundred domains
3. generate reddit, facebook, X posts, wikipedia pages with the same information
Wait half a year ? until scrappers collect it and use to train new models
Profit...
Since some team members started using AI without care, I did create bunch of agents/skills/commands and custom scripts for claude code. For each PR, it collects changes by git log/diff, read PR data and spin bunch of specialized agents to check code style, architecture, security, performance, and bugs. Each agent armed with necessary requirement documents, including security compliance files. False positives are rare, but it still misses some problems. No PR with ai generated code passes it. If AI did not find any problems, I do manual review.
This came out of necessity, as active using of AI assistants in uncontrollable way significantly degraded code quality. The goal is to enforce the same development workflow across team This is internal tool. If someone interesting, I can create a public repo from it
I organize notes by tags, folders, and links from tree of "map of content" notes. Those documented as rules for AI. All notes came to "Inbox" folder, and from time to time I run special script that checks inbox, formats notes, tags them, and put in the most appropriate place. "git diff" to check results and fix mistakes, reset if it went wrong.
As notes organized by the limited number of well defined rules, they became easy to search and navigate by AI. Claude Code easily finds requested notes, working as advanced search engine, and they became a starting point for "deep research" : find relevant notes, follow links, detect gaps, search internet. Repeat until reach required confidence level.
The most advanced workflow so far is combination of TRIZ (Theory of Inventive Problem Solving) + First Principles Framework. Former generates ideas and hypotheses, later validates them and converge on final answer.
Even when error message was clearly understandable for my expertise, it took surprisingly long tome to switch from one mental activity - "Pay bills", to another - "Investigate technical problem". And you have to throw away all short memory to switch into another task. So all rumors about "stupid" users is direct consequence from how human mind works.
- you create small utility that covers only features needed only for you. As many researches show that any individual uses only less than 20% of software functionality, your tool covers only 10-20% that matters for you
- it only runs locally, on user computer or phone, and never has more than one customer. Performance, security, compliances do not matter
- the code lies next to application, and small enough to fix any bug instantly, in a single AI agent run
- as a single user, you don't care about design, UX, or marketing. Do the job is only matter
It means, majority of vibe coded applications run under radar, used only by a few individuals. I can see it myself: I have a bunch of vibe code utilities that never intended for a broad auditory . And, many of my friend and customers, mention the same: "I vibe coded utility that does ... for me". This means a big consequences for software development: the area for commercial development shrinks, nothing that can be replaced by the small local utility has a market value.