162 karma · joined October 1, 2016
I’ve stopped using llms to generate architecture, which i design and write myself and let the machine pattern match the gaps. I also use it to review issues which I lot of the times push back against.
I’m working on a stateful application sitting on top of a data warehouse and have to implement a stream of messy half defined feature requests and navigate on top of an ever changing infrastructure layer. LLMs rarely get the infra layer even if it is written as code and have hard time grasping how to deal with tech debt, when and how to re-architecture parts of the stack or even implement stuff based on a detailed openspec design.
With that said I think doing some kind of workout even on vacation is important.
we will pick up that book about the monster, it is really scary (slippers on already) and we will sit on the sofa (already carrying the child). Are you cold? Let’s find that pink sweater…
Why I engage with this comment? I truly believe that trusting your own knowledge and skills is the way forward. If you don’t have the skill? Build it. Don’t outsource thinking and learning.
def my_swe_percentile(best_agent_swe_percentile):
return min(100, best_agent_swe_percentile * 1.25)Does this need an agent though is my question? Maybe generating a test case and a loop doing git bisect but why on earth would we want to run it through the internet and gpus and whatnot when it can be run on a single core celeron.
There is nothing on horizon which automates a programmer’s work. Typing in code is faster now, and some things “only need pointing out” like an existence of a “bug” which an llm + harness might be able to mitigate. Automated tests might capture regressions and possibly written by llm + harness. If you replicate this in other professions what will you get?
From this I assume you think that what the llm has generated is as valuable as your own work generally is. How do you even calculate this?
I’d never just turn on the heater silently if someone said this to me. I think it means something else.
The issue is that validation needs presence and it is the limiting factor - common knowledge, but is part of the “physics”. Also maintenance gets really tricky if the codebase has warts in it - which it will have. I get much more easy to understand architecture out of an LLM driven code generation process if I follow it and course correct / update the spec process based on learnings.
Example: yesterday I’ve introduced a batch job and realized during the implementation phase that some refactoring is needed so the error boundary can be reused in the batch application from the main backend. This was unplanned and definitely not a functional requirement - could be documented as non-functional. There was a gap between the agent’s knowledge and mine even though the error handling pattern is well documented in the repository. Of course this can be documented better next time if we update the process of openspec writing but having these gaps is inevitable unless formal and half-formal definitions are introduced - but still there needs to be someone with “fresh eyes” in the loop.
Have you built anything purely with LLM which is novel and is used by people who expect that their data is managed securely, and the application is well maintained so they can trust it?
I have been writing specifications, rfcs, adrs, conducting architecture reviews, code reviews and what not for quite a bit of time now. Also I’ve driven cross organisational product initiatives etc. I’m experimenting with openspec with my team now on a brownfield project and have some good results.
Having said all that I seriously doubt that if you treat the english language spec and your pm oversight as the sole QA pillars of a stochastic model transformer you are making a mistake.
I think I have a relatively good life, but I still have hard times. I had circa 6 months long depression streak after my child was born (I'm male).
For me the best mood fixer is a walk still. Super small commitment, great with a dog too. For a weekend the best is a longer hike. I practice yoga and train my body - great mood boosters. I've trained my body to be able to sit comfortably on the ground so I can work from anywhere - sunshine in park hellooo.
Hope you find your rhythm soon!
I’d use a fast model to create a minimal scaffold like gemini fast.
I’d create strict specs using a separate codex or claude subscription to have a generous remaining coding window and would start implementation + some high level tests feature by feature. Running out in 60 minutes is harder if you validate work. Running out in two hours for me is also hard as I keep breaks. With two subs you should be fine for a solid workday of well designed and reviewed system. If you use coderabbit or a separate review tool and feed back the reviews it is again something which doesn’t burn tokens so fast unless fully autonomous.
Some improvement ideas:
A prototype can help in the “Better communicate the idea/feature” part but it is even better if you let engineers do this as learning by doing is better than just being shown the result.
Vibe coding doesn’t help in “Understand the systems” - on the contrary, this is already a well known fact that vibecoding has negative effect in understanding the underlying system. It should be hardboiled documentation reading, trial and error which helps, otherwise you get only the illusion of competence.
- ragebait them by saying AIs don’t think
- …