100 karma · joined October 19, 2016
GitHub: panda01 medium: khalah.medium.com
My reasoning is people aren’t happy with many of the titans we have for certain apps, uber is my favorite example. Everyone hates the price it charges thinks it’s unfair, but to go back to a time before uber is unfathomable.
That to me is a ripe market to be overtaken. What it really needs is a taxi driver or operator so fed up with the current system they spend their energy organizing and making their own uber with less greed.
For sure not every tool will become huge, but most of the tools we use nowadays started with someone having a problem, being frustrated with the status quo, and having the will and means to execute.
The means necessary to execute get lower and lower everyday.
The goal of writing the git messages is to slow down and prevent a worst case scenario where work is lost. Given how powerful git is I think giving control over is like doing a sure fire migration without a backup in the way that it can lead to easily preventable problems that are difficult to fix afterwards, it’s just a bad paradigm.
I can appreciate some parts of this, like keeping a painfully detailed record of the changes written by ai, I have this too, but it is separate for the content meant for actual human eyes.
Like others here I find a hard time finding specific evidence or reasons for some of op’s thoughts, but in general it just seems like a recipe for problems when you’re too trusting with ai with all the processes
So workflow for a full web app is make e2e tests for all use cases. Then add a very strict duplication checker, and linter, and then just tell the ai to hit a certain duplication limit like 3%, check the linter, and add unit tests to ~95% or greater of the code.
With the right CI and other checks that are deterministic you can really do a lot with a codebase.
Honestly I don’t even write tests manually because of coverage checks. Being that the coverage check is not something easily manipulated, I always tell the ai, don’t ever change configs, and make the coverage pass whatever I set it to, most times > 95%. I just tell the AI, make this coverage pass.
I find tremendous success with this technique, or anytime really I can find an objective way for the ai to test its work.
The same output that is such a bad thing in this article can also be used to gain context, by making a thorough plan with your ai first, reading through the plan and proposing changes just like you would with a real developer.
You can also use this output to have the ai write a journal as well. The journal can be as detailed as possible and essentially a ledger of all of the changes your ai has made to the code. This allows not only for your teammates reviewing your pr to gain greater context, but also can be used by yourself, or even the ai itself to figure the why behind a particular implementation was done the way it was, far into the future even.
Lastly how many of us ever deploy code without actually checking the feature works e2e? I would gather not many of us do, I don’t, because even though we may have a greater understanding of the code, we can make mistakes in the code or in our logic. And I keep coming back to why would we treat llms any differently? I believe we should be spending our energy thoroughly manually testing a feature to make sure when we brainstormed we actually did get every edge case, and it works well.
I now on all of my projects have an ai journal that stands as a ledger for every change the ai has made, and why it was made. I don’t read it that hard personally because I spend so much time planning with my agent before letting it code. However I have found it very useful in sharing code between people, or having Claude look through the journal to gain context when modifying or adding a feature.
Yes I agree for sure llms write terrible code when left to their own devices, but so do most engineers. Which is why we have so many tools to help keep a certain level of quality. Duplication checks, tests, linters, other engineers.
I find whenever you make an llm repo without these checks, and more, it will write like an enthusiastic junior engineer, wrong and strong. However a junior engineer would be hard pressed to get 95% coverage on a codebase, the ai is more than willing and does it in a few minutes. We can use things like this to our advantage, how many people have ever seen a repo with 100% test coverage? With ai this is very possible, with people not so much.
LLM’s writes terrible code, we know this, but when dealing with humans that write terrible code we have many techniques. We should be using those same techniques to keep the llms honest, but more importantly verifiable.
I think most problems with ai tend to be around can you deterministically test the thing you are asking it to do?
How many of us would never ever show work, without going to check the thing we just built first?
Also shameless plug, I wrote something about this very thing. https://khalah.medium.com/getting-ai-to-work-right-27b750dba...
IDEs are made for just a human to interact with code. I think the paradigm of forcing these tools that weren’t built for this to do this, is us trying to fit a square peg in a round hole.
Call me old, but don’t put ai in my ide. My ide was made for a human, not an ai. For the established players for sure it makes sense since they already have space on our machines. But for the new ones imo terminal, or dedicated llm interfaces are where it’s at.
If I’m writing code sure suggest the next line. If the machine is writing code, let it, and just supervise properly. and have the proper interface that allows the strength of each
Social media was so powerful because it convinced people their world view was popular and correct even if it definitely wasn’t, but at least then it was actually happening somewhere. This is just completely made up, to feed into people’s world view, and they don’t even have to find the perfect figure they can make one up.
We’re definitely going to suffer a lot before it gets better. Interesting times we’re living in.
If you really want to see what I was messing with email me. I’ll share on non public forums.
The reason for the post is that even without the actual website one should be able to envision the technique and how it may or may not work. Also if you look above recently I added links to the Claude.md for another thing I was working on for a friend that also had to solve this problem.
Just want to give people the tools to use ai well from my own findings
I am still trying to learn how to wrangle Claude properly, but I have this Claude.md[1] for that I used to make the website. In particular one of the last rules about using imageMagick for comparison.
I haven’t touched this website in a bit (waiting on client) so now I use playwright mcp for the screenshots and the browser interactions.
[1] https://github.com/panda01/hartwork-woocommerce-wp-theme/blo...
Here is the line from my Claude code to get something like this. Keep in mind I didnt use mcp for playwright with this particular implementation but it is my preferred method currently. Tha
CRITICAL - When implementing a feature based off of an image mockup, use google chrome from the applications folder set the browser dimensions to the width and height of the mockup, capture a screenshot, and compare that screenshot directly to the mockup with imagemagick. If the image is less than 90% similar go back and try and modify the code so that way the website matches the mockup closer. If a change you make makes the similarity go down, undo it, and try something else. be mindful the fonts will never be laid out exactly like the mockup, please use blur at a max of 10% to see if the images are closer matching. If you spend more than 10 cycles screen-shotting and comparing, stop and show the user how similar they are mentioning any problems
The more text the harder it becomes and it’s why we really need the blue because fonts are almost always rendered differently.
Yes ai can’t see, it only understands numbers. So tell it to use image magick to compare the screenshot to the actual mockup, tell it to get less than 5% difference and don’t use more than 20% blur. Thank me later.
I built a whole website in like 2 days with this technique.
Everyone seems to have trouble telling ai how to check its work and that’s the real problem imho.
Truly if you took the best dev in the world and had them write 1000 lines of code without stopping to check the result they would also get it wrong. And the machine is only made in a likeness of our image.
PS. You think Christian god was also pissed at how much we lie? :)
Thorough CLAUDE.md, that makes sure it checks the tests, lints the code, does type checks, and code coverage checks too. The more checks for code quality the better.
It’s just a bowling ball in th hands of a toddler, and needs to ramp and guide rails to knock down some pins. Fortunately we get more than 2 tries with code.
This even if their was a gain from watching others suffer, the lack of discipline, guidance, sternness, is way more detrimental than the positives of fearing the consequences
My favorite instruction is using component A as an example make component B
Clearly the comma thing is a bug, it's the lack of wanting to fix it actually that is a bit disheartening, and why I think it is a deadish repo
1 such bug, find a foreign language with commas in between numbers instead of periods, like Dutch(I think), and a lot of prices on the page. It’ll think all the numbers are relevant text.
And of course I tried to open a pr and get it merged, but they require tests, and of course the tests don’t work on the page Im testing. It’s just very snafu imho