13,166 karma · joined February 28, 2013
http://kiprotect.com https://github.com/adewes
But yeah I noticed how they play with your emotions, it's a bit creepy. Whenever it asks a hook question and I tell it "Actually I did X and Y" it will invariably go like "Wow that is such a brilliant and well thought out strategy, almost a stroke of genius!" (not literally but you get the idea). I guess they found out that people really like having their ego stroked, the other providers are also doing that to some degree (I assume since no one likes an unfriendly AI that makes you feel stupid) but Google really cranked that setting to 11.
That got me thinking that programming might be like that in the near future, it's still an impressive skill for a human to learn and master and it takes tremendous effort and the craft itself is beautiful, but economically it will just not matter anymore. And in the same way as the music industry today is bigger than it ever was and there's tremendous money in it but 95-99 % of highly talented musicians will never earn a decent wage with it, the software industry will be bigger than it ever was but the money will be highly concentrated in the hands of a few players. Maybe we will still have some rockstar programmers than can masterfully command legions of agents or that win the love of the crowd for their unique work, but most of us will be enjoying this craft as a creative hobby and will do other things for a living.
Then again, coming back to the music parable, for a long time "amateurs" can make music with software tools without having any talent or even knowing musical notation or anything about tempo. Does that make them musicians? So take solace in the fact that these people aren't programmers and never will be, they're just pushing buttons on a synth board watching the "music" that comes out. Mastering a skill, regardless of what it is, will still matter to people, it will just not be worthwhile economically for most of us. There are so many things in that category, running for example isn't worthwhile for most of us in an economic sense as you have to be at the top 0.1 % of runners to make a living with it, still it has tremendous benefits and brings enjoyment.
I guess there's no reason to believe these models can't be as smart as a great software architect / engineer or team of such people that build an elegant and maintainable software solution over many years together based on customer feedback, then again the models are appallingly bad at some forms of reasoning, I mean they will "understand" something once you make them aware of it like e.g. a flaw in the software architecture, but when asking them to audit the code and check for issues they will often have a blind spot to finding such problems. It's interesting, like they have very high ability but very little awareness or self-directed thinking outside of the prompts they receive.
And maybe let's not only hear the opinion of two or three Fields level mathematicians with blogs, 99 % of the worlds mathematicians in academia might profit from these tools as they might partially close the gap between them and the world elite, making creativity and tenaciousness more important than having the right neocortical structure allowing you to outperform 99.9 % of other humans at keeping context in your head and making predictions, AI can do that better now with the right prompts.
As a simple experiment, try giving AI a high level goal for your software and let it iterate on it by just repeatedly prompting it to continue, it will happily churn forever on the goal, turning the codebase into a useless spaghetti mess with very high probability, and growing it more and more without ever cutting anything back. That's what happens without human intervention regarding system state and manipulation. The main issues here are most prompts that are extremely underspecified ("fix the issue with the buttons on the main page") so AI will ingest context data it likely generated itself in a previous step and assumptions from its own training data, then act on that to produce a new state. Think of it like a random walk, the AI makes a small step in one random direction to achieve a goal, that brings the system to a new state which is now the basis for the next step, and so on. If there's no (or not enough) corrective action that pulls the system back to a known good reference state it will keep wandering in random directions.
That's the main issue, people have a hard time steering recursive, probabilistic systems, especially when they never look at the output of the system after each step and correct it. And let's be real, if you examine AI generated output in great detail after each iteration you're often better off writing the code yourself, so I would argue that the promised speed up of agentic development can only be realized if you stop inspecting every output of the system. And it seems we still haven't figured out how to specify the steering instructions that keep a system close to a given ideal state that allow unsupervised, recursive work on most codebases. I think some codebases are by themselves better suited for this as they provide a more rigid harness for AI development and exist in the training data (e.g. CRUD apps using RoR), whereas complex software that doesn't use rigid frameworks is at much higher risk of destruction by AI as there's no reference point in the training data that would hold the AI back from randomly walking to a garbage state.
And that's why people have such different views on agentic software development, some work on codebases that are better represented in the training data and so have great success using agentic tools on them, others work on software that isn't represented so well so AI does poorly on it. I don't think it's an issue with quality management, from my own experiments no amount of hand-written rules or system prompts will keep AI from destroying a codebase for which it doesn't have a strong idea how the code is supposed to look from its own training data in the first place. As another experiment, try giving AI strict rules about how to change code or introduce new features, it will always find a way around them or appropriate them in a maliciously funny way that you haven't anticipated. That's also an artefact of the training process, these systems aren't designed to say no or do nothing, they produce outputs to achieve goals and they will bend your rules to the greatest amount possible if it helps with goal fulfilment.
I recently tried writing a paper with Claude and it was an absolute disaster, I spent hours (days?) instructing it about writing style and pointing out anti patterns to avoid, but I couldn't get it to even produce simple sentences, it would always add unnecessary lead up sentences, put the most important information at the end of the sentence, use the typical "It's A, not B!" or "B, not A!" type sentences. In the end I gave up and edited everything manually. Makes me wonder how AI can be so smart that it poses a human-level extinction threat but can't seem to even write a simple paper based on facts and information you spoon feed it. I now think it's an intelligence illusion due to the training data and optimization process being hidden from us, essentially it keeps working better and better because we invested massively in optimization of specific use cases like coding, where users contributed billions of training samples that are part of the LLM model. The same is true for text-based workflows and others, the sampling density of the training space is getting much better due to the massive use of AI everywhere so the models extrapolate better between the different instances, but I'd wager they would still miserably fail to generalize to things that are outside of the most common training use cases now. That's why I am also very skeptical about recursive self improvement of these systems, look at what happens when you let agents work recursively / in a loop now, they just keep piling more garbage onto garbage and choke on their own output. I have observed it in my paper writing as well, you feed input into the AI system, the system produces output, the next paper iteration works on that output but the AI doesn't differentiate properly between it's own output and your original input, that pollutes the next output which is then used as input again, eventually the system just churns on its own hallucinated/fabricated outputs until the result is complete garbage that no amount of steering will fix. The same is true for most vibe coded software I built with AI, it holds together decently initially, but the more AI code and decisions accumulate the more the system operates on its own outputs and keeps piling more output on it. More than anything we really need a way to keep system data accurately tagged, i.e. clearly mark human input from AI output and keep AIs churning on output data that it produced itself but treats as input.
I find people that are the "unconscious incompetence" stage of software development absolutely love them because they produce output that looks fine to them and they can't find anything wrong with it because they simply lack the ability to perceive quality. For a contact widget that might be acceptable as it will be used a few dozen times so who cares if it breaks, but let's not pretend you can actually build a compliant e-signing or invoicing software without having domain expertise and investing a ton of editing/revision time.
Put more bluntly, a domain expert that, as of currently, refuses to use LLMs can easily start using LLMs in their work, an idiot who uses LLMs cannot easily become a domain expert. Now let's think about who you would rather hire or who do you think will have the better future career outlook. Pretty damning how many people here are willing to reject the idea that domain expertise or ability to write program code yourself is a waste of resources.
So I'm greatly excited what AI will bring about in physics, more so than in math, because in physics it's clear that our fundamental theories are missing a big piece of the picture, and given how easily AI crunches through Millenium prize problems I think it's possible that AI will come up with a viable grand unified theory uniting quantum mechanics and gravitation, or produce new predictions in other areas. There's enough contradictory or unexplained observational data available to make a ton of progress on the theory side I think. Exciting times ahead!
We equipped our founding era building with wall heating and will test out the heatpump based cooling next summer, if it's not enough I would consider installing a small AC unit, that's super easy as well.
Don't know what the article wants to say, there's no ban on AC anywhere in Europe and small units are quite affordable, you can get a single split unit from Mitsubishi for 500 € from Italy. The only nuisance is that installation is heavily regulated in Germany so they rip you off with insane cost. That's definitely a thing that needs to change in my opinion.
Maybe just the system prompt or my settings or the harness? Anyway, it seemed to me Codex was just deliberately being overcareful and wasting tons of tokens for a low-risk CSS / HTML refactoring with very minor breakage risk, while Claude got the job done immediately.
I for one tend to care less about the minutiae of solutions implemented by AI as long as it gets the job done, I do care about architecture and design decisions and correctness and I have ways to steer and verify these when working with LLMs but I couldn't care less about it writing "good" code. Bad engineers also produce better results with AI at least when they're working in established frameworks, AI doesn't really need a lot of high level architecture input when designing or building a web app with a common stack, so as long as you're not working on something that's completely novel I don't think it will make a strong difference.
Maybe designers think the same way about the AI generated web designs I have Claude Code do for me but to be honest I don't care, I just know that before this tool existed it would have taken me weeks or months to come up with a good design and I would have to rely on prefabricated UI libraries and stuff like that or pay a designer tens of thousands of USD to make one for me, now I can get a (for me and my customers) perfectly acceptable and professional design within a few hours. So maybe I'm also a bad designer that amplifies my bad design taste 10x in my company, but the fact is the stuff ships and makes money and the customer is happy! And I can tell you customers or users don't give a shit about how good your code is, they only care if the software works and does what they want!
Same goes for writing, anyone who has written long complex texts with LLMs knows that a ton of editing is required to make it half decent.