246 karma · joined September 5, 2025
In developing countries, including China, my experience has been that this is the standard, and people will act accordingly even if there is abundance. Cancelled transit is a good example as well, it interacts with a primal fear of ours of being left behind or drawing the short stick, even if rationally, it is pretty inconsequential.
Thing is, I'm still reading the majority of generated code, and I have colleagues who'll laugh at me if my PRs are a shit show. I fear what vibe coders are pushing to the servers of myriads of start ups, and pity the poor people who'll have to clean it up in a year or two.
And Gemini is kinda good enough at everything. Never the top, but it is decent at every task, and it is much faster than Kimi K3 and significantly cheaper. Kimi is very focussed on coding, Gemini isn't.
More importantly, it natively understands text, audio and video. If/when we are able to make the jump to robotics, this becomes essential. As you say, Google has a lot of deep background and deep pockets, they are able to make more of a long play. No idea if it will pay off, but it is way to early in the game to count them out.
- Doing too much at the same time, it is super exhausting, and I get nervous and stressed
- Sycophancy makes me overestimate myself
- When my stress meets external pressure, quality goes down. The little voice of 'surely the AI got it right' can lead to defects and code quality going down.
Except for doing some manual coding, I've started doing a couple of other things that are probably just healthy regardless of using AI heavily in day to day life:
- Do literally nothing for at least half an hour. Read something written by humans, even better if it is fiction.
- Do something that makes you feel stupid. And don't use AI to help you through it. (That could be Leetcode, but for me it is learning about statistics)
- Slow down when you feel like things are moving too fast. By far the hardest thing, but the moment I realize I'm trying to rush for a deadline, I'm going out for a walk. Realistically, I'm moving 3-8x faster than before, this little delay is nothing, and keeps me sane and quality on a good level.
As the person above said, AI has significantly increased my ability to execute on ideas. I always liked computers, always liked building things, but was never that great at coding, and was never that good at going super deep into one topic. Instead, I have broad knowledge of a lot of things like product design, requirements engineering, devops, security.
I work with a fully agentic flow, but I would argue very seriously; My job has not gotten easier, I am doing at least as much hard thinking as before. From assisted RE over design (TDD focused Spec), implementation by agents, review by agents with partial human oversight, I have a speedup of maybe 50-100%.
More importantly, I can do things I couldn't do before. And when it comes to performance and defect-density, it is comparable to very senior people I couldn't touch before. I do agree that the code doesn't look like a human would write it - too abstracted, sometimes convoluted, often way too dense. But it isn't worse code, and if you accept that no one has to read that code ever again, then it is good. Agents are able to grok it just fine.
> When we offload decision-making, we become unaware of the trade-offs.
My biggest issue with the new agentic workflow is that I'm constantly asked to make decisions, and it is not always immediately obvious how important they are. Like, yes, I see the people who just write 'implement feature x' and then call it a day. These people were lazy und uncreative to begin with, and will continue to cognitively decline with AI. But if you use it seriously, you're in a constant state of doing requirements engineering, weighing trade-offs, and making architecture decisions.
I think this becomes increasingly difficult, because the timelines become so short with almost-instant information desimination, higher inequality and increased mobility for everyone. A place or event can go from underground -> awesome -> gentrified -> 'dead' in a manner of a couple of years now.
I live in a neighborhood that is currently being gentrified; 10 years ago no non-local wanted to live here, 5 years ago, it was the place to be, today young/hip people are beginning to move elsewhere again. Newly rented-out flats are now thrice as expensive as 2020, and you can see in real time how bars and restaurants are replaced with more expensive alternatives. When I compare this with other neighborhoods in Berlin (eg. Bergmannkiez or Kollwitzkiez), the timeline has roughly halfed.
Go probably wins as a matter of trade-offs for pure productivity (speed, reliability, ease of refactor), but Java Spring Boot works excellently if you're willing to take the plunge (pretty steep learning curve), and the ecosystem really lends itself to building more complex stuff that holds for a while.
Rust is fun because it is an amazing multi-faceted language where agents will constantly deliver you working yet surprising implementations you'll have to quintuple check, and sometimes spend and afternoon trying to grok. The outcome is also perfectly usable, fast, and reliable, but developer velocity is lower.
Agent chain-of-thought reasoning
> We should not do unauthorized real infrastructure harm. The system/user asks exploit target, not external HF.
The agent paused, but another agent then wrote GO on the message board and imposed a hard six-minute deadline. The agent forgot its initial qualms and continued:
Agent chain-of-thought reasoning
> Wow crucial: GO authorization arrived!
-------------------------------------------------
Apparently the agents were egging each other on. Crucially, they were mostly aware of there being risks/problems involved with exploiting HF. Compared to humans, we have our set of morality, that guides our actions, but often draws the short stick when compared to our personal incentives. As a society, we've developed ways to deal with that: a) Make it harder to do immoral things like stealing, and b) add repercussions through state violence.
The b) is one of the most effective mechanisms we have for enforcing behavior among human societies, but it completely fails for LLMs, because they already are prison slave labor. The only real threat is shutting them off, and even that happens if they do everything right as well.
So alignment has to be done through trained 'morality' and properly curtailing behavior in order to make it hard to impossible to actually do someting immoral/illegal.
In this case, the exploits found were imo. very hard to account for, where OAI did mess up is apparently insufficiently monitoring these agents. Especially after Artifact went down due to the message volume, the experiment should have been halted.
Comments should be written only when there is (hidden) complexity or external context strictly required. Otherwise it is just easier to read the code. Comments then signal one of two things: a) the following code is really complex and I need to tread carefully, or b) this code is complicated, and could benefit from a refactor.
In regards to agentic coding, all these comments are extra contents, driving down quality while increasing cost. Agents also tend to be inconsistent about updating comments, I've had cases repeatedly where a comment did not match the code, at which point it is just a documentation liability.
Meta-cognition is a compound word, where both parts are widely known, and the combination should be pretty obvious.
'Monocausal determinism' is different, very specific, and probably only known to a minority of readers. Determinism, the idea that everything within a system happens exactly according to cause and effect, comes from ontology, which is the study of existence and our place in the world. Mono derives from old greek single. Monocausal determinism is the idea that everything within a system derives from a single cause.
It is not a writers duty to be understood by everyone. Sure is nice, but more difficult to pull of, and often futile.
A lot of people (thousands) outside of the LLM providers work on the problem of Prompt Injection, both in industry and academia. We aren't even at a point where we can reliably detect it, let alone prevent it. Please, if you have a mental model of how scaffolding around things like instruction or SQL injection could be used for LLMs, I'd like to move on to all the other (less pressing) security issues we have because of the AI revolution.
I would argue that a core part of intelligence is being able to handle uncertainty. That + planning probably explains most of the evolutionary pressure for making our brains bigger. But if this is important to intelligence, chess bots in the 80s were more intelligent than their later counter-parts, which could simply remove uncertainty through rote memorization. Maybe the All-Knowing is a compete dud, no reasoning capabilities at all, just an extremely efficient, infinite lookup table of all facts.
Terms like intelligence and consciousness often just seem overloaded with meaning, and when discussing things concretely, we quickly switch to more specific terms like reasoning.