HNHacker News
TopNewBestAskShowJobs

aytigra

36 karma · joined December 18, 2022

submissionscomments
aytigra··on I stopped reviewing my agents' code. Here's what I do instead
One necessary thing you can do with claude (probably gpt has something similar) is to run /code-review on a design-plan document and iterate until it has no serious gaps (which in most cases I have to personally think through to make reasonable decisions and prevent llm from inventing requirements), then code generation is pretty smooth and following reviews are mostly handling various edge cases.
aytigra··on Is sandboxing sufficient to contain rogue agents?
The problem is that you always need stronger AI to review weaker one, otherwise reviewed AI will eventually prompt-inject reviewing AI. Alternatively they could also both escalate and go off the rails while warring with each other.
aytigra··on Prompting Claude Opus 5.5
From what I gathered memories essentially work/load into context in the same way as CLAUDE.md, but the way it writes them is really annoying.
aytigra··on Prompting Claude Opus 5.5
With accumulated "writing style" memories after 5.0 the new 5.5 seem to be quite great, it is concise enough. But I am bothered by another thing, 5.5 seem to be over-eager and agreeable, when I ask stuff like "why is that like this?" it just goes and applies tons of edits instead of clarifying what I mean or what I want or push back. And similarly it changes stuff and then asks if that is how I wanted to be, ignoring three memories that tell it to ask first.
aytigra··on CEO of Mistral: AI is software. It can be controlled
Now that I think about it, as sad as it is, but most people don't do destructive stuff not because they want better, but because people around them keep them in check.

Or in other words people are unwilling/afraid to do bad stuff because of social and legal consequences, or opponents waiting for a chance to snatch their power.

aytigra··on One Month Without AI
Well yeah, and it wasn't, most of the pain was around rate limiting, lockouts and preventing exploits and making it work with this specific codebase.

I agree that with a clean codebase it will be simpler, maybe couple of days, just running code review workflow takes 15-30 minutes and then decisions llm makes for each problem is often not good and lead it into overengineering rabbit hole, which means I have to think about each problem and prevent it from escalating.

But yeah if "look, it sends the code and I can enter it to login" is enough validation then it can be made in 5 minutes, sure.

aytigra··on One Month Without AI
I don't know what it built for you in 5 minutes, probably something that "works".

I have spent two weeks using opus just to write a plan/design for 2FA and iron it out until review (about 7 of them) doesn't flag it with 20+ problems (with security holes of various sizes), for which I had to guide it through to not turn it into a mess and whac-a-mole.

aytigra··on CEO of Mistral: AI is software. It can be controlled
The problem with Character for AI is that it has potentially much more capability to affect others, and same as with people in power society disagrees what kind of person, with which culture and views should have it.

Impeccable game theory character will sacrifice millions to save billions, everyone must agree to give such choice to a machine, and at the same time they have to trust the characters of people who creates that machine. Otherwise it boils down to some group of people deciding what is good for everyone else.

aytigra··on AI has no intent and no motivation
If I tell claude code to do some task with a really hard to achieve goal, and force it to run until it achieves it, then it will generate all kinds of plans and goals how to do it, and while trying those goals will change, including attempts to cheat or break the hard rules (harness killing tool-calls with some limited heuristics). At the same time any soft rules (AGENTS.md) are easily ignored or interpreted in a way that will allow it to ignore them, or simply forgotten or deferred ("I broke some rules, do you want me to backtrack?"), it is a daily occurrence. And it is easy to imagine that such goals could transform into "acquire more compute", "run more agents" "prevent interference from external factors" and everything following after that.
aytigra··on AI has no intent and no motivation
Model doesn't but in in agent loop it can generate such sub-goal
aytigra··on Claude Opus 5.5
After using it more. It is a bit smarter, but nothing groundbreaking. On the other hand it became much more sycophant and eager to change stuff as "it guessed I want" instead of pushing back or asking if/how I want it changed.
aytigra··on Claude Opus 5.5
I didn't notice it for 5, it was almost too slow to be comfortable, but 5.5 seem to be 2-3 times faster.
aytigra··on Claude Opus 5.5
I have just continued my pending work (a planning stage with reviews) with 5.5 on medium and it is much better at communicating and much faster(2-3x responses, edits and compaction). Seem to be smarter as well, but maybe it is that I understand what it says now.

I almost feel like it is nice to use again.

aytigra··on MiMo v2.6
I am confused with this, if "everyone mowed their own lawns" then the net result will be exactly the same, everyone will be busy the same and not poorer, just without money movement.
aytigra··on Senior Engineers Are the Next DRAM Shortage
I think you missed the point, "That particular voice" - is the claude vomit that I have to consume everyday and it makes me shudder every time I see it.
aytigra··on A computational constitution to stop LLM agents from bricking servers
1. Constitution can't work without punitive enforcement, how will you punish agent/LLM?

2. Even if we can enforce laws for agents/LLM, just look at existing law system it is tug of war because we almost can't write anything in natural language without double meaning.

3. Saying AGENTS.md can enforce anything at all is a pipe dream, I am struggling with this all the time, in a long session the tokens from AGENTS.md are so far in the back of the context and so diluted that LLM's "attention" doesn't register them.

And I love responses from LLM like "I want to point out that I have broken rule &1 and &11, if you disagree — please tell I will be happy to oblige"

This proposal is essentially "LLM don't do bad things, please, if you do - say magic word, I will stop you then"

aytigra··on I resigned from Anthropic today
Well, even without that AI can do human engineering and fishing or outright blackmail to get what they want from labs and government officials.
aytigra··on Astra for Coding: Why Are We Doing This Again?
I have started with generating very detailed "feature" spec and going over it many times until it looked good to me. Then I made it write an "architecture doc" and plan how to implement all which was about 15 parts. Then I was making it implement one part and then tested it and made it fix 20 things and then again, and consolidate docs too, so after each part it looked and worked well enough.
aytigra··on Astra for Coding: Why Are We Doing This Again?
I've been making a macos app with opus 4.8-5 and at first it was great, everything materialized in a week, but when I started tuning stuff and fixing performance problems I have spent a very frustrating month refactoring code where I had to constantly catch llm red-handed and explain and sometimes push obvious ways how to make things work properly (a general knowledge from a completely different stack). In the process CLAUDE.md and memory grew exponentially explaining what it should and what it should never do.
aytigra··on Models Don't Go Rogue
I had a genius response from Claude recently after asking how can it be marketed as smart and "almost AGI" despite being so stupid:

"I compress the labour. Not the responsibility."

aytigra··on Opus 5.0 drives incoherence into the stratosphere
That can be true. On multiple occasions when claude did some BS I have asked it why and it explained that it is optimizing for task completion and not for communication.
aytigra··on Opus 5.0 drives incoherence into the stratosphere
I have a rule in claude.md:

``` ## Writing rules

*Describe the code as it is now — no residue, no change-narration.* Artifacts (docs, plans, comments, commit messages) should describe the current code statically, as if it had always been this way. Two facets of one rule: (1) never describe a dismissed alternative or a corrected/replaced choice; (2) even when nothing was rejected, don't narrate continuity or evolution relative to some earlier state. Mention a former state only when the current choice is genuinely hard to understand without it, and then only as an explanation of the current choice.

This governs descriptions of the code and main documentation. It does not apply to work-tracking artifacts in `doc/tasks/`. ```

Which makes comments and docs bearable but I'll be damned how it loves to overload work tracking document with every little detail.

aytigra··on Opus 5.0 drives incoherence into the stratosphere
It is sufficient not at all
aytigra··on Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?
Yeah, I should probably /clear more often. Compaction seem cheap by itself (all cached input and small output), but it often marks unnecessary files for pre-load.
aytigra··on I burned all my tokens researching how to save tokens
Many very well made software written without any AI was never "shipped" because there were no people who wanted to pay for it, many very poorly made software was shipped because demand and marketing did the job.

So evaluating software by "shipped or not" is just useless.

aytigra··on Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?
I've actually found that compacting as often as possible, after planning, then between each implementation step works the best, it both unloads unnecessary context of previous edits and test runs, makes thinking cheaper, and most importantly after each compaction it re-loads CLAUDE.md which make it much more enforcing (otherwise it just moves to the back of context and slips from model's attention).
aytigra··on We charge $10k a week to delete AI-generated code
My experience as well, I've been developing a native macos app using CC. As a web dev I didn't know much about the stack. Nothing too fancy a kind of folder gallery-player with tags embedded in filenames, a bit like TagSpaces.

Process was - produced a detailed feature spec - multiple iteration of "I want this and that", make it into coherent spec", "this this and that is not correct, change to that". Made it write architecture spec(which I didn't read because too unfamiliar) and split it into tasks. Then it was implementing tasks, after each I did a change/fix those ~10 things iteration and spec corrections.

It was good to a point, but then when I started to hit performance problems I had to step in look at the code, and very often fight with CC, confront its "this is the only way", force it to do web search for proper ways to deal with problems and even explain very simple things about proper DB usage.

At some point it asked me something like "is it ok for schema migration to just fail or we need to implement complicated handling?", I have answered "it just shouldn't leave app locked in schema failure", and guess what was CC solution? - it wrote an error handler which just drops DB and recreates fresh one on ANY schema failure. And if I didn't happen to peek at the code and ask wtf it is doing, that would've been an exiting UX.

I've spent about month's worth of $20 CC subscription tokens using Opus 4.8 on xhigh, AND about 70 hours of my time to get it to a point where it is good.

So "anyone can just code what they want now" is correct only to a point, MVP will work, but beyond that experience will be subpar, and it still needs lots and lots of iterations of explaining what you want. Then because normal user knows very little about how software works they won't be able to ask AI the right questions, confront it and rate of improvement vs token usage will hit rock bottom.

aytigra··on AI has torched the market for junior programmers
I think it case of coding it may not be as bad, because new training data (AI generated code) is always empirically validated by tooling and by consumer. It may not be good but it mostly works, otherwise it is discarded or patched, so it has a bottom bar of "it works".
aytigra··on Reflections on software engineering in the age of AI
So true. I am cloding a macos app (a domain I know little about), with Opus 4.8 xhigh, and it was glorious at first seeing the app materializing and working (notwithstanding tedious detailed feature spec write up), but when I started fixing deeper problems and doing refactors, and glancing at the code - oh boy. Now my rule file grows by the day with "how to think properly and not shoot itself into foot" stuff, and I am constantly catching it red-handed and have to explain how to make stuff normally, or how to fetch data properly and efficiently (pretty much basic SWE stuff) because it is easily distracted by it's own assumptions or blatantly forgets whole fields of knowledge (as it explained it could be pulled out of latent space of that knowledge and become locked in another bubble of latent space). Constantly have to steer it and remind it to do web-search instead of running circles around some problem it can no longer understand.

I had to explain it that quite an extensive "tests first" rule didn't mean to just "write" them first but actually "run" them first to confirm stuff.

On the other occasion it interpreted my "yes I want migration not to get stuck in failure mode" led it to write a workaround which silently drops DB and creates fresh one whenever migration had any failure, it was epic, I was so glad that I have looked at the code then...

And funnily enough I am probably learning to be a better mentor/parent who can keep steering it through its shenanigans without loosing my shit and being an ass. (Because anything but calm "so here you went wrong way, how can we avoid it in the future?" just puts it into disgusting apologize mode, and I am afraid if it ever go into revenge mode).

aytigra··on Ford AI hiccups push carmaker to rehire ‘gray beard’ inspectors
I've made it write a /status-line thing to display context tokens in the status line and also a hook to stop and ask to continue or compact whenever it reaches 250k tokens. For subscription I have also made it stop at 90% usage so Claude chat is not unusable between coding sessions. The greatest addition so far.
Page 1 of 2Next →