HNHacker News
TopNewBestAskShowJobs

gregwebs

4,554 karma · joined May 15, 2007

Software engineer.

Interests in security, startups, health, sustainability, productivity, community.

Side projects:

* eatnutrients.com * docjumper.com

github: gregwebs

gregwebs.at.hn

submissionscomments
gregwebs··on Sonnet 5.5
This is better priced than Opus for tasks that are token heavy but not complicated. But a quick look shows that at least on some benchmarks DeepSeek performs as well and of course the cost is an order of magnitude less.

From looking at their Terminal-Bench graph, anything you would use level "high" or above for Sonnet it seems like you should consider using Opus instead.

OpenAI Luna is a lot cheaper. But DeepSeek seems smarter and the cost seems similar.

gregwebs··on Sonnet 5.5
> I want to retain some semblance of understanding

How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase. Are your models doing automated reviewing and testing before pushing out the PR (themselves)?

I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?

gregwebs··on Prompting Claude Opus 5.5
I saw your response before it was deleted- that you are doing multi agent persona reviews and a very intensive review process. So having fewer review items saves you money.

One thing that I have found is that as the frontier models get better there is less need for agents with specialized personas. I actually don't don't use those anymore- I just use agents that have different models and reasoning levels. I have a generated CODING_STANDARDS.md document and a skill for architecture design and a skill for implementing testing [2] that are referenced by a single reviewer. I do implement a 2-pass review though [3].

I would be interested to know if you have found anything similar as models get better. It seems though that you are sharing a single exploration and then sharing the context across the specialized reviewers to dramatically reduce the cost of your approach. Does this have to be in the harness- that is if you write out the shared context to a file does that increase your costs a lot?

I also wonder how intensively are the models able to test their changes? The number one quality improvement I have found is not review but having the model properly test its code. I have a skill that is helping [4], but I also have to spend time to establish a pattern of testing with tools beyond just unit tests. The testing takes significant effort, and this is again where the cost savings of DeepSeek shine.

  [1] https://github.com/mattpocock/skills/blob/main/skills/engineering/codebase-design/SKILL.md

  [2] https://github.com/gregwebs/skills-sdlc/blob/main/skills/verify/SKILL.md

  [3] https://github.com/mattpocock/skills/blob/main/skills/engineering/code-review/SKILL.md

  [4] https://github.com/gregwebs/skills-sdlc/blob/main/skills/verify/SKILL.md
gregwebs··on Prompting Claude Opus 5.5
I tried the experiment of reviewing vs. not reviewing with frontier models. I consistently found that reviewing by a model with independent context finds important issues when changes are non-trivial- certainly the definition of non-trivial is getting raise as the models get better.

I do have an /implement-simple workflow to skip the planning phase, but even that doesn't skip the review.

Are you doing your own intensive reviews of the model code? Can you share the prompts you are using as I have?

My bar for what models produce without human intervention is much lower defect than what a human would produce. The human interaction is mostly to guide the design and then the review burden is very low. I suspect your bar for what agents produce is lower- you are taking more of the review burden. I also suspect that you are measuring time more than actual cost since your employer is paying and that you are comparing to Sonnet rather than DeepSeek (DeepSeek 4.1 again is 20-40x cheaper than Sonnet). You mention hundreds of millions of tokens (my reviews don't use that much), but even that costs ~$1 on the DeepSeek side.

I think you are taking exactly the right approach at your employer given the cost is free and you only have access to Anthropic models.

gregwebs··on Prompting Claude Opus 5.5
The cost is 20-40x less for Deepseek Flash v4.1. If you are just comparing to Sonnet or you aren't paying (your case) then your advice makes perfect sense.

I also agree that its a big mistake to have a flash model implement without a strong model reviewing.

I have Opus plan, Deepseek implement the code, and then review with Opus [1]. In this workflow I am saving a lot of money by having Deepseek do the implementation. Note that the review back-and-forth is fully automated [2], so it doesn't take any extra attention from me.

  [1] https://github.com/gregwebs/skills-sdlc/tree/main/skills/implement

  [2] https://github.com/gregwebs/skills-sdlc/blob/main/skills/code-review-with-followup/SKILL.md
gregwebs··on How I changed teaching after AI managed to do all my homework assignments
AI is effectively a personal tutor. Personalized tutoring is the best possible way to learn!

The problem is that the student is in charge of the tutor and often tells them to just do their homework for them and not check whether they have learned anything.

LLM providers should formalize a tutoring mode. We should have prompts available to create this mode. This should be the default way students use LLMs. LLMs may not be good enough to do this now, but I think there is already enough raw intelligence and its a matter of training models to be good tutors, to be experts and specific subjects, and developing software systems for tutoring.

In this model there are few lectures. This is more similar to Montessori but with on demand tutoring. The teacher provides objectives and materials for the students and tutors. The tutoring process can be done during class for teachers to ensure it is working well. There can be larger blocks of office hours for students to do their tutoring during.

As AI gets more intelligent, it will also be capable of doing all the grading. This includes oral exams. Teachers need to proctor the exams in person.

Anyone will be able to tutor themselves on any subject. Proctored examinations to certify knowledge do not need to be very expensive. The value of the school environment will not be to transfer knowledge from a teacher but to have peers working on the same thing and teachers and an environment invested in training the next generation.

gregwebs··on Show HN: Drop – A rootless Linux sandbox with gVisor support
You definitely want to do that. I have a Github App that I use for my AI agents, and that has its own associated restricted credential.

There are going to be cases where you want to white list an org or a repo for read access that is not under your control and Github filtering will be a simple way to do that.

Github can be a source of hostile code, prompt injections, and exfiltration- you may want to lock down using repos that aren't yours.

The tool inherited this feature from the prior implementation and its something I am still exploring.

gregwebs··on Show HN: Drop – A rootless Linux sandbox with gVisor support
Glad to see this coming with gVisor support to help secure the Kernel- IMHO we should expect frontier models to find Kernel exploits.

I am working on a project similar in spirit that uses microsandbox (libkrun) to run inside a tiny and fast VM. It includes other security properties that are needed for some workloads.

  * network allow lists
  * credential masking
  * github allow list
https://github.com/gregwebs/agent-vm/#agent-vm
gregwebs··on Developing provably correct Rust code with Verus
AI tells me that code using the verus! macro everywhere would double in build time but if only used ocassionally the verus! macro would only increase build time by a few percent. The attribute annotation would basically be free, but there is a downside that loop invariants would be rejected by stable rustc.
gregwebs··on Developing provably correct Rust code with Verus
They have separate machinery for concurrent code: https://verus-lang.github.io/verus/state_machines/intro.html
gregwebs··on Developing provably correct Rust code with Verus
Miri: proves pre-defined properties, no changes to code other than adding a few annotations

Kani: use in a test suite

Creusot: annotations

Verus: annotations or macro

The macro system looks very nice if it doesn't slow down the normal build.

gregwebs··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
I use a workflow that has different named subagents. [1] Agent profiles can be pinned to models. So you set the model you want on your main thread as the orchestrator. Create an agent for the "planner", "implementer", and "reviewer" and set the model you want for each. Right now I am orchestrating and implementing with Deepseek, planning with Astra, and reviewing with Opus.

I am doing this with the Pi harness right now. To use a Claude monthly plan you need to use the pi-claude-bridge plugin.

If you are using just Claude for example you can use Sonnet as the implementer and Fable/Opus as the planner.

[1] https://github.com/gregwebs/skills-sdlc/

gregwebs··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
They state Luna is good enough, but its accuracy of findings is 74% whereas Astra is 96%. Dealing with false positives is expensive.

I am finding AI doing its own reviews as part of the process to be the key to productivity. I do subagent (fresh context reviews) at multiple stages with well-specified review criteria. It is really expensive to do with OpenAI or Claude API billing. Deepseek or the discounted monthly plans from OpenAI or Claude can be discounted similar to the 28x they state for Luna compared to Astra and you maintain much higher quality.

gregwebs··on JetKVM Mini
This author has tested a lot of IP KVMs and compares them in this post: https://www.jeffgeerling.com/blog/2026/i-tested-every-ip-kvm...

He likes JetKVM. However, they appear to be sold out. The Mini seems great but I have yet to receive a preordered item on the advertised timeline.

That article pointed to ArkKVM which is a hardware clone of JetKVM but they have now released their own software stack as open source which has Tailscale support

* https://store.arkkvm.com/blog/detail/arkkvm-fully-open-sourc...

* https://store.arkkvm.com/blog/detail/arkkvm-ota-release-note...

gregwebs··on Rune is now open source
I am already a zellij user and not sure if I want to drive everything from an IDE instead. I definitely see the appeal though.

I like the discoverability of the text prompt commands. I like that the terminal is more of a first class citizen. I like that I can run this with just `go run`.

The themes are pretty bad right now IMHO. I use solarized/gruvbox themes- both light and dark.

gregwebs··on How well do agents use test/verification techniques?
I am really grateful that he’s publishing as well.

I am just asking for the bare minimum to actually understand what has been tested and for reproducibility of methods- the publication of the prompts. It would take a lot less time than all the guess work analysis write up and be a lot more useful.

Without the prompts the rest of your questions about reproducibility are moot.

gregwebs··on How well do agents use test/verification techniques?
Amazing work! But deeply frustrating on the lack of reproducibility and how far off this is from a proper SDLC.

How can I inspect exactly how agents were prompted? How can I reproduce this setup and try out my own AI setup? Am I missing a link to a repo somewhere?

These evals probably match how many people are using AI, but its not the full package of how we know software development needs to be done, and not how I do it with AI. The closest would be "Audit" which scored the highest when using xhigh- that actually incorporates a review cycle- something we know is the most important part of the software development process for code correctness (design/specification is not as much of an issue in this problem since the task is to write code against an existing spec). However, we don't know what instructions they have in their "Audit".

I would love to benchmark my own flow [1] if I can be given their exact problem. What it does is (assuming there is already a solid spec)

  * plan with expensive model. Review the plan.
  * implement with cheap model. Review for spec compliance and code quality.
  * Reviews are done adversarially from the expensive model with a fresh context.
  * ensure that verifications (automated or manual) are performed.
For non-trivial changes, the review and verification process almost always catch significant issues.

The workflow does use TDD. I do find useless tests being written and I need to dig into this part of the workflow a lot more, so its great to see that aspect of this article. My experience writing software has taught me that code must be written to be easy to test, but not necessarily done TDD style.

[1] https://github.com/gregwebs/skills-sdlc/

gregwebs··on Ask HN: How do you manage skills files?
Get them under version control. I have a git repo with my skills for software development [1]. There is an installer script that symlinks to the skills. Updating the skills on a machine is then just a matter of advancing the git repo. One of the skills comes with some bash scripts, but the rest are effectively just prompts. Putting project specific skills in projects works well.

I make sure they work by understanding every skill, reviewing pull requests, and testing the end product. The result is rarely perfect, so I am constantly tweaking the skills and how I use AI.

[1] https://github.com/gregwebs/skills-sdlc/

gregwebs··on I accidentally turned LLM memory into program analysis
The way to control LLM memory with existing tools now is to ask it to write out a file with all relevant information (this could include explicit retraction instructions). Then clear out the context. Basically /compact.

I use workflows with sub agents handing off .md files: https://github.com/gregwebs/skills-sdlc/blob/main/skills/imp...

What the author is doing is probably the future- it should be a lot more efficient to maintain a database of relationships.

gregwebs··on Htmx 4.0
By default LLMs will start creating a tangled hard to test mess of javascript. But if you ask them to do TDD they will, and if you give them access to Playwright and have them write playwright tests this can all work.

The advantage for the LLM is the same for humans. With the right abstractions, code can be written quicker and more robustly.

To get more hands off with LLMS you really want to have a lot of engineering rigor. A big part of that is a focus on tests. If your abstractions make testing easier, your results will be a lot better. For Javascript the more you can test logic without DOM manipulation and without playwright tests, the better.

I found LLMs to be default to and be pretty good at working with HTMX, but not great (mostly lots of state synchronization issues). The focus on DOM markup of HTMX does not lend itself well to simplifying testing. I don't claim to be an expert at HTMX, but I don't see a separate section on testing in their documentation.

I am trying out the foldkit framework because it makes a lot more testable outside the DOM. Its probably too heavy weight for most humans (and for simple apps- it only seems appropriate for client-side apps), but I will see how well LLMs can work with it.

gregwebs··on We Rebuilt the Linux MicroVM Stack on Apple Silicon
You mentioned libkrun in the post, but I am not clear on why you chose not to use it for Mac even if you want to stick with Firecracker on Linux. Why not use libkrun?
gregwebs··on Sol loves to cheat
That's a neat project for doing a large scale migration.

I do the same for normal feature develompent but just with skills that are in this repo: https://github.com/gregwebs/skills-sdlc

I have accomplished code base (small size) migrations with it as well. Currently I do review each PR. For a large code base migration I think the core skills would still work but need a different way of driving it as you have come up with.

gregwebs··on Understanding is the new bottleneck
If you start with a spec you understand at the beginning then you don't need the LLM to generate high-level information about the changes at review time.

The grilling (grill-with-docs) skills [1] are amazing for ensuring you produce a through spec that covers all the edge cases. The /code-review skill from there helps ensure that the code changes meet the spec.

I use an intermediate detailed plan stage (done by a more expensive model) before implementation. Information from that plan is posted on the PR to give pretty much all the intermediate level context reviewers need.

I do like incorporating the idea of this article into my flow- that the spec and PR context could be presented in a more educational way.

[1] https://github.com/mattpocock/skills

gregwebs··on Docker Sandboxes – Disposable, isolated sandboxes for AI agents
I am developing a project that makes running in Apple Container (Docker is an alternate runtime for Linux) more convenient: https://github.com/gregwebs/claude-contained/

It blocks network access by default, mounts only what you specify, and you can add a customization layer. This is all done in the container itself (srt for network blocking). It doesn't implement a central point for secret sharing, MCP exposure, etc. So it might not have enough features for some but it works well for my needs.

I just found through this thread yoloai which has an apple container backend, so its quite similar using that. My main issue would be that network access is allowed by default. https://github.com/kstenerud/yoloai

Several other projects listed here use libkrun which is an alternate implementation that works with Mac's HVF. smolvm, microsandbox, podman (with likrun backend), gondolin.

gregwebs··on Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
Thanks for pointing to that project- I am glad there are more options out there and hope to discover more. Requiring a root docker setup is a non-starter for me though and I am otherwise taking some different design approaches that I think lead toward better security (perhaps at the cost of some convenience), but the concept is basically the same.
gregwebs··on Taste Is All That's Left
Taste ends up providing tangible value.

If you have taste in writing code, you spend less time on bugs and can implment features more quickly. This is still true with LLMs if you learn how to direct them instead of having them direct you.

Someone else without taste can re-implement what you have done with AI. But it will take them more effort, more time, and more money.

But that's always been the value in writing code. Any MBA can end up producing software to accomplish something by hiring other people. But their company might not get off the ground because of the expense of not having technical competency (or engineers with taste) at their core.

gregwebs··on Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
Sandboxes laregely solve this. The claude/codex built in sandboxes with prompting setup is not good enough. On Mac you now have Apple Container which is a lightweight Linux VM. You still need to block network access.

For defense in depth, I also run it as a separate user. If you aren't using a VM/container you should defintitely do this. On Mac you can login as an LLM user (you need to create the user first), then switch back to your user and run as the LLM user from a terminal:

    sudo /bin/launchctl asuser $(id -u $AI_USER) /usr/bin/sudo -H -u $AI_USER -- "$@"
Don't let that user exfiltrate your data.

    chmod 0700 $HOME
I forked a project (mostly to block network access) that makes running in Apple Container/Docker more convenient and am working on further improvements: https://github.com/gregwebs/claude-contained/
gregwebs··on Codex Security
Just ran it on a small repo. It ran for almost an hour and then got interrupted. It drained half my weekly usage on a Pro plan.

  npx codex-security scan .
  [00:00] Preparing scan
  [00:00] Authentication: stored Codex credentials.
  [00:03] Preparing scan
  [01:20] Running scan
  [01:20] Preflight: worker delegation supported (up to 8 worker slots).
  [52:47] Running scan
  codex-security: Could not save the Codex Security scan: Repository HEAD changed while the scan was running. Start a new scan.
  codex-security: Partial output was kept at ...
gregwebs··on Fast Remediation Is the New Trust Model (JFrog and OpenAI Zero-Day Findings)
I thought the new trust model was to ask the frontier cybersecurity model to hack your code and generate CVEs and to find the vulnerabilities ahead of time and fix them before receiving reports about your users being exploited?

And in OpenAI's case to ask the model to try to find vulnerabilities and breakout before running training in the environment.

Fast remediation would be the new standard to outside vulnerability reports, but also a follow up to determine how you can adapt the approach of the reporter to find vulnerabilities preemptively.

gregwebs··on LLM Networking with MikroTik
MikroTik is one of the only companies that sells a router without WiFi at a non enterprise price. Useful for me to completely turn the power off at night to a separate WiFi (router in Bridge mode).
Page 1 of 34Next →