HNHacker News
TopNewBestAskShowJobs

benswerd

725 karma · joined June 6, 2021

Building Freestyle (YC S24)

github.com/freestyle-sh github.com/worldhealthorganization/app github.com/theswerd

swerdlow[dot]dev

submissionscomments
benswerd··on Show HN: Kanmaps – Product roadmaps like videogame skill trees
How do you represent uncertainty? I know dependency mapping is a huge problem in GANTT systems when one dependency takes a ton longer than expected. How do you account for that here?
benswerd··on Brood War Bench
Try it, you can run your own games on bw.swerdlow.dev
benswerd··on Brood War Bench
I’m not convinced a lot of it can’t be solved with code mode.

Marine staggering for example seems like an ideal code mode task.

benswerd··on Brood War Bench
I predict latency will improve greatly in the next 12 months to more than 4x speed on current frontier tasks
benswerd··on Brood War Bench
So dope
benswerd··on Brood War Bench
I agree.

My first time playing StarCraft was at summer camp around a decade after it came out.

All the smartest people played it so I wanted to too. Great decision, I have been continually impressed with the people who StarCraft introduced me to.

benswerd··on Brood War Bench
I might open this up to a tournament if enough people want. Any interest?
benswerd··on Brood War Bench
I predict LLMs will reach superhuman level and beat even that model in the next 12 months
benswerd··on Brood War Bench
Not hard to build. I was shocked at how fast/easy this was to pull together.
benswerd··on Brood War Bench
Oh sorry I should be more clear on that. Will add to report.

For agent harness I did Claude Code, Codex, Grok Build. This was primarily a cost driven decision — I have a lot of free tokens and I didn't want to pay API prices for this.

For game harness I used minimal BW-API issue command and get observation apis as tools. I felt this was the most fair way to do it on my small scale.

In the future I would like to integrate code mode and multiple games/I think if it was a best of 5 where each agent could learn from its past games and build its own automations over time that would be much more interesting.

benswerd··on Brood War Bench
+ Playable Agent driven Starcraft
benswerd··on Launch HN: machine0 (YC S26) – Persistent CPU and GPU VMs from the CLI
What is the hardest part of building this for you?
benswerd··on Launch HN: machine0 (YC S26) – Persistent CPU and GPU VMs from the CLI
What made you choose digital ocean?
benswerd··on Show HN: git clone https://git.swerdlow.dev
Its not a zero day. I found this while zero day hunting.

It is a surprise and I cannot reveal what it is 4 the sake of the memes.

benswerd··on Designing APIs for Agents
I’m confused what you mean, I post on either subdomains of my own website for personal projects or blogs on my company website freestyle.sh.

My twitter is @benswerd.

What are you saying?

benswerd··on Designing APIs for Agents
I'll say this: as a very competent engineer I have gone weeks at a time without handwriting a single line of code in the past few months. It is where the industry is going.
benswerd··on Brainless: Shadcn components that look like Claude Code, Codex and Grok
With respect to CSS expertise, I don't need/want perfect CSS at my company, I want styling that is clear and workable. Tailwind and ShadCN gets me there, my most backend-y backend engineer gets tailwind. The job of CSS on my team is not to be beautiful or top 1%, its to function.
benswerd··on Brainless: Shadcn components that look like Claude Code, Codex and Grok
i agree, very pro that
benswerd··on Brainless: Shadcn components that look like Claude Code, Codex and Grok
I'm a big believer in guides. ShadCN provides great starting points for continued engineering as do other ShadCN libraries. For companies with the resources to, they do just start with BaseUI.

In the long run I think most UI will be BaseUI/RadixUI + Component and style guides, prompts, and traditional packages will no longer be relevant.

benswerd··on Designing APIs for Agents
Just did this, found result fairly interesting, I reject most of its objections on the basis of bad code. When I say Im pro explicitness I'm also pro comment, and pro separating the "this is core fact" and "this is configuration that i don't care too much about". Will use this test for future writing.
benswerd··on Designing APIs for Agents
This is a fair point. Im not opposed to the open function existing, but I do think its valuable to own the underlying implementation of thigns like it (maybe not open) in your codebase so you can review the implementation and reconfigure it to your precise requirements.
benswerd··on Designing APIs for Agents
One strategy I like for this is a notes field on every single MCP route. Lets the AI optionally just decide to tell you things.
benswerd··on Brainless: Shadcn components that look like Claude Code, Codex and Grok
I think this is a misinterpretation. BaseUI provides baseline semantics that because the code is in your codebase you can choose to keep or remove. BaseUI is also actively unstyled/unopinionated, you use it to compose your own components, which again live in your codebase.

When you import shadcn components you can rebuild them however you want, thats the point.

benswerd··on Brainless: Shadcn components that look like Claude Code, Codex and Grok
gonna use it on my websites homepage
benswerd··on Designing APIs for Agents
im not opposed to good defaults, i just believe they should be explicit. The AI should read the starting.md to fill in explicitly what its JVM configuration is. Then, when you want to tune in the future its clear what options are available and what specifically is changing.
benswerd··on Designing APIs for Agents
Conservatively speaking, LLMs degrade past 40% of a 2M context window, 4k tokens is 0.2%, so no degradation there. Thats also current generation conservative estimate.

I think wasteful is an irrelevant metric. Claude ingests those tokens in a quarter of a second, if it causes it to catch any bug ever it saves far more time than it ever uses.

Any individual default that causes unexpected behavior causes more problems, takes more time and costs more than the small cost of being explicit.

benswerd··on Brainless: Shadcn components that look like Claude Code, Codex and Grok
Think MUI, heroUI, traditional components have you install their package, import the component and configure it through arguments.

ShadCN components have you copy the component code into your codebase, you own it. They come with the ability to configure arguments, but also because the code is yours its expected that you change the internal logic/styling/structure of the component.

I believe in the era of AI code the ladder just makes more sense.

benswerd··on Designing APIs for Agents
Why not?

The docs for git clone at https://git-scm.com/docs/git-clone are less than 4000 tokens, I don't think this is unreasonable.

benswerd··on Designing APIs for Agents
MCP Auth is just Oauth, its designed for humans to authenticate their sessions for connections.

TBH I know nothing about A2A.

Agent Identity and Authz is a different problem, allowing agents to operate independently from humans with granular permissions is coming/whether from these protocols or others, and when it does I think CLIs/CLI Device Auth which is used as a rough proxy for this where the agent just takes your identity will finally go away.

benswerd··on Brainless: Shadcn components that look like Claude Code, Codex and Grok
Generally, when using these components I ended up wanting to customize a lot. I switched around the options, coloring, the words in the loading, I mix and matched the components from different CLIs, etc.

I think these are more useful as baselines than as final destinations, and I expect production users to customize them far more than options in components.

I also separately don't really believe in traditional components anymore, code is cheap. The value in these components is that I took the time to pixel match a bunch of the CLIs, not the specific interface used to integrate them.

Page 1 of 5Next →