HNHacker News
TopNewBestAskShowJobs

alexhans

147 karma · joined April 11, 2023

alexhans.github.io (blog)

https://ai-evals.io (community site for eval-driven development as a shared language for product building)

submissionscomments
alexhans··on Ask HN: Who's still keeping a DOS machine up because the business depends on it?
I've not virtualized things in that way for ages. I think I was using VirtualBox years ago and com0com for serial communication.

What do you use for VMs these days?

alexhans··on Pion, an agent designed to run any company autonomously
> We have decided to organize these like departments similar to the way you might hire out humans. I'm not 100% sure if that's the best approach, but I will say it's been easier for people to understand because they're more mentally easily able to traverse the bot org chart if it somewhat reflects a traditional business org chart

I think this way of thinking is going to be very important to drive adoption of AI systems because the human analogies benefit from the pre-existing domain knowledge and expectations of people.

I like using the exam analogy for evals as a qualifier for work for your "AI hires" so you can trust them to work on a specific domain.

I'd be quite curious to see what your approach to evals/testing/tracing and agent/system mutation is.

alexhans··on Ask HN: How do you manage skills files?
Since I install them with symlinks in the tools "global" locations I get access to them across projects.

Think ~/.codex/skills/<symlink-to-myskill-a/

Same for ~/.Claude or any other tool that supports skills.

alexhans··on Ask HN: How do you manage skills files?
When I say evals I mean the evals you write that verify that your use cases are upheld. Think of it like a regression test for different behaviours/user stories.

The idea would be that if you already know what you want from an autonomous system, you don't need to verify manually every time and instead just run these tests to see if there's any regression of any kind. Generally I recommend structure output and evals that are just a plain assertion, if possible. Cheaper, faster, deterministic assertions.

Does that make more sense?

alexhans··on Ask HN: How do you manage skills files?
- I don't find skills, I create them

- Keep them organised in software repos that you install with symlinks for all coding harnesses that you have. Progressive disclosure based on the frontmatter does the rest.

- I make sure they work with AI evals. Think of them like integration tests to prove behaviour. They're useful to optimize your flows. I try to make my skills be mostly a translation between natural language and good small fast tools that they call.

- I change them as a new problem arises. Not just because.

Skills can't be eaten by model capabilities if skills represent a workflow that is custom to my team or my person.

I wrote about a good mental model in the past:

https://alexhans.github.io/posts/series/evals/building-agent...

alexhans··on Pi's Minimalism Is Its Advantage
You can also use PI with local models like Qwen3.5-35B-A3B [1] and they can be surprisingly good if you development in minimalistic/simple ways.

One game changer when it comes to tweaking configs that are optimized for your use case is that you can easily use a "more powerful" cloud model to identify a good enough config for your local server/pi settings combination [2] in a pattern that applies pretty much anywhere.

- [1] https://huggingface.co/Qwen/Qwen3.5-35B-A3B

- [2] https://alexhans.github.io/posts/find-the-loop-story-first.h...

alexhans··on Ask HN: I still don't understand why AI agents need "skills"
Skills represent a pattern many of us converged to before it was named as such (very useful for comms): Progressive disclosure, determistic scripts bundled (not just markdowns). See past comment:

https://news.ycombinator.com/item?id=48132477

Don't underestimate the value of "Skill builder" skills too. Great UX

alexhans··on AIs don't do what you want. This is bad
I could but I'd probably miss the angle that is helpful to you. Do you want to instead share what your pain points are, what are you trying to delegate to AI and what you aren't? Based on that I can expand my answer better with a thing or two you can try.
alexhans··on AIs don't do what you want. This is bad
I'm a broken record but with:

- evals

- limiting AIs to tool calling, bounded planning, interpreting/producing natural language.

- bounding non determinism

- investing in small tools/security (If something shouldn't happen, then it shouldn't not be possible, RBAC style).

They can be good enough for a massive amount of contexts.

alexhans··on China’s open-weights AI strategy is winning
If you actually are interested in understanding what I said, you can look at these links:

- https://en.wikipedia.org/wiki/Fear,_uncertainty,_and_doubt

- https://www.theregister.com/software/2001/06/02/ballmer-linu...

alexhans··on China’s open-weights AI strategy is winning
Every Linux user or FOSS enthusiast knows the acronym FUD: Fear, Uncertainty and Doubt, which were a set of techniques commonly used to disparage efforts of open source communities. Linux was evil and anticapitalist and we needed to use "CorporateTool" and ban/restrict Linux.

The same companies later would be running their entire infrastructures on it and on open source.

With AI, open weights and local models, we will see the same claims, even if the named fears change.

The end users and humanity are better served by collaboration and openness than by creating oligarchies.

alexhans··on I burned all my tokens researching how to save tokens
Many of us were saying it a year ago but now with model restrictions (e.g. fable) and pricing changes it should be obvious to people that part of the economics is avoiding vendor lock-in with evals (so you can easily switch providers/models/optimize) and increasing control by investing in local models which could be good enough for your tasks, at whatever the price for your cloud compute is. Eventually consumer hardware will also be able to run good enough.

You can use Big/Cloud LLMs to help you "find good enough configs" for your local/small llms [1] and stay quite nimble in the face of rapid change.

- [1] https://alexhans.github.io/posts/find-the-loop-story-first.h...

alexhans··on Show HN: Learn by rebuilding Redis, Git, a database from scratch
Got it. Different audience. Cool stuff. I'll take a better look as soon as I can.
alexhans··on Show HN: Learn by rebuilding Redis, Git, a database from scratch
I've helped people get into programming face to face and also in a site I liked called exercism which also had a multi language track unit test passing style which I really value and it was purely command line, and I can't stress enough how important the command line is for me for people who want to dabble. Nowadays it's easier to get people into the command line because of Claude/codex.

I only have browsed your site from a phone and looks interesting but I wanted to ask if you had particular insights around getting people to approach learning, design through tests, breaking down problems, without having someone to guide them. Have you had a chance to observe people using your tool and adjust or it's been mostly dog fooding something you would've loved to have.

alexhans··on Show HN: Getting GLM 5.2 running on my slow computer
Having a thin python/ts orchestrator and workers that pick up tasks from the directories like events and decide whether to make deterministic calls and wait is pretty standard albeit custom way of doing things in this space where you're bottlenecked by the concurrent call your workers/agents can make.

The hard thing is always keeping complexity low and being ZeroOps.

alexhans··on Opinionated and easy Pi.dev configuration
I can understand someone being overwhelmed and not wanting to configure and "build your own pi" which is really one of the beautiful points of pi, but like with vim, I do recommend that after playing with this for a while you go back to pure pi and then decide what do you really need and incrementally add it.

The power of the incremental in control approach is huge. It allows you to keep moving in whatever direction you want instead of taking yet another dependency.

alexhans··on GPT-5.6
> What's the consensus today on codex vs claude code, does it really matter anymore?

Consensus is probably the wrong word for the popular opinions reflected in HN that you might get.

I would recommend that you have 2 of each at all times when it comes to AI so you don't necessarily become overly locked to quirks of one thing. You'll soon realize that things move so fast that you just start internalizing common patterns instead of depending on one specific vendor.

I recommend that you try pi and codex besides claude, to get your own feel for it.

alexhans··on Learn Vim motions with an ice-cream van
Touch typing essentially. It's such a comfortable way to work. Remapping mode switching to something like jk instead of Esc is vital to stay comfortably in the home row.

I always liked this site to grok some of those vim fundamentals [1] and the touch typing part was going to touch typing exercise webpages and getting pure practice.

- [1] http://learnvimscriptthehardway.stevelosh.com/

alexhans··on Show HN: Chess-Inspired Roguelike
Fun, responsive and intuitive. Congrats.
alexhans··on Extreme Heat conference cancelled due to extreme heat warning
I guess it depends on what your exact hypothesis is Vs mine.

For at least 6 years we've had AC worthy temps.

alexhans··on Why eval startups fail (2025)
The way eval startup is defined here is very specific and doesn't cover successful eval farmwork/SaaS vendors like Arize, Promptfoo, deepeval, etc

The author does have a point around generic benchmarks not being super valuable for companies. But evals should be seen as verifying design/behaviour constraints and can greatly aid product building, golden dataset creations and good software practices.

It's just that the aim should be "how to generate your own good evals, even if it's hard" as not so much "here's some generic evals about models".

alexhans··on Extreme Heat conference cancelled due to extreme heat warning
I'd argue that these temperatures are not a new thing. It's just a contradiction that is almost a tradition at this point, imho.
alexhans··on Extreme Heat conference cancelled due to extreme heat warning
Living in London and Dublin, what I've observed is that we get the following contradictory statements:

- "We don't need AC, It's only hot a few times during the year." - "Oh what a terrible heat, global warming is getting worse every year."

Pair to that the fact that in many places windows don't open all the way due to bureocratic regulations and many interior designs are very questionable in terms of air flow and you get some unpleasant scenarios.

alexhans··on There is minimal downside to switching to open models
Exactly. I'm very happy the discourse has moved on from "but X model is the best" to "you can use open models".

Whether you're using SDK or harness based agents, having evals means you're able to modify any part of your agent and still know what satisfies your "good enough".

It's great for designing products that are easy to change as well.

alexhans··on I told them forced consent was unlawful. 5 years later it cost Elkjop €1.8M
It's always satisfying when customer rights stories have a known positive outcome. The timeline is unfortunately quite slow and bureocractic but I'm glad OP managed to find out about it.
alexhans··on I Could've Rickrolled the FIFA World Cup. All I Needed Was My ID
I think the criticism is constructive. It's really not about hating. I'd wager many of the people who convey this criticism do use AI to aid their writing as well.

It's just that this one in particular lacks one more edit pass removing some of the AI noise on branding-speak and needless repetition (AI tends to list things and beat the point).

alexhans··on I Could've Rickrolled the FIFA World Cup. All I Needed Was My ID
Agreed. The post looks great. The story is great but the AI style in this case does distract.

I'm not against using AI for writing at all but you want to be careful that the output doesn't contain too much of this noise over signal type of wording that repeats and wants to just sell you something.

alexhans··on RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
I think it's important to be able to do both so you can stay in control of the price to value created relationship.

In last year, some people were publishing aider /ollama/open router [1] and now thankfully people are publishing all around about pi/qwen/llama.cpp/openrouter. It's widespread.

[1] https://alexhans.github.io/posts/aider-with-open-router.html

alexhans··on Ryanair dark UX patterns summer 2026 refresher
The Foil Arms & Hog 12 year old skit [1] is still relevant.

- [1] https://youtu.be/Id-zzOGnN6A (Website part at 1:42 calling out the insurance example).

alexhans··on Learn SQL Once, Use It for 30 Years
I was a fan of Seven Languages in Seven Weeks [1] because it exposed you to different paradigms which you could then try to apply where they made sense on whatever tools you were using or building: prototype based, fault tolerante, funcional, logical. Very fun book when used right.

The point being that sometimes the tools themselves don't need to survive because you take the lessons from one thing to another (e.g. move semantics and rust/modern c++)

[1] - https://pragprog.com/titles/btlang/seven-languages-in-seven-...

Page 1 of 4Next →