What do you use for VMs these days?
147 karma · joined April 11, 2023
https://ai-evals.io (community site for eval-driven development as a shared language for product building)
What do you use for VMs these days?
I think this way of thinking is going to be very important to drive adoption of AI systems because the human analogies benefit from the pre-existing domain knowledge and expectations of people.
I like using the exam analogy for evals as a qualifier for work for your "AI hires" so you can trust them to work on a specific domain.
I'd be quite curious to see what your approach to evals/testing/tracing and agent/system mutation is.
Think ~/.codex/skills/<symlink-to-myskill-a/
Same for ~/.Claude or any other tool that supports skills.
The idea would be that if you already know what you want from an autonomous system, you don't need to verify manually every time and instead just run these tests to see if there's any regression of any kind. Generally I recommend structure output and evals that are just a plain assertion, if possible. Cheaper, faster, deterministic assertions.
Does that make more sense?
- Keep them organised in software repos that you install with symlinks for all coding harnesses that you have. Progressive disclosure based on the frontmatter does the rest.
- I make sure they work with AI evals. Think of them like integration tests to prove behaviour. They're useful to optimize your flows. I try to make my skills be mostly a translation between natural language and good small fast tools that they call.
- I change them as a new problem arises. Not just because.
Skills can't be eaten by model capabilities if skills represent a workflow that is custom to my team or my person.
I wrote about a good mental model in the past:
https://alexhans.github.io/posts/series/evals/building-agent...
One game changer when it comes to tweaking configs that are optimized for your use case is that you can easily use a "more powerful" cloud model to identify a good enough config for your local server/pi settings combination [2] in a pattern that applies pretty much anywhere.
- [1] https://huggingface.co/Qwen/Qwen3.5-35B-A3B
- [2] https://alexhans.github.io/posts/find-the-loop-story-first.h...
https://news.ycombinator.com/item?id=48132477
Don't underestimate the value of "Skill builder" skills too. Great UX
- evals
- limiting AIs to tool calling, bounded planning, interpreting/producing natural language.
- bounding non determinism
- investing in small tools/security (If something shouldn't happen, then it shouldn't not be possible, RBAC style).
They can be good enough for a massive amount of contexts.
- https://en.wikipedia.org/wiki/Fear,_uncertainty,_and_doubt
- https://www.theregister.com/software/2001/06/02/ballmer-linu...
The same companies later would be running their entire infrastructures on it and on open source.
With AI, open weights and local models, we will see the same claims, even if the named fears change.
The end users and humanity are better served by collaboration and openness than by creating oligarchies.
You can use Big/Cloud LLMs to help you "find good enough configs" for your local/small llms [1] and stay quite nimble in the face of rapid change.
- [1] https://alexhans.github.io/posts/find-the-loop-story-first.h...
I only have browsed your site from a phone and looks interesting but I wanted to ask if you had particular insights around getting people to approach learning, design through tests, breaking down problems, without having someone to guide them. Have you had a chance to observe people using your tool and adjust or it's been mostly dog fooding something you would've loved to have.
The hard thing is always keeping complexity low and being ZeroOps.
The power of the incremental in control approach is huge. It allows you to keep moving in whatever direction you want instead of taking yet another dependency.
Consensus is probably the wrong word for the popular opinions reflected in HN that you might get.
I would recommend that you have 2 of each at all times when it comes to AI so you don't necessarily become overly locked to quirks of one thing. You'll soon realize that things move so fast that you just start internalizing common patterns instead of depending on one specific vendor.
I recommend that you try pi and codex besides claude, to get your own feel for it.
I always liked this site to grok some of those vim fundamentals [1] and the touch typing part was going to touch typing exercise webpages and getting pure practice.
For at least 6 years we've had AC worthy temps.
The author does have a point around generic benchmarks not being super valuable for companies. But evals should be seen as verifying design/behaviour constraints and can greatly aid product building, golden dataset creations and good software practices.
It's just that the aim should be "how to generate your own good evals, even if it's hard" as not so much "here's some generic evals about models".
- "We don't need AC, It's only hot a few times during the year." - "Oh what a terrible heat, global warming is getting worse every year."
Pair to that the fact that in many places windows don't open all the way due to bureocratic regulations and many interior designs are very questionable in terms of air flow and you get some unpleasant scenarios.
Whether you're using SDK or harness based agents, having evals means you're able to modify any part of your agent and still know what satisfies your "good enough".
It's great for designing products that are easy to change as well.
It's just that this one in particular lacks one more edit pass removing some of the AI noise on branding-speak and needless repetition (AI tends to list things and beat the point).
I'm not against using AI for writing at all but you want to be careful that the output doesn't contain too much of this noise over signal type of wording that repeats and wants to just sell you something.
In last year, some people were publishing aider /ollama/open router [1] and now thankfully people are publishing all around about pi/qwen/llama.cpp/openrouter. It's widespread.
[1] https://alexhans.github.io/posts/aider-with-open-router.html
- [1] https://youtu.be/Id-zzOGnN6A (Website part at 1:42 calling out the insurance example).
The point being that sometimes the tools themselves don't need to survive because you take the lessons from one thing to another (e.g. move semantics and rust/modern c++)
[1] - https://pragprog.com/titles/btlang/seven-languages-in-seven-...