HNHacker News
TopNewBestAskShowJobs

visarga

13,867 karma · joined September 4, 2011

submissionscomments
visarga··on Gemini 4 Argon
I built a harness that externalizes state into files so you can swap providers. I regularly move between claude and gpt, but it works with all providers including pi and local models.

I log everything - user messages, tasks, project memory, even bash commands for forensics. As a consequence you can do reflection where you analyze past work and extract refinements for the harness and realign the project when it diverged from user intentions. I don't have to do this manually, it reduces steering work.

https://github.com/horiacristescu/playbook-harness

visarga··on What reversing, modernising old games tells us about the economic impact of AI
When they diverge you use the agent to classify bug one one of them or spec is unclear. You can both fix bugs and refine the spec iteratively.
visarga··on Show HN: Raven – The harness of harnesses, built for RSI
You can't fool me - this prompt refinement turtle stack only has 3 turtles. Not much way down.
visarga··on Show HN: Raven – The harness of harnesses, built for RSI
> But surely in the real world, anyone who's got a real product to make is going to want to steer what's happening.

In my own harness I log each and every user message by hook and use the model to extract user intent by re-reading the chat log from time to time. The raw messages are very important, they contain information that can be used to refine the harness on the one hand, and to validate if the agent still follows user intent on the other. Models tend to get lost in the details and forget the big picture.

visarga··on What reversing, modernising old games tells us about the economic impact of AI
> But this verification loop exists in a closed world. No equivalent exists for most programming disciplines.

You can do N-version programming with coding agents and you get tests for free. It is easier to implement something than to test the same thing. This approach turns implementing into automated oracle for testing, takes less time (can implement in parallel), and every divergence between versions is either a bug fix or a requirement clarification. In order to make the N-versions more diverse we can use different model providers, programming language or libraries.

visarga··on The problem is not AI code, but not knowing about system architecture or intent
I have been trying to define "understanding". Is it when you can predict something that you understood it? Or maybe when you can explain it? Or how about when you can control it? Or invent it.

This time I add another definition "when you can own it".

visarga··on The problem is not AI code, but not knowing about system architecture or intent
The invisible hand of the market is in control, and we don't understand it either.
visarga··on Imp is a full port of DSPy to the BEAM
We moved from structured decoding to tool calls and agents a long time ago. It was more useful when models were too weak to hold syntax reliably.
visarga··on How to keep enjoying programming in a world of LLMs
> After just a few weeks of not coding and handing everything to agents you’ll notice that you have a hard time returning to coding yourself.

The trend is in the opposite direction though, keeping your capability for hand coding is not going to help as much as you imply. How many people know how to ride a horse today, or routinely multiply large numbers by hand.

The skill we need today is to compensate for coding agent blindspots and limitations, know their problems, have ways to approach those problems and still get code you can trust to be reliable and aligned with your intent.

visarga··on We're gonna need a lot more mathematicians
I agree we can't predict, and I explain it this way - it is like a football game, where the ball will be 10 seconds later depends on what every player is doing. It is the same with AI, what we will do with it depends on what others will do with it, the reason we can't foresee it.
visarga··on Goodbye Google
> We can also choose not to work towards expanding the capacity of said symptoms.

Not if we need to eat and raise our kids. Then we need careers, which depend on what everyone else is doing, so we can't stop, our own future depends on how competitive we remain in the new context. I see an idealist streak here, some kind of naive platonism that ignores everything runs at a cost and nothing executes for free like platonic heaven. We can't take decisions that contradict our own needs.

visarga··on Contrastive Language Models
My environments don't score actions, they score state, so no matter how the agent chooses to solve a task it gets scored correctly. It's called Potential Based Reward Shaping and its main benefit is that it does not introduce reward hacking.
visarga··on Contrastive Language Models
Yes, Jev is calibrated "from factory" on a bunch of tasks, but we can be almost sure our own bespoke tasks are not covered. So the model does not really know how to produce calibrated confidence scores.

What it would need is a calibration dataset on which to align. There is no calibration in the abstract, only relative to a set of test examples. A model with an uncalibrated output probability can be recalibrated using conformal prediction. You run the model over your calibration examples, get the probabilities.

Assume the new example's answer is y, and calculate its nonconformity score, higher means a worse fit. Count how many calibration examples have a score at least as high as that. Add one to this count, then divide by the total number of calibration examples plus one.

visarga··on Contrastive Language Models
> The latter is a strategy: there usually isn’t a correct answer, now or in the future. The model is playing a game consisting of repeated rounds, and the only way to evaluate it is to see how well it plays

This is self contradictory. You can tell a correct answer as you said, by looking at the score. In my $DAYJOB I am making hundreds of RL environments that produce a score for each intermediate state.

visarga··on Jev in 25 Lines of Python
Yes, it is at https://github.com/horiacristescu/semlabel
visarga··on Jev in 25 Lines of Python
I also built one, but mine uses embeddings. It classifies concepts defined by a collection of positive and negative examples. The classifier model is trained in <1 second using ridge regression. The model itself is exactly the same shape as the embedding, so it works as a concept embedding. Since I already have a dataset, I can use it to do conformal prediction in order to calibrate confidence scores. Jev, on the other hand, has a generic model, not trained on in-domain examples, so its confidence scores are uncalibrated for any non-generic task.

So you might ask: how do I obtain the training examples? Just collect samples and use a coding agent to classify them as match or no match. From time to time you can add more examples to the dataset to have your concept adapt to changes in input distribution. It's all automated, but it only uses LLMs to train concept vectors, after that it works like a regular embedding model with a calibrated classifier on top. It's also 20-30x faster than Jev, free, and runs on CPU.

An illustration of how it defines a concept as opposed to simple cosine similarity: https://github.com/horiacristescu/semlabel/raw/main/images/c...

visarga··on I don't want to read what you didn't write
I tried a lot to get rid of Claude vibes and I failed. I made blacklists, I used prompt evolution to test against good human writing samples, even created a dynamic policy for dialogue, a cli tool used to steer agents dynamically by well chosen questions. The poor writing style made even models like Fable look worse than Sonnet 3.7 to me.
visarga··on I don't want to read what you didn't write
> So yeah, if you proofread and iterate on your AI's output until you feel you'd be proud if you had written it yourself, I'm happy to read it, too.

I agree, using LLMs is disproportionately frowned upon, but what matters is if you invested your own attention in the process. I use LLMs a lot for sparing, usually ask it to assume some opposing persona or use web search, not relying on its defaults.

visarga··on I asked Meta’s Muse for its filesystem and it sent me 6.8GB
It might not be apparent from the start what are the best demands to put inside a skill, you can only know by evals. There are whole papers dedicated to changing a few details in a coding harness. https://arxiv.org/abs/2609.20519
visarga··on MCP was always a bad idea?
What matters is local server or not. Using a remote server means exposure.
visarga··on Prompts Aren't Real
They can want anything they like, if customers want to use agents and they don't provide APIs they will lose out.
visarga··on Prompts aren’t Real
> the textual nature of prompts leads us to take the intentional stance towards systems which aren’t conscious, and thus miss the essential nature of their non-meaning

I see LLMs as being capable of making useful distinctions and having a rich action space. They are widely used because their operation is useful, and that can only happen when semantics work well in practice. But useful things that pay for themselves don't need our "essential nature" blessing, they already have persistence by mutual entanglement with us.

visarga··on Spain orders blocks on Archive.today and its mirrors
You can't host the internet on your machine, or download a "Google" or "Meta" but you can download a local model. So I don't think AI will lead to more centralization. Besides local models, the interface itself - prompt based steering - is eminently more open ended that apps or websites ever were. Cloud or local, you get more control now as a user.
visarga··on If math is more than proof, we need to better celebrate the rest of it
Not every problem will receive $20M in funding to be solved by AI; for the rest, good human guidance will have to suffice. Labs only pulled this stunt because they wanted to show investors how powerful their models are on their own. But look again at the cost of that army of 10,000 SOTA agents. At the very least, I foresee a need for humans to decide when costly AI resources should be committed to a specific search plan. Grant review remains irreducibly human because it involves choosing which directions to fund and weighing opportunity costs: taking one path forecloses others.
visarga··on How to Write with an LLM
> If you can't spend the time to write it, why should anyone read it?

I don't think writing it is the whole process, research could be 10 or 100x longer.

visarga··on HarnessTax: How Much Does the Harness Matter for Coding Agents?
A big part of that responsibility can be put in code tests.

Ensuring good test coverage and quality is how you purchase trust in the work agents do. This also reduces the context problem - a test collection has no recall issues, it just runs every time you call it, the whole battery, checks all the things we could check by code in one fast tool call. For the rest, the things we can't test by code, I use manual testing.

A large class of problems are intent divergence, when the model passes tests but it didn't do what I asked. For that I keep a log of all user messages in the project history and review it with agents. This intent alignment is repeated from time to time to catch drift.

So I see the "why should I trust the work agent did?" problem as a combination of 1. ensure good testing 2. review intent alignment.

visarga··on ImpactGate: A merge gate that scores the structural decay AI adds
I think agents are pretty capable of reading a log when given the explicit task of extracting the latest version of what the user wants. In general, they work well for direct tasks like this. They don't forget and do something else the way they do when they are deep into development work or debugging.

Besides intent, I also mine signs of "user friction," which I use as input for the agent to come up with new tests. What I complain about is one of the signals driving testing.

visarga··on OpenSpec – A lightweight and configurable AI spec framework
I just rely on a log of user messages, all messages the user typed in a project as raw data and do a pass with agents to synthesize intent. Then use this for planning and validation of code. I think the user messages are the most valuable data in a project for this reason. Doing this reflection pass on messages takes just a few minutes even for thousands of messages. It keeps global perspective which is often lost in local work.
visarga··on ImpactGate: A merge gate that scores the structural decay AI adds
You usually don't know what you want upfront, in real life it is a stream of specification and steering.
visarga··on Why I'm still bearish on LLMs after Navier-Stokes
You don't need to describe it; just show samples of the style you want to achieve. Of course it's not perfect, but it's easier than describing it.

I have my own theory about why it's impossible to remove the human from the loop:

1. Any task emerges from a need, from a human context. We need the human to pay and assume the risks and costs of using the model. So intent emerges from context.

2. While the task is being worked on, constant interaction with the context is needed, for action, for feedback, and for steering.

3. At the end of a task, consequences accumulate in the context, they don't fly to the model provider. The cost, risk, liability, gains and losses remain there.

So the LLM is great except for the start, middle and end of a task. Contexts are humans, teams, projects, and they are maximally distributed, you can't copy a context, it is indexical and relational, just as you can't copy my phone number or eat for me.

Page 1 of 34Next →