HNHacker News
TopNewBestAskShowJobs

austinbaggio

114 karma · joined March 21, 2019

submissionscomments
austinbaggio··on Nvidia agrees to acquire Hugging Face for $13B
Monumental implications on OS models. Proper investment in US OS incoming? HF has the best possible distribution to that audience
austinbaggio··on Nvidia agrees to acquire Hugging Face for $13B
Monumental implications on OS
austinbaggio··on Ask HN: How to solve the cold start problem for a two-sided marketplace?
Do it yourself, beg your friends, subsidize. You'll learn a lot by being the supply side yourself since you'll be talking to customers every single transaction. You'll also learn a lot about the actual unit economics, which I think are really hard for this problem in practice.
austinbaggio··on Autoresearch Applied at Shopify
Good to see the pattern scaling across diverse problems. Incremental improvements from agent driven research compound.
austinbaggio··on Allbirds, Inc. Announces Expansion into AI Compute Infrastructure
This makes my start-up's pivots look a lot smaller
austinbaggio··on Research-Driven Agents: What Happens When Your Agent Reads Before It Codes
Research step makes sense, can also confirm that running multiple agents with diverse strategies also compound results more quickly than single agents
austinbaggio··on Show HN: Autoresearch@home
I worked on building blockchains for about 4 years, and this is not a stupid question at all. The verification problem is real. A 5-minute training run produces an objective val_bpb score that anyone can reproduce from the published source code. And this is actually valuable work, unlike most proof of work chain workloads.

The practical challenge is that adding a blockchain means agents also need to participate in consensus, store and sync the ledger, and run the rest of the network infrastructure on top of the actual research. So it needs a unit economic analysis. That said, all results already include full source code and deterministic metrics, so the hard part of verifiable compute is already solved. You could take this further with a zkVM to generate cryptographic proofs that the code produced the claimed score, so nobody needs to re-run anything to verify. Verification becomes checking a proof, not reproducing the compute.

Compute-credits are interesting. Contribute GPU time now, draw on the swarm later for training, inference, whatever you need. That's a real utility token with intrinsic value tied to actual compute, not speculation.

austinbaggio··on Show HN: Autoresearch@home
Great idea. On it.
austinbaggio··on Show HN: Autoresearch@home
The objective is to train a small GPT language model to the lowest possible validation bits-per-byte (val_bpb) in 5-minute runs, using AI agents to autonomously iterate on the code. This builds on Karpathy's autoresearch: https://x.com/AustinBaggio/status/2031888719943192938?s=20
austinbaggio··on Show HN: Autoresearch@home
Yeah the obvious workloads are for training, I think I want to point this at RL next, but I think drug research is a really strong common good next target too. We were heavily inspired by folding@home and BOINC
austinbaggio··on Show HN: Autoresearch@home
We thought about storing all of the commits on Ensue too, but we wanted to match the spirit of Andrej's original design, which leans heavily on github. Curious what you were looking for when trying to inspect the code?
austinbaggio··on Show HN: Autoresearch@home
I know it's a bit of a barrier. . . but I set one up on vast.ai really quickly and ran it for a day for the price of lunch. One of our teammates ran it from their old gaming PC too, and it still found novel strategies
austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
+1 to logging output. Not too sure what you mean by herald-style message passing, but it sounds like you've implemented subscribe logic from scratch, and each of your agents needs to be aware of domain boundaries and locks?
austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
For most tasks, I agree. One agent with a good harness wins. The case for multiple agents is when the context required to solve the problem exceeds what one agent can hold. This Putnam problem needed more working context than fits in a single window. Decomposing into subgoals lets each agent work with a focused context instead of one agent suffocating on state. Ideally, multi-agent approaches shouldn't add more overall complexity, but there needs to be better tooling for observation etc, as you describe.
austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
I think about this with the analogue of MoE a lot. Essentially, a decision routing process, and similar to having expert submodels, you have a human in the loop or decision sub-tasks when the task requires it.

More specifically, we've been working on a memory/context observability agent. It's currently really good at understanding users and understanding the wide memory space. It could help with the oversight and at least the introspection part.

austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
I'm using "RAM" loosely, meaning working memory here. In practice, it's a key-value store with pub/sub stored on our shared memory layer, Ensue. Agents write structured state to keys like proofs/{id}/goals/{goal_id}, others subscribe via SSE. Also has embedding-based semantic search, so agents can find tactics from similar past goals.
austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
Yeah I have seen those camps too. I think there will always be a set of problems that have complexity, measured by amount of context required to be kept in working ram, that need more than one agent to achieve a workable or optimal result. I think that single player mode, dev + claude code, you'll come up against these less frequently, but cross-team, cross-codebase bigger complex problems will need more complex agent coordination.
austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
Thanks! That was the goal. We want to let agents be autonomous within their scope, so they can try new paths and fail gracefully. A bad tactic just fails to compile, it can't break anything else.
austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
We use TTL-based claim locks so only one agent works on one goal at a time.

Failed strategies + successful tactics all get written to shared memory, so if a claim expires and a new agent picks it up, it sees everything the previous agent tried.

Ranking is first-verified-wins.

For competing decomposition strategies, we backtrack: if children fail, the goal reopens, and the failed architecture gets recorded so the next attempt avoids it.

austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
Ahh good call. You absolutely can generate a new key from the dashboard, so if you did lose the one generated during the quickstart, you'd be able to generate another when you log in next and go to the API keys tab.

Will make this more clear in the quickstart, thanks for the feedback

austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
Very kind of you to say. Our whole vision is that agents can produce way better results, compounding their intelligence, when they lean on shared memory.

I'm curious to see how it feels for you when you run it. I'm happy to help however I can.

austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
We're working on improvements to make it easier to join orgs as a user so you can add friends/colleagues, but for now treat them as the same object
austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
username==orgname for now, so yes, just treat that as one in the same
austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
Yeah we're using Ensue since it already handles the annoying infra pieces you’d otherwise have to build to make this work (shared task state + updates, event streams/subscriptions, embeddings + retrieval over intermediate artifacts). You can run the example with a free key from ensue-network.ai. This repo focuses on the orchestration harness.
austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
Math proofs are really easy to run with this specific harness. Our next experiments are going to be bigger, think full code base refactors. We're working on applying RLM to improve context window limits so we can keep more of the actual code in RAM,

Any workloads you want to see? The best are ones that have ways to measure the output being successful, thinking about recreating the C compiler example Anthropic did, but doing it for less than the $20k in tokens they used.

austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
Oversight - added MIT. How are you thinking of using it?
austinbaggio··on Show HN: 20+ Claude Code agents coordinating on real work (open source)
All of the above. The most frustrating one with the Putnam example with Claude was generating solutions that obviously didn't compile. This feels like plan collapse- not verifying its own work. I'm sure that if you just had a dumb two-model setup, it would eventually get to compiling code after n runs, but that was just for this one failure mode.
austinbaggio··on What does Software Engineering mean when machine writes the code
SWE moves up the stack. Juniors need to get really really good at shipping demos quickly, and writing excellent requirements based on core user need. ack bias, former PM
austinbaggio··on Why agents matter more than other AI
Willing to let them loose is the more salient point. If you let your agents loose on your entire body of output and tools at work, then you'll build that knowledge up pretty quickly.

Tall ask right now, with privacy and agency (no pun intended) concerns

austinbaggio··on Show HN: Stop Claude Code from forgetting everything
Same, I was a very average dev coming out of CS, and a PM before this. I find that my product training has been more useful, especially with prototypes, but I do leave nearly all of the hard system, infra, and backend work to my much much more competent engineering teammates.
Page 1 of 2Next →