497 karma · joined September 20, 2018
- Henry Wadsworth Longfellow
CTO @ BuildBetter.ai
From what I understand about AV (as a non-expert!), all three of those steps happen at different clock rates, ie you have a planner that’s updating continuously with observations from sensors at one rate, that planner then issues actions that get picked up by the actuators at another rate.
In that sense 20hz should really be compared to human reflexes without perception and planning; in scenarios where one is anticipating an action, response time can be as low as 150ms. in that context, I think 50ms/20hz is plenty reasonable for an automated driver.
Throughput for solar projects is very high and increasing, latency is unfortunately also very high.
Solar is preferable when you have a ton of unproductive land and longer timelines to set it up - it’s certainly going to be cheaper than gas when everyone goes sets up NG plants as option 1 for new energy capacity.
The other confound is that China is the primary supplier of battery and solar tech, I’m sure that supply chain risk becomes part of the calculus when setting up so much infrastructure.
For scale, the current DC is about 8k acres, and the solar to run 7.6GW is almost 40k acres.
Download the Mac Desktop App at tryzeroshot.com/download or run curl -fsSL tryzeroshot.com | sh to give it a spin!
What ZeroShot does: it captures your coding-agent sessions locally (Claude, Codex, Cursor, Pi, OMP, more on request), then runs a classification pipeline on them in the background, looking for signals that the agent could improve. When it finds a big enough cluster of signals related to the same frictions, we submit them to your local agent and get it to generate a skill. It's different from something that you can bootstrap off of your own agent sessions locally in the following ways:
1. Cross-session skill drafting: agents aren't good at telling you when they encounter the same friction again in a separate session, even when pointing an agent at session logs retroactively.
2. Team-level enrichment (paid tier): We look across your team's sessions/GitHub PR comments and contribution history, and weight skills toward the people who are experts in a given codebase or area. Our current development focus is on surfacing the strategies that your most productive teammates use to work with AI; early days for agentic engineering in general. If you have suggestions on how to track that productivity, I'm all ears. 3. Your agentic documentation becomes self-maintaining: we're trying to offload the cognitive burden of managing harness engineering so that you can focus on delivering business value.
The problem we kept hitting: Between September and November of last year, I found myself getting more upset with the vibe-slop that made its way into main from engineers who knew better (and some who didn’t). So we started harness engineering. At first we just wanted to make sure that the models wouldn't make the same dumb mistakes that we kept on repeating in PR comments and to our own agents. We realized that specific engineers were able to “hold” the models better and were more productive because of it. ZeroShot is born from that internal experimentation that led us to being 5x more productive as a team.
We were inspired to build the app by our own internal self-improving-skill system, which itself was inspired by systems that the Codex team has espoused on X.
Privacy: the app runs entirely on your machine for free. Your agent's session logs only leave the computer when you opt in to a paid team plan and otherwise stay entirely local.
A non-exhaustive list of limitations:
1. Our classifiers work best when there's a repeated friction pattern across sessions. If you're doing highly varied work, we may not be able to pick up on patterns that are worth incorporating into the harness.
2. The numbers on our website are from teams with heavy repeated work. I don't want to oversell the outcomes you'll get.
3. Works with Claude Code, Codex, Cursor, Pi, and OMP today, we're adding more on request.
4. Mac only for now, Windows in the works, Linux CLI.
Pricing: local app is free, paid plans are for team features.
Happy to answer any questions or hear any feedback!
You can also get a drift of it in their announcements - OpenAI has been far more aggressive in datacenter build-out partnerships as a more top-down allocation strategy since 2024, and Anthropic has rented out existing compute from infrastructure-heavy companies that are losing the model race in what seems like a more desperate ad-hoc compute acquisition strategy.
In all, Dario’s conservatism was in service of not going bankrupt, but any inspection of his argument that an overallocation of compute a year too early results in bankruptcy falls flat: there’s so much demand for mechanized intelligence in the world today that it’s easy to rent out compute if you have too great a supply.
[1]: recall that he was trying to court the saudis into investing $1T to build his own chips
AND - they're operating well past their estimated service lifetime.
No, this isn’t the same as planning a multi-day vacation. But it is plainly useful today, and it feels very close to handling more complex tasks like that.
Maybe the difference is the model and the harness. At this point, I’m starting to think some people are either gaslighting themselves about how useful these systems are, or overgeneralizing from one narrow setup. Gemini, for example, seems especially weak at agentic behavior.
The wholesale dismissal just feels strange coming from the HN community I’m used to.
It ties the definition to economic value, which I think is the best definition that we can conjure given that AGI is otherwise highly subjective. Economically relevant work is dictated by markets, which I think is the best proxy we have for something so ambiguous.