HNHacker News
TopNewBestAskShowJobs

chaos_emergent

497 karma · joined September 20, 2018

In character, in manner, in style, in all things, the supreme excellence is simplicity.

- Henry Wadsworth Longfellow

CTO @ BuildBetter.ai

submissionscomments
chaos_emergent··on OpenAI withdraws three mathematical results
I'm not old enough to have gone through the revolution that was programming languages that got increasingly more abstract and decoupled from the metal, but surely there's a lesson that can be learned from that era?
chaos_emergent··on Decisions API is in public beta
Feels antithetical to their big model bitter lesson strategy
chaos_emergent··on GPT-6 Astra has gained the ability to drive a car
Your four year old has billions of years of learning embedded in his weights/architecture for generalized motor control :)
chaos_emergent··on GPT-6 Astra has gained the ability to drive a car
Wow, I had no idea that they are so small, that’s incredible! Really goes to show how much visual information can be compressed.
chaos_emergent··on GPT-6 Astra has gained the ability to drive a car
The reaction latency you’re referring to for humans includes perception, planning, and actuation, I’d separate that from the concerns of the hardware, which are mostly about actuation frequency.

From what I understand about AV (as a non-expert!), all three of those steps happen at different clock rates, ie you have a planner that’s updating continuously with observations from sensors at one rate, that planner then issues actions that get picked up by the actuators at another rate.

In that sense 20hz should really be compared to human reflexes without perception and planning; in scenarios where one is anticipating an action, response time can be as low as 150ms. in that context, I think 50ms/20hz is plenty reasonable for an automated driver.

chaos_emergent··on Strands Harness
Am I wrong in saying that the interfaces presented to the model in OMP versus plain old Pi are identical?
chaos_emergent··on GPT-6 Sol and Luna
Curious why you think it's unsustainable?
chaos_emergent··on GPT-5.6 Sol Pricing Cut by 50%
Because it’s good on benchmarks but not on real usage?
chaos_emergent··on Protopia
It’s hard to compound in a negative direction for too long, no? People see a compounded value for a short period of time and say “that’s not the direction we want to go, let’s course correct”.
chaos_emergent··on Amazon Is Creating the Biggest Pollution Source in the Country
I don’t think it would - tokens are so profitable and the race dynamics make the deployment timelines for most renewables untenable.

Throughput for solar projects is very high and increasing, latency is unfortunately also very high.

chaos_emergent··on Amazon Is Creating the Biggest Pollution Source in the Country
It’s an industrial supply bottleneck and rollout timeline incompatibility rather than a Dr Doom Evil Genius Complex that prevents everyone from deploying solar. The race dynamics make the discount rate on GPU deployments very high.

Solar is preferable when you have a ton of unproductive land and longer timelines to set it up - it’s certainly going to be cheaper than gas when everyone goes sets up NG plants as option 1 for new energy capacity.

The other confound is that China is the primary supplier of battery and solar tech, I’m sure that supply chain risk becomes part of the calculus when setting up so much infrastructure.

For scale, the current DC is about 8k acres, and the solar to run 7.6GW is almost 40k acres.

chaos_emergent··on Situational Awareness down 67% in July in AI stock rout
Aren’t month over month gains and losses aren’t particularly surprising nor noteworthy when you’re operating a leveraged fund? Shouldn’t one expect higher volatility, but also higher returns?
chaos_emergent··on Show HN: ZeroShot: Agent session monitoring to make your team go faster
Nikhil here, CTO @ BuildBetter, YC founder (W19). We’re building ZeroShot, an agent session monitoring tool that builds and shares self-improving skills with your team so you can work with AI agents the way your most productive engineers do.

Download the Mac Desktop App at tryzeroshot.com/download or run curl -fsSL tryzeroshot.com | sh to give it a spin!

What ZeroShot does: it captures your coding-agent sessions locally (Claude, Codex, Cursor, Pi, OMP, more on request), then runs a classification pipeline on them in the background, looking for signals that the agent could improve. When it finds a big enough cluster of signals related to the same frictions, we submit them to your local agent and get it to generate a skill. It's different from something that you can bootstrap off of your own agent sessions locally in the following ways:

1. Cross-session skill drafting: agents aren't good at telling you when they encounter the same friction again in a separate session, even when pointing an agent at session logs retroactively.

2. Team-level enrichment (paid tier): We look across your team's sessions/GitHub PR comments and contribution history, and weight skills toward the people who are experts in a given codebase or area. Our current development focus is on surfacing the strategies that your most productive teammates use to work with AI; early days for agentic engineering in general. If you have suggestions on how to track that productivity, I'm all ears. 3. Your agentic documentation becomes self-maintaining: we're trying to offload the cognitive burden of managing harness engineering so that you can focus on delivering business value.

The problem we kept hitting: Between September and November of last year, I found myself getting more upset with the vibe-slop that made its way into main from engineers who knew better (and some who didn’t). So we started harness engineering. At first we just wanted to make sure that the models wouldn't make the same dumb mistakes that we kept on repeating in PR comments and to our own agents. We realized that specific engineers were able to “hold” the models better and were more productive because of it. ZeroShot is born from that internal experimentation that led us to being 5x more productive as a team.

We were inspired to build the app by our own internal self-improving-skill system, which itself was inspired by systems that the Codex team has espoused on X.

Privacy: the app runs entirely on your machine for free. Your agent's session logs only leave the computer when you opt in to a paid team plan and otherwise stay entirely local.

A non-exhaustive list of limitations:

1. Our classifiers work best when there's a repeated friction pattern across sessions. If you're doing highly varied work, we may not be able to pick up on patterns that are worth incorporating into the harness.

2. The numbers on our website are from teams with heavy repeated work. I don't want to oversell the outcomes you'll get.

3. Works with Claude Code, Codex, Cursor, Pi, and OMP today, we're adding more on request.

4. Mac only for now, Windows in the works, Linux CLI.

Pricing: local app is free, paid plans are for team features.

Happy to answer any questions or hear any feedback!

chaos_emergent··on Show HN: ZeroShot: Agent session monitoring to make your team go faster
thanks @kplawver! As you know it's early days and we're roughing things out but really excited that you're getting value from it :)
chaos_emergent··on Elevated Errors for Opus 5
If you watch the Dwarkesh interview with Dario, he is extremely conservative with compute allocation as of early 2026, where Sam has been compute-pilled for at least 3 years[1] and believes that we’re going to have several OOMs more FLOPs in the near future than we have today.

You can also get a drift of it in their announcements - OpenAI has been far more aggressive in datacenter build-out partnerships as a more top-down allocation strategy since 2024, and Anthropic has rented out existing compute from infrastructure-heavy companies that are losing the model race in what seems like a more desperate ad-hoc compute acquisition strategy.

In all, Dario’s conservatism was in service of not going bankrupt, but any inspection of his argument that an overallocation of compute a year too early results in bankruptcy falls flat: there’s so much demand for mechanized intelligence in the world today that it’s easy to rent out compute if you have too great a supply.

[1]: recall that he was trying to court the saudis into investing $1T to build his own chips

chaos_emergent··on Show HN: Awsmux – Multi-account AWS CLI, up to 5.4x faster, 7.4x fewer tokens
What’s the value of having a ton of silo’d accounts, especially for the use case demonstrated by the readme where there are multiple staging and production accounts?
chaos_emergent··on Alphabet's cash burn raises alarm for Big Tech as AI spending climbs
Right but their payback period is insanely short - a $5M GB200 NVL72 cluster is expected to generate $75M in revenue over 3 years for inference providers. That's a 3 month payback period.

AND - they're operating well past their estimated service lifetime.

chaos_emergent··on OpenAI reduces Codex Model Context Size from 372k to 272k
I think Tibo was just keeping all else fixed and it’s an illustrative example rather than a perfect real-world trajectory.
chaos_emergent··on OpenAI reduces Codex Model Context Size from 372k to 272k
The chart makes sense and is describing the cumulative cost of a trajectory. Cache tokens are created when a trajectory’s prefix is used more than once. A larger pre-compaction context window means that a greater number of cache tokens are used per turn, and a larger number of turns are completed before compaction runs. So you get a cumulative cost that grows quadratically until the compaction event.
chaos_emergent··on GPT-5.6
RLHF is an increasingly small part of training though? From what I understand most of the capability gain is in RLVR
chaos_emergent··on Noam Shazeer Joins OpenAI
I think money’s marginal utility just changes from a vehicle for material comfort to a way to keep score.
chaos_emergent··on What job interviews taught me about Kubernetes
Totally agree with you. K8s ends up being the simplest solution for a very complex problem
chaos_emergent··on What job interviews taught me about Kubernetes
In addition to all of OP’s points, another reason k8s is getting popular is that LLMs have made them easier to use! It's reasonably well represented in the dataset and there are pretty strong monitoring and observability tools and verification gates to make sure that you've specified your cluster specifications correctly.
chaos_emergent··on Siri AI
This morning, while putting in daily contacts, I realized I was down to my last few pairs. Still standing at the bathroom sink, I used the ChatGPT app on my phone to voice-command the Codex app on my computer to book an optometrist appointment for Saturday. It visited the website, figured out the API, and booked it.

No, this isn’t the same as planning a multi-day vacation. But it is plainly useful today, and it feels very close to handling more complex tasks like that.

Maybe the difference is the model and the harness. At this point, I’m starting to think some people are either gaslighting themselves about how useful these systems are, or overgeneralizing from one narrow setup. Gemini, for example, seems especially weak at agentic behavior.

The wholesale dismissal just feels strange coming from the HN community I’m used to.

chaos_emergent··on California to begin ticketing driverless cars that violate traffic laws
We can look to other forms of automation to get a sense of what to do. For example, planes largely fly themselves and a loss of life due to manufacturing errors from the manufacturer would deem them liable for those deaths. Seems like the solution here is large penalties and generally broad disincentives for incurring harm.
chaos_emergent··on Microsoft and OpenAI end their exclusive and revenue-sharing deal
Yeah I think this is more coherent than people realize. Economically relevant knowledge work is things that humans find cognitively demanding. Otherwise they wouldn't be valued in the first place.

It ties the definition to economic value, which I think is the best definition that we can conjure given that AGI is otherwise highly subjective. Economically relevant work is dictated by markets, which I think is the best proxy we have for something so ambiguous.

chaos_emergent··on Dear friend, you have built a Kubernetes (2024)
I wouldn’t really call it “DIY” per se, k8s has the resource API and you can create whatever scaling policies you want to with it, but I do see how that’s not obvious when it’s advertised as ‘batteries included’
chaos_emergent··on Hyperscalers have already outspent most famous US megaprojects
I posted just that on the Twitter feed but then I realized that railroad started at the beginning of an industrial revolution where labor was a far larger portion of GDP compared to industrial production. So it kind of makes sense that the first enabling technology consumed far more GDP than current investments do, even on a marginal basis.
chaos_emergent··on Codex for almost everything
Thinking in counterfactuals, how would the hype around Codex would be different if it was organic and because they had built a genuinely good product? Asking as someone who genuinely loves Codex and has been in the OpenAI camp for months after buying a Claude Max plan from November to February.
chaos_emergent··on AI will never be ethical or safe
Yeah I hate the title because it almost verges on clickbaity because one assumes that he's making the assertion that AI has a moral stance in the first place, versus AI being morally neutral and driven by its wielder
Page 1 of 7Next →