HNHacker News
TopNewBestAskShowJobs

aryamanagraw

7 karma · joined July 2, 2025

submissionscomments
aryamanagraw··on Show HN: Hearth – a shared family workspace where an agent can build apps
This looks cool! wondering how this can connect to the different kinds of devices in my home? and what does it look like interfacing with with Hearth for different member's of the family?
aryamanagraw··on Why everything breaks at 150 people
Interesting point. Just read about Coase's theorem of the firm. Isn't it slightly narrower than the ceiling metaphor? a firm expands "until the costs of organizing an extra transaction within the firm become equal to the costs of carrying out the same transaction by means of an exchange on the open market." That's a price related argument. The Dunbar limit doesn't.

This is where I would push, if Dunbar's number is a manifestation of the same thing, shouldn't it move when the price moves? Coase's ceiling has moved over time, while Dunbar's limit isn't. What's happening there?

The intersection in this case comes from source of truth information getting out of date with moving code. You can't tell whether its wrong until its been acted upon. So people stop trusting the docs, which is where Dunbar's number comes up again

aryamanagraw··on LLM-as-a-Courtroom
How so? Care to elaborate?
aryamanagraw··on LLM-as-a-Courtroom
That's really cool! That's actually the standpoint we started with. We asked what a collaborative reconciliation of document updates looks like. However, the LLMs seemed to get `swayed` or showed `bias` very easily. This brought up the point about an adversarial element. Even then, context engineering is your best friend.

You kind of have to fine-tune what the objectives are for each persona and how much context they are entitled to, that would ensure an objective court proceeding that has debates in both directions carry equal weight!

I love your point about incentivization. That seems to be a make-or-break element for a reasoning framework such as this.

aryamanagraw··on LLM-as-a-Courtroom
That's the thing with documentation; there are hardly any situations where a simple true/false works. Product decisions have many caveats and evolving behaviors coming from different people. At that point, a numerical grading format isn't something we even want — we want reasoning, not ratings.
aryamanagraw··on LLM-as-a-Courtroom
The funnel is the answer to this. We're not running four agents on every PR — 65% are filtered before review even begins, and 95% of flagged PRs never reach the courtroom. This is because we do think there's some value in a single agent's judgment, and the prosecutor gets to make a choice when to file charges vs not.

Only ~1-2% of PRs trigger the full adversarial pipeline. The courtroom is the expensive last mile, deliberately reserved for ambiguous cases where the cost of being wrong far exceeds the cost of a few extra inference calls. Plus you can make token/model-based optimizations for the extra calls in the argumentation system.

aryamanagraw··on LLM-as-a-Courtroom
We kept asking LLMs to rate things on 1-10 scales and getting inconsistent results. Turns out they're much better at arguing positions than assigning numbers— which makes sense given their training data. The courtroom structure (prosecution, defense, jury, judge) gave us adversarial checks we couldn't get from a single prompt. Curious if anyone has experimented with other domain-specific frameworks to scaffold LLM reasoning.