HNHacker News
TopNewBestAskShowJobs

CharlieDigital

6,567 karma · joined July 2, 2021

If you are an experienced C# dev and would like to work at a well-funded, profitable, series-C, former YC startup building green-field products, check out: https://www.usemotion.com/careers and apply directly

In my free time, I'm making this [https://turas.app] and this [https://coderev.app]

More: [https://chrlschn.dev] [https://github.com/CharlieDigital] [https://charliedigital.com] [https://www.linkedin.com/in/charlescchen] [https://www.youtube.com/@chrlschn]

Let's connect!

submissionscomments
CharlieDigital··on A Million Agents Is a Distributed System Problem
Interesting.

I wanted to test the veracity of this result so I tested some of my own 100% hand written pieces and what do you know? It came back 100% hand written.

This is a better ad for Pangram than for OP :D

CharlieDigital··on Jev is 13.6x faster, 2.7x cheaper than GPT Luna 6
- This is only true if your constraint is online

- If it is offline processing (e.g. classifying text snippets stored in a DB (from call transcript, from logs)), you can batch to Luna

- If you batch to Luna and then divide both cost and time by the batch size, you will find that Luna beats Jev.

- Now tune your batch size for your test dataset and see where the batch size causes accuracy falloff.

Tested this approach with GPT-6 Luna. It was 1.6x cheaper, 1.2x slower, within margin of error performance vs Jev with batch size 20 using a CFPB complaint dataset. Tuning batch size up to even double would likely yield similar accuracy while reducing both the per-record run time as well as per-record token cost.

The methodology of comparing single record is only valid for the on-line, real-time use case. For every other case, Luna can match or beat Jev by simply batching. If you find no dropoff at larger batch sizes, you will be significantly cheaper than Jev.

* Batching here does not mean the native batch API but actually placing 20 records (batch size) into one prompt and getting 20 results back in one response.

CharlieDigital··on Strands Harness

    > The model matters more than the harness anyway
    > 
    > Everyone is benchmaxxing
    > 
    > ...harnesses tend to be chosen on voodoo and hunches...
I get what you're saying, but their graphic on performance here uses the exact same model with different harnesses and definitively shows that there is a significant difference in both accuracy and cost. The whole point of their technical implementation and design decision here is to highlight that it's not "voodoo and hunches", but observable data.

Fable 5 on Claude Code scored 61.8% at a cost of $248.05 while Fable 5 on OpenCode beat it at 66.3% at $73.42. The same model, the same benchmark; only the harness is different with a ~5 point difference in accuracy while costing significantly less. So if we are to believe the author and these results are repeatable, then it would seem that the harness matters.

The point of this framing here is specifically to address 1) benchmaxxing by using the same model, 2) NOT choose a harness on "voodoo and hunches" by using actual data to back the assertions. Your comment feels misguided and completely hand waves the actual data points here.

CharlieDigital··on Strands Harness
It won't be Pi because there's really no singular Pi that is broadly useful without explicit configuration of plugins. Pi requires plugins to do lots of things that are OOB on other harnesses that people care about. MCP, OpenTelemetry, etc. It may be some offshoot or something built on top of Pi that is more standardized, but it won't be Pi.
CharlieDigital··on 28% of job postings on company career sites have been open over 90 days
Company's own website job listings are likely actually in an applicant tracking system (ATS) like Greenhouse or Ashby because they need to manage the pipeline, not just list the job.
CharlieDigital··on 28% of job postings on company career sites have been open over 90 days
You pay for job listings so it would be silly to not take them down.
CharlieDigital··on GPT-6 Sol and Luna
I can already see it. 7-Nebula, 8-Galactic, 9-Cosmos; The size inflation is real.
CharlieDigital··on OpenAI is well positioned to fast-follow Jev
Why would it need to? There is a deterministic flow here for the actions that are allowed. AGI isn't needed for this at all if you can map out the flow and use a classifier to decide which route to follow.
CharlieDigital··on OpenAI is about to eat Jev's lunch – Arcturus Labs

    > so why invest into a more complex solutions
Not sure what's more complex about one REST API call versus another REST API call...
CharlieDigital··on OpenAI is well positioned to fast-follow Jev
It's just another tool. Luna exists for a reason: it's the right tool for the job. If they release AGI and it costs $1 and 5 seconds to decide "is the customer asking for a refund", then that's a terrible use case for AGI if another tool can do it with 95% accuracy for $0.002 and 50ms.
CharlieDigital··on AI Has No Wisdom and Neither Will You

    > "the scaffolding produced by the team" also includes the scaffolding produced by past AIs.
This statement is also true of humans. Everything you've stated here is also true for human engineers.

    > "the scaffolding produced by the team" also includes the scaffolding produced by past engineers.
But the agent can be instructed reliably to keep documentation accurate and up to date and will then do so dutifully. Put it in AGENTS.md that it must always update the /docs directory by creating a new doc or updating an existing doc and it will do it. (Yes, adherence may be 95% of the time, but that is likely several points higher than with most non-NASA human teams)

Better yet, extract docs from code comments. Even better when the docs are spatially co-located and line of sight as the agent crawls through code.

CharlieDigital··on AI Has No Wisdom and Neither Will You

    > After a certain point, people would be forced to refactor, because they find themselves unable to handle the complexity.
This is a fallacy; this is why legacy code exists that teams just work around. They lack the tests to verify it, the person that wrote it is long gone, it's handling some mission critical dataflow so no one touches the code and just builds around it.
CharlieDigital··on AI Has No Wisdom and Neither Will You
AI is not better nor worse at producing code rot; just faster at it.

AI produced code is a function of the team driving and instructing the agents along with the scaffolding produced by the team (skills, examples, docs, comments); same with human teams.

A team that cannot guide a human team to produce better code will not be able to guide an AI team to produce better code because it's the same skillset: being able to write good docs, create constraints structurally in code, produce core architecture that enforces good behavior.

CharlieDigital··on AI Has No Wisdom and Neither Will You

    > aka it is the matter of long term memory that currently LLM architecture is not capable of
Long term memory is easier than you think when you consider what an agent has to do when it is reading and editing code: instruct the agent to leave comments on its rationale and reasoning directly in the code. This is infrastructure free memory that every agent that then sees the code will read. Your code review agent will see the reasoning and decision making your coding agent formulated. When an agent comes and refactors this code in 6 months, the comments will be there (and it will update it!). When an agent is trying to troubleshoot an issue, it will read the comment. No infrastructure needed! Don't overthink it; use comments.

Code comments are line-of-sight for agents and one of the cheapest, highest leverage ways to get better coding performance from AI because unlike skills that may or may not activate, comments end up in context as long as they are well placed and carry the right instructions.

Best places to have it leave comments: 1) start of the file because it frequently uses `sed -n 1,200p` to read files and 2) inside the body of the method because it may find by keyword and read a few lines past. If your harness is set up with an LSP, language standard comments are also useful because then it can read comments on the member.

Tips for comments: point it to other, related members or artifacts; point it to external canonical docs; point is to a specific issue number or PR; have examples directly in the comment using your language's example markers; point it to example, reference usages in code. Use AGENTS.md to tell your agents how you want it to leave comments and to specifically read, follow, and maintain comments.

You don't need infrastructure or special architecture; Every coding agent is text-in, text-out. You need comments that get carried with text-in and a bit of guidance to the agent on how to use comments effectively.

CharlieDigital··on AI Has No Wisdom and Neither Will You
AI written code is a function of the human created constraints around it.

That is why I believe the most senior engineers on the team with the most scars and most experience need to shift into writing those constraints instead of writing code.

In writing those constraints, they can multiply their effect across a tireless fleet of agents that generally want to copy existing patterns and can be guided to use skills.

CharlieDigital··on AI Has No Wisdom and Neither Will You

    > because you have time to reflect
It doesn't mean that people do. This is a false narrative we tell ourselves. Yes, there are craft-oriented devs and teams, but these are the exception rather than the rule because in the end, it is the GTM and business teams that define what, when, how and rarely the engineering teams.

There is no team without tech debt because there is no "golden" project where every decision has been made right because of reflection on decisions made wrong.

CharlieDigital··on AI Has No Wisdom and Neither Will You

    > Fact is, vibe-coded projects devolve over time into an unmaintainable mess. The reason is simple, yet hard to fix: code maintainability and good architecture don’t have good measurements that we can apply, because it takes months, years even, to notice the effects of bad architecture or of unmaintainable code.
    > 
    > For one, AI is not trained on what it means for code to be maintainable. For instance, any reinforcement learning done needs a reward signal that can be measured immediately, not in months or years.
Sad to say, but this is no different from human written code. Human written code just takes even longer to realize the mistakes because the pace is slower.

I think at the end of the day, it is not impossible to have AI write "good" or "high quality" code. If anything, once the patterns are established, AI will be more likely to adhere to the patterns and rules than any human team. It requires the most experienced engineers on the team to split their time writing the core patterns and documenting them in references/skills.

But it takes a lot of "taste" and a willingness to slow down a bit with AI (to create necessary artifacts), something teams find hard to do when you can ship so fast now.

My experience has been that there is a camp of very senior engineers that are unwilling to adapt to reality and focus on documentation and writing (effectively producing skills and agent guidance which multiplies their effectiveness); they will cling to their knowledge thinking coding a sacred art.

CharlieDigital··on AI coding has made CI a bottleneck, so we reworked ours to keep up

    > If there is one universal amongst organizations, its that they love walls and silos.
What's true for human organizations isn't necessarily true for agent-driven engineering. People and human teams struggle with contracts because there's always human negotiation involved. If the decisions are instead made by a team of agents, there's no more ego, ownership, miscommunications; just decisions based on whatever rules have been given to the orchestrator.

    > Your agents may be happy in their tiny walled kingdom...
Yes indeed; the agents will always be happier if they can iterate faster, lint faster, build faster, test faster, ship faster, with smaller context.

That would be the point of using contracts as boundaries so the agent can iterate more autonomously so long as it maintains the externally facing contract or version the contract if it needs to.

CharlieDigital··on AI coding has made CI a bottleneck, so we reworked ours to keep up
My take: it seems like systems should become smaller, more isolated, and contract-oriented.

I have been a long time proponent of monoliths, but it seems like agents would be happier with smaller, more isolated services. The more isolated, the better. Contracts between the service components only. Then it can iterate internally as long as it satisfies the contract. If it needs to, it can version the contract and keep iterating.

CharlieDigital··on AI coding has made CI a bottleneck, so we reworked ours to keep up
Question here:

    > where sub agent orchestration is done through agent to agent messaging
How do you expect to see the history/record of what the agents did and why? Is it enough to see it in PRs? Do you expect tickets that have the design and history? How are you thinking of agents being able to historically resolve reasoning/why/decisions made in earlier passes?

Genuine open question here. My assumption is that a GH or Linear or Jira is still useful as a decision store. It may as well be a custom app over Postgres, but it seems like something is needed to store this and for observability. A GH/Linear/Jira is nice if only because of standard APIs and integration points (whatever you build would likely end up duplicating a subset of those).

CharlieDigital··on Show HN: Lossless-memory – a personal AI memory that never summarizes
The point is to build systems that don't require the agent to "ask".

The human becomes the bottleneck as systems become more agentic.

OP's system is first ingesting and storing the individual messages and then recompiling it into "working memory".

CharlieDigital··on Show HN: Lossless-memory – a personal AI memory that never summarizes
The specific design of this system uses the raw memories and allows rebuilding the operational memory from the raw memory (the LLL "Left Leg Layer").

To do this requires that there is an ordering of which memory came last.

    "always do X before committing Y"
    "always do X before committing Y except after Z"
Which of these is the current state? Without the date, it is not possible to rebuild the operational state of the rule from the raw records.
CharlieDigital··on MCP was always a bad idea?
I also take a lot of the guidance from influencers with a grain of salt for that reason. What works for the solo dev on greenfield projects rarely scales to the enterprise 1:1. Brownfield. Security policies. Enterprise controls. Legacy systems.

There is usually a kernel of value that has to be extracted and translated to apply what works for the solo AI engineer and what works for the team.

CharlieDigital··on MCP was always a bad idea?
I think many devs misunderstand MCP because they work in solo mode. In solo mode, you just have your secrets local. You don't care about auditing access. There's no IT managing the infra and third party secrets. You don't have to account for different harnesses and tooling; you just use your own harness and adapt your tooling to it. You're not thinking about revoking/rotating secrets when someone leaves your team. You don't have to deliver capabilities to many different runtimes and stacks.

Back in March, everyone was already pronouncing it dead[0] when in fact, it has only proliferated and become even more essential for both 3rd party systems as well as platform level capabilities[1] as agentic tooling has moved a bit more slowly into the enterprise. It was apparent even back then that enterprises will need MCP.

In a team context? Enterprise? Building web server backed or in-process agents? Not sure how you replicate the control, auditability, accessibility, composability, and security boundary that you get with MCP over HTTPS without a lot of bespoke, point solutions; in the end, protocols almost always win.

Could you do it with just REST APIs? I mean, MCP is just JSON-RPC over HTTP with standardized auth, schemas, and agent specific exchange flows (tools, prompts, resources, etc.). Could you do it with just CLIs? You lose a lot of the control mechanisms offered by MCP (auditing, security, centralized auth, etc.). The context savings are overblown except with CLIs that have good representation in training (curl, jq, cat, sed, etc.; your custom CLI is going to need to produce instructions and add to context all the same)

Heuristic is simple: don't use MCP for local, solo dev. As soon as you need MCP, you'll know it and you'll understand why it exists. Many of the harness level capabilities themselves are implemented as first party MCP (more apparent in the CLIs). MCP is basically REST for the agentic era.

[0] https://chrlschn.dev/blog/2026/03/mcp-is-dead-long-live-mcp/

[1] https://blog.cloudflare.com/mcp-v2/

CharlieDigital··on Sherline Tools Is Going Out of Business
The US is kind of "collapsing" in third spaces both public and private, but especially public due to funding cuts.
CharlieDigital··on Towards Self-Driving Codebases
I would certainly consider Roslyn analyzers capable of covering "a tiny fraction" of possible errors :) They are quite capable of covering for many common types of structural coding mistakes.
CharlieDigital··on Towards Self-Driving Codebases
You would be surprised. Humans will do human things like be extremely inconsistent, ignore warnings (if they are not enforced as errors), skip steps because they are lazy (devs often chose to skip our pre-push hooks and preferred to run in CI and babysit the PR).

Agents can also do all of those things, but they are generally more compliant to instruction.

CharlieDigital··on OpenJev
Yes, agree, but also limited to specific types of use cases.
CharlieDigital··on Towards Self-Driving Codebases
I don't think I implied that they could.

Each of these are just layers of control at different lifecycles of agent code generation. Analyzers are nice because it gives targeted, static analysis that can prevent certain classes of errors very early and at lower iterative cost (e.g. a build)

CharlieDigital··on An Empirical Study of Harness Design for Coding Agents

   > In other words, MCP was just a bunch of bullshit
This is a complete misunderstanding of why a team would want MCP.

If you're just using bash scripts, where are you putting your enterprise secrets for external systems? How do you cleanly revoke them when a developer leaves your team?

MCP moves execution into a remote environment where it is easy for enterprises to secure access to internal and external systems. OAuth based access makes it easy to audit and revoke tokens. Central HTTP interface makes it trivially easy to monitor and audit.

They solve different problems.

← PreviousPage 2 of 34Next →