HNHacker News
TopNewBestAskShowJobs

CharlieDigital

6,609 karma · joined July 2, 2021

If you are an experienced C# dev and would like to work at a well-funded, profitable, series-C, former YC startup building green-field products, check out: https://www.usemotion.com/careers and apply directly

In my free time, I'm making this [https://turas.app] and this [https://coderev.app]

More: [https://chrlschn.dev] [https://github.com/CharlieDigital] [https://charliedigital.com] [https://www.linkedin.com/in/charlescchen] [https://www.youtube.com/@chrlschn]

Let's connect!

submissionscomments
CharlieDigital··on AI Has No Wisdom and Neither Will You

    > because you have time to reflect
It doesn't mean that people do. This is a false narrative we tell ourselves. Yes, there are craft-oriented devs and teams, but these are the exception rather than the rule because in the end, it is the GTM and business teams that define what, when, how and rarely the engineering teams.

There is no team without tech debt because there is no "golden" project where every decision has been made right because of reflection on decisions made wrong.

CharlieDigital··on AI Has No Wisdom and Neither Will You

    > Fact is, vibe-coded projects devolve over time into an unmaintainable mess. The reason is simple, yet hard to fix: code maintainability and good architecture don’t have good measurements that we can apply, because it takes months, years even, to notice the effects of bad architecture or of unmaintainable code.
    > 
    > For one, AI is not trained on what it means for code to be maintainable. For instance, any reinforcement learning done needs a reward signal that can be measured immediately, not in months or years.
Sad to say, but this is no different from human written code. Human written code just takes even longer to realize the mistakes because the pace is slower.

I think at the end of the day, it is not impossible to have AI write "good" or "high quality" code. If anything, once the patterns are established, AI will be more likely to adhere to the patterns and rules than any human team. It requires the most experienced engineers on the team to split their time writing the core patterns and documenting them in references/skills.

But it takes a lot of "taste" and a willingness to slow down a bit with AI (to create necessary artifacts), something teams find hard to do when you can ship so fast now.

My experience has been that there is a camp of very senior engineers that are unwilling to adapt to reality and focus on documentation and writing (effectively producing skills and agent guidance which multiplies their effectiveness); they will cling to their knowledge thinking coding a sacred art.

CharlieDigital··on AI coding has made CI a bottleneck, so we reworked ours to keep up

    > If there is one universal amongst organizations, its that they love walls and silos.
What's true for human organizations isn't necessarily true for agent-driven engineering. People and human teams struggle with contracts because there's always human negotiation involved. If the decisions are instead made by a team of agents, there's no more ego, ownership, miscommunications; just decisions based on whatever rules have been given to the orchestrator.

    > Your agents may be happy in their tiny walled kingdom...
Yes indeed; the agents will always be happier if they can iterate faster, lint faster, build faster, test faster, ship faster, with smaller context.

That would be the point of using contracts as boundaries so the agent can iterate more autonomously so long as it maintains the externally facing contract or version the contract if it needs to.

CharlieDigital··on AI coding has made CI a bottleneck, so we reworked ours to keep up
My take: it seems like systems should become smaller, more isolated, and contract-oriented.

I have been a long time proponent of monoliths, but it seems like agents would be happier with smaller, more isolated services. The more isolated, the better. Contracts between the service components only. Then it can iterate internally as long as it satisfies the contract. If it needs to, it can version the contract and keep iterating.

CharlieDigital··on AI coding has made CI a bottleneck, so we reworked ours to keep up
Question here:

    > where sub agent orchestration is done through agent to agent messaging
How do you expect to see the history/record of what the agents did and why? Is it enough to see it in PRs? Do you expect tickets that have the design and history? How are you thinking of agents being able to historically resolve reasoning/why/decisions made in earlier passes?

Genuine open question here. My assumption is that a GH or Linear or Jira is still useful as a decision store. It may as well be a custom app over Postgres, but it seems like something is needed to store this and for observability. A GH/Linear/Jira is nice if only because of standard APIs and integration points (whatever you build would likely end up duplicating a subset of those).

CharlieDigital··on Show HN: Lossless-memory – a personal AI memory that never summarizes
The point is to build systems that don't require the agent to "ask".

The human becomes the bottleneck as systems become more agentic.

OP's system is first ingesting and storing the individual messages and then recompiling it into "working memory".

CharlieDigital··on Show HN: Lossless-memory – a personal AI memory that never summarizes
The specific design of this system uses the raw memories and allows rebuilding the operational memory from the raw memory (the LLL "Left Leg Layer").

To do this requires that there is an ordering of which memory came last.

    "always do X before committing Y"
    "always do X before committing Y except after Z"
Which of these is the current state? Without the date, it is not possible to rebuild the operational state of the rule from the raw records.
CharlieDigital··on MCP was always a bad idea?
I also take a lot of the guidance from influencers with a grain of salt for that reason. What works for the solo dev on greenfield projects rarely scales to the enterprise 1:1. Brownfield. Security policies. Enterprise controls. Legacy systems.

There is usually a kernel of value that has to be extracted and translated to apply what works for the solo AI engineer and what works for the team.

CharlieDigital··on MCP was always a bad idea?
I think many devs misunderstand MCP because they work in solo mode. In solo mode, you just have your secrets local. You don't care about auditing access. There's no IT managing the infra and third party secrets. You don't have to account for different harnesses and tooling; you just use your own harness and adapt your tooling to it. You're not thinking about revoking/rotating secrets when someone leaves your team. You don't have to deliver capabilities to many different runtimes and stacks.

Back in March, everyone was already pronouncing it dead[0] when in fact, it has only proliferated and become even more essential for both 3rd party systems as well as platform level capabilities[1] as agentic tooling has moved a bit more slowly into the enterprise. It was apparent even back then that enterprises will need MCP.

In a team context? Enterprise? Building web server backed or in-process agents? Not sure how you replicate the control, auditability, accessibility, composability, and security boundary that you get with MCP over HTTPS without a lot of bespoke, point solutions; in the end, protocols almost always win.

Could you do it with just REST APIs? I mean, MCP is just JSON-RPC over HTTP with standardized auth, schemas, and agent specific exchange flows (tools, prompts, resources, etc.). Could you do it with just CLIs? You lose a lot of the control mechanisms offered by MCP (auditing, security, centralized auth, etc.). The context savings are overblown except with CLIs that have good representation in training (curl, jq, cat, sed, etc.; your custom CLI is going to need to produce instructions and add to context all the same)

Heuristic is simple: don't use MCP for local, solo dev. As soon as you need MCP, you'll know it and you'll understand why it exists. Many of the harness level capabilities themselves are implemented as first party MCP (more apparent in the CLIs). MCP is basically REST for the agentic era.

[0] https://chrlschn.dev/blog/2026/03/mcp-is-dead-long-live-mcp/

[1] https://blog.cloudflare.com/mcp-v2/

CharlieDigital··on Sherline Tools Is Going Out of Business
The US is kind of "collapsing" in third spaces both public and private, but especially public due to funding cuts.
CharlieDigital··on Towards Self-Driving Codebases
I would certainly consider Roslyn analyzers capable of covering "a tiny fraction" of possible errors :) They are quite capable of covering for many common types of structural coding mistakes.
CharlieDigital··on Towards Self-Driving Codebases
You would be surprised. Humans will do human things like be extremely inconsistent, ignore warnings (if they are not enforced as errors), skip steps because they are lazy (devs often chose to skip our pre-push hooks and preferred to run in CI and babysit the PR).

Agents can also do all of those things, but they are generally more compliant to instruction.

CharlieDigital··on OpenJev
Yes, agree, but also limited to specific types of use cases.
CharlieDigital··on Towards Self-Driving Codebases
I don't think I implied that they could.

Each of these are just layers of control at different lifecycles of agent code generation. Analyzers are nice because it gives targeted, static analysis that can prevent certain classes of errors very early and at lower iterative cost (e.g. a build)

CharlieDigital··on An Empirical Study of Harness Design for Coding Agents

   > In other words, MCP was just a bunch of bullshit
This is a complete misunderstanding of why a team would want MCP.

If you're just using bash scripts, where are you putting your enterprise secrets for external systems? How do you cleanly revoke them when a developer leaves your team?

MCP moves execution into a remote environment where it is easy for enterprises to secure access to internal and external systems. OAuth based access makes it easy to audit and revoke tokens. Central HTTP interface makes it trivially easy to monitor and audit.

They solve different problems.

CharlieDigital··on OpenJev
OP's point here is that the overall approach of restricting output token space and using parallel prompts to produce concurrent results and taking the most relevant ones isn't something novel to Jev (not saying there's nothing novel, but a facsimile can be created at the application layer using any small, fast model)
CharlieDigital··on Towards Self-Driving Codebases
You can express them as tests, but you also need a feedback mechanism that creates the rule that when the LLM generates some net new code or performs some refactor, that there are these CAPAs that it needs to cover with test cases.

The CAPA is a learning that sits outside of the mechanism of verification; it is a record of problem:root_cause:preventative_action. I see it as the instruction that would be required to generate the test case to prevent the next occurrence of a class of failures.

In a real-world process, for example, there is usually a QA lead that is verifying that the process is followed by looking at the paperwork and evidence.

CharlieDigital··on Towards Self-Driving Codebases
This may be platform dependent.

C# Roslyn Analyzers[0], for example, are quite powerful and can identify complex patterns in code. One approach to deterministic enforcement would be to ensure that the project is set up with an analyzers library and mistakes that can be deterministically flagged are

[0] https://learn.microsoft.com/en-us/visualstudio/code-quality/...

CharlieDigital··on Towards Self-Driving Codebases
This is a question of context management and I suppose some would classify this as "harness engineering" as the trend of the moment.

One approach, for example, might be to have the a standalone code reviewer agent that is solely responsible for interfacing with the CAPA system (e.g. via a tool, via MCP) and acts as a back stop. When it finds a new type of CAPA, it stores it (and the backend indexes it with enough metadata to support broad types of retrieval). When it reviews a piece of code, it finds past CAPAs. By file locality. By business domain in the application. By keywords.

Same tool and repository available to both building agents and review agents, but use the review agent as a dedicated back stop as part of the verification process.

CharlieDigital··on Towards Self-Driving Codebases

    > It’s actually fine if agents make a lot of boneheaded mistakes. What’s not ok is if they keep making the same mistakes. 
I worked in life sciences for a bit. There is a process in clinical trials called corrective and preventative actions (CAPA). You'll also find this in other areas where failure tolerance is low (e.g. aircraft).

It's simple: when a mistake happens, you run your CAPA process (Google CAPA form and see examples to extrapolate what that process might look like) and determine the root cause and the correction to the process that allowed the mistake to happen in the first place.

(At least as a SaaS vendor in life sciences, when we had a CAPA (e.g. after a SEV0 failure), it would be folded into our SOPs and then we would be required to retrain on the SOP. Auditors would want to see our evidence of CAPAs, the versions of our SOPs, the records of training. All to extreme for most shops, but I add this for context/color)

This is something most eng shops do not have the discipline for since it requires some diligence.

Should it be fully agentic? Should there be human intervention here to approve the CAPA? Open questions to be answered.

CharlieDigital··on Rate limits on GitLab.com are changing
This is not fully true.

Most agents will use curl | jq to slice what they need (assuming a known API)

CharlieDigital··on HarnessTax: How Much Does the Harness Matter for Coding Agents?
Claude Code (CLI) has been really bad with handling window resizing and layout changes. For long sessions, it will freeze while redrawing. On an M3 Max. With 64 GB or memory. Codex CLI does not have this problem, nor does OpenCode. Every time I split the pane, I'm dumbfounded by how bad this is.
CharlieDigital··on Japan's book scene is moving from bookstores to libraries
Not the case!
CharlieDigital··on OpenSpec – A lightweight and configurable AI spec framework
This section https://openspec.dev/docs/setup links to "Concepts"

Concepts links here: https://github.com/Fission-AI/OpenSpec/blob/main/docs-lab/gu...

All the docs here are the templates rather than the actual file (I presume: https://github.com/Fission-AI/OpenSpec/blob/main/docs/concep...)

Somehow not very confidence inspiring...

CharlieDigital··on Performance Improvements in .NET 11
Flip side: LLMs have to be very explicitly told to code in "modern" .NET and C# due to lack of representation in the training set.

Case in point: extension members from C# 14 is one that LLMs commonly stumble on and requires an explicit example. It still sometimes says that this is not valid syntax.

    extension(SomeType instance)
    {
        public OtherType DoSomething() { ... }
    }
Agents really struggle on this one for some reason.

Even older releases have a few that I notice LLMs making mistakes on like use of `System.Threading.Lock` over `Object` when locking.

CharlieDigital··on Japan's book scene is moving from bookstores to libraries
I recently had a similar episode as the author in another east Asia country: Taiwan.

Found myself in Taitung (along the much less densely populated, remote east coast) with several inches of rain forecast. Spent the day at the Taitung County Public Library [0] and it was a very pleasant public space which had more tourists than patrons!

Shots of the interior for anyone interested (surprisingly Scandinavian): https://youtu.be/gq8xlk_JEG4?t=108

[0] https://maps.app.goo.gl/tnEvHQLMNBGAnMWF9

CharlieDigital··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
This is only true if you only run the code reviewer in the final PR.

The way we set it up was that the code review responded to two signals:

1. GH PR webhook

2. The exact same review agents running as an MCP tool that the local agent can invoke before pushing.

Practically, what this means is that it's OK to have a false positive because the local agent can make the check with the full context.

This would be the same as if the entire team used Codex and, for example, had a sub-agent configured to run code reviews using a smaller model. In this case, the benefit to this tool-based approach is that the exact same agent works for all harnesses across a team and also works in the PR itself.

CharlieDigital··on Java 27
C# gRPC story is really, really good, too.
CharlieDigital··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
Because your local agent can already see the full codebase; the code review only needs to see what's changing and evaluate the change.
CharlieDigital··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
I found Luna and even 5.4-mini to be quite good at code review provided a few things:

1. Run it in multiple cycles, only on the diff, and only emit a few findings at a time.

2. Give it a memory so each cycle, it knows the previous finding to check if it's been fixed.

3. Give it access to canonical docs that encode your human reviewer heuristics. I exposed these as tool calls so they could be tracked via telemetry.

4. Run multiple reviewers, each with a tight focus. Security, performance, structural, database, etc. Each a separate prompt and persona. Additionally, we had file activation filters so the FE React reviewer didn't activate on BE only changes.

Luna and 5.4-mini with no reasoning were exceptionally fast and almost always found issues with code produced by Opus and Fable.

Default prompts for the curious (these are templates deployed by default, but customizable).

Performance: https://github.com/zeeq-ai/zeeq-app/blob/main/src/backend/Ze...

Structural: https://github.com/zeeq-ai/zeeq-app/blob/main/src/backend/Ze...

(Keep in mind each agent also has tools to access and reference external docs.)

← PreviousPage 3 of 34Next →