HNHacker News
TopNewBestAskShowJobs

CharlieDigital

6,609 karma · joined July 2, 2021

If you are an experienced C# dev and would like to work at a well-funded, profitable, series-C, former YC startup building green-field products, check out: https://www.usemotion.com/careers and apply directly

In my free time, I'm making this [https://turas.app] and this [https://coderev.app]

More: [https://chrlschn.dev] [https://github.com/CharlieDigital] [https://charliedigital.com] [https://www.linkedin.com/in/charlescchen] [https://www.youtube.com/@chrlschn]

Let's connect!

submissionscomments
CharlieDigital··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
I found Luna and even 5.4-mini to be quite good at code review provided a few things:

1. Run it in multiple cycles, only on the diff, and only emit a few findings at a time.

2. Give it a memory so each cycle, it knows the previous finding to check if it's been fixed.

3. Give it access to canonical docs that encode your human reviewer heuristics. I exposed these as tool calls so they could be tracked via telemetry.

4. Run multiple reviewers, each with a tight focus. Security, performance, structural, database, etc. Each a separate prompt and persona. Additionally, we had file activation filters so the FE React reviewer didn't activate on BE only changes.

Luna and 5.4-mini with no reasoning were exceptionally fast and almost always found issues with code produced by Opus and Fable.

Default prompts for the curious (these are templates deployed by default, but customizable).

Performance: https://github.com/zeeq-ai/zeeq-app/blob/main/src/backend/Ze...

Structural: https://github.com/zeeq-ai/zeeq-app/blob/main/src/backend/Ze...

(Keep in mind each agent also has tools to access and reference external docs.)

CharlieDigital··on Introduction to AI Engineering
Way too shallow; was expecting a lot more meat here...
CharlieDigital··on Nike exits the S&P 100 after 18 years and a $200B market-cap wipeout
Nowadays, it seems that catering to any audience other than 1) white 2) male is "culture war".
CharlieDigital··on Show HN: Authorize MCP tool calls without giving agents the credentials

    > How do you let an agent use a tool that needs credentials without giving the credentials to the agent?
I took a different approach [0].

    1. Register a secret and get back an opaque identifier (out-of-band, human action; can extend and add an API for listing by endpoint)
    2. Generate a JavaScript API call that includes the opaque identifier as header (for example)
    3. Send the JavaScript to a server that runs a JS interpreter
    4. Pre-process the inbound JS and replace the opaque identifier with the actual token
    5. Run the JS to invoke the API, make transforms, process the result, etc.
Effectively using agent generated JS script to make the actual API call where the agent only sees the opaque ID of the secret.

[0] https://github.com/CharlieDigital/runjs

CharlieDigital··on Making Startups Powerful
I think this is a very classic "waterfall" world view.

In many cases, the problem can only be understood through many failed solution attempts. It could be because the actual problem is novel. It could be because there is a different way of thinking about the problem that is not apparent studying only what a customer does today with the limitations of the software that they have right now.

The point of "move fast, break things" is to address this fallacy that anyone can fully understand a problem before attempting the solution; it is often the case that in attempting a solution, one comes to understand the actual, valuable problem.

CharlieDigital··on Don't be the out of touch Kung Fu master

    > Then you look at the code. Code that has passed review, generally. You realize that the database schema has been broken silently, and that the agent has rewritten the tests or the golden fixtures to match. 
I find that the code review leg of this is critical to invest a lot of energy into hardening.

First is don't trust the code review from your local harness, even if it uses sub-agents; externalize it into another system.

Second is to get your most critical human code reviewers to encode their heuristics in markdown files and feed those to the code review agents.

Third, if possible, is to bring the code review "into the loop" so that it's not only running in the PR, but also running in the coding loop so the coding agent has immediate, external feedback. Final PR code review is a backstop.

This pattern [0] works well because it solves for some team level problems where folks are using different harnesses or different models (consistency issues) and it means that code reviews don't just sit at the end of the loop; it actively alters the code production cycle.

[0] https://zeeq.ai/docs/key-features/code-review-tool

CharlieDigital··on Scheduling is NP-Hard. We need responses in milliseconds
I'm not sure about the specifics of what Cal.com is trying to do, but interval trees [0] are a good fit and generally simplify a lot of the otherwise iterative or recursive logic for scheduling. There are some other related algorithms in 3D to calculate occulusion that can be used.

I have a practical writeup here as well that demonstrates pulling multiple calendars, building an interval tree, and then checking for overlaps/finding free slots [1]. It is possible to modify the algorithm so that the free slots are also tracked so no iteration is required to detect conflicts (trading space for time).

There are some other algorithms for bin-packing (calendar scheduling over a closed interval is just a bin-packing problem) that can also be used for this type of solving. Google OR-Tools, for example [2].

However, my sense is that interval trees or bin-packing solvers are what teams building schedulers really want, but teams building schedulers usually start with simple. Hope the Cal.com team sees this and gives it a go!

[0] https://en.wikipedia.org/wiki/Interval_tree

[1] https://chrlschn.dev/blog/2022/11/concurrent-processing-dotn...

[2] https://developers.google.com/optimization/pack

CharlieDigital··on Open Source Durable Objects for Postgres
It is perhaps more accurate to say that they can meter by the second, but the billing calculation is a trailing activity.
CharlieDigital··on DeepSeek v4.1 Flash
Politics and profits.

Deepseek delivers 1 product; Microsoft delivers dozens (or hundreds depending on how you want to count it) across various domains.

CharlieDigital··on Claude, change the "Add to Cart" button to blue
Downside would seem to be that CI tends to increase the cycle time and feedback loop and add their own cost into the equation.
CharlieDigital··on How well do agents use test/verification techniques?
I worked for a long stint building SaaS for life sciences/pharma.

A big, big part of this space is "traceability" through the SDLC. If something goes wrong, there is a long chain of liability from the final effect (a subject with an adverse event) to the root cause (system recorded the wrong data in a clinical trial) all the ways to a 3rd party vendor. We call software "validated" if it has gone through the GAMP 5 V-model of software development and produced the necessary artifacts that correlate verified behavior to specification.

While this is way too much rigor for many scenarios, I think more folks should be familiar with the GAMP 5 V-model in the agentic era. The left side of the V defines the requirements moving from high level business requirements to low level technical requirements, the right side of the V defines the verification artifacts corresponding to the technical and business requirements (in that order). In this structure, testing covers both the functional requirements as well as the technical requirements.

I think this mental model is extremely useful and reflects the interaction model with agents: humans now focus on the requirements and less on the implementation details. It is useful to separate the business from the technical and therefore, the two types of artifacts that need to be produced to satisfy verification thresholds.

This style of structured software development (pre-agents) is very expensive. With a 2-3 weeks planning/specification phase on the front-end (since the test specification has a dependency on the approved requirements), a 4 week implementation phase (this is iterative with the test team, but informal; changes from the approved plan are documented as deviations), and then a 2-3 week formal test/verification phase. But in a post-agent world, this model feels like it is 1) more reasonable, 2) perhaps more effective, 3) more manageable.

While the teams I've now been on in startups have focused heavily on fast iteration with AI, a big downside I've noted is that it seems that we keep building the wrong things or things that are not useful because we are no longer verifying the business requirements before building. We assume the cost of building is low so we build and the iterate; testing in this reality becomes ad-hoc and full of gaps because the functional behavior of the software is no longer scoped before building; it's "vibes" based on each iteration because the iterations are cheap. The result? The teams I've been on feel like we're going nowhere in many cases, just spinning faster and producing more throwaway code (not a comment on whether this is good or bad; it may vary by domain and nature of the company).

If one is looking to build a software factory or pipeline, I think it behooves an architect to examine the GAMP 5 V-model and consider how to adapt ideas from this model to agent systems.

CharlieDigital··on We are not going anywhere
In my 30's I really hit my stride as a developer and system architect. Enough experience, seniority, and autonomy to own and build out large complex systems.

Many, many mid career devs, I fear, will miss this window. They will become reliant on the LLMs more and more. 2 months back, I was asked to backtest an interview question and 2 out of our 3 most senior engineers (both in their 30's) could no longer write a generic method.

I liken this experience to learning cursive as a kid. It wasn't about writing cursive; it was about developing dexterity and hand-eye coordination. Getting the reps in, so to speak. Even if the future is all AI, that window of expanding one's knowledge and understanding of system design and architecture through hands-on experience (and failure!) facilitates the formation of "taste" through reps: why A over B or C. When B over A or C?

Many, many devs will end up "going nowhere". They will be able to prompt and push code with the façade of productivity, but I think building stable, scalable, complex systems requires knowing which angles to probe and which questions to ask; things learned via reps of trying, failing, learning, failing some more, thinking hard, drawing it out, and finally hitting the breakthrough.

I recently published a series of blog posts that focuses on the underlying architecture decisions that I think can help teams set a solid foundation for building with AI [0]. I think the guidance and patterns in it are unlikely to be emergent from an LLM without very explicit prompting. The goal is to share the thought process and intent for each technical decision. I think this type of thinking may become more rare as folks surrender their reps to LLM defaults.

[0] https://chrlschn.dev/blog/2026/08/the-unexpected-ai-stack-cs...

CharlieDigital··on The Unexpected AI Stack: C# + .NET (Part 1)
This is a 5 part series that is intentionally (hand) written to help dev teams understand how to scaffold a codebase for agentic engineering by focusing on key, underlying technical decisions and manual wiring before building with AI. Getting the foundations right helps provide the tools and safeguards for coding agents to iterate more efficiently while reducing slop.

Specifically:

- Giving agents access to programmable runtime orchestration (Aspire.dev[0]) - Empowering agents to iterate rapidly with runtime mutability (using CSharpRepl[1]) to dynamically modify code at runtime while retaining full application state - Using the GitHub Copilot SDK to build an agentic core with a multi-platform harness, BYOK, any model provider - Testcontainers[2] with automatic transactions to streamline and isolate integration tests - A well-documented, AI-friendly UI framework (Nuxt UI[3]) - Logging and telemetry to give agents insights and visibility into the runtime state of the application

The core setup is used at a series C, post-YC startup to ship fast with AI while maintaining high quality standards (in combination with other tools facilitating code review and context management)

Part 1 (https://chrlschn.dev/blog/2026/08/the-unexpected-ai-stack-cs...) is an intro into a few key parts of this stack.

Part 2 (https://chrlschn.dev/blog/2026/08/the-unexpected-ai-stack-cs...) is focused on walking through the hands on scaffolding.

Part 3 (https://chrlschn.dev/blog/2026/08/the-unexpected-ai-stack-cs...) covers wiring GitHub Copilot SDK as an agent runtime and incorporating CSharpRepl to allow agents to dynamically work with the runtime DI container

Part 4 (https://chrlschn.dev/blog/2026/08/the-unexpected-ai-stack-cs...) wires up the test harness using Testcontainers to give agents isolated test environments

Part 5 (https://chrlschn.dev/blog/2026/08/the-unexpected-ai-stack-cs...) wires up logging and telemetry to give agents visibility into runtime state and I start to build the prototype application now that the foundations are ready.

---

The project repo is here: https://github.com/zeeq-ai/zeeq-tmpl (be sure to check the branches; main is currently the base code only)

I encourage working through the posts since the goal is to underscore the platform level decision making process and assembly of the foundational core.

[0] https://aspire.dev/

[1] https://fuqua.io/CSharpRepl/

[2] https://testcontainers.com/

[3] https://ui.nuxt.com/

CharlieDigital··on DeepSeek Harness developer preview
Everyone thinks their workflow is a special snowflake.

Truth is that useful dev workflows and tooling probably coalesces in a tight band. There's really no point in re-inventing the wheel over and over again at this level (the raw tooling).

CharlieDigital··on Ask HN: What are you working on? (August 2026)
I'm (solo) spinning off the internal agent knowledge and observability layer used at Motion (YC W20, series C, ~40 devs). It was born out of a need to lift all devs using heterogeneous models, harnesses, and prompting styles to a higher level of consistency and quality while surfacing the visibility to see what's working and what's not. The knowledge layer is integrated into the code review layer and all of it tied together with telemetry to the document section level to see how each piece of context is affecting code generation.

Open source repo (.NET 10 + C# + Postgres): https://github.com/zeeq-ai/zeeq-app

Docs and more info: https://zeeq.ai

CharlieDigital··on Stateless MCP has recaptured my interest
MCP is even more important now in enterprise and thus the shift towards the streamable, stateless HTTP server implementation.
CharlieDigital··on Stateless MCP has recaptured my interest
It isn't.

MCP over streamable HTTP is just JSON RPC payload over HTTP.

The internal interface for the ChatGPT app and Coded are all MCP based (stdio on local).

All enterprise AI rollout is eventually going to converge on MCP as a key piece of that infra because it can mask credentials behind one gateway that's easier to control and more secure.

CharlieDigital··on Pi's Minimalism Is Its Advantage
how many extensions do you need for browsing the web? Adblocker. Not much else because the browser already does the important work.

Same with any of the off the shelf harnesses.

Go build something useful instead of tweaking the minutiae of the harness.

CharlieDigital··on Stateless MCP has recaptured my interest
Stateless MCP was already possible before this and made sense for whole classes of use cases where it helps to have a remote fleet of servers.

Wrote about this back in March: https://chrlschn.dev/blog/2026/03/mcp-is-dead-long-live-mcp/

MCP is going to be a foundational piece of enterprise agent infra.

CharlieDigital··on Pi's Minimalism Is Its Advantage
End up building Pi extensions instead of actual product.

Doesn't make sense to me, but to each their own.

CharlieDigital··on Codex logging bug may write TBs to local SSDs
> ...Do they just want to force you to keep busy

Given functionally unlimited access to tokens with frontier models, there is really no "force you to keep busy"; it should just bake overnight. We're talking about a rather simple and well-defined specification; not something novel and complex.

CharlieDigital··on Codex logging bug may write TBs to local SSDs
Anthropic were the progenitors of the Model Context Protocol. Claude Code does not fully implement the client end of the protocol. A protocol; a literal pre-defined spec that an agent should be able to one-shot. Neither does Codex. Codex does not implement MCP Prompts.

(I want Codex to implement MCP Prompts because then we have one central way to ship skills from a server).

The fact that neither platform can implement a protocol given what is functionally infinite frontier model tokens really says a lot. I do not care what kind of random project some influencer can ship with a swarm of 1000 agents. If you cannot make the basics work, it is a farce.

CharlieDigital··on Fruit Is Too Sweet
Pink Lady, Snapdragon, Sweetango are all probably closer to alanced compared to Cosmic Crisp. Cosmic Crisp being more tart than Sugarbee, but still definitely more sweet than tart in flavor profile IMO.

Sweetango and Pink Lady are probably what I would consider balanced sweet and tart.

CharlieDigital··on The only scalable delete in Postgres is DROP TABLE

    > And you cannot keep doing high concurrent DROP TABLEs to run your large scale CRUD app
In this kind of use case/design, I would assume it would make use of partitions to make this more palatable in which case it would seem that you would bypass this issue of "high concurrent DROP TABLE". Large scale CRUD app just points to recent-ish partitions. Old partitions are either going to be low or on access and can be dropped easily or transformed/transferred into some long term/cold storage.
CharlieDigital··on pg_durable: Microsoft open sources in-database durable execution
A few things are not clear to me from reading through docs and examples:

    df.wait_for_schedule()
How does this call work? Is it idempotent if I call it from an application? If I run it 2x with the same parameters, does it double tick? Am I invoking this manually from a query console to only do this one time? Am I running this as part of a migration script?

For this[0]:

    -- Wait for human signal (5 minute timeout)
    ~> (df.wait_for_signal('approval', 300) |=> 'sig')

    ~> df.if(
        $$SELECT NOT ($sig::jsonb->>'timed_out')::boolean
            AND ($sig::jsonb->'data'->>'approved')::boolean$$,
Is the `timed_out` a fixed constant that is returned on timeout?

Also not immediately clear: how to handle errors/exceptions?

[0] https://github.com/microsoft/pg_durable/blob/main/examples/i...

CharlieDigital··on VoidZero Is Joining Cloudflare
Really good talk that goes over this: https://corecursive.com/vue-with-evan-you/

Totally worth the listen.

CharlieDigital··on Uber's $1,500/month AI limit is a useful signal for AI tool pricing

    > We are shipping more features
That's not really the important question; the important question: is it generating revenue.

If you increase your spend -> ship more features -> no correlated increase in revenue, that's just burning money.

If a team of 10 spends 1 extra headcount ($180k/year) and ships features with no corresponding growth in revenue, what does that mean?

There was probably a reason it was on the backlog (because it didn't really have value).

CharlieDigital··on Uber's $1,500/month AI limit is a useful signal for AI tool pricing
Even if the laptop costs $5k and you upgrade it every year with the latest hardware and run local models (assuming your workload can tolerate smaller models at slower tok/s), you win.
CharlieDigital··on Uber's $1,500/month AI limit is a useful signal for AI tool pricing
That's because for some of these folks, the cost of the tokens doesn't have to match the value of the output; the hype from the story is all they need.

Normal people have to produce something of value from that spend. So starting 100 agents and then waking up to something cool but useless just means you spent a few thousand dollars and created nothing of value............

CharlieDigital··on Uber's $1,500/month AI limit is a useful signal for AI tool pricing
You're a content creator; you define your revenue stream.

Uber engineers do not define their revenue stream; the product leadership team does.

$1500/mo of AI spend by engineers does not equate to revenue. They need to figure out revenue first before zeroing in on AI spend.

← PreviousPage 4 of 34Next →