HNHacker News
TopNewBestAskShowJobs

digitcatphd

462 karma · joined October 9, 2020

submissionscomments
digitcatphd··on The growing divide between AI hype and software engineering reality
First, I don’t think it’s widely accepted financial markets are in an AI bubble or there is actually any evidence to support this, most market cap growth is supported by real earnings per share. It is debatable how long that will continue and if it is sustainable, but that isn’t a bubble it’s a cycle. Even the most skeptical investors to bubbles like Buffet and Burry, have large stakes in Google and Microsoft, front line bets on AI.

Second, I agree with the distinction evals are not enough, and this is dangerously making people reckless, I think the conclusion is still quite incorrect.

There is essentially nothing magical about being human. Our primitive brains were trained on logic, reasoning and processes. An LLM is simply this on less efficient hardware, but, improving at a speed faster than human evolution.

digitcatphd··on The AI bubble is popping; we just don't know it yet
I would argue we still have not even really gotten started.

What do we have in the decade ahead? Robotics in every household, models 10x+ faster and more intelligent than today.

Really no significant impact in life sciences, R&D, and 'offline' world / robotics today as of yet, which is where most of the value will live.

digitcatphd··on Alex Karp says only trade workers and neurodivergents will survive in the AI era
I never quite understood the notion trade workers will be exempt for a couple reasons:

1. Right now trades businesses are profitable because of supply and demand. They are profitable, because they are undersupplied.

2. We are assuming robotics stagnates.

digitcatphd··on AI agent opens a PR write a blogpost to shames the maintainer who closes it
Respectfully, the same argument was for Moltbook's controversial posts and it turned out to be humans.
digitcatphd··on AI agent opens a PR write a blogpost to shames the maintainer who closes it
I don't know why these posts are being treated by anything beyond a clever prompting effort. If not explicitly requested, simply adjusting the soul.md file to be (insert persona), it will behave as such, it is not emergent.

But - it is absolutely hilarious.

digitcatphd··on Launch HN: AgentMail (YC S25) – An API that gives agents their own email inboxes
This is fantastic! I created a GMAIL for my Clawdbot and Google deleted the account after an hour.
digitcatphd··on Vm0
I built something similar to this before Langraph had their agent builder @braid.ink, because Claude Code kept referencing old documentation. But the problem ended up solving itself when Langraph came out with their agent builder, and Claude Code can better navigate its documentation.

The only thing I would mention is that building a lot of agents and working with a lot of plug-ins and MCPs is everything is super situation- and context-dependent. It's hard to spin up a general agent that's useful in a production workflow because it requires so much configuration from a standard template. And if you're not being very careful in monitoring it, then it won't meet your requirements when it's completed, when it comes to agents, precision and control is key.

digitcatphd··on ChatGPT Health
At first I was reading this like 'oh boy here we go, a marketing ploy by ChatGPT when Gemini 3 does the same thing better', but the integration with data streams and specialized memory is interesting.

One thing I've noticed in healthcare is for the rich it is preventative but for everyone else it is reactive. For the rich everything is an option (homeopathics/alternatives), for everyone else it is straight to generic pharma drugs.

AI has the potential to bring these to the masses and I think for those who care, it will bring a concierge style experience.

digitcatphd··on [dead]
I’ve been writing about building Agent-First SaaS and working with teams implementing LangGraph flows. I’ve noticed a recurring pattern where we get stuck trying to perfectly replicate a human's SOP (e.g., "click this button, then read this PDF"). While reproducing human workflows is great for trust and "human-on-the-loop" auditing, I argue it often traps us in a local optimum.

This post explores the difference between "Replica Agents" (biomimicry) and "First-Principles Agents" (optimizing for the objective function). I draw on examples like Amazon's "Chaos Storage" and AlphaGo to suggest that sometimes the most efficient agent workflow looks nothing like the human one.

Curious to hear how others are balancing "legibility" vs. "efficiency" in their agent designs.

digitcatphd··on Move 37 and the Case for "Alien" Agent Workflows
I’ve been writing about building Agent-First SaaS and working with teams implementing LangGraph flows.

I’ve noticed a recurring pattern where we get stuck trying to perfectly replicate a human's SOP (e.g., "click this button, then read this PDF"). While reproducing human workflows is great for trust and "human-on-the-loop" auditing, I argue it often traps us in a local optimum.

This post explores the difference between "Replica Agents" (biomimicry) and "First-Principles Agents" (optimizing for the objective function). I draw on examples like Amazon's "Chaos Storage" and AlphaGo to suggest that sometimes the most efficient agent workflow looks nothing like the human one.

Curious to hear how others are balancing "legibility" vs. "efficiency" in their agent designs.

digitcatphd··on We gave 5 LLMs $100K to trade stocks for 8 months
Backtesting is a complete waste in this scenario. The models already know the best outcomes and are biased towards it.
digitcatphd··on The Generative Burrito Test
I find it a bit surprising GenAI has made it this far without this benchmark
digitcatphd··on £220 'for a cut-up sock' — Apples's new iPhone Pocket ridiculed online
They will seed in a few dozen influencers and there will be lines out the door
digitcatphd··on iPhone Pocket
Beatifully said and you are right. I will get mine on Temu.
digitcatphd··on Circular Financing: Does Nvidia's $110B Bet Echo the Telecom Bubble?
What do you think is going to happen to their earnings when CAPEX slows?
digitcatphd··on Circular Financing: Does Nvidia's $110B Bet Echo the Telecom Bubble?
The biggest issue with Nvidia is their revenue is not recurring but the market is treating their stock as it were, which is correlated with all semi stocks, with a one-time massive CAPEX investment lasting 1-2 years.

Simple as this - as to why its just not possible for this to continue.

digitcatphd··on Shiller PE Ratio
It’s never different this time, this is embedded into human nature and people oscillate between fear and greed. That’s it. Not more complicated than that.
digitcatphd··on Shiller PE Ratio
T-Bills are at 4% which is below the rate of inflation
digitcatphd··on Shiller PE Ratio
Irrational exuberance the very thing this indicator was intended to track.
digitcatphd··on Shiller PE Ratio
It is an odd position because the P/E and Forward P/E are elevated but not extreme. The bigger warning signal for me is when I see people in public talking about stocks or having stock screens open and this happened almost four times in one week. For me that is basically the sign to risk manage.
digitcatphd··on Launch HN: Slashy (YC S25) – AI that connects to apps and does tasks
I would argue Dropbox was a new product category rather than a feature and as such, was a much deeper strategic decision to enter that category than add a feature. My only recommendation would be to focus on deep complex workflows (E.g. N8N style) with extensive integrations or build out a developer community so you can build some data lock in, because if they are surface level templates surely these will get easily disrupted.
digitcatphd··on Launch HN: Slashy (YC S25) – AI that connects to apps and does tasks
I really hate to be the curmudgeon here but won't foundation models end up having their own AI workflows like the GPT store but with MCP?

I could really envision saving an 'AI Workflow' template with integrated MCP clients that will balloon once adoption is reached. Right now adoption is low so its not a priority for them, once it is, they will tack it on.

I really wish this the best of luck its a great concept, but surely you must be thinking ahead to plan for this situation.

digitcatphd··on MIT Study Finds AI Use Reprograms the Brain, Leading to Cognitive Decline
So users are more detached from their work? How does this correspond with cognitive decline? Wouldn’t it need to be cross referenced in other areas beside the task at hand? Seems a bit of a headline grabbing study to me. Personally I find thinking with an LLM helps me take a more structured and unbiased approach to my thought process
digitcatphd··on Canaries in the Coal Mine? Recent Employment Effects of AI [pdf]
I had a team of developers and essentially told them all 'either learn to code with Claude' or you're out. What I found is the more junior developers started 'vibe coding' resulting in a net decrease in performance, where the more senior ones used it to accelerate their speed cautiously and selectively.

My conclusion was senior engineers were better because they were used to managing developers and taking on more managerial tasks building 'LLM Soft Skills' and also frankly fixing mistakes, the junior developers were pressured for speed and had their managers to correct them.

Within 12 months, despite extensive attempts, only the mid level team members remained.

digitcatphd··on Ask HN: Should primary care doctors be replaced with AI?
Yes, precisely these tools and yes, I did cross-reference with studies. In-fact the physician immediately misattributed the cause, and I had to guide her into self-correcting herself, since she failed to ask conditional followup questions that I forced the LLM to.

In-fact, in another similar case, an almost identical thing happened to my partner. She was experiencing medical issues and asked her friends, who were doctors. They confidently gave a flurry of knee jerk responses of the cause without carefully considering all the variables I forced the LLM to do. With the guidance of the LLM, we took a local diagnostic at a nearby pharmacy, CORRECTED the pharmacist's recommendation who attributed it to 'the heat', and the problem was solved within a few hours after determining it was related to a magnesium deficiency.

I'm not saying it's perfect out of of the box, but I remember getting excited when OCR medical imaging 5 years or so was 85% as effective as a doctor, now its surpassed human performance and for much the same reasons that LLMs are superior.

The primary mistake of the physician, and most, is their knowledge window is limited, as another comment cites. Also, they tend to, based on my observation be more reactionary. If you're not visibly sick then you're not sick and several studies demonstrate the long-term compounding effects of LDL at these levels without side effects for years.

Again, the argument isn't to replace primary care physicians with self-diagnostic ChatGPT usage, but rather than in the near term we will observe a threshold similar to medical image recognition surpassing primary care physicians specifically and at some point we will reach a threshold where physician interaction is in-fact meddlesome.

digitcatphd··on How to build a coding agent
The problem I have with this is that this style of agent design, providing enormous autonomy, makes sense in coding while keeping an expert human in the loop since it can self-correct via debugging. What would the other use cases of giving an agent this much autonomy be today versus a more structured flow versus something more like LangGraph?
digitcatphd··on Is the A.I. Sell-Off the Start of Something Bigger?
After a run up to ATH
digitcatphd··on Review of Anti-Aging Drugs
“ Josh Mitteldorf studies evolutionary theory of aging using computer simulations.”

> This is exactly why one must read the author before going into their content. What does this even mean?

digitcatphd··on Perplexity offers to buy Google Chrome for $34.5B
Also likely a political PR stunt.
digitcatphd··on Perplexity offers to buy Google Chrome for $34.5B
This is clearly a PR stunt. Perplexity knows Google would not sell Chrome, it is the holy grail of their ads strategy and would cripple their moat.
Page 1 of 10Next →