HNHacker News
TopNewBestAskShowJobs

supermdguy

1,293 karma · joined February 20, 2017

Github: https://github.com/supermdguy LinkedIn: https://www.linkedin.com/in/matthew-dangerfield/

Authentication Stuff: hnchat.com:0m9YzMwf4kdLeEnrHt2n 547e7b5026bc47ac8302bf3d213909be [ my public key: https://keybase.io/supermdguy; my proof: https://keybase.io/supermdguy/sigs/fHS654nzPCSpfpoOPNrHUJrgc2jOy21zNsnFwyeLDfs ]

submissionscomments
supermdguy··on There's no point at which turning your brain off will work
I've noticed some tasks (80-90%) can be accomplished by quickly typing out a feature/bug fix, accepting plans with ~10% brain power, and moving on. But there are other tasks where the LLMs struggle, and it requires significantly more time and mental effort to think through the architecture and requirements before it's to a point where the LLMs can process it.

What's difficult, is I now spend most of my time on these high-throughput tasks, where my brain isn't exactly off, but focused on quantity over quality. It's often difficult to identify when a task requires "real" brain-power, and when it does, it adds a lot of friction to dig into the code and orient myself. So it's hard to know when it's the right time to do that.

supermdguy··on Matt Mullenweg tells Automattic staff in Slack he's back in control after ouster
Probably has an HDR layer, kind of surprised social networks let that happen since it's really annoying imo. There was a discussion about it here a couple months ago: https://news.ycombinator.com/item?id=49402521.
supermdguy··on How accurate have Ed Zitron's AI skeptic predictions been?
Wait are the hyperscalers booking unrealized gains as income? Or are they selling their positions?
supermdguy··on Path to Astra: critical capabilities and frontier safeguards
> As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities.

Funny to read this in the wake of the HuggingFace hack. I'm sure this is based on a clean run, but I can't help thinking PHASEONE[big] would be proud.

supermdguy··on Claude Fable 5.1 and Claude Mythos 5.1
They originally released it at a "temporary discounted price", then made it permanent (probably due to competitive pressure). It's still way more expensive per task, due to tokenizer changes and general verbosity.
supermdguy··on Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
Yes! I was also reminded of the truth mines in Diaspora.
supermdguy··on Tell HN: Cloudflare silently injects its analytics when you switch nameservers
Cloudflare sets up a reverse proxy as part of their core offering, so by default they can MITM your proxy. The “orange cloud” by a DNS record means it points to their proxy instead of your server.
supermdguy··on Meta releases Muse Glimmer, 30B parameter open-weight edge model
Dupe: https://news.ycombinator.com/item?id=49241679
supermdguy··on Timeline of the OpenAI accidental attack against Hugging Face
Here's the performance of frontier models without reasoning, to more directly address the claim that raw performance is plateauing:

https://artificialanalysis.ai/evaluations/artificial-analysi...

I don't have any insider info, but if model sizes actually have increased exponentially since GPT 4.1, there's an argument to be made that there are diminishing returns in scaling pretraining alone.

Also interesting thing I haven't noticed before, Opus models have followed a really consistent linear improvement, while it looks like OpenAI struggled with base model performance until 5.5/5.6 (EDIT - 5.5 was their first new pretraining run in over a year).

supermdguy··on Prime Agent: A self-improving RLM agent
Interesting, what did you use for the data? And do you have a write-up anywhere?
supermdguy··on Prime Agent: A self-improving RLM agent
It'll be really interesting when they run RL training on the harness self-improvement loop. I've tried using LLMs for harness engineering, but it often creates too much bloat that weighs things down in the end. Guessing it's just not something the models are tuned to do by default.

Curious if anyone's tried using RL for harness engineering? I think we're still pretty far away from the optimal harness, especially when it comes to long-context memory management.

supermdguy··on qm – Multiplayer agent harness for work
It has read only access, so it’s able to check query statistics, find slow queries, then run EXPLAIN ANALYZE to find the root cause and either tweak the query or suggest indexes. A lot of it is low hanging fruit, I just haven’t put in the time to fix it (startup).
supermdguy··on qm – Multiplayer agent harness for work
Still figuring it out, but it's been really convenient to have an always-on agent that has access to internal systems and can be triggered by webhooks. Some examples of what we use it for:

- automatically fixing simple CI failures

- getting production alerts and automatically creating RCAs and a fix PR

- periodically checking slow DB queries and finding ways to speed them up.

- creating charts to answer one-off questions about our data

I've tried using it as an on-the-go coding agent as well, but found I prefer more interactive agents, so I can see what the code looks like.

supermdguy··on OpenAI’s accidental attack against Hugging Face is science fiction that happened
It would require collusion with HuggingFace, including getting them to release their disclosure blogpost a week in advance. Huggingface is primarily a hub for open models, so there's not really an incentive for them to jump through hoops/lie to provide marketing for OpenAI's (closed) models. So it's highly unlikely this is fabricated.
supermdguy··on No link between acetaminophen use during pregnancy and adverse birth outcomes
The administration promised they'd find a cause by September 2025. I guess that was the best they had at that point?

https://www.cbsnews.com/news/rfk-jr-cause-of-autism-research...

supermdguy··on Why write code in 2026
“Do you know what the industry term for a project specification that is comprehensive and precise enough to generate a program?

Code. It’s called code.”

- CommitStrip (https://www.reddit.com/r/ProgrammerHumor/comments/1p70bk8/sp...)

I think if you’re doing it right, the core of your code should be the simplest expression of the underlying business logic. Of course there’s always going to be supporting layers, and maybe those don’t need to be reviewed. But if you haven’t read the code, there’s an extent to which you don’t know the business logic.

supermdguy··on Previewing GPT‑5.6 Sol: a next-generation model
> We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed.

This is really exciting. I work on voice AI, and we're still using 4.1/4.1 mini since none of the frontier models come close on latency. I'm excited to be able to have more interactive experiences, I think it'll unlock new ways of working with these models.

supermdguy··on How many of the 170k English words do you know?
Agreed, there were also a few where I deduced the correct definition by comparing the options.
supermdguy··on .gitignore Isn't the only way to ignore files in Git
One trick I’ve used is creating a folder and then adding a .gitignore inside it with *. Then nothing in that folder gets tracked, without needing to add anything to the public gitignore. Didn’t know about .git/config though!
supermdguy··on Why LLMs still lack taste
I agree that current memory systems are pretty bad, and I think that’s because memory is a prompted behavior instead of a learned one. In theory, if memory was an emergent behavior instead of a prompted one, it would be a lot better.

I think you’re right that changing its own harness would be bad and skew towards prompt engineering instead of learned behavior. So maybe instead it could start with a harness with memory CRUD tools and then learn how to most effectively use them.

supermdguy··on Why LLMs still lack taste
I think good taste is objective to an extent, especially given a particular production context.
supermdguy··on Meta's ships facial recognition on smart glasses
Someone posted a tool that does this recently: https://news.ycombinator.com/item?id=47140042
supermdguy··on The Public Should Own Half of the Big A.I. Companies
There was a show HN for something similar a couple months ago[0]. Looks like they shut it down. Probably too difficult/low-margin to run as a business, but I think the co-op model you mentioned has potential.

[0]: https://news.ycombinator.com/item?id=47639779

supermdguy··on MAI-Thinking-1
It's interesting because their last model series (Phi) was based around the thesis that high-quality synthetic data is better than a large pre-training corpus.
supermdguy··on 2026 HIPAA Security Rule Update
I think the conduit exception still applies for analog faxes. Which makes no sense, since tapping a fax line is probably way easier than compromising a data center.
supermdguy··on Temu is advertising filet mignon on X
Apparently from a third party seller in New York....and tastes really bad. I was surprised steak could be safely mailed in such normal looking packaging!

https://www.delish.com/food/a70539084/temu-meat-review/

supermdguy··on Local AI needs to be the norm
Interesting to see this after the recent post about Chrome’s on-device model using up 4gb of storage, which frustrated a lot of people [1].

I agree local models are great, and it’s cool that Apple has models built in now. But I feel like it basically has to be an OS level feature or users are going to get upset. I’d certainly rather have a small utility call out to OpenAI than download its own model.

[1]: https://news.ycombinator.com/item?id=48019219

supermdguy··on From Supabase to Clerk to Better Auth
Better auth is great! I love how it's way more hackable than using a something like Clerk. We were able to add a plugin to allow auth via iframe postMessage (embedded in a CRM) and everything worked seamlessly.
supermdguy··on Agent Skills
I've used it off and on over the last month or so. For more complicated tasks (30+ minutes) it works well, and seems to replace a lot of prompting that I'd normally need to do (e.g. asking questions about requirements, creating specs and implementation plans, staying on task). For simple tasks, it tries to do too much and gets in the way.
supermdguy··on Apocalypse Early Warning System
There’s a really good analysis here: https://www.lesswrong.com/posts/LBC2TnHK8cZAimdWF/will-jesus...

> The Yes people are betting that, later this year, their counterparties (the No betters) will want cash (to bet on other markets), and so will sell out of their No positions at a higher price.

Page 1 of 11Next →