HNHacker News
TopNewBestAskShowJobs

supermdguy

1,444 karma · joined February 20, 2017

Github: https://github.com/supermdguy LinkedIn: https://www.linkedin.com/in/matthew-dangerfield/

Authentication Stuff: hnchat.com:0m9YzMwf4kdLeEnrHt2n 547e7b5026bc47ac8302bf3d213909be [ my public key: https://keybase.io/supermdguy; my proof: https://keybase.io/supermdguy/sigs/fHS654nzPCSpfpoOPNrHUJrgc2jOy21zNsnFwyeLDfs ]

submissionscomments
supermdguy··on Apocalypse Early Warning System
There’s a really good analysis here: https://www.lesswrong.com/posts/LBC2TnHK8cZAimdWF/will-jesus...

> The Yes people are betting that, later this year, their counterparties (the No betters) will want cash (to bet on other markets), and so will sell out of their No positions at a higher price.

supermdguy··on Looking for feedback on a paper about a revision-capable language model [pdf]
Overall, I'm really impressed by what you accomplished! I'm not a researcher, so not sure if this is that helpful, but here are some thoughts:

- I wonder if the "move" action is difficult for the model to learn to use well. The model sees token location as positional encodings in the embedding, not sparse character offsets. Would be interesting to see something more like "jump to next/previous [token or set of tokens]". Or maybe a find/replace like most coding harness edit tools use?

- I'd move the exact training data generation details to an appendix. Could be summarized to improve the flow of the paper.

supermdguy··on You have to rank 5 projects before you can post your own
Cool concept! I think the hardest part will be getting people in the target audience to use it. A lot of indie hackers make software for other indie hackers, but that isn't true of most other verticals. And honestly building software for indie hackers feels like a losing battle. Any ideas of how to incentivize none-builders to rank projects?
supermdguy··on Does Gas Town 'steal' usage from users' LLM credits to improve itself?
From the most recent comment, looks like this is a bug, triggered by the system inadvertently activating an internal release tool [0]. Still a pretty wild bug, but not as dramatic as the title suggests. Which is kind of unfortunate honestly, the chaos of every gas town instance automatically contributing to itself would be beautiful to see.

- https://github.com/gastownhall/gastown/blob/main/internal/fo...

supermdguy··on All elementary functions from a single binary operator
Next step is to build an analog scientific calculator with only EML gates
supermdguy··on Anthropic downgraded cache TTL on March 6th
Bizarre reading the thread, it feels like their Claude responding to the other posters’ Claudes
supermdguy··on Will I ever own a zettaflop?
If all LLM advancements stopped today, but compute + energy got to the price where the $30 million zettaflop was possible, I wonder what outcomes would be possible? Would 1000 claudes be able to coordinate in meaningful ways? How much human intervention would be needed?
supermdguy··on Muse Spark: Scaling towards personal superintelligence
And also OpenAI’s codex spark?
supermdguy··on OpenAI Codex Moves to API Usage-Based Pricing for All Users
Headline/article is extremely misleading. They still have subscription plans with included usage, but those usage limits are now based on tokens instead of messages.

https://help.openai.com/en/articles/20001106-codex-rate-card

supermdguy··on "Good Taste" Is Just Experience
I like this, and think it's true for how humans learn. What's interesting to me is that it seems LLMs are significantly smarter than they were two years ago, but it doesn't feel like they have better "taste". Their failure modes are still bizarre and inhuman. I wonder what it is about their architecture/training that scales their experience without corresponding improvements in taste.

In theory, RLVR should encourage less error-prone code, similar to a human getting burned by production outages like the article mentioned. Maybe the scale in training just isn't big enough for that to matter? Perhaps we need better benchmarks that capture long-term issues that arise from bad models and unnecessary complexity.

supermdguy··on Does Syntax Matter?
Correct URL: https://www.gingerbill.org/article/2026/02/21/does-syntax-ma...
supermdguy··on AI isn't killing jobs, it's 'unbundling' them into lower-paid chunks
I’ve tried having one “big” task that I’m focusing on with active back and forth while letting other Claude instances handle easier back-burner type tasks that it can effectively one-shot. But I’ve noticed that often turns into me spending more time/focus than I’d want on tasks that aren’t actually that impactful. I still think I get more done than I would otherwise, but I still haven’t found the best management strategy.
supermdguy··on [dead]
Yeah that confused me, but the compression paper also doesn’t make a ton of sense since I doubt Google would have released it if it was actually such a competitive advantage compared to what other labs are doing. So I wonder what’s actually causing the price decrease.
supermdguy··on [dead]
Any other sources on the OpenAI claim? Regardless, it’ll be nice to have cheaper RAM
supermdguy··on Show HN: Git bayesect – Bayesian Git bisection for non-deterministic bugs
Okay this is really fun and mathematically satisfying. Could even be useful for tough bugs that are technically deterministic, but you might not have precise reproduction steps.

Does it support running a test multiple times to get a probability for a single commit instead of just pass/fail? I guess you’d also need to take into account the number of trials to update the Beta properly.

supermdguy··on HyperAgents: Self-referential self-improving agents
It's surprising that this works so well considering that AI-generated AGENTS.md files have been shown to be not very useful. I think the key difference here is that the real-world experience helps the agent reach regions of its latent space that wouldn't occur naturally through autoregression.

I wonder how much of the improvement is due to the agent actually learning new things vs. reaching parts of its latent space that enable it to recall things it already knows. Did the agent come up with novel RL reward design protocols based on trial and error? Or did the tokens in the environment cause it to "act smarter"?

supermdguy··on China is mass-producing hypersonic missiles for $99,000
> Nobody knows yet the true capabilities of the missile, but it doesn’t matter. The accuracy doesn’t matter very much, the payload doesn’t matter very much. If it’s launched at a certain target in Tel Aviv, it still is going to hit something in Tel Aviv. The Israelis have no choice but to attempt an intercept, and will spend millions to do so

Sounds like the massive price disparity more than makes up for any accuracy issues

supermdguy··on If you thought code writing speed was your problem you have bigger problems
> Experimenting with ideas/refactors to see how they'll play out (often the agent can just tell you how it's going to play out)

This has helped me a lot. Normally I'd feel really attached to big refactors because of sunk costs, but when AI does a huge refactor it's easier to honestly decide that it wasn't worth it and unnecessarily increased complexity.

supermdguy··on Tree Search Distillation for Language Models Using PPO
> One might note that MCTS uses more inference compute on a per-sample basis than GRPO: of course it performs better

This part confused me, it sounded like they were only doing the MCTS at train time, and then using GRPO to distill the MCTS policy into the model weights. So wouldn’t the model still have the same inference cost?

supermdguy··on Cloudflare crawl endpoint
I’ve actually written a crawler like that before, and still ended up going with Firecrawl for a more recent project. There’s just so many headaches at scale: OOMs from heavy pages, proxies for sites that block cloud IPs, handling nested iframes, etc.
supermdguy··on Tell HN: I'm 60 years old. Claude Code has re-ignited a passion
I’ve been trying to learn a lot about domain driven design, I think knowledge crunching will be a huge part of the new software development role.
supermdguy··on Two kinds of error
I'd also add errors with third-party systems, which aren't the developer's or the user's fault, but which are probably worth handling nicely (e.g. retry with backoff).
supermdguy··on Verified Spec-Driven Development (VSDD)
Probably referencing this: https://news.ycombinator.com/item?id=47034087
supermdguy··on Sandboxes won't save you from OpenClaw
Yeah, OpenClaw agents have a full set of tools to interact with a browser in arbitrary ways. My idea was to instead give it a tool for a browser wrapper with a limited API surface. And that tool could use LLMs internally in specific contexts.
supermdguy··on Sandboxes won't save you from OpenClaw
I like simonw's definition: "An LLM agent runs tools in a loop to achieve a goal."

I guess agent isn't the best term here since the LLM wouldn't be driving the logic in the daemon. Using an LLM to select which item to add to the cart would mimic the behavior of full agentic loop without the risk of it going off the rails and completing the purchase.

supermdguy··on Sandboxes won't save you from OpenClaw
One promising direction is building abstraction layers to sandbox individual tools, even those that don't have an API already. For example, you could build/vibe code a daemon that takes RPC calls to open Amazon in a browser, search for an item, and add it to your cart. You could even let that be partially "agentic" (e.g. an LLM takes in a list of search results, and selects the one to add to cart).

If you let OpenClaw access the daemon, sure it could still get prompt injected to add a bunch of things to your cart, but if the daemon is properly segmented from the OpenClaw user, you should be pretty safe from getting prompt injected to purchase something.

supermdguy··on The only moat left is money?
I think there's still value in building quality products, but AI makes it easy to build something that appears good but doesn't actually work that well. It's very difficult to communicate the thought and intentionality that went into a well-designed product in a way that stands out amongst the noise.
supermdguy··on PsiACE/Skills – A small, shared skill library
Has anyone had success using skills like these without an agent that supports them? I’m using the Zed agent, which doesn’t have skills support, and was thinking of just adding in a summary of the skills directory and how to use it inside my AGENTS.md.
supermdguy··on Murder-suicide case shows OpenAI selectively hides data after users die
Looks like this would affect around 4.3% of chats (the "Self-Expression" category from this report[0]). Considering ChatGPT's userbase, that's an extremely large number of people, but less significant than I thought based on all the talk about AI companionship. That being said though, a similar crowd was pretty upset when OpenAI removed 4o, and the backlash was enough for them to bring it back.

[0]: https://www.nber.org/system/files/working_papers/w34255/w342...

supermdguy··on The compiler is your best friend
That's a good point, thinking about it some more, I think the business logic feels so trivial that it would make the code harder to reason about if it were separated from the effects. Currently, I have one giant function that pulls data, filters it, conditionally pulls more data, and then maybe has one line of effectful code.

I could have one function that pulls the wallet balance for all users, and then passes it to a pure function that returns an object with flags for each user indicating what action to take. Then another function would execute the effects based on the returned flags (kind of like the example you gave of processing a pending charges table).

The value of that level of abstraction is less clear though. Maybe better testability? But it's hard to justify what would essentially be tripling the lines of code (one function to pull the data, one pure function to compute actions, one function to execute actions).

Additionally, there's a performance cost to pulling all relevant data, instead of being able to progressively filter the data in different ways depending on partial results (example: computing charges for all users at once and then passing it to a pure function that only bills customers whose billing date is today).

Would be great to see some more complex examples of "functional core imperative shell" to see what it looks like in real-world applications, since I'm guessing the refactoring I have in my head is a naive way to do it.

← PreviousPage 2 of 11Next →