HNHacker News
TopNewBestAskShowJobs

dbreunig

2,877 karma · joined March 24, 2008

submissionscomments
dbreunig··on How I use LLMs to learn complex topics
Wrote about this awhile ago, and it hasn’t changed:

> Time and time again, when talking to people who rely on ChatGPT, Claude, Perplexity, and other general AI tools, I hear them say, “AI is incredible. It handles nearly everything I throw at them.”

> “What does it fumble with?” I’ll ask.

> “Well, it still gets things wrong when it comes to my line of work.”

https://www.dbreunig.com/2025/04/08/on-ai-observational-comi...

dbreunig··on How I use LLMs to learn complex topics
> What you get is a beautiful animation that is 100% accurate and free of hallucinations

How does he know?

dbreunig··on Show HN: I left Figma to build a diffusion-based UI design tool
I got an early tip on this and used it to build a brand mood board and it crushed previous attempts with Claude. Highly recommend.
dbreunig··on Sam Altman's Business Dealings Under GOP Scrutiny Ahead of OpenAI's IPO
Elizabeth Lopatto at The Verge makes a strong case we _do_ have proof that Musk is actively gathering and throwing fuel on the fire: https://www.theverge.com/ai-artificial-intelligence/929129/s...

> But the thing is, Molo doesn’t actually have to be good at this job, because the point of this trial isn’t to win — though I’m sure Musk wouldn’t mind a win. The point is to punish Altman, Brockman, and OpenAI. Musk has done that pretty thoroughly — reinforcing in the public’s mind that Altman is a liar and a snake. This morning, I read an exclusive in The Wall Street Journal that assorted Republican AGs and the House Oversight committee wanted to look into Sam Altman’s investments. References to the trial are peppered throughout the article.

dbreunig··on Accelerating Gemma 4: faster inference with multi-token prediction drafters
Among benchmarkers its a frequent topic. Qwen BURNS reasoning to get its scores.
dbreunig··on If DSPy is so great, why isn't anyone using it?
Model testing and swapping is one of the surprises people really appreciate DSPy for.

You're right: prompts are overfit to models. You can't just change the provider or target and know that you're giving it a fair shake. But if you have eval data and have been using a prompt optimizer with DSPy, you can try models with the one-line change followed by rerunning the prompt optimizer.

Dropbox just published a case study where they talk about this:

> At the same time, this experiment reinforced another benefit of the approach: iteration speed. Although gemma-3-12b was ultimately too weak for our highest-quality production judge paths, DSPy allowed us to reach that conclusion quickly and with measurable evidence. Instead of prolonged debate or manual trial and error, we could test the model directly against our evaluation framework and make a confident decision.

https://dropbox.tech/machine-learning/optimizing-dropbox-das...

dbreunig··on Learnings from a No-Code Lib: Keep the Spec Driven Development Triangle in Sync
No reason it can't. I know people currently generating specs from existing code; just gotta write the pipeline.
dbreunig··on We gave terabytes of CI logs to an LLM
"Think step by step," was just a sentence you appended to your prompt.

It ended up kicking off reasoning training which enabled the massive gains in coding, tool use, and more over the last 18 months.

So yeah, it's "just using LLMs in a specific way."

dbreunig··on Meta’s AI smart glasses and data privacy concerns
Last year they pushed out an update stating if any “Meta AI” is left on, they can access image data for training,

I turned the AI off and used them as headphones and taking videos while biking. After a couple rides, I couldn’t bring myself to put them on because people started to recognize them and I realized I didn’t want to be associated with them (people are right to assume Meta has access to what they see).

Meta Ray Bans, if kept simple, could have been a great product. They ruined them.

dbreunig··on We gave terabytes of CI logs to an LLM
Check out “Recursive Language Models”, or RLMs.

I believe this method works well because it turns a long context problem (hard for LLMs) into a coding and reasoning problem (much better!). You’re leveraging the last 18 months of coding RL by changing you scaffold.

dbreunig··on Why is Claude an Electron app?
Author of the post here.

I didn’t say AI was bad and I acknowledged the benefits of Electron and why it makes sense to choose it.

With 64gb of RAM on my Mac Studio, Claude desktop is still slow! Good Electron apps exist, it’s just an interesting note give recent spec driven development discussion.

dbreunig··on Why is Claude an Electron app?
I keep saying this, it’s my new favorite metaphor.
dbreunig··on A OSS Library with No Code, Only Specs
That's cute.
dbreunig··on Three kinds of AI products work
Agree. I bucket things into three piles:

1. Batch/Pipeline: Processing a ton of things, with no oversight. Document parsing, content moderation, etc.

2. AI Features: An app calls out to an AI-powered function. Grammarly might pass out a document for a summary, a CMS might want to generate tags for a post, etc.

3. Agents: AI manages the control flow.

So much of discussion online is heavily focused towards agents so that skews the macro view, but these patterns are pretty distinct.

dbreunig··on The World's 2.75B Buildings
There was a good study on this a few years ago that ran the numbers on this and landed on white paint for residential homes as the best option, for a few reasons, if I remember correctly:

- Installation, maintenance and transmission costs are lower when solar is aggregated on farms - Solar offsets air conditioning, but that moves the heat outside. White roofs reduce the need for AC, which helps significantly with urban heat scenarios

A quick search yields a UCL study, which supports the lower claim: https://phys.org/news/2024-07-roofs-white-city.html

dbreunig··on DeepSeek writes less secure code for groups China disfavors?
Yes, if you put unrelated stuff in the prompt you can get different results.

One team at Harvard found mentioning you're a Philadelphia Eagles Fan let you bypass ChatGPT alignment: https://www.dbreunig.com/2025/05/21/chatgpt-heard-about-eagl...

dbreunig··on We're Joining OpenAI
Yeah, I agree. Almost mentioned in the post how I imagine an ad PM at OpenAI is jealous of an ad PM at Perplexity.
dbreunig··on The AI Job Title Decoder Ring
I also dislike the term. It feels concocted to evoke “tacticool” vibes.

Unless you’re pushing new firmware onto a drone in Ukraine, FDE is stolen valor.

dbreunig··on The AI Job Title Decoder Ring
You should read the post. You might find the “domain” discussion interesting.
dbreunig··on The AI Job Title Decoder Ring
I will be thinking about this comment for a bit. Thanks for this perspective!
dbreunig··on Claude Sonnet 4 now supports 1M tokens of context
The team at Chroma is currently looking into this and should have some figures.
dbreunig··on The surprise deprecation of GPT-4o for ChatGPT consumers
I’m wondering that too. I think better routers will allow for more efficiency (a good thing!) at the cost of giving up control.

I think OpenAI attempted to mitigate this shift with the modes and tones they introduced, but there’s always going to be a slice that’s unaddressed. (For example, I’d still use dalle 2 if I could.)

dbreunig··on FLUX.1-Krea and the Rise of Opinionated Models
I’m aware of LoRA, Civitai, etc. I don’t think they are “widely known” beyond AI imagery enthusiasts.

Krea wrote a great post, trained the opinions in during post-training (not during LoRA), and I’ve been noticing larger labs doing similar things without discussing it (the default ChatGPT comic strip is one example). So I figured I’d write it up for a more general audience and ask if this is the direction we’ll go for qualitative tasks beyond imagery.

Plus, fine-tuning is called out in the post.

dbreunig··on An LLM does not need to understand MCP
Came here to say this: people present MCP’s verbosity as all the context the LLM needs. But almost always, this isn’t the case.

I wrote recently, “ Connecting your model to random MCPs and then giving it a task is like giving someone a drill and teaching them how it works, then asking them to fix your sink. Is the drill relevant in this scenario? If it’s not, why was it given to me? It’s a classic case of context confusion.”

https://www.dbreunig.com/2025/07/30/how-kimi-was-post-traine...

dbreunig··on Ask HN: Best AI Automation Platform
I really like Relay.app for non-coders. People can get wrapped around the wheel with n8n and co.
dbreunig··on Irrelevant facts about cats added to math problems increase LLM errors by 300%
A similar, fun case is where researchers inserted facts about the user (gender, age, sports fandom) and found alignment rules were inconsistently applied: https://www.dbreunig.com/2025/05/21/chatgpt-heard-about-eagl...
dbreunig··on Irrelevant facts about cats added to math problems increase LLM errors by 300%
Wrote about this about a month ago. I think it’s fascinating how they developed these prompts: https://www.dbreunig.com/2025/07/05/cat-facts-cause-context-...
dbreunig··on The new skill in AI is not prompting, it's context engineering
I studied linguistic anthropology, in addition to CS. Been at it since 2002.

And I wrote the first post before the meme.

dbreunig··on The new skill in AI is not prompting, it's context engineering
Sometimes buzzwords turn out to be mirages that disappear in a few weeks, but often they stick around.

I find they takeoff when someone crystallizes something many people are thinking about internally, and don’t realize everyone else is having similar thoughts. In this example, I think the way agent and app builders are wrestling with LLMs is fundamentally different than chatbots users (it’s closer to programming), and this phrase resonates with that crowd.

Here’s an earlier write up on buzzwords: https://www.dbreunig.com/2020/02/28/how-to-build-a-buzzword....

dbreunig··on The new skill in AI is not prompting, it's context engineering
While researching the above posts Simon linked, I was struck by how many of these techniques came from the pre-ChatGPT era. NLP researchers have been dealing with this for awhile.
Page 1 of 9Next →