HNHacker News
TopNewBestAskShowJobs

gdudeman

660 karma · joined May 4, 2010

submissionscomments
gdudeman··on Anthropic appears to be A/B testing reduced effort levels in Claude Code
There is so obviously competition in this market it’s astounding to me that people attribute all this malicious behavior to the model companies.

They’re growing over 10x a year. They want users and revenue. In order to get users and revenue, they want to provide the smartest models at affordable prices. If they unnecessarily burn tokens, users will get less value and switch.

This thread is filled with competing comments about their monopolistic power and how when one model provider was no longer doing a good job people switched to a different one.

The competition in this market is ferocious!

gdudeman··on GPT 5.6 Sol is the best "vision" model OpenAI ever released
In my experience starting with Gemini 2.5 Pro, moving to 3 and 3.1, 3.5 Flash, 3.6 Flash, and finally 3.7 Flash, 3.7 Flash is just as good if not better than 3 especially on high resolution mode (same token count per page as 3.1).

I run complicated, messy PDFs through these models. 2.5 Pro required a lot of kludgy hacks to get it to fully "see," but from 3.1 pro on I've removed many of them and haven't spotted problems.

3.7 Flash scores better than 3.1 pro on most benchmarks, leading me to believe that even if your OCR requires reasoning to interpret text or data, 3.7 Flash is probably going to be better.

gdudeman··on Claude Opus 5
Yes. I have Claude route to codex all the time as part of a dev process where fable is manager and it oversees the work of sidekicks and subagents. It’s happy to comply.
gdudeman··on The shingles vaccine may reduce the risk of dementia
Sorry - the older group has higher rates of dementia than the younger group when they reach the same age - so when they reach 75, they're more likely to have dementia.
gdudeman··on The shingles vaccine may reduce the risk of dementia
I believe that there are studies that show that merely getting very sick increases your chance of dementia - essentially it ages you faster or brings chronic disease forward. If that's the case, vaccines for things like the flu - a disease you're likely to get - are probably good overall.

They also activate the immune system in generaly, which could probably go either way in terms of longevity.

In general I don't think vaccines are preventing so much as delaying dementia, but if they stop chronic infection they might be.

gdudeman··on The shingles vaccine may reduce the risk of dementia
If you haven't seen the chart from the UK study, I highly recommend checking it out: https://erictopol.substack.com/p/the-shingles-vaccine-and-re...

When the vaccine came out in the UK, they included a hard age cutoff: Above a certain age, you weren't eligible. Below that age, you were eligible.

They looked at the probability of a dementia diagnosis over the 7 years after the vaccine was introduced.

People who were born in the "can get the vaccine" group have markedly lower rates of dementia. People in the "too old" group have higher rates. It's cut and dry. The researchers didn't separate out the people who actually got the vaccine.

It's one of those studies where you don't even need to look at the p-value to see the difference between the cohorts.

gdudeman··on The shingles vaccine may reduce the risk of dementia
People who are immune compromised qualify for this. I'm not sure how much your insurance company would push back on that though.
gdudeman··on The shingles vaccine may reduce the risk of dementia
This really isn't the case in the Shingles vaccine unless the UK study is flawed a way that isn't clear to me.

The study looked at the effect of being eligible for a vaccine and the results were clear. (see chart below the fold here: https://erictopol.substack.com/p/the-shingles-vaccine-and-re...)

There was a hard age cutoff in the UK study. Above a certain age, you weren't eligible. Below it, you were. People who were born in the "can get the vaccine" group have markedly lower rates of dementia. People in the "too old" group have higher rates. It's one of those studies where you don't even need to look at the p-value to see the difference.

I'm very open to being wrong about this!

gdudeman··on Computer use in Gemini 3.5 Flash
Computer use is a great idea. It gets the job done when nothing else will.

If you're a person trying to get their job done at a big company, but half your job is in 1-2 proprietary tools or is stuck behind an API you can't program against, computer use can allow you, a non-techie, to do your job more efficiently.

I think it's an awesome way to circumvent gate keepers and the IT department to let people accomplish their goals.

gdudeman··on A walking tour of surveillance infrastructure in Seattle (2020)
For anyone trying to figure out how to build a society where no one wants to be a criminal, I highly recommend When Brute Force Fails: How to Have Less Crime and Less Punishment by Mark Kleiman.

There are evidence-backed ways of reducing criminality.

One counterintuitive way of reducing crime is to increase the likelihood of being caught, to have small-but-increasing consequences for committing crimes, and to increase the swiftness of sentencing.

For example, if you are caught drinking and driving, you immediately spend 1-2 days in jail.

Long sentences are not very productive at reducing crime or at least are a very inefficient way to do so.

gdudeman··on Sleep research led to a new sleep apnea drug
Curious how you tested this - did you do repeated home tests to figure this out?
gdudeman··on Tesla's lithium refinery discharges 231,000 gallons of polluted wastewater a day
I think it does. Crops pull up a set amount of water. If it's concentrated, then they'll pull up a lot of heavy metals. If it's at very low levels, then they won't.
gdudeman··on Tesla's lithium refinery discharges 231,000 gallons of polluted wastewater a day
Is there an alternative?

We live in a much, much cleaner world than we did 50 years ago. Legislation and environmental rules have worked. There are some areas where it could obviously be better, but also some areas where regulation is too strict (blocking housing, renewables, transit) and the system is evolving to address those.

I think the loss of local media has made it harder for misdeeds to come to light, but I don't want to throw up my hands and cede everything to commercial interests et al.

gdudeman··on Tesla's lithium refinery discharges 231,000 gallons of polluted wastewater a day
Then we should work together as a society to fix the law and make sure it's applied evenly. Hard to do, but necessary.

Is there an alternative?

gdudeman··on Anthropic's Prompt Engineering Tutorial (2024)
This is written for the 3 models (Sonnet, Haiku, Opus 3). While some lessons will be relevant today, others will not be useful or necessary on smarter, RL’d models like Sonnet 4.5.

> Note: This tutorial uses our smallest, fastest, and cheapest model, Claude 3 Haiku. Anthropic has two other models, Claude 3 Sonnet and Claude 3 Opus, which are more intelligent than Haiku, with Opus being the most intelligent.

gdudeman··on Claude Code 2.0
Ooo - good question. I'm unsure on this one.
gdudeman··on Claude Code 2.0
This enables the /resume command that lets you start mid-conversation again.

Storing the data is not the same as stealing. It's helpful for many use cases.

I suppose they should have a way to delete conversations though.

gdudeman··on Claude Code 2.0
Resume is now a drop down menu at the top in the new VS Code plugin and it's much easier to read.
gdudeman··on Claude Code 2.0
> New native VS Code extension

Looks great, but it's kind of buggy:

- I can't figure out how to toggle thinking

- Have to click in the text box to write, not just anywhere in the Claude panel

- Have to click to reject edits

gdudeman··on Sampling and structured outputs in LLMs
I was burned by this for a while because I assumed structured output ordering would be preserved.

You can specify ordering in the Gemini API with propertyOrdering:

"propertyOrdering": ["recipeName", "ingredients"]

gdudeman··on Fartscroll-Lid: An app that plays fart sounds when opening or closing a MacBook
Even farther off topic, but this reminds me of the time my friends and I recorded a 3 minute long wav file that ended with a quiet “this is god. Can you hear me? I’d like to talk with you,” and set it to be the error sound on a friend’s PC.

Much hilarity ensued.

gdudeman··on Claude Sonnet 4 now supports 1M tokens of context
You don't say that - you instruct the LLM to read files about X, Y, and Z. Putting the context in helps the agent plan better (next step) and write correct code (final step).

If you're asking the agent to do chunks of work, this will get better results than asking it to blindly go forth and do work. Anthropic's best practices guide says as much.

If you're asking the agent to create one method that accomplishes X, this isn't useful.

gdudeman··on Claude says “You're absolutely right!” about everything
Warning: A natural response to this is to instruct Claude not to do this in the CLAUDE.md file, but you’re then polluting the context and distracting it from its primary job.

If you watch its thinking, you will see references to these instructions instead of to the task at hand.

It’s akin telling an employee that they can never say certain words. They’re inevitably going to be worse at their job.

gdudeman··on Claude Sonnet 4 now supports 1M tokens of context
I'll note this saves a lot of wait time as well! No sitting there while a new Claude builds context from scratch.
gdudeman··on Claude Sonnet 4 now supports 1M tokens of context
Claude is very agreeable and is an eager helper.

It gives you the benefit of the doubt if you're coding.

It also gives you the benefit of the doubt if you're looking for feedback on your developers work. If you give it a hint of distrust "my developer says they completed this, can you check and make sure, give them feedback....?" Claude will look out for you.

gdudeman··on Claude Sonnet 4 now supports 1M tokens of context
With an empty text box, double escape shows you a list of previous inputs from you. You can go back and fork at any one of those.
gdudeman··on Claude Sonnet 4 now supports 1M tokens of context
Git or my favorite "Undo all of those changes."
gdudeman··on Claude Sonnet 4 now supports 1M tokens of context
A tip for those who both use Claude Code and are worried about token use (which you should be if you're stuffing 400k tokens into context even if you're on 20x Max):

  1. Build context for the work you're doing. Put lots of your codebase into the context window.
  2. Do work, but at each logical stopping point hit double escape to rewind to the context-filled checkpoint. You do not spend those tokens to rewind to that point.
  3. Tell Claude your developer finished XYZ, have it read it into context and give high level and low level feedback (Claude will find more problems with your developer's work than with yours).
If you want to have multiple chats running, use /resume and pull up the same thread. Hit double escape to the point where Claude has rich context, but has not started down a specific rabbit hole.
gdudeman··on Claude Sonnet 4 now supports 1M tokens of context
Yes: type /model and then pick Opus 4.1.
gdudeman··on Do LLMs identify fonts?
Missing from the methodology: - was thinking on or off? (At least for Gemini) - was web search allowed? - was tool use allowed?

It’s quite likely LLMs don’t “know” the fonts in the dataset, but they could figure many of them out.

Page 1 of 5Next →