HNHacker News
TopNewBestAskShowJobs

dekhn

33,458 karma · joined January 12, 2013

submissionscomments
dekhn··on Improper redaction reveals Google Data Center water and electricity usage
The big hyperscalers always knew that eventually the tide would turn against them (for example when local towns noticed that the DCs didn't actually need very many locals to run the facilities).

They carefully hid things so they could expand for as long as possible, knowing that having the necessary footprint for services, cloud, and AI would be key to winning the race. When I worked at Google, our leaders were always "proud" they had built such a huge fleet (originally for search + ads but rapidly growing to a wide range of products and research) without competitors knowing just how large, because they figured eventually AWS and MSFT would realize they had to build out massively to handle the coming loads.

So they hide these things to stave off complaints for as long as possible.

dekhn··on Improper redaction reveals Google Data Center water and electricity usage
Different state, but: almond trees growing in California consume 1 trillion gallons a year (estimated).

So probably we could grow fewer almonds in California.

dekhn··on Automating my 35mm film scanning pipeline
It's not a hill I want to die on either, ultimately the moderators run the site and make the rules and we comply with them at their pleasure.

dang believes that the "transcription and formatting" adds a signature. I entered into Gemini Pro (at the highest level of reasoning) and it calls it human written text and provides a far more detailed description of why: "If an LLM was used here, it was likely just used as a grammar checker or "polisher" rather than a generator. If we had to flag anything, it would be this section:... .his reads like a standard, high-quality comment from a technical forum like Hacker News or Reddit's r/analog. If the user utilized an LLM, they likely wrote the entire comment themselves and merely passed it through a tool like Grammarly or ChatGPT with a prompt like "Fix my typos" without letting it change their voice or vocabulary."

I tried ChatGPT and Claude (all the most recent models I have access to) and they all agree: human text, slightly modified to be more readable.

Looking at Pangram's output, it only mentioned the last sentence as AI generated:

"Would you recommend it, and roughly what did you pay? Have you tried camera scanning? I’ve been considering automating film advance and capture on my setup, but consistent color conversion is the part I’m least confident about."

That's absolutely a sentence a human expert might write. Even if it wasn't, the content contributes to the discussion. Seriously, why would grammar checkers for voice dictation even be remotely an issue?

dekhn··on Automating my 35mm film scanning pipeline
I ended up discussing it with the moderators and they interpret "transcribe and formatting" as contributing enough that they want to "educate" the user to not do that.

In my mind, using an LLM to lightly format a legitimate questions falls below any threshold I'd apply to AI-written writing. I often dictate HN messages to my phone, and in my case, personally edit the message (tediously), but if there was a tool that produced text that looked lightly LLM, I wouldn't care about it at all.

It's not a big deal because in five years, the idea of chiding a user for using an LLM to enhance their text is going to be considered ridiculous. I've lived through enough computer revolutions to know this; once, I didn't get a job (in ~1989) because I could type 80wpm but made a bunch of errors... on a typewriter! The agency told me they needed people who could type 80wpm on a typewriter.

dekhn··on Treachery in the Rodin Museum 3D scan verdict
I used to want programmatic law for exactly the reason that I could determine what was illegal so I didn't break the law.

After a while and talking to smart people I became convinced that was impossible; instead, we have a judicial system filled with experts who make heuristic decisions, and (ideally) it's biased towards not finding people guilty of breaking complex laws they couldn't have figured out. Life requires flexible thinking.

dekhn··on Automating my 35mm film scanning pipeline
It wasn't a generated comment at all! Read the comment, and then what the author wrote elsewhere saying they spoke it into their phone. https://news.ycombinator.com/item?id=49948614

"To reiterate, those were legit questions. I used voice dictation to my phone, with an app that transcribed and formatted my message."

harassing a user for doing that is just unncessary and I dont't think it violates the site's guidelines. Further, the comment itself was high quality, and contributed to the discussion.

dekhn··on Automating my 35mm film scanning pipeline
Who cares? The comment stands on its merit.
dekhn··on Automating my 35mm film scanning pipeline
Dang, would you reconsider your message here? I think it was uncalled for and chilling to legitmate discourse.
dekhn··on Automating my 35mm film scanning pipeline
It is a very sensible and reasonable comment- literally all questions I'd want to ask to replicate the setup, in a format that's easy to parse. If the author did use LLM it did not substantially alter the content of their question.
dekhn··on Treachery in the Rodin Museum 3D scan verdict
Are laws expected to be completely self and cross consistent?

I wanted programmatic law in the past and then after thinking and talking a bit, concluded that self and cross consistency in the law is not considered necessary.

dekhn··on Things that apparently cause cancer
Well, I wouldn't categorically exclude radiation exposure, but yes, I strongly agree that it's unlikely to be radiation exposure.
dekhn··on Things that apparently cause cancer
Oh- I shoudl have said earlier, yes, the original paper is likely nonsense (based on my priors of epi work from Harvard Public Health), but I'm not qualified to read and critique it in a detailed way.
dekhn··on Things that apparently cause cancer
Where did I say that? What came out the cooling stack was radioactive? You're presuming a mechanism. My whole point here is that the original paper just described an association. We don't know what the cause of that was or if the authors even did a competent study.
dekhn··on Things that apparently cause cancer
he was a construction engineer, not a policymaker.
dekhn··on Things That Apparently Cause Cancer
Attributable risk is a specific term in epi, the FDA uses it, and it doesn't imply causality.
dekhn··on Things that apparently cause cancer
Who is the quack here? Did you take this criticism article as being absolutely true? I found their argument unconvincing (I am not really qualified to judge the original paper; running a good epi study is hard, and interpreting the results even harder.
dekhn··on Things that apparently cause cancer
See- you immediately went to a mechanism that justified the result. It's all too easy to convince yourself something is true because you can see a pathway- yet that pathway might not matter.
dekhn··on Things that apparently cause cancer
I could go on a long rant on how sociologists, pyschologists, and epidemiologists all play games with data to support their own pet theories and causes, but I won't.

To answer your question: without other data, generally I would expect a person who works for an industry supporting institute to have greater vested interest than academics working at a university- academics mainly just want to get more funding for their individual research, while the industry is dealing with multi-billion-dollar industries and they get paid well to write articles like this. But now that I look carefully, I don't think their institute is funded directly by the nuclear power/plant industry.

dekhn··on Things that apparently cause cancer
Cancer risk associations studies are not normally described as "proving" anything. The article being criticized uses 'risk' and 'association'. There is a long history of argument around these sorts of studies, there was a previous series of these around people living near power plants (where it seems like SES, not plant proximity, was the strongest explanatory variable, see https://www.aps.org/archives/publications/apsnews/200710/ele... ). I've seen similar "living near a freeway causes cancer" arguments. There was a massive court case by flight attendants, who have a higher rate of cancer than the regular population.

In reading the criticism, I noticed the authors keep using the term "prove" and they also keep trying to come up with mechanisms ("refueling of the plant"). Even the argument about plant worker exposure compared to people living far away doesn't completely work, because the plant workers are taking all sorts of precautions to minimize exposure to radiation, but there are still mechanisms where something could go out the cooling stacks and deliver something harmful downwind.

The biophysics of cancer causation is entirely nontrivial and looking for the actual sources of the cancer risk is challenging, and watching physics people argue with epidemiologists gets old quickly (my field is biophysics, and I've had a few physics people insist that non-ionizing radiation couldn't possibly cause cancer, "because it doesn't damage DNA". Unfortunately, that argument isn't good, because it presupposes a mechanism (DNA damage due to radiation); we know now that non-ionizing radiation causes cellular heating, stress response, and more, which are all associated (based mostly on in vitro studies) with increased rate of cancer.

Note the authors (and the institute they work for) have vested interest: Dr. Adam Stein is the Director of the Nuclear Energy Innovation program at the Breakthrough Institute, where his work centers on the technology, regulation, economics, and risk governance of advanced nuclear energy.

Deric Tilson is a Senior Nuclear Energy Innovation Analyst at The Breakthrough Institute, where he focuses on advancing nuclear energy as a critical pathway to a carbon-free and energy-abundant future.

dekhn··on A 12-year sequence of telescope images of a star and four planets orbiting
The term I would use instead is that the data provided observational support of a hypothesis. It didn't "prove" anything- proofs only exist in math. (yes, I know people use "prove" is colloquial way, but it's misleading, especially in observational work where you can't control variables to find causality.
dekhn··on Frog and Toad and the Increasingly Capable Machines
Thanks for mentioning this- I completely forget this toad is from a different book.
dekhn··on Nicholas Polson has authored 258 academic papers in 2026 so far
Haha, Sean Eddy used to complain about reproduction efforts in bioinformatics (specifically, people trying to benchmark HMMER and doing a bad job). Fortunately, my advisor helped me learn the techniques and Sean approved (he was also happy that HMMER beat BLAST for remote homolog detection).

I also left a postdoc position over my professor's decision to rewrite my paper to juice all the stats- not a specific error, but selectively interpreting the data to make the results look better than state of the art, when they were not.

I would absolutely love to have my "paper correctness AI" mark all those bad reproductions in the literature and outright misrepresentations- it's all too easy to rush to publish and get a lot of attention- especially if your advisor or coauthors are prestigious and mildly unethical.

dekhn··on Nicholas Polson has authored 258 academic papers in 2026 so far
I can't speak to math and theoretical CS, because both of those fields now have methods to make proofs that can be verified (and math is probably the only situation where you can "prove" something true; all other fields are effectively probabilistic, not logical in nature).

I disagree with your premise. Part of our job as scientists (thankfully no longer mine) is to reduce the irrelevant and incorrect noisy as early as possible. I have seen so many grad students get excited by a paper and put enormous effort into reproducing somethign that was a false or fake result.

dekhn··on Nicholas Polson has authored 258 academic papers in 2026 so far
I'm not aware of any fatally flawed paper that leads to the claimed result when done properly. Could you share an example?
dekhn··on Nicholas Polson has authored 258 academic papers in 2026 so far
Change the word LLM to "teacher" and I think the statement is still true.
dekhn··on Nicholas Polson has authored 258 academic papers in 2026 so far
You can get a good idea by reading Retraction Watch and following a few papers/scientists who show up in it. That defines the current norms.

To me, a single significant error of any kind brings the entire paper into question. If a less important figure contains an image duplication, that makes me wonder if I can trust any of the images.

dekhn··on Nicholas Polson has authored 258 academic papers in 2026 so far
The most common and obvious example is image duplication. This is used to invalidate large numbers of paper (I was absolutely shocked at the observed rate of image duplication). I am not sure I would call that a model.

The next example I can think of- I am not sure it qualifies. I read a paper where they deleted one gene at a time in yeast (it has 6000 genes) and determined whether the mutated yeast could live or not. For each gene where the yeast died, they added that to a list of "essential for life" genes. The paper concluded they had found some interesting proteins that should be studied. I read the paper and the first thing that sprang to mind, are any of these genes overlapping? Because we know (somebody already demonstrated in a lab) that genes do overlap (which is truly weird!)

I wrote a script and showed that every gene they reported as essential for life overlapped an already known gene that was essential for life. I wrote the authors, who never responded, but wrote a followup paper where they acknowledged they probably had a high false positive rate due to overlapping genes with known fatal effects. My guess is you'd say that either I used a model (existing literature) or I compared against the world, but realistically, what I did was trivially come up with a better explanation than the authors. T hat's what I want LLMs to do for me, and it seems like the direction LLMs are going will fulfill my desires.

dekhn··on Nicholas Polson has authored 258 academic papers in 2026 so far
I think you misunderstood. I'm not asking for an oracle that can determine whether a paper is correct, I want an oracle that can find real mistakes in papers (thus invalidating them).

I work full time on "lab in the loop" AI, so I'm pretty familiar with the need for real-world experiments. I am not proposing a fully autonomous scientist that could read an arbitrary paper and emit whether it's universally true without some verification method.

Also, to your statement: " Because by the time you are a practicing scientist, you've developed a feel for what constitutes a satisfying explanation."

I'm a practicing scientist (well, ex-scientist) and it seems like most "satisfying explanations" end up being wrong or incomplete simply because they seem so satisfying.

dekhn··on Nicholas Polson has authored 258 academic papers in 2026 (so far)
This is realistic. We already do this today: it's called "journal club". A bunch of grad students read the same paper and then criticize it. I've read papers that I thought were amazing only to. have somebody else notice a key issue in a method, or a conclusion that didn't follow, or outright omission of an important detail, in a way that could be verified by both the students and the authors of the paper.
dekhn··on Nicholas Polson has authored 258 academic papers in 2026 so far
I think one of the end-game components for AI is to be able to read all the literature and generate a reliable list of which ones are not correct and a convincing reason why.
← PreviousPage 3 of 34Next →