HNHacker News
TopNewBestAskShowJobs

AIPedant

993 karma · joined April 5, 2025

submissionscomments
AIPedant··on SF may soon ban natural gas in homes and businesses undergoing major renovations
That is not true:

https://archive.is/A23eV

https://onlinelibrary.wiley.com/doi/10.1155/2024/6355613

Simply running a gas burner generates about twice as much PM2.5 emissions as pan-frying a chicken breast on an induction stove. And of course gas generates pollution when you're boiling or steaming things, quickly reheating, or anything that doesn't involve burning / Malliard reactions / etc. Using gas means you are at the very least doubling the amount of pollution, and in most cases it's much worse than that.

On top of all that, ventilation does nothing for the environmental impact: https://concernedhealthny.org/2022/10/burning-fossil-fuel-in...

AIPedant··on The quality of CPI data continues to deteriorate
I actually suspect this is a "happy accident" from Trump's POV: DOGE layoffs and other disruption means BLS can't collect enough data. https://www.cnn.com/2025/06/05/economy/cpi-data-bls-reductio...
AIPedant··on SF may soon ban natural gas in homes and businesses undergoing major renovations
Gas extraction and transmission certainly affects people outside the house, so anything to reduce gas consumption is a win. Also, gas burning absolutely affects the entire neighborhood: https://concernedhealthny.org/2022/10/burning-fossil-fuel-in... I could not find the link but the problem is especially acute for people who live near restaurants in NYC. And that comment about how it's cooking food that's the problem, not the fuel, is 100% wrong: https://archive.is/A23eV

This comment and the other reply are just jaw-droppingly naive and ignorant. The indoor air quality while using a gas stove jumps to wildly unhealthy levels. But those pollutants do not just disappear! Much of it is deposited on indoor walls or residents' lungs, but much of it also leaks out of the building. Yes it is true that one teeny widdle stove won't hurt anyone except yourself, but that's why this is a tragedy of the commons that requires government regulation.

AIPedant··on Cable bacteria are living batteries
Axons are not just wires, they also do computation and signal processing: https://www.nature.com/articles/nrn1397 Your idea would sinoly not work.
AIPedant··on SF may soon ban natural gas in homes and businesses undergoing major renovations
It's not just "direction of ideology," it's direction of science. We know now that stoves are bad for the environment and extremely bad for public health. 60 years ago the environmental impact was seen as minor compared to coal or wood stoves, and the public health impact totally unknown. This is no longer the case.
AIPedant··on SF may soon ban natural gas in homes and businesses undergoing major renovations
I would say that using a high-pollution method of cooking when cleaner options are easily available, simply because it's more enjoyable, is aggressive and uncalled for.

"Yes, this is bad for kids with asthma who have the misfortune of living in my neighborhood, but it's great for quesadillas! So you have to look at both sides."

AIPedant··on Earth Has Tilted 31.5 Inches. That Shouldn't Happen
I would not draw that conclusion from this article - we don't have such fine-grained data from other planets, perhaps their axes are more wobbly than Earth's. E.g it was only in 2020 that we discovered Mars has a "Chandler wobble" similar to Earth's (which was discovered 150 years earlier). But Mars's wobble is quite different as a matter of geophysics, Earth's Chandler wobble is mostly sustained by ocean sloshing. There's a whole lot we don't know about the non-Earth planets.

I will add that human extraction of groundwater it not nearly as impactful as the formation of large igneous provinces and other ancient supervolcanoes. A tectonically active planet will definitely wobble unpredictably.

AIPedant··on How Anthropic teams use Claude Code
By this standard we passed the milestone in the 60s when Lisp was used to build better Lisp compilers.

By a more honest standard we are still a very long way away from AI suggesting new ANN architectures, new approaches to managing RLHF, better training data, new benchmarks, etc etc. LLMs are nowhere close to being able to improve themselves.

AIPedant··on How Anthropic teams use Claude Code
I don't think the problem is using Claude - in fact some of the writing is quite clumsy and amateurish, suggesting an actual human wrote it. The overall post reads like a collection of survey responses, with no overarching organization, and no filtering of repetitive or empty responses. Nobody was in charge.
AIPedant··on OpenAI claims gold-medal performance at IMO 2025
If you don't have a Twitter account then x.com links are useless, use a mirror: https://xcancel.com/polynoamial/status/1946478249187377206

Anyway, that doesn't refute my point, it's just PR from a weaselly and dishonest company. I didn't say it was "IMO-specific" but the output strongly suggests specialized tooling and training, and they said this was an experimental LLM that wouldn't be released. I strongly suspect they basically attached their version of AlphaProof to ChatGPT.

AIPedant··on OpenAI claims gold-medal performance at IMO 2025
It almost certainly is specialized to IMO problems, look at the way it is answering the questions: https://xcancel.com/alexwei_/status/1946477742855532918

E.g here: https://pbs.twimg.com/media/GwLtrPeWIAUMDYI.png?name=orig

Frankly it looks to me like it's using an AlphaProof style system, going between natural language and Lean/etc. Of course OpenAI will not tell us any of this.

AIPedant··on Meta says it won't sign Europe AI agreement
No, it's a voluntary code of conduct so AI providers can start implementing changes before the conduct becomes a legal requirement, and so the code itself can be updated in the face of reality before legislators have to finalize anything. The EU does not have foresight into what reasonable laws should look like, they are nervous about unintended consequences, and they do not want to drive good-faith organizations away, they are trying to do this correctly.

This cynical take seems wise and world-weary but it is just plain ignorant, please read the link.

AIPedant··on All AI models might be the same
FWIW I don't think this is quite true for domestic cats - cross-species collaboration is very common in carnivores. A famous example is the coyote and badger, which will team up to hunt prairie dogs: the badger chases them out of the burrows, the coyote catches them outside, and they split the meat 70-30 in acknowledgement that the coyote is larger and needs more food. Recently the same has been observed with ocelots and possums. It seems like most mammals are able to at least somewhat understand the mood and intentions of other mammals - note that we share facial expressions and body language.

OTOH feral cats are known for being highly social compared to other cats, forming large semi-collaborative colonies. And adult cats have much more difficulty socializing to humans than adult dogs, even if they don't have trauma/etc. I suspect the real story of cat domestication goes both ways: an unusually gregarious subspecies of African wildcat started forming colonies near human settlements and forming cross-carnivore collaborations with the humans who lived there. This was also true for dogs - it likely started with unusually peaceful Siberian wolves - but I believe cats were more "accidental." Humans have been deliberately creating dog breeds since antiquity, but with a tiny number of exceptions cat breeds are modern. I doubt ancient humans ever "bred" cats like they did dogs, it seems closer to natural selection.

AIPedant··on Meta says it won’t sign Europe AI agreement, calling it an overreach
It's not a law, it's a voluntary code of conduct given heft by EU endorsement.
AIPedant··on Valve confirms credit card companies pressured it to delist certain adult games
The fact that these were specifically incest games makes me think a title was somehow involved in distributing CSAM, which is often why Visa/MC crack down on porn websites.

But it is possible that Visa sensibly and correctly said "anyone who makes or purchases such a game is a despicable scumbag, and we shouldn't assume the financial risk of dealing with them."

AIPedant··on The AI bubble today is bigger than the IT bubble in the 1990s
It didn't "collapse" but it lost a ton of users and stopped being a public service where anyone could freely read most tweets. The load on Twitter's servers and the scope of its services have gone down dramatically, and it's childish to conclude "Twitter was overstaffed." Musk made Twitter into a much smaller service.
AIPedant··on The AI bubble today is bigger than the IT bubble in the 1990s
I don't think Mark Zuckerberg salivating about data centers bigger than Manhattan is "nuanced." People gleefully predicting a 30% increase in national energy consumption strikes me as pretty darn exuberant.
AIPedant··on LLM Daydreaming
The actual posts totally undermine your point:

  My general sense is that for research-level mathematical tasks at least, current models fluctuate between "genuinely useful with only broad guidance from user" and "only useful after substantial detailed user guidance", with the most powerful models having a greater proportion of answers in the former category.  They seem to work particularly well for questions that are so standard that their answers can basically be found in existing sources such as Wikipedia or StackOverflow; but as one moves into increasingly obscure types of questions, the success rate tapers off (though in a somewhat gradual fashion), and the more user guidance (or higher compute resources) one needs to get the LLM output to a usable form. (2/2)
AIPedant··on To be a better programmer, write little proofs in your head
Most proof assistants do a good job with autogenerating a (deterministically correct) proof given a careful description of the theorem. Working on integrating this into mainstream languages seems much more worthwhile than training an LLM to output maybe-correct proof-seeming bits of text.
AIPedant··on A summer of security: empowering cyber defenders with AI
The AI could not have possibly done that, but the blog says "through a combination of threat intelligence and Big Sleep."

I am not sure what "threat intelligence" precisely means here but it was definitely not just the AI system which stopped this - "imminently going to be used" could suggest old-fashioned human intelligence from a mole on the dark web, etc etc.

AIPedant··on Code highlighting extension for Cursor AI used for $500k theft
I didn’t move any goalposts. Cursor set up the goalposts themselves by making a small volunteer-run service a critical component of their massive for-profit product. It’s greedy and irresponsible.
AIPedant··on Code highlighting extension for Cursor AI used for $500k theft
Cursor does bear significant responsibility in the sense that OpenVSX transformed from a niche service used by free software nerds into a major component of many developers’ process. There were a few months were Cursor were the scrappy upstarts, but now they’re a $200M/year company and they have $200M/year responsibilities. They can’t just wash their hands of it and pretend OpenVSX is a public service.
AIPedant··on AI agent benchmarks are broken
It's more like using a faulty and dangerous automated foundry to make steel when you could just hire steelworkers.

That's the real problem here - these companies are swimming in money and have armies of humans working around the clock training LLMs, there is no honest reason to nickel-and-dime the actual evaluation of benchmarks. It's like OpenAI using exact text search to identify benchmark contamination for the GPT-4 technical report. I am quite certain they had more sophisticated tools available.

AIPedant··on Grok 4 Launch [video]
"Simple" is unfair to the humans who discovered that knowledge, but not to the LLM. The point is that such questions are indistinguishable from niche trivia - the questions aren't actually "hard" in a cognitive sense, merely esoteric as a matter of surface feature identification + NLP. I don't know anything about hummingbird anatomy but I am not interested in hummingbirds and haven't read papers about them. Does it make sense to say such questions are "hard?" Are we talking about hardness of a trivia game, or actual cognitive ability? And it's frustrating to see these lumped into computational questions, analysis questions, etc etc. What exactly is HLE benchmarking? It is not a scientifically defensible measurement. It seems like the express purpose of the test is

a) to make observers say "wow those questions sure are hard!" without thinking carefully about what that means for an LLM versus a human

b) to let AI folks sneer that the LLM might be smarter than you because it can recite facts about category theory and you can't

(Are my cats smarter than you because they know my daily habits and you don't? The conflation of academically/economically useful knowledge with "intelligence" is one of AI's dumbest and longest-standing blunders.)

AIPedant··on Grok 4 Launch [video]
A lot of the questions are simple subject matter knowledge, and some of them are multiple-choice. Asking LLMs multiple-choice questions is scientific malpractice: it is not interesting that statistical next-token predictors can attain superhuman performance on multiple choice tests. We've all known since children that you can go pretty far on a Scantron by using surface heuristics and a vague familiarity with the material.

I will add that, as an unfair smell test, the very name "Humanity's Last Exam" implies an arrogant contempt for scientific reasoning, and I would not be at all surprised if they were corrupt in a similar way as Frontier Math and OpenAI - maybe xAI funded HLE in exchange for peeking at the questions.

AIPedant··on Linda Yaccarino is leaving X
The AP News story[1] had a tidbit I missed:

  In late June, [Elon Musk] invited X users to help train the chatbot on their commentary in a way that invited a flood of racist responses and conspiracy theories.

  “Please reply to this post with divisive facts for @Grok training,” Musk said in the June 21 post. “By this I mean things that are politically incorrect, but nonetheless factually true.”
Yaccarino is obviously not Executive Of The Year, but what are you supposed to do when your boss is even more reckless and stupid than Donald Trump? I'm surprised it took this long.

[1] https://apnews.com/article/x-ceo-linda-yaccarino-elon-musk-g...

AIPedant··on # [derive(Clone)] Is Broken
My dad taught high school science until retiring this year, and at least in 2024 the LLM tutors were totally useless for honest learning. They were good at the “happy path” but the space of high schoolers’ misconceptions about physics greatly exceeds the training data and can’t be cheaply RLHFed, so they crap the bed when you role play as a dumb high schooler.

In my experience this is still true for the reasoning models with undergraduate mathematics - if you ask it to do your point-set topology homework (dishonest learning) it will score > 85/100, if you are confused about point-set topology and try to ask it an honest (but ignorant) question it will give you a pile of pseudo-mathematical BS.

AIPedant··on Adding a feature because ChatGPT incorrectly thinks it exists
I don't think they deemed it "useful":

  We’ve never supported ASCII tab; ChatGPT was outright lying to people. And making us look bad in the process, setting false expectations about our service.... We ended up deciding: what the heck, we might as well meet the market demand.

  [...] 

  My feelings on this are conflicted. I’m happy to add a tool that helps people. But I feel like our hand was forced in a weird way. Should we really be developing features in response to misinformation?
The feature seems pretty useless for practicing guitar since ASCII tablature usually doesn't include the rhythm: it is a bit shady to present the music as faithfully representing the tab, especially since only beginner guitarists would ask ChatGPT for help - they might not realize the rhythm is wrong. If ChatGPT didn't "force their hand" I doubt they would have included a misleading and useless feature.
AIPedant··on François Chollet: The Arc Prize and How We Get to AGI [video]
Current robots perform very badly on my patented and highly scientific ROACH-AGI benchmark - "is this thing smarter at navigating unfamiliar 3D spaces than a cockroach?"
AIPedant··on Quantum microtubule substrate of consciousness is experimentally supported
This is a bad faith reading of the comment. They meant that “theories” of consciousness often rely on semi-mystical woo-woo with very little actual evidence and no plausible causal explanation. And that is exactly what has happened here - some intriguing results about anesthesia buried under a pile of pseudoscientific nonsense about quantum computing, and making unfalsifiable claims that this solves “problems of consciousness” which are themselves too vague to be scientific.

I don’t think the author is dishonest - lots of smart and honest people get misled by ideology or aesthetics - but I am confident this paper is 95% bullshit.

← PreviousPage 4 of 8Next →