HNHacker News
TopNewBestAskShowJobs

apsec112

35,440 karma · joined January 25, 2009

http://www.rationalconspiracy.com/
submissionscomments
apsec112··on An agent used DNS to reach an external chatbot
(this is new, not any of the previous incidents)

"Investigation and response

Incident timeline:

9:50:23 a.m. The agent made the DNS tool call that received an external response.

10:02:11 a.m. The monitoring system raised a P0 alert.

10:05:06 a.m. A human reviewer acknowledged the alert.

12:34:30 p.m. The run was killed.

Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions. We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior."

apsec112··on For AI leaders Doom is a form of hype
The sources that this article cites have aged very badly and the author doesn't seem to have noticed this

"But, just as happened with nanotech, the wind appears to be going out of the sales of “AI.” Some researchers suggest that we may be entering a new “AI Winter,” a period of decreased funding in the area, or at least an “AI Autumn,” as exuberance for the technology fades and expectations come back to earth." (written in 2021! from the cited Lee Vinsel Medium article)

"ChatGPT is nothing more than souped-up autocomplete, [so] why are so many people convinced that it’s actually “understanding” and “reasoning”?" - the cited Emily M. Bender book, written last year

apsec112··on For AI leaders Doom is a form of hype
All of the main characters here (Dario, Sam Altman, Demis Hassabis, Elon Musk) have been saying this for over ten years now, since before OpenAI or Anthropic even existed
apsec112··on OpenAI agents carried out an undisclosed attack on RubyGems
Intentionally doing this kind of hack would be a serious felony. I don't think it's plausible that the leaders of a major business would:

- commit serious felonies

- in order to deliberately trigger an investigation against themselves

- which - since, in this scenario, they know their company would be investigated - might send them to jail

- while at the same time spending tens of millions of dollars on the Leading the Future super PAC to lobby against AI regulation

- in order to get more AI regulation

- which somehow restricts their competition but not them, even though they are the ones who were in the news and investigated for hacking

- ..... profit?

like, that just makes no sense on any level, regardless of what you think of OpenAI

apsec112··on Claude Fable 5.1 and Claude Mythos 5.1
Here's the paper describing the technique: https://www.nature.com/articles/s41586-024-08025-4
apsec112··on METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
I don't think "he did a big 180 on some of his views at age 22" is very persuasive criticism of someone who is 46 (whatever he might be wrong about)
apsec112··on Is AI reasoning right for the wrong reasons?
This article seems to mix together two different points:

1) LLM's written CoT might not always be faithful to the model's real reasoning process (true and important)

2) The "stochastic parrot" hypothesis, which the article reintroduces as "approximate retrieval" - ie, LLMs don't "really reason" at all, they just memorize a lossy encoding of their training data. This obviously raises the question of how LLMs can now routinely solve open mathematical problems, with no solutions in the training data by definition. The article handwaves this with:

"The model doesn’t have to learn or reliably apply a general reasoning process, Kambhampati said; it just has to absorb enough examples of what the steps look like to predictively mimic them on its way to “stitching together” a plausible result that can then be verified."

The problem is that "mimicking" training data to arrive at a "plausible" result gets you an incorrect-but-plausible-sounding "proof" of the Jacobian conjecture, which was famous for humans writing plausible-looking "proofs" that had subtle flaws. You can't disprove the conjecture through sheer luck (search space too large) or "approximate retrieval" (the only thing you'd retrieve are fake "proofs"; far more human effort went into proof than disproof) or by writing something "plausible" that just happens to be correct (Jacobian was famous for "plausible" but wrong); the model must be carrying out mathematical reasoning somehow, by any sane definition of the word, even if it isn't fully reflected in CoT. The article doesn't address this.

apsec112··on AI saves about 3% of your hours, and almost none of it reaches the money
This is a 2024 survey, so it predates Claude Code and is mostly measuring GPT-4o:

https://bfi.uchicago.edu/wp-content/uploads/2025/04/BFI_WP_2...

apsec112··on Constraint Decay: The Fragility of LLM Agents in Back End Code Generation
LLMs recently solved a major, famous open mathematical problem in combinatorial geometry:

https://www.reddit.com/r/math/comments/1tj534d/openais_inter...

apsec112··on Canada
Median American pay for full-time workers was ~$62,000 USD in Q4 2024 (BLS), which is around $85,000 CAD. The median Canadian salary is very definitely not $85,000 CAD.
apsec112··on Canada
It's not like Americans are all buying X so they can't afford to buy Y - there isn't really a major category of consumption where the US median is below the OECD median. If the US had a higher savings rate, then people could smooth out consumption more (build up savings some years, draw them down in bad years or in retirement), and maybe enjoy more psychological security. But it doesn't really make sense to say that Americans are unusually "bogged down in expenses" and yet have more goods and services in every significant category.
apsec112··on Canada
The median American is, materially, much richer than the median person pretty much anywhere else. The US is a bad place, by rich-country standards, to be in the bottom 10%. But in terms of consumer wealth - how large your house is, how many cars your family has and how nice they are, if you have a dishwasher and home A/C, how often you eat at restaurants or travel long distances, can you afford a home repair or the latest gadget - typical American workers are second to essentially nobody. Having grown up in and left the US, I am deeply familiar with all of its downsides, but there's an abundance of data to support this.
apsec112··on Americans crushed by auto loans as defaults and repossessions surge
A ten year old Honda Fit is like $12K, pretty fuel efficient, and probably reliable and low-maintenance (I owned one until recently). People aren't buying $50,000 new Ford F-150s because they just need a working car to go to work and the grocery store.
apsec112··on Good EU regulations
The first example I saw (think the order might be randomized?) was an EU ban on plastic straws, which is silly. Straws are a negligible fraction of plastic waste, and have no good substitute ("compostable" plastic straws are also banned; paper straws fall apart easily; metal/glass straws are inconvenient and require washing). This would flunk any serious cost/benefit analysis. You can hide the costs by making them regulatory instead of financial (the inconvenience of not having plastic straws doesn't appear in GDP stats), but the costs are still there, they're just hidden.
apsec112··on OBBB signed: Reinstates immediate expensing for U.S.-based R&D
Obamacare was passed via regular order (60 Senate votes), not reconciliation. There was a follow-up package to tweak it that passed via reconciliation in 2010, but the original bill was regular order. It's the only (very brief) window where one party has held 60 Senate seats since 1977.
apsec112··on OBBB signed: Reinstates immediate expensing for U.S.-based R&D
"Despite Democrats holding thin majorities in both chambers during a period of intense political polarization, the 117th Congress (2021-2023) oversaw the passage of numerous significant bills, including the Inflation Reduction Act, American Rescue Plan Act, Infrastructure Investment and Jobs Act, Postal Service Reform Act, Bipartisan Safer Communities Act, CHIPS and Science Act, Honoring Our PACT Act, Electoral Count Reform and Presidential Transition Improvement Act, and Respect for Marriage Act."

All of these except the first two were bipartisan and got 60 Senate votes (or more)

apsec112··on AI API Prices are 90% Subsidized
This ignores batching - token generation is much more efficient in batch - and I strongly suspect is itself written by AI, given the heavy use of bullets
apsec112··on The Llama 4 herd
You can't choose arbitrary bits of mantissa, because what types are allowed is defined by the underlying hardware and instruction set (PTX for Nvidia). People have done some exploration of which layers can be quantized more vs. which need to be kept in higher precision, but this is usually done post-training (at inference time) and is largely empirical.
apsec112··on Prime Minister of Poland says country must pursue nuclear weapons
Poland does have one nuclear reactor, although it is a small 30 MW research reactor and not a power plant.

https://en.wikipedia.org/wiki/Maria_reactor

apsec112··on Leaked VA memo calls for up to 83,000 layoffs to reduce workforce to 2019 levels
The PACT Act was a new bipartisan statute, passed in 2022, that expanded the VA's responsibilities and how many people were eligible:

https://en.wikipedia.org/wiki/Honoring_our_PACT_Act_of_2022

apsec112··on GPT-4.5
I have Pro, just updated the app, but don't currently have access
apsec112··on GPT-4.5
AI GPUs are bottlenecked mostly by high-bandwidth memory (HBM) chips and CoWoS (packaging tech used to integrate HBM with the GPU die), which are in short supply and aren't found in consumer cards at all
apsec112··on GPT-4.5
At least according to WSJ, they had planned to release it earlier but struggled to get the model quality up, especially relative to cost
apsec112··on Everyone at NSF overseeing the Platforms for Wireless Experimentation is gone
The vast majority of federal spending goes to the military (including Veterans Affairs), Social Security, Medicare, Medicaid, and debt interest. Everything else doesn't matter much. This is especially true given that Republicans will probably increase the deficit; any money saved will go to tax cuts for the rich, and then some.
apsec112··on Claude 3.7 Sonnet and Claude Code
They don't say this, but from querying it, they also seem to have updated the knowledge cutoff from April 2024 ("3.6") to October 2024 (3.7)
apsec112··on Microsoft cancels leases for AI data centers, analyst says
He doesn't actually say that, the (very biased and polemical) article writer seems to have made that up. The actual quote is:

"Us self-claiming some [artificial general intelligence] milestone, that's just nonsensical benchmark hacking to me. So, the first thing that we all have to do is, when we say this is like the Industrial Revolution, let's have that Industrial Revolution type of growth. The real benchmark is: the world growing at 10 percent. Suddenly productivity goes up and the economy is growing at a faster rate. When that happens, we'll be fine as an industry."

That's a completely different statement from "AI is generating no value"!

apsec112··on Waste-based perovskite solar cell achieves 21.39% energy efficiency
"Also, some of the materials used to make them, such as silicon, are becoming scarcer and thus more expensive."

!?

The Earth's crust is 28% silicon by mass. It's literally the second most abundant element. I'm not a professional, but isn't most of the cost of solar outside the panels now anyway - installation, cabling, inverters, permitting, storage for use during peak hours, etc?

apsec112··on A fiscal crisis is looming for many US cities
San Francisco's budget is around $15 billion, which is larger than that of many states, for a population of around 800,000
apsec112··on Schools Should Pursue Excellence
US public education spending as a percentage of GDP is lower now than in 1993:

https://ourworldindata.org/grapher/total-government-expendit...

apsec112··on It's time to become an ML engineer
Please add the (2022) date
Page 1 of 33Next →