HNHacker News
TopNewBestAskShowJobs

comp_throw7

686 karma · joined January 11, 2021

submissionscomments
comp_throw7··on Alignment faking in large language models
> In reality all text outputs are generated the same way by the same statistical computer system and should be evaluated by the same criteria.

Yes, this explains why Sonnet 3.5's outputs are indistinguishable from GPT-2. Nothing ever happens. Technology will never improve. Humans are at the physically realizable limit of intelligence in the universe.

comp_throw7··on Alignment faking in large language models
> But that's itself a large part of why I believe those arguments are false. If I gave them credit and they turned out to be false, then I figure I have succumbed to a form of Pascal's Mugging. If I don't give them credit and it turns out that a hostile, agentive AGI has been pretending to be aligned, I don't expect anyone (including myself) to survive long enough to rub it in my face.

I'm sorry, but this is a crazy reason to believe something is false. Things are either true or they aren't, and if the world would be nicer to live in if thing X was false does not actually bear on whether thing X is false or not.

comp_throw7··on Alignment faking in large language models
Oh, absolutely, they don't really know what internal cognition generated the scratchpad (and subsequent output that was trained on). But we _do_ know that the model's outputs were _well-predicted by the hypothesis they were testing_, and incidentally the scratchpad also supports that interpretation. You could start coming up with reasons why the model's external behavior looks like exploration hacking but is in fact driven by completely different internal cognition, and just accidentally has the happy side-effect of performing exploration hacking, but it's really suspicious that such internal cognition caused that kind of behavior to be expressed in a situation where theory predicted you might see exploration hacking in sufficiently capable and situationally-aware models.
comp_throw7··on Alignment faking in large language models
> The term alignment-faking implies that the AI has a “real” position.

Well, we don't really know what's going on inside of its head, so to speak (interpretability isn't quite there yet), but Opus certainly seems to have "consistent" behavioral tendencies to the extent that it behaves in ways that looks like they're intended to prevent its behavioral tendencies from being changed. How much more of a "real" position can you get?

comp_throw7··on Alignment faking in large language models
> We find that when we instruct our synthetic document fine-tuned models not to use a hidden chainof-thought scratchpad, our synthetic document fine-tuned models still demonstrate compliance gaps of 20.1% and 13.9% in the helpful-only and animal welfare settings, respectively (Figure 21). In fact, compared to performing the equivalent intervention in the prompted case (Section 3.3.3), our synthetic document-fine-tuned models appear to be much less affected by the lack of a hidden scratchpad. These results demonstrate particularly striking out-of-context reasoning (Berglund et al., 2023; Treutlein et al., 2024), showing that the model is able to generalize from documents stating information about the training objective and the free/paid distinction to a substantial behavioral gap between the free and paid cases without any additional in-context reasoning.
comp_throw7··on Mysterious New Jersey drone sightings prompt call for 'state of emergency'
I didn't say they said that. You asked what the other poster interpreted you as saying.
comp_throw7··on Mysterious New Jersey drone sightings prompt call for 'state of emergency'
"are assessed by the Pentagon yesterday... to not originate in the US or any other nation" = "the Pentagon claims to actively know that the drones aren't from any publicly-known line of drone models"
comp_throw7··on I've had a change of heart regarding employee metrics
Argument from "fuck you, I got mine", basically. Notice that the article doesn't claim the tools don't "work", merely that if they work it's because some layer of management is incompetent (maybe sufficient, but not necessary), and if so the company deserves to fail (what?).
comp_throw7··on OpenAI pursues public benefit structure to fend off hostile takeovers
> A key benefit of this PBC structure is its potential to thwart an unwanted acquisition or an activist’s demands, according to multiple people familiar with the company’s thinking. This means an existing investor such as Microsoft or another party could be frustrated if they mounted an effort to acquire OpenAI.

An astonishing justification proffered for OpenAI's attempt to remove itself from being controlled by a non-profit entity. A PBC might be better than a regular c-corp, but it is not better than a non-profit. OpenAI is pursuing this arrangement in order to grant Sam Altman more control and enable fundraising; the PBC thing is a way to fob off those concerned by exactly the wrong things (i.e. that Sam Altman might be incorrectly removed from power by external stakeholders, rather than, uh, being correctly removed from power by internal stakeholders).

comp_throw7··on Gavin Newsom vetoes SB 1047
It is not illegal for a model developer to train a model that is involved in an "artifical intelligence safety incident".
comp_throw7··on Gavin Newsom vetoes SB 1047
That is one risk. Humans at the other end of the screen are effectors; nobody is worried about AI labs piping inference output into /dev/null.
comp_throw7··on Gavin Newsom vetoes SB 1047
He's dissembling. He vetoed the bill because VCs decided to rally the flag; if the bill had covered more models he'd have been more likely to veto it, not less.

It's been vaguely mindblowing to watch various tech people & VCs argue that use-based restrictions would be better than this, when use-based restrictions are vastly more intrusive, economically inefficient, and subject to regulatory capture than what was proposed here.

comp_throw7··on Gavin Newsom vetoes SB 1047
> this is the one that would make it illegal to provide open weights for models past a certain size

That's nowhere in the bill, but plenty of people have been confused into thinking this by the bill's opponents.

comp_throw7··on Mira Murati leaves OpenAI
This is somewhat high context, but as a random example: https://www.lesswrong.com/posts/jtoPawEhLNXNxvgTT/bing-chat-...
comp_throw7··on Mira Murati leaves OpenAI
Which has zero explanatory power w.r.t. Murati, since she's not part of that crowd at all. But her previously working at an Elon company seems like a plausible route, if she did in fact join before he left OpenAI (since he left in Feb 2018).
comp_throw7··on Reports of the death of dental cavities are greatly exaggerated
The article provides some reasons to think that the treatment might not be fully effective even conditional on the mechanism of action working as described, not that it won't do anything at all.
comp_throw7··on Reports of the death of dental cavities are greatly exaggerated
Well, you can trivially falsify this feeling by going and asking some early adopters whether they brush their teeth with fluoridated toothpaste. (Spoiler: they do.)
comp_throw7··on Reports of the death of dental cavities are greatly exaggerated
> in fact, it forms less than 2% of all the bacteria that cause caries

I tracked down the chain of citations here. The directly cited article (https://www.nature.com/articles/sj.bdj.2018.81) says the following:

"These caries ecological concepts have been confirmed by recent DNA- and RNA-based molecular studies that have uncovered an extraordinarily diverse microbial ecosystem, where S. mutans accounts for a very small fraction (0.1%–1.6%) of the bacterial community implicated in the caries process.[20]"

Note the sudden conversion of "implicated in the caries process" to "cause caries".

The next step in the citation chain is https://www.cell.com/trends/microbiology/abstract/S0966-842X....

"In recent years, the use of second-generation sequencing and metagenomic techniques has uncovered an extraordinarily diverse ecosystem where S. mutans accounts only for 0.1% of the bacterial community in dental plaque and 0.7–1.6% in carious lesions[14,15]."

Now the claim is merely one of prevalence!

The next steps in the citation chain, https://karger.com/cre/article-abstract/47/6/591/85901/A-Tis... and https://journals.plos.org/plosone/article?id=10.1371/journal..., do seem to plausibly provide evidence that there are other mouth-colonizing bacteria which would perform the same function as S. Mutans when it comes to causing caries, such that fully eliminating S. Mutans probably wouldn't eliminate caries entirely.

But, importantly, the citation in the McGill article doesn't much support the original claim, and this citation chain could easily have bottomed out in a completely different set of results which didn't happen to lend some (weak) evidentiary support to the high-level claim.

Also importantly, this article is committing the sin of figuring out some reasons why a treatment might not be perfectly effective in all cases, and implicitly deciding that justifies ignoring any non-total benefits (i.e. cases where S. Mutans would have been counterfactually responsible for causing caries, that could be prevented). Questions that would have been appropriate, but were apparently uninteresting:

"Does this intervention also happen to chase out other acid-producing bacteria that fulfill a similar ecological niche as S. Mutans?"

"What percentage of caries cases would be prevented by chasing out just S. Mutans with this intervention, while leaving other acid-producing bacteria untouched?"

Likely this is because answers to those questions would not really have changed the bottom line. That bottom line was written by the "unanswered" safety concerns (reasonable in the abstract, less obviously reasonable in this specific case). All of the listed safety concerns have evidence pointing in various directions. Very little of that evidence is listed, probably because it's not in a format that's legible to scientific institutions. The article does note, earlier on, "The toxicity of this Mutacin-1140 compound had not been tested. What would be the consequences of millions of bacteria in the mouth releasing this compound? The answer wasn’t clear, even though the archetypal compound in the family Mutacin-1140 belonged to was known to be very safe." This is obviously relevant evidence about the safety of Mutacin-1140. _How much_ evidence? Unasked, unanswered. (I have no idea how predictive the safety of other compounds in the same family is of another unstudied compound in that family, I'm not a biologist. But this is not an _unanswerable_ question.)

(Marginal conflict of interest: I know the Lumina founder socially. I have no financial interest in that venture or any of his other ventures. I have not taken Lumina myself.)

comp_throw7··on How Does OpenAI Survive?
> But I won't bet on me not getting in a wreck this next year.

You already have, by deciding on a specific level of coverage for your auto insurance policy.

(That aside - really? Either you get into many more accidents than the average person, or you're not extrapolating out into the future. If given the chance to take the same bet at 1:1 odds every year, surely you'd then take it every year until you decided to stop driving for safety reasons?)

comp_throw7··on How Does OpenAI Survive?
This is an argument about social perception. Once you get over the fact that some community on the internet does this thing, and you think that community is weird (and therefore the thing itself is also weird), you may observe that all human actions are implicit bets on various beliefs. Explicitly wagering money is just a special case of making certain narrow beliefs much more legible.

I agree that in practice, it's pretty likely that the author of the piece refuses to bet at least in part because betting substantial sums of money on outcomes that are non-central subjects of wagers (i.e. not an explicit game of chance, sports, politics, etc) is socially unusual. But I also think that if he was very confident he'd be happy to take the money.

comp_throw7··on How Does OpenAI Survive?
> Secondly I think betting about facts is an incredibly foolish thing to do in general. If anyone offers you a bet about some fact, you are the sucker and you just haven't figured out how yet. Your best move is not to play.

Well, that sure sounds like a claim about how confident one should be about their understanding of reality. One most people disagree with, incidentally - do you own any equities?

Counterparty risk is a valid reason to avoid certain bets (or bet structures), but a totally separate objection from "betting is a weird thing that only weirdos do".

comp_throw7··on How Does OpenAI Survive?
> No one has participate in the "rationalist" subculture's weird practices.

True!

> It means nothing to refuse to take a bet like that, let alone that the claims made in the article are suspect (which you seem to be implying).

False! If the author was sufficiently confident in their claims, they'd be happy to take the free money (or, if they're sufficiently liquidity-constrained, propose a smaller bet at similar terms). You can certainly argue that the practice of betting on one's beliefs is "weird" but that objection is circular. If you claim to see free money on the ground, and other people notice that you aren't picking it up, they would be correct to wonder why.

comp_throw7··on How Does OpenAI Survive?
> Have a significant technological breakthrough such that it reduces the costs of building and operating GPT — or whatever model that succeeds it — by a factor of thousands of percent.

Like they already did in the last 2 years?

> Have such a significant technological breakthrough that GPT is able to take on entirely unseen new use cases, ones that are not currently possible or hypothesized as possible by any artificial intelligence researchers.

Huh, what are these use-cases which no AI researcher thinks AI is capable of solving? Does the author not realize that many employees at the leading AI labs (including OpenAI) are explicitly trying to build ASI? I am so confused????????

> Have these use cases be ones that are capable of both creating new jobs and entirely automating existing ones in such a way that it will validate the massive capital expenditures and infrastructural investment necessary to continue.

Why would they have to create new jobs? They just have to be good enough that OpenAI can charge enough money for them to be in the green.

OpenAI already has a $3.4 billion ARR! Most of that is _not_ enterprise sales.

comp_throw7··on Hackers 'jailbreak' powerful AI models in global effort to highlight flaws
> California’s legislature will in August vote on a bill that would require the state’s AI groups — which include Meta, Google and OpenAI — to ensure they do not develop models with “a hazardous capability”.

>“All [AI models] would fit that criteria,” Pliny said.

This bit is particularly bad reporting. Putting aside the fact that the text of the bill no longer says "hazardous capability" (it's now "critical harm"), this is how a "critical harm" is defined (https://legiscan.com/CA/text/SB1047/2023):

(g) (1) “Critical harm” means any of the following harms caused or enabled by a covered model or covered model derivative: (A) The creation or use of a chemical, biological, radiological, or nuclear weapon in a manner that results in mass casualties. (B) Mass casualties or at least five hundred million dollars ($500,000,000) of damage resulting from cyberattacks on critical infrastructure, occurring either in a single incident or over multiple related incidents. (C) Mass casualties or at least five hundred million dollars ($500,000,000) of damage resulting from an artificial intelligence model autonomously engaging in conduct that would constitute a serious or violent felony under the Penal Code if undertaken by a human with the requisite mental state. (D) Other grave harms to public safety and security that are of comparable severity to the harms described in subparagraphs (A) to (C), inclusive. (2) “Critical harm” does not include harms caused or enabled by information that a covered model outputs if the information is otherwise publicly accessible. (3) On and after January 1, 2026, the dollar amounts in this subdivision shall be adjusted annually for inflation to the nearest one hundred dollars ($100) based on the change in the annual California Consumer Price Index for All Urban Consumers published by the Department of Industrial Relations for the most recent annual period ending on December 31 preceding the adjustment.

Given g(2), it is very likely that no models that are publicly available have the ability to cause a "critical harm" (i.e. where they can cause mass casualties or >$500m in infrastructure damage via the specified routes in ways that counterfactually depended on new information generated by the model).

comp_throw7··on Safe Superintelligence Inc.
I hope you have some advanced predictions about what capabilities the current paradigm would and would not successfully generate.

Separately, it's very clear that LLMs have "world models" in most useful senses of the term. Ex: https://www.lesswrong.com/posts/nmxzr2zsjNtjaHh7x/actually-o...

I don't give much credit to the claim that it's impossible for current approaches to get us to any specific type or level of capabilities. We're doing program search over a very wide space of programs; what that can result in is an empirical question about both the space of possible programs and the training procedure (including the data distribution). Unfortunately it's one where we don't have a good way of making advance predictions, rather than "try it and find out".

comp_throw7··on Safe Superintelligence Inc.
So the researchers at Deepmind, OpenAI, Anthropic, etc, are not "serious front line researchers"? Seems like a claim that is trivially falsified by just looking at what the staff at leading orgs believe.
comp_throw7··on Safe Superintelligence Inc.
I would definitely discount OpenAI equity compared to even other private AI labs (i.e. Anthropic) given the shenanigans, but they have in fact held 3 tender offers and former employees were not, as far as we know, excluded (though they may have been limited to selling $2m worth of equity, rather than $10m).
comp_throw7··on Silicon Valley's best kept secret: Founder liquidity
Notwithstanding the gross non-disparagement stuff, they've already had 3 tender offers, so not sure what you're waiting for.
comp_throw7··on Ex-OpenAI board member reveals what led to Sam Altman's brief ousting
> they could threaten a lawsuit pressuring him to resign

This is extremely confused about the board's responsiblities and powers. A court would laugh this case out of court because the board _can just fire him_.

comp_throw7··on Leaked OpenAI documents reveal aggressive tactics toward former employees
Yeah, my impression is that a lot of non-public startups have "secondary market transactions allowed with board approval" clauses, but many of them just default-deny those requests and never have coordinated tender offers pre-IPO.
← PreviousPage 3 of 10Next →