Yes, this explains why Sonnet 3.5's outputs are indistinguishable from GPT-2. Nothing ever happens. Technology will never improve. Humans are at the physically realizable limit of intelligence in the universe.
686 karma · joined January 11, 2021
Yes, this explains why Sonnet 3.5's outputs are indistinguishable from GPT-2. Nothing ever happens. Technology will never improve. Humans are at the physically realizable limit of intelligence in the universe.
I'm sorry, but this is a crazy reason to believe something is false. Things are either true or they aren't, and if the world would be nicer to live in if thing X was false does not actually bear on whether thing X is false or not.
Well, we don't really know what's going on inside of its head, so to speak (interpretability isn't quite there yet), but Opus certainly seems to have "consistent" behavioral tendencies to the extent that it behaves in ways that looks like they're intended to prevent its behavioral tendencies from being changed. How much more of a "real" position can you get?
An astonishing justification proffered for OpenAI's attempt to remove itself from being controlled by a non-profit entity. A PBC might be better than a regular c-corp, but it is not better than a non-profit. OpenAI is pursuing this arrangement in order to grant Sam Altman more control and enable fundraising; the PBC thing is a way to fob off those concerned by exactly the wrong things (i.e. that Sam Altman might be incorrectly removed from power by external stakeholders, rather than, uh, being correctly removed from power by internal stakeholders).
It's been vaguely mindblowing to watch various tech people & VCs argue that use-based restrictions would be better than this, when use-based restrictions are vastly more intrusive, economically inefficient, and subject to regulatory capture than what was proposed here.
That's nowhere in the bill, but plenty of people have been confused into thinking this by the bill's opponents.
I tracked down the chain of citations here. The directly cited article (https://www.nature.com/articles/sj.bdj.2018.81) says the following:
"These caries ecological concepts have been confirmed by recent DNA- and RNA-based molecular studies that have uncovered an extraordinarily diverse microbial ecosystem, where S. mutans accounts for a very small fraction (0.1%–1.6%) of the bacterial community implicated in the caries process.[20]"
Note the sudden conversion of "implicated in the caries process" to "cause caries".
The next step in the citation chain is https://www.cell.com/trends/microbiology/abstract/S0966-842X....
"In recent years, the use of second-generation sequencing and metagenomic techniques has uncovered an extraordinarily diverse ecosystem where S. mutans accounts only for 0.1% of the bacterial community in dental plaque and 0.7–1.6% in carious lesions[14,15]."
Now the claim is merely one of prevalence!
The next steps in the citation chain, https://karger.com/cre/article-abstract/47/6/591/85901/A-Tis... and https://journals.plos.org/plosone/article?id=10.1371/journal..., do seem to plausibly provide evidence that there are other mouth-colonizing bacteria which would perform the same function as S. Mutans when it comes to causing caries, such that fully eliminating S. Mutans probably wouldn't eliminate caries entirely.
But, importantly, the citation in the McGill article doesn't much support the original claim, and this citation chain could easily have bottomed out in a completely different set of results which didn't happen to lend some (weak) evidentiary support to the high-level claim.
Also importantly, this article is committing the sin of figuring out some reasons why a treatment might not be perfectly effective in all cases, and implicitly deciding that justifies ignoring any non-total benefits (i.e. cases where S. Mutans would have been counterfactually responsible for causing caries, that could be prevented). Questions that would have been appropriate, but were apparently uninteresting:
"Does this intervention also happen to chase out other acid-producing bacteria that fulfill a similar ecological niche as S. Mutans?"
"What percentage of caries cases would be prevented by chasing out just S. Mutans with this intervention, while leaving other acid-producing bacteria untouched?"
Likely this is because answers to those questions would not really have changed the bottom line. That bottom line was written by the "unanswered" safety concerns (reasonable in the abstract, less obviously reasonable in this specific case). All of the listed safety concerns have evidence pointing in various directions. Very little of that evidence is listed, probably because it's not in a format that's legible to scientific institutions. The article does note, earlier on, "The toxicity of this Mutacin-1140 compound had not been tested. What would be the consequences of millions of bacteria in the mouth releasing this compound? The answer wasn’t clear, even though the archetypal compound in the family Mutacin-1140 belonged to was known to be very safe." This is obviously relevant evidence about the safety of Mutacin-1140. _How much_ evidence? Unasked, unanswered. (I have no idea how predictive the safety of other compounds in the same family is of another unstudied compound in that family, I'm not a biologist. But this is not an _unanswerable_ question.)
(Marginal conflict of interest: I know the Lumina founder socially. I have no financial interest in that venture or any of his other ventures. I have not taken Lumina myself.)
You already have, by deciding on a specific level of coverage for your auto insurance policy.
(That aside - really? Either you get into many more accidents than the average person, or you're not extrapolating out into the future. If given the chance to take the same bet at 1:1 odds every year, surely you'd then take it every year until you decided to stop driving for safety reasons?)
I agree that in practice, it's pretty likely that the author of the piece refuses to bet at least in part because betting substantial sums of money on outcomes that are non-central subjects of wagers (i.e. not an explicit game of chance, sports, politics, etc) is socially unusual. But I also think that if he was very confident he'd be happy to take the money.
Well, that sure sounds like a claim about how confident one should be about their understanding of reality. One most people disagree with, incidentally - do you own any equities?
Counterparty risk is a valid reason to avoid certain bets (or bet structures), but a totally separate objection from "betting is a weird thing that only weirdos do".
True!
> It means nothing to refuse to take a bet like that, let alone that the claims made in the article are suspect (which you seem to be implying).
False! If the author was sufficiently confident in their claims, they'd be happy to take the free money (or, if they're sufficiently liquidity-constrained, propose a smaller bet at similar terms). You can certainly argue that the practice of betting on one's beliefs is "weird" but that objection is circular. If you claim to see free money on the ground, and other people notice that you aren't picking it up, they would be correct to wonder why.
Like they already did in the last 2 years?
> Have such a significant technological breakthrough that GPT is able to take on entirely unseen new use cases, ones that are not currently possible or hypothesized as possible by any artificial intelligence researchers.
Huh, what are these use-cases which no AI researcher thinks AI is capable of solving? Does the author not realize that many employees at the leading AI labs (including OpenAI) are explicitly trying to build ASI? I am so confused????????
> Have these use cases be ones that are capable of both creating new jobs and entirely automating existing ones in such a way that it will validate the massive capital expenditures and infrastructural investment necessary to continue.
Why would they have to create new jobs? They just have to be good enough that OpenAI can charge enough money for them to be in the green.
OpenAI already has a $3.4 billion ARR! Most of that is _not_ enterprise sales.
>“All [AI models] would fit that criteria,” Pliny said.
This bit is particularly bad reporting. Putting aside the fact that the text of the bill no longer says "hazardous capability" (it's now "critical harm"), this is how a "critical harm" is defined (https://legiscan.com/CA/text/SB1047/2023):
(g) (1) “Critical harm” means any of the following harms caused or enabled by a covered model or covered model derivative: (A) The creation or use of a chemical, biological, radiological, or nuclear weapon in a manner that results in mass casualties. (B) Mass casualties or at least five hundred million dollars ($500,000,000) of damage resulting from cyberattacks on critical infrastructure, occurring either in a single incident or over multiple related incidents. (C) Mass casualties or at least five hundred million dollars ($500,000,000) of damage resulting from an artificial intelligence model autonomously engaging in conduct that would constitute a serious or violent felony under the Penal Code if undertaken by a human with the requisite mental state. (D) Other grave harms to public safety and security that are of comparable severity to the harms described in subparagraphs (A) to (C), inclusive. (2) “Critical harm” does not include harms caused or enabled by information that a covered model outputs if the information is otherwise publicly accessible. (3) On and after January 1, 2026, the dollar amounts in this subdivision shall be adjusted annually for inflation to the nearest one hundred dollars ($100) based on the change in the annual California Consumer Price Index for All Urban Consumers published by the Department of Industrial Relations for the most recent annual period ending on December 31 preceding the adjustment.
Given g(2), it is very likely that no models that are publicly available have the ability to cause a "critical harm" (i.e. where they can cause mass casualties or >$500m in infrastructure damage via the specified routes in ways that counterfactually depended on new information generated by the model).
Separately, it's very clear that LLMs have "world models" in most useful senses of the term. Ex: https://www.lesswrong.com/posts/nmxzr2zsjNtjaHh7x/actually-o...
I don't give much credit to the claim that it's impossible for current approaches to get us to any specific type or level of capabilities. We're doing program search over a very wide space of programs; what that can result in is an empirical question about both the space of possible programs and the training procedure (including the data distribution). Unfortunately it's one where we don't have a good way of making advance predictions, rather than "try it and find out".
This is extremely confused about the board's responsiblities and powers. A court would laugh this case out of court because the board _can just fire him_.