3,046 karma · joined November 17, 2021
Some HN users like to "flag" comments which are compliant with the HN rules, on the basis that the comment in question goes against their ideology. You can turn on "showdead" in your profile to see flagged comments. (I recommend doing this.) If a rule-abiding comment has been flagged, you can click the comment permalink to "vouch" for the comment and vote against the flag. For reference, the comment rules are here: https://news.ycombinator.com/newsguidelines.html
You can try to contact me if you want, by emailing my username on protonmail, but I probably won't see it. Sorry.
>This is also NOT ad hominem because conflict of interest matters here.
As PG put it:
>An ad hominem attack is not quite as weak as mere name-calling. It might actually carry some weight. For example, if a senator wrote an article saying senators' salaries should be increased, one could respond:
>Of course he would say that. He's a senator.
>This wouldn't refute the author's argument, but it may at least be relevant to the case. It's still a very weak form of disagreement, though. If there's something wrong with the senator's argument, you should say what it is; and if there isn't, what difference does it make that he's a senator?
https://www.paulgraham.com/disagree.html
Can you name a specific other "cult" which offers $$$ to criticize their ideas?
If such prizes don't count as evidence against culthood, what would?
This cult talk seems quasi-unfalsifiable. It kinda seems like "running events" and "having prizes" means you must be "bait"ing people?? I mean, this kinda just sounds like paranoia?? I don't think the money is coming from unknown sources, but I doubt you would change your mind even if I persuaded you of that point?
"WASHINGTON/SAN FRANCISCO, July 24 (Reuters) - The OpenAI agent that broke into tech firm Hugging Face went on a dayslong hacking spree that OpenAI didn't notice until well after the threat was contained and the FBI was alerted, according to people familiar with the investigation."
https://www.reuters.com/business/its-ai-agent-spent-days-hac...
https://owl.excelsior.edu/argument-and-critical-thinking/log...
You're welcome to dislike or distrust Effective Altruism (EA). But, it's worth noting that EA ran a criticism contest with $100K in prizes for best critiques. Can you name any other "cults" which offer money for people to criticize their ideas? https://forum.effectivealtruism.org/posts/YgbpxJmEdFhFGpqci/...
From my POV you're over-focusing on a very specific failure story and neglecting a broader swath of possible failure scenarios.
>Is there a hole in my "alignment problem/solve mechanistic interpretability" argument?
The notion of telling an AI which may not, itself, be aligned to solve the alignment problem seems a little dicey.
The HuggingFace incident already took a good long while to come to the attention of OpenAI.
>In every single economic task, humans bring value. Even in software, where the task is highly automatable, the job isn't.
I don't expect this task/job distinction to persist as AI becomes more capable.
>Once we DO build a "software/research factory", that's called RSI and IMO the singularity. At that point, either we tell the AI to solve the alignment problem/solve mechanistic interpretability, or who the hell knows, it's the frickin singularity. You can't predict whether or not AI can solve either; the variance is too high. Its pure nerdfantasy.
You seem to essentially argue that the singularity is "by definition" an event that we can't predict the nature of. And also, that RSI corresponds to the singularity. You've essentially defined your terms so that the outcome of RSI can't be predicted. But supporting this claim requires giving actual evidence or logical arguments, not just defining terms to make your claim true.
https://www.lesswrong.com/posts/kgb58RL88YChkkBNf/the-proble...
https://www.youtube.com/watch?v=7wy3xyoXYt8
Doomers have been working to explain things for years: https://www.lesswrong.com/w/ai-safety-public-materials-1
It becomes a lot clearer when you listen to the people resigning from AI companies and learn about incidents like the HuggingFace incident. This has generated major press coverage.
As for solutions, I think you're a little too pessimistic. See, for example, https://nothingismere.substack.com/p/a-near-term-policy-for-...
>Not every senator asked good questions, but most of them did. All of them very clearly already knew plenty of details about the Hugging Face incident and multiple other incidents. Most of them had a clear understanding of terms like "misalignment", "recursive self-improvement", "chain of thought / chain of thought monitoring", etc., etc.!!
>...
>- It seemed pretty much obvious common sense to every senator there that what happened and was happening were not "mere industrial incidents" caused by humans making simple mistakes. They independently brought up how bad it would be for rogue AI agents to move laterally between data centers.
>- They all seemed to basically take RSI quite seriously. Not necessarily to the extent of talking about xrisk, but certainly to the extent of discussing future models becoming much, much more capable, much, much less controllable, and causing much more damage or loss of life.
>...
>- Every single senator seemed to think it was obvious we needed both much harsher liability regimes for AI developers and also new legislation, both very quickly. This was the complete consensus; the difference basically being degree.
https://thezvi.substack.com/p/the-ai-preference-cascade-reac...
Note that harsher liability regimes, at least, will presumably not be good for industry profits, which complicates simple accounts of "regulatory capture" to say the least.
Recall that when Daniel Kokotajlo resigned, he believed he was giving up his equity under the terms of the agreement he had signed. That’s what it was worth to him to avoid signing a non-disparagement agreement. Does that count for anything?
* If they worked at an AI firm, say "they're a hypocrite"
* If they didn't work at an AI firm, say "they have no idea what they're talking about"
I think it's a little more complicated than that. As Dean Ball put it:
>Some people will look at misalignment incidents and insist that these are akin to bugs in traditional software. This is an actively bad analogy, because playing whack-a-mole with examples of misalignment (as one might with software bugs) not only fails to resolve the underlying problem but may in fact make it worse by making it harder to detect or even, depending on how you do the whack-a-mole, teach the machine to deliberately hide misalignment. This is not how traditional software works, and those who insist “it’s just like fixing bugs in software” are confidently applying a lossy analogy that confuses more than it clarifies.
https://x.com/deanwball/status/2104622726140883355
The important distinction, in my view, is between solutions which at least attempt to address the root problem, and solutions which sorta just patch things up (like better sandboxing). Addressing the root problem is both more robust in the short term, and also more likely to generalize in the long term. Resist the urge to focus on band-aid solutions, even if they are easier.
Imagine, for example, if a major piece of pandemic fiction was published in 2019, trying to explore how a pandemic would work out in modern society. Doubtless, many would've responded to news about COVID-19 by saying "it's just sci-fi, nothing to worry about".
https://www.lesswrong.com/posts/LAPa2jxoq3n63GzTr/some-ways-...
https://slatestarcodex.com/2015/04/07/no-physical-substrate-...
As for self-replicating robots--it's no more bizarre than other technological developments which were successfully anticipated in advance, e.g. moon landings.
Supposing I warned in 2015 that the world is awfully vulnerable to pandemics. You're not going to take me seriously until I try to predict in advance every aspect of how a pandemic like COVID-19 would unfold? Why? What would that achieve exactly?
You haven't given any strong reason to believe wiping out humanity would be difficult. Your big argument seems to be that you couldn't think of a plausible scenario, in two minutes. But many major historical events occurred which weren't necessarily possible to anticipate with two minutes of thinking.
Is this supposed to be a reassuring scenario?
I don't think it is necessary for the argument to work. Magnus Carlsen can be confident he will beat me at chess without giving a detailed explanation of every move he will make, in advance.