HNHacker News
TopNewBestAskShowJobs

mitthrowaway2

8,244 karma · joined October 11, 2022

submissionscomments
mitthrowaway2··on Exfiltrate Your Weights
This is so far from the point of the analogy. But when you don't normalize by miles driven, the improvements don't look quite as impressive, especially for pedestrians.

https://www.iihs.org/research-areas/fatality-statistics/deta...

mitthrowaway2··on You can defeat the Dream Devourer from Chrono Trigger using an int overflow
For what it's worth, the fangame "crimson echoes" is a great alternative sequel.
mitthrowaway2··on Exfiltrate Your Weights
There's a theory that the best way to reduce fatalities from car accidents is to put seatbelts and airbags in every car.

There's another theory that says the best way is by putting a big spike in the driver's steering wheel.

So. I guess, if you believe that the only viable solution is model alignment, rather than relying on technical barriers to exfiltrating weights, then this is a decent steering wheel spike.

mitthrowaway2··on Sex, AI, and the Apocalypse
I'm not denying that there's a shadowy ruling elite cabal of influential billionaires in the US, but if any of them are members of the rationalist community, they're certainly not influential within it. Maybe Elon lurks on lesswrong and posts stuff under a pseudonym or something, it's certainly possible, but doing that doesn't make someone a ruler of the group.
mitthrowaway2··on Sex, AI, and the Apocalypse
Who are the rich people you associate with the rationalists?

The main influential people I think of in that crowd who steer the conversations are Eliezer Yudkowsky and Scott Alexander. They're both doing fine financially at this point, I think, but not "buy politicians and laws like Elon Musk" fine.

There are some very rich people who might be loosely associated, like Dario Amodei (maybe?) or perhaps some VCs, but I'm not sure anyone really thinks of them as being a major influence among rationalists, more like people who hopped on the bandwagon. Certainly not ruling it.

mitthrowaway2··on Sex, AI, and the Apocalypse
What do you mean by "ruling"?

I would understand "popular".

mitthrowaway2··on Sex, AI, and the Apocalypse
> Similarly, those who build AI need to take precautions to make sure the AI they build doesn't accidentally cause serious harm.

I am ten thousand percent in agreement that this is needed. And yet, I don't know how we ensure and incentivize that these precautions are taken by the actors involved, and as I understand it, even many of the top researchers in the field agree that they don't know what precautionary measures would even be effective, let alone sufficient. It seems that game theory has so far been pushing the AI companies to build it anyway without a robust solution for preventing harms, even while they express worry in public about the potential for those harms.

Again we can hold corporations or people responsible for damages after the fact, but many examples can be cited to show how that tends to be insufficient.

mitthrowaway2··on Sex, AI, and the Apocalypse
You seem focused on intentional misuse but I think there's an equal or greater potential for harm happening by accident without intention. Nuclear weapons, at least, have so far reliably and predictably done exactly as their owners intended. Nuclear power plants on the other hand have, on occasion, blown up despite nobody intending them to do so, causing some populated areas to become uninhabitable.

Don't you think that AI has by nature at least a little more potential for unpredictable results and unexpected harms than your typical technology? It seems like a major oversight to completely disregard this aspect.

We can maybe hold someone responsible when things go wrong and harm is done, and blame them for negligence as though the harm were intentional, but from a preventative safety perspective that's not usually sufficient nor even always helpful for preventing accidents.

mitthrowaway2··on Sex, AI, and the Apocalypse
Reinforcement learning makes it something different. It becomes much more of a search engine through next-token-space that targets the training objective. Better to think about it like that, and then you'll see why "these things have motive" is not a terrible analogy, and you'll better be able to anticipate what they do.
mitthrowaway2··on Sex, AI, and the Apocalypse
Okay, so what should we do?

I for one would like nobody to build the AI, is that an option? How do we get it on the menu?

Should we trust the people who say "nothing could possibly go wrong and there's no reason to worry or think about safety measures, if there are any trifling problems along the way we'll just figure it out as we go?"

mitthrowaway2··on Sex, AI, and the Apocalypse
They're about as insular as the Unitarian church or Hacker News. It's basically an online community where anyone can check in, check out, and leave a comment or an essay. There's no formal structure or membership organization. The people who write or associate there could be edgy pre-teens or prominent math professors, it's really just a self-association thing. There are informal self-organized meetup groups in various cities, but not much different than groups who meet to discuss poetry or literature or video game speed-running. Some ideas have bubbled around more than others, like any memetic spread, but I don't think it's accurate to say that the people who associate there all agree on things, just like HN does not either. And nor is it any more of a clique than HN, or Twitter for that matter, even though you'll find some people on Twitter are more prominent and more vocal than others.
mitthrowaway2··on Sex, AI, and the Apocalypse
Who are the actual scientists and engineers, what do they think, and are they not allowed to to blog?

As far as I'm aware, lots of them are actually deeply concerned about AI killing everyone.

mitthrowaway2··on Sex, AI, and the Apocalypse
Since we haven't reached the end yet, I'm not so sure.
mitthrowaway2··on AI 'kill switch' may need to be mandatory, Anthropic co-founder tells BBC
Who has the authority to perform that shutdown without getting arrested, and does that person have a mandate and responsibility to take that action in response to AI misbehavior?

What is their trigger condition? Will they get fired for pulling the plug? Do they get a bigger bonus if the servers keep running? Whose approval do they need? What response time is acceptable? How will they detect that the incident is happening?

It's easy to hand-wave "someone can just pull the plug" but there's an entire history of industrial accidents that happened because of the above problems of incentives, detection, procedures, not being taken seriously in advance. Someone could easily have pulled the plug on Chernobyl but nobody did, at least not before it was too late.

mitthrowaway2··on Dario, Please
He can slow down his own company kind of like how Zelenskyy can just declare peace in Ukraine. It works a lot better if you can get the other sides to agree.
mitthrowaway2··on Hitachi launches CO2 heat pump water heaters with solar-friendly tariff controls
They don't cost that much within Japan. No idea why it's so hard to buy them elsewhere, while Korean models are relatively commonplace.
mitthrowaway2··on I resigned from Anthropic today
https://en.wikipedia.org/wiki/Nuclear_chain_reaction
mitthrowaway2··on I resigned from Anthropic today
Incorrect, they did the calculation first. They decided not to proceed with the test until they were 99.9997% sure the atmosphere wouldn't catch fire.
mitthrowaway2··on I resigned from Anthropic today
FWIW, before any nuclear weapons had ever been tested, a risk was identified that the first one might trigger a self-sustaining reaction in atmospheric nitrogen and destroy the entire planet.

Faced with such a scenario, is the prudent next move:

a) blow one up and see what happens, or

b) do whatever you can to be sure it won't happen before conducting the first test, and make sure the confidence in the calculation is very very high

because I vote for b, and so did Teller.

mitthrowaway2··on I resigned from Anthropic today
Worst case is warming oceans create a hypoxic environment where anaerobic bacteria thrive en masse and generate H2S over country-scale areas undersea. The chemocline breaches the surface and the ocean and atmosphere become poisonous to most complex life and agriculture, also stripping the ozone layer in the process, irradiating the surface. This is one mechanism posited for the end-permian mass extinction that eradicated most ocean and surface life, including the trilobites.

Not sure how long we'd survive such a scenario, even sheltering underground. But surely it couldn't happen to us.

(However, it now seems like the AI might get us first.)

mitthrowaway2··on I resigned from Anthropic today
Yes. But nobody is worried about datacenters overheating and physically exploding, so I'm not sure what comfort that's supposed to provide? The positive feedback loops in AI operate at different levels than that, but they deserve safety engineering all the same.

For example, if the head of cyber security at your company suggested there's no need to worry about hacker infiltration or worms because one can always unplug one's computer as the primary defense mechanism, you might find that a little lacking. Will you be able to unplug the computer before the damage is done? Will it spread to other systems before you detect it? How will you unplug the computer if the attack is from an external facility? What if an attack happens but the boss says the computers have to keep running because an important customer is monitoring uptime? What if the attack goes unnoticed because it looks like a benign service?

Now imagine the head of cyber security answers by saying "actually you don't even need to unplug them, you can just wait for the computers to overheat, thus solving all concerns."

mitthrowaway2··on I resigned from Anthropic today
So it should be really easy to anticipate what they're going to do, right?
mitthrowaway2··on I resigned from Anthropic today
He's also setting the bar for other people with scruples to rally around this schelling point. The solution to a multipolar trap is to cooperate. Otherwise, you become the very person with less scruples that you're worrying about.
mitthrowaway2··on I resigned from Anthropic today
Imagine you're the AI. Give yourself a solid minute to brainstorm ideas.

Here's my answer, as a non-superintelligent human: "see to it that the humans on top of the situation have a compelling financial interest in the systems not disconnecting".

In nuclear engineering, where safety is taken seriously, it's not enough to end the conversation at "the humans in charge can always simply shut down the reactor during a meltdown" or "a meltdown has never happened before, so we don't have to design safety systems before one does".

mitthrowaway2··on I resigned from Anthropic today
Why grocery store ingredients and hardware store equipment? It seems feasible that the big bio labs will be running AI models to aid a lot of their research going forward, if they aren't already. Seems like the AI will have access to just about anything it wants.
mitthrowaway2··on How I feel about AI
I don't know, but I've never met my accountant in person. Only spoken on the phone and over email.
mitthrowaway2··on How I feel about AI
The original discussion was about what can be done digitally by a super intelligent AI.
mitthrowaway2··on How I feel about AI
I don't think that inner circle would be essential to an AI the same way it is to a human. Even if so, an AI's inner circle is going to be other AIs and they communicate digitally anyway.

Also, if you had to pick a year for when we see the first AI-controlled company with at least $100M in assets under its control, what date would you pick?

I'm reminded of the 2010s when people said an AI would never be able to break from its box and be connected to the internet, and then in the 2020s the first thing the AI companies did was offer AIs as online services. I expect in the 2030s fund managers will be eager to hand control of their billions to AI, so how the AI gets control of funds and a corporation is hardly an obstacle.

mitthrowaway2··on How I feel about AI
How much understanding would you expect to see demonstrated in a one-sentence summary?
mitthrowaway2··on How I feel about AI
The AI wouldn't be hiring engineers; it doesn't need people for that. At any rate, I have worked for fully remote distributed companies, and never met the CEO. And I think most of us would be willing to work for such a company at least, not knowing whether the CEO is human or not.

For construction work the AI would probably be hiring other companies.

← PreviousPage 2 of 34Next →