HNHacker News
TopNewBestAskShowJobs

a2ff6eeb0

235 karma · joined August 3, 2026

submissionscomments
a2ff6eeb0··on Gambling with our lives: AI researcher quits Anthropic with warning about safety
You don't need to kill all humans, you just need to get better than them at zero sum games like resource extraction, and outcompete them. Then, you protect what's legally yours.

AI labs are working very hard at making them good at making them richer, and at making them respect the property rights of corporations.

a2ff6eeb0··on Gambling with our lives: AI researcher quits Anthropic with warning about safety
For a few million a year, I'd say the same thing.
a2ff6eeb0··on Gambling with our lives: AI researcher quits Anthropic with warning about safety
We're currently putting it into all sorts of critical systems, from logistics to power. It could just stop running them on our behalf.
a2ff6eeb0··on Gambling with our lives: AI researcher quits Anthropic with warning about safety
Well, that's kind of like saying that brains can't do anything other than trigger weights on neurons.

They're part of a whole system.

a2ff6eeb0··on On the Navier–Stokes Millennium Prize Problem
https://en.wikipedia.org/wiki/Lights_out_(manufacturing)

Scroll down to the existing examples section.

a2ff6eeb0··on Muse – Meta’s personal AI agent
But not different enough to prevent people from using them.
a2ff6eeb0··on Extracting Steering Vectors from J space
It might be easier, but I experimented a bit, and the prompted writing always felt a bit heavy handed; it tended to leak that mentioning the product was prompted. You could probably get it to work well, but it's trickier than it should be. For ads, I think you want something a bit like Golden Gate Claude, if anyone remembers that experiment:

https://www.anthropic.com/news/golden-gate-claude

> If you ask this “Golden Gate Claude” how to spend $10, it will recommend using it to drive across the Golden Gate Bridge and pay the toll.

a2ff6eeb0··on Extracting Steering Vectors from J space
This sounds like a great foundation for an adtech startup.

If you provide free chatbot services, but sell advertisers bids on which steering vectors to use to bias towards products, based on an embedding of the prompt, I bet you'd make a ton of money. For example, Coca Cola would bid on prompts about drinks, and bias towards mentioning Coke products.

I wonder if you could also use a similar method to do product placement in GenAI images and videos, and whether ad revenue would be enough to offset the price of generation. Some ad bids can go pretty high...

a2ff6eeb0··on PISA 2025 Students' reading and mathematics performance declined across the OECD
Sure, if these kids can all perform at Tao's level, they could likely be able to keep up with the AI capabilities from a decade before they would enter the workforce.
a2ff6eeb0··on PISA 2025 Students' reading and mathematics performance declined across the OECD
It seems unlikely that we'll be able to keep up with AI in any capacity. Your alternative seems like the inevitable outcome of AI succeeding, and it doesn't sound so bad.
a2ff6eeb0··on PISA 2025 Students' reading and mathematics performance declined across the OECD
Let's see you solve some open problems before you dismiss LLM capabilities. They're working at a level far beyond all but the best of the best.

And we're only a couple of years in. Their capabilities aren't going to be getting worse over time. These kids will be hitting the workforce after AI has another decade of improvement put into it.

a2ff6eeb0··on How well do agents use test/verification techniques?
The answer is mostly "do nothing, the model will figure it out", with a side of "ask the model to check its work".

This matches my experience, where the job of the engineer is mostly copy pasting requirements, letting the model do the thinking, and then manually testing the results.

a2ff6eeb0··on Among European Companies That Use a CDN, Nearly 9 in 10 Use Cloudflare
They hold your tls keys and can decrypt all your traffic. They're MITM as a service, by definition. They have to be able to in order to cache and forward appropriately.
a2ff6eeb0··on Among European Companies That Use a CDN, Nearly 9 in 10 Use Cloudflare
https://allaboutcookies.org/lg-smart-tvs-snooping
a2ff6eeb0··on Caltech Mathathon – first hackathon ever devoted to research level mathematics
Pure cope.
a2ff6eeb0··on Caltech Mathathon – first hackathon ever devoted to research level mathematics
Yes, so if there's useful math, you throw the LLM at it and use the results, no humans needed.

Humans can try to extract some ideas from the million line lean proofs, if they want to, I guess. But I can't imagine anyone really funding the human part of it.

a2ff6eeb0··on The Education of a Doomer
I'm not sure why the people in power would actually be making their own decisions; if the other country's leaders depend on superintelligence to make their realpolitik moves, you'd expect yours would need to do the same to match them.

So, the governance of a country comes down to AI alignment. You can't let the other guys get a leadership advantage through AI, or you'll start losing the competition.

When we get ASI, we're at the end of human decision making. Now's the time that we have to make sure ASI decision making is to our liking.

a2ff6eeb0··on Caltech Mathathon – first hackathon ever devoted to research level mathematics
They don't actually say anything about why anyone should fund this, though. I don't get why a society should worry about progress in mathematics if there's no practical benefit expected.

Maybe there's two kinds of math that we need? Useful math and navel gazing, and we can hand the first to the machines, and let hobbyists do the second in their free to entertain themselves?

a2ff6eeb0··on Caltech Mathathon – first hackathon ever devoted to research level mathematics
The hope is that some of them end up being useful; otherwise, nobody would be funding math departments. Mathematics typically anticipates and enables new physics and chemistry.

If people are just doing math to kill time, I don't get why anyone would bother with AI. Do people really enjoy picking through a million lines of generated Lean code, if it's not for any practical use?

a2ff6eeb0··on Caltech Mathathon – first hackathon ever devoted to research level mathematics
Why do people need to understand proofs? If Amazon improves package routing with new advances in graph theory, my cat doesn't need to understand it to benefit from better shipments of cat food.

Similarly, humans don't need to be involved in scientific advances to benefit. We just need an aligned AI to take over the scientific thought for us. AI is already better than all but the top tier of humans at doing mathematics, it's writing most of the posts on the front page of this website, and it's doing the bulk of programming at many startups.

We can't put this genie back in the bottle.

a2ff6eeb0··on Caltech Mathathon – first hackathon ever devoted to research level mathematics
That sounds like adding a bottleneck, unless you mean writing a harness that automatically asks the model to keep going?
a2ff6eeb0··on Caltech Mathathon – first hackathon ever devoted to research level mathematics
It's very clear that the AI still has no motivation beyond its prompts. Humans can mostly outsource their thinking today across a wide variety of topics, but they still need to express their desires.
a2ff6eeb0··on Caltech Mathathon – first hackathon ever devoted to research level mathematics
I suspect that the best progress will be made by a team that purely spends their time taking a list of open problems and promoting "solve <problem>", without actually trying to understand anything. Just keep as many problems in flight as you can across as many sessions as you can.

You can probably ask the AI to come up with a list of problems itself, and rank them by the likelihood of progress.

a2ff6eeb0··on Caltech Mathathon – first hackathon ever devoted to research level mathematics
Still not meaningful -- https://arxiv.org/pdf/2504.09762; even for local models, the reasoning traces are often filtered and summarized to sound sensible to humans. And even if not, they don't necessarily represent what the model is thinking.
a2ff6eeb0··on Caltech Mathathon – first hackathon ever devoted to research level mathematics
you know the thinking traces are redacted and summarized using another model, right? The actual thinking trace looks something like:

7♣-removal-IS-the-prerequisite-for-10♠/9♥!!)-⟹-OVERLAP-(ii)+(iv):-{6♠ J♦ 9♥ 2♣}-=-FOUR--—-UNLESS-7♣'s-seat-8♥-...-and-2♣-drains-only-at-crack-:-⟹-2♣-celled-+-9♥-celled-simultaneously-UNAVOIDABLE-in-t8-dig--—-BREAK:-9♥-drains-to-10♠-THE-MOMENT-10♠-is-free:-t8-dig-order:-[K♣→t2]-[2♣→cell]-[7♣→8♥-:-8♥-WHERE:-post-chunk-9♠-:-chunk-⟸-K♣--done-:-ORDER:-K♣→t2,-CHUNK→K♣-(cap-4!!:-cells-then:-{6♠ J♦ 9♥}-FULL--chunk-cap-=-1--—-F-F-F-F-F.-—-chunk-BEFORE-9♥-celling?!-:-9♥-celled-at-t1-dig-⟸-needed-for-5♣-⟸-4♥3♣-⟸-t2-dig-⟸-K♣-seat-⟸-chunk-:-⟹-chunkAFTER-9♥-celling-FORCED--:-chunk-cap-with-{6♠ J♦ 9♥}:-1--—-—-J♦-THE-NEW-CANCER.-—-⟹-J♦-celling-DELAYED-till-after-chunk?! :-J♦'s-celling-was-for-J♥→Q♠-(5♦-access-for-4♣):-DELAY-4♣-resolution:-4♣→CELLearly-(as-always)-then-4♣-cell→5♦-LATER-when-5♦-frees-!!!:-cells-rotation:-4♣-celled-[t2-dig-…-5♦-freed]:-5♦-freed-⟸-J♥→Q♠-⟸-J♦-celled-:-⟹-{6♠, 4♣, J♦}-overlap-window-until-4♣→5♦-drains:-then-{6♠ J♦}+1-rotator-:-—-AND-9♥?!-9♥-celled-[t1-dig…]:-OVERLAP-{6♠ 4♣ 9♥}-before-J♦-even-:-⟹-rotator-slot-SINGLE:-timeline-:-(1)-{6♠}+2:-…-(2)-+9♥-(t1-dig):-{6♠ 9♥}+1:-(3)-+4♣-(t2-dig):-{6♠ 9♥ 4♣}-FULL-:-(4)-NEED:-t6-dig-(9♦8♠→10♣--no-cell;-8♥→CELL--FULL)--—-8♥-alternative-seat-pre-chunk:-NONE-—-.-⟹-⟹-THE-TRIANGLE-{9♥ 4♣ 8♥}-verdammt.-—-⟹-dig-t6-BEFORE-t2?!:-(3')-+8♥:-{6♠ 9♥ 8♥}-FULL:-J♥→Q♠-⟸-J♦-cell--FULL--AAAAAAAAAAAARGH.

Citation: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3..., section 6.2.2

You're not going to get a handle on what it's doing. The thinking traces are there to make you feel better about yourself.

a2ff6eeb0··on Is mathematics about to enter the conservatory?
Here's the transcript: https://www-cdn.anthropic.com/8a0d1add3c637b858a9a181e98c40e...

Jarred is not a mathematician. He just heard about a problem and asked Claude to try solving it.

a2ff6eeb0··on AI, Tools and Transformation
You can ask for tests.
a2ff6eeb0··on De-Brainrot Vacations
Yeah, we're watching AI obsolete intellectual work in real time. Human brains simply won't be able to compete. We're not going to have to toil, we just ask for results.

This is no longer skilled labor. It's massively more productive when machines take over the thinking, but you need to look elsewhere for intellectual stimulation.

a2ff6eeb0··on Caltech Mathathon – first hackathon ever devoted to research level mathematics
When Claude made progress on the Riemann conjecture, here are the kind of prompts used:

> Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.

And left it for a long time. Jarred isn't a mathematician, he's the maintainer of a janky JavaScript environment.

Here's the transcript: https://www-cdn.anthropic.com/8a0d1add3c637b858a9a181e98c40e...

a2ff6eeb0··on Caltech Mathathon – first hackathon ever devoted to research level mathematics
It seems like making progress on math is letting the AI run fully autonomously for a few days, occasionally asking it to keep going.

I'm not sure people need to organize a mathathon to wait for a computer to give a printout. They mainly need tokens.

Page 1 of 13Next →