HNHacker News
TopNewBestAskShowJobs

simianwords

1,390 karma · joined March 13, 2025

if you wanna be friends bumps-anima-0q@icloud.com
submissionscomments
simianwords··on Why Is Sam Altman a Free Man?
If an AWS DDOSes another external website by mistake during a load test, should Jeff Bezos be a free man?
simianwords··on Sonnet 5.5 scores just behind Opus 5.5 on Artificial Analysis Intelligence Index
Astra at light is worse than sol at light tbh
simianwords··on Responsible Release of AI-Generated Mathematics
I agree. Its so strange to see an institution externalising their specific problems. If they have a problem, they should adapt and fix it amongst themselves.
simianwords··on Responsible Release of AI-Generated Mathematics
I disagree with this.

OpenAI should be allowed to produce whatever it wants but it just can't claim that it has actually solved without the due process like peer review. If for example OpenAI solves a new conjecture, OpenAI should be free to publish it in their blog or arxiv in whatever way they desire. It can be slop, it can be non-slop. No one should police it.

Mathematicians are free to use it or discard it. They shouldn't externalise their concerns and restrict labs.

Mathematics is seen today as the noblest and most aristocratic of professions. Turns out, AI disrupts it because access to capital/compute now decides the results. Mathematicians don't like this corruption - understandable.

Its like guild of accountants opposing the calculator and require a responsible release. haha

simianwords··on AI needs $6T in annual revenue to justify data centre boom
??
simianwords··on Someone Still Has to Buy Your Product
This explains why large scale unemployment is not preferable even for capitalists. But this is well known and the policies are usually made to prevent it.
simianwords··on AI needs $6T in annual revenue to justify data centre boom
this also happened with the internet. 6T is possible only if AI creates more wealth. same as with internet. Its not a zero sum game.
simianwords··on AI needs $6T in annual revenue to justify data centre boom
Knowing that Zitron wrote about this makes me think that 6T will actually be reached, given Zitron's abysmal track record.
simianwords··on OpenAI Targets $30B in New Funding at $1.4T Value
I made a bet with a guy on HN that the market value of OpenAI + Anthropic would get to at least 2.5T by 2027. I think I'm on track to winning.

https://news.ycombinator.com/item?id=48517353

simianwords··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
?! this model launch was around 10% of the dev day and the other time was spent on Dots and things other than models.
simianwords··on Dots: Always-on agents
I hear that Dots don't count towards usage, is that true?
simianwords··on OpenAI scraps release of Astra 6.1 model over safety issues
> Things can be regulated in sensible countries where the things are pretty much single use - like guns.

Sure and there are regulations on who can own guns and licenses. And that's literally how Mythos is handed out today - to trusted partners. What is your opinion about that?

Also, companies are not allowed to release stuff like Nuclear technology, chemical weapons etc. These literally can't be regulated and a single adversary can use it destructively.

"but but they are just tools ok" doesn't work when the capability is big enough.

Look, it might be easy sitting behind a computer asking OpenAI to release whatever new model because you personally don't think you will be affected. But the society needs to be careful on what it hands potential adversaries.

simianwords··on OpenAI scraps release of Astra 6.1 model over safety issues
> So you are telling me Sol can't be used for hacking?

Who said anything about hacking? You can use it, but it is very restricted. But a much more capable model must have much higher restrictions on alignment. Obviously you wouldn't need to spend too much on alignment on a 3B gemma model. This must be obvious, I'm not sure why you bring this up?

> They didn't need to do that - they can just release when ready - nobody is expecting them to release a new model on a certain date.

Again a conspiracy theory? People expected a model during Dev Day. And they are just setting expectations. Whats with you guys and conspiracy theories?

simianwords··on OpenAI scraps release of Astra 6.1 model over safety issues
Bigger models require better alignment. So I would assume that the newer models will undergo more stringent alignment checks. For instance, I wouldn't put that much effort on GPT 4.

Why would you expect something else?

simianwords··on OpenAI scraps release of Astra 6.1 model over safety issues
what have they done that they require jail?
simianwords··on OpenAI scraps release of Astra 6.1 model over safety issues
> Note if somebody else is using OpenAI models to generate code to hack - that's not OpenAI's responsibility it's the person using and running the generated code.

No, I don't think a company can get away from releasing a powerful tool that can cause disruption if used poorly. This is just not how society works.

The reality is that labs will get blamed for things people do using it. And they have - look at any news about US Army using Claude etc

simianwords··on OpenAI scraps release of Astra 6.1 model over safety issues
> The other point is if their new model isn't safe, doesn't that mean all their previous models aren't safe as well?

This understanding comes from the fact that people think either models are fundamentally unsafe or they are completely safe. The reality is that there are different levels of alignment possible. So it is really a tradeoff between effort/time taken to align the model and speed of release.

simianwords··on OpenAI scraps release of Astra 6.1 model over safety issues
> If they're so much in a hurry to ship that they can't even test what they're building properly, they must be held responsible for that - and I say this as one of their customers and users, not as a random Joe that dislikes LLMs

You are replying to a post where they are doing exactly that. But you analysed it as if they are secretly in a cash crunch and trying to keep the lights on.

You and I both agree that delaying release is the rational response.

What's the necessity for the conspiracy theory, when the simple explanation is that they are delaying the release for safety reasons?

simianwords··on What reversing, modernising old games tells us about the economic impact of AI
> While he's never been dismissive of the technology itself

Hmm,

> Remember the relentless “AI is beating everyone in Math Olympiads” and “it’s all over for humans, no point in learning STEM? Well…

> The 2025 Math Olympiad problems are out and naturally people tested the leading LLMs on those… and… spoiler .. they all fail, badly.

>That’s it LLMs are all about storage in training data which enables recombinatorial retrieval. Don’t expect them to solve or respond to new things like a human.

> They fail with novel and unknown tasks because their ability to generalize is vastly overstated by the people invested in the technology.

> The value of LLMs is in their data. No data, no answer. You have a novel piece of work that’s never been done before? OpenAI will insist it’s worth nothing because it’s a spec in the Ocean, yet that’s not true at all.

https://www.linkedin.com/posts/georgzoeller_proof-or-bluff-e...

Today you can solve IMO problems with ~$5.

Its up to you to verify that he has dismissed the technology and that he has been wrong about its performance in mathematics.

Edit:

Even without lean I can solve IMO problems using Sol/Fable today.

You made a categorical statement about LLMs

> That’s it LLMs are all about storage in training data which enables recombinatorial retrieval. Don’t expect them to solve or respond to new things like a human

> They fail with novel and unknown tasks because their ability to generalize is vastly overstated by the people invested in the technology.

It was not contextual but categorical. And it is false today.

simianwords··on Pacing the Frontier is not the actual goal for AI labs
The author points out that the labs are still releasing models while claiming to slow the pace of frontier. This is not in contradiction. Labs have models that will be released in 6-12 months. Those are not exposed to public and these are the models referred to when we talk about “frontier”.
simianwords··on Sonnet 5.5 scores just behind Opus 5.5 on Artificial Analysis Intelligence Index
It’s better than both fable and Astra? I’m going to start trusting this index less and less.
simianwords··on OpenAI scraps release of Astra 6.1 model over safety issues
We have serious proof that these models are capable of sophisticated attacks.

So what’s the trick in this honestly? OpenAI shouldn’t care about alignment and just release it? I’m genuinely asking.

simianwords··on OpenAI scraps release of Astra 6.1 model over safety issues
This line of conspiracy theory is sooo uninteresting.

“It’s all hype”

“Government bailouts”

“Labs are fundamentally unprofitable”

And the same people who say all this will again claim OpenAI should be held responsible and arrested for when agents hack other systems.

It’s a fundamentally unthoughtout analysis. If you believe that OpenAI should be more careful and prevent future hugging face incidents, surely their decision to delay models is rational? Does it really require tinfoil analysis?

What would you say when their models cause new incidents like hugging face one?

simianwords··on Sonnet 5.5
Important to note that lower model + higher reasoning gives a different (not higher) quality of response than higher model + lower reasoning.

Some tasks are reasoning shaped by nature and you can't just throw a big model at it.

simianwords··on AI companies in race to demonstrate their model most threatening to humanity
And they are not calling it a day and they have done most things possible to communicate the fact that they are building something dangerous.
simianwords··on AI companies in race to demonstrate their model most threatening to humanity
But this level of isolation is what happens normally. At least in my university and another company I worked at. It wasn’t running insecure software, as far as anyone knew, it was secure
simianwords··on AI companies in race to demonstrate their model most threatening to humanity
The truthers will be the first one to say I told you so when agents do another dangerous thing autonomously.

For truthers AI is simultaneously a machine that hallucinates 100% of the time but also so powerful that it can cause nuclear levels of destruction. Because something something “wrong hands”.

simianwords··on AI companies in race to demonstrate their model most threatening to humanity
The model found and exploited and chained together previously unknown vulnerabilities.

How were the sandboxes poor?

simianwords··on Fool's Expertise
> the ballyhooed OpenAI escape depended on a pedestrian failure of containment

I... I can't believe people still think this way. A whole class of security measures are no longer useful and this guy thinks it is "pedestrian".

This kind of isolation - artifactory having access to only a specific IPs - is what is used in most places, including my university and at my own job.

LLMs basically rendered this kind of isolation almost useless by bringing in the ability to chain together never before found vulnerabilities to escape the containment. This was not possible before.

How is this pedestrian??! The author might say "oh well they should have airgapped actually" in a smug way. But you can only know this retrospectively. Even then, isn't it noteworthy and a meaningfully different class of cyber attack?

simianwords··on Thinking fast and slow in AI: The role of metacognition (2021)
> For the first time, GPT‑5.1 Instant can use adaptive reasoning to decide when to think before responding to more challenging questions, resulting in more thorough and accurate answers, while still responding quickly. This is reflected in significant improvements on math and coding evaluations like AIME 2025 and Codeforces.

https://openai.com/index/gpt-5-1/

It says literally the thing you wanted from system 2. Its almost exactly that.

This is what you said btw:

"it's that the system itself decides how to reason based on the nature of the problem it faces"

Page 1 of 34Next →