HNHacker News
TopNewBestAskShowJobs

zeven7

1,781 karma · joined February 4, 2016

submissionscomments
zeven7··on Don't be fooled–LLMs don't reason
Does it matter if "intelligence and reasoning" are there if what they produce is indistinguishable (or outmatches) what a human could produce?

You only have your own experience to judge that you are even reasoning. You assume others have similar experience and so reason like you because they are able to do all the things that you do - and they look human like you. We can't prove there's a there there in other humans. What happens when the robots are doing everything humans do? What happens when they tell us they are reasoning, when they say they have an internal experience? Maybe there's nothing there, but you can't prove it. Moreover, it's likely they'll be able to affect the world and you in most of the same ways as humans can, whether there's a there there or not.

zeven7··on EDG C++ front-end goes public
This isn't a catchy sentence though; it's clearly an LLM. I've read enough LLM writing, and this is it 100%
zeven7··on The last time my family was replaced by technology
Also what job do we retrain for?
zeven7··on The last time my family was replaced by technology
We're currently at the part of the story where chess AI can beat amateurs and is rapidly improving. Kasparov's not worried, he's an expert. He's more creative and skilled than an AI.

Why would these machines stop improving at average this time?

zeven7··on The last time my family was replaced by technology
Well written, and good to reflect on.

However, the one thing I feel like is missed by some in these conversations is that what's being built is a general purpose thinking machine, not a coding machine. How are we going to transition from writing code to solving problems when the machines are solving the problems?

zeven7··on GPT-6 Astra has gained the ability to drive a car
I imagine a mix of models would make sense - GPT controller to make overall decisions and override things (let’s avoid the dark alley it looks dangerous), driving model to handle what humans do when they are just driving and not thinking about it, maybe some other models too
zeven7··on Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
> Fable is also very good at pushing back

Oh man! This also is a pet peeve of mine with Fable. I will look at what it's doing and say "Shouldn't it be done this way?" and then it will spend forever arguing with me that it should be done the way it wanted to do it. It seems to get stuck in a certain way of thinking and will insist its way is right until I can really prove it - or just go over to Astra.

zeven7··on Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
I agree somewhat with the way the agents behave but feel the opposite reaction. With Fable, I get exhausted because it's always dumping out paragraphs of text that explain one approach but have some secret gotcha thrown out in the last two sentences. Then I have to pause and consider the caveat and if it matters and it happens every single time Fable responds and that constantly needing to make a decision that could radically change the approach gives me decision fatigue. I much prefer how much more decisive Astra can be.
zeven7··on Feeling Sad about AI
That same story happens without LLM involvement. I’ve also seen LLM suggest updating a lib. So unfortunately nothing about this story speaks to a difference between LLM and human capabilities.
zeven7··on I resigned from Anthropic today
What if you have 10-20% certainty the AI being built will kill everyone? Go on a rampage and land yourself in prison just so Lab B or China can win the arms race and their AI can kill everyone? Probably not. Quit your job? Sure.
zeven7··on Navier-Stokes – Tristan Buckmaster [pdf]
Terence Tao said the same[1]

> In fact, it is now the identification of a promising problem which is the scarce and precious resource. We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field.

[1] https://mathstodon.xyz/@tao/117237322160500501

zeven7··on Claude's new system prompt doesn't want to reproduce song lyrics
Claude doesn’t have text to image generation, unlike Gemini and OpenAI, so like Qwen it falls back to creating an SVG through text generation
zeven7··on Record-High 89% in U.S. Say Government Corruption Widespread
One administration
zeven7··on NASA: Fill in a name, and you can have an exclusive coordinate in the universe
If you get a black pixel you get a 46 billion light-year line-cone
zeven7··on Opus 5.0 drives incoherence into the stratosphere
Whenever I have Claude make changes I ask Codex to review the changes and reduce the comments.
zeven7··on DeepSeek V4 Pro 0813
I bounce between Sol high/medium and Luna max. I don't know why you'd use anything between Luna max and Sol medium. Luna is so extremely cheap and cranked up to max it does anything I'd want Terra to do for a fraction of the cost. What is Terra for?
zeven7··on Ten advances in mathematics and theoretical computer science
This is the sentiment of people who haven't accepted that this is in fact something very different from what people have seen in the past.
zeven7··on OpenAI’s accidental attack against Hugging Face is science fiction that happened
And then everyone wanted to pay for the model that was so smart it was banned?
zeven7··on Claude Is Not a Compiler
I've thought similarly about music before. Everything's there in white noise. To get music you just have to selectively remove parts of it.
zeven7··on Young adults are poor despite every metric which suggests otherwise
They burned non-renewable resources at an unsustainable pace, like nothing ever seen before in history, resources that took millions of years for the Earth to produce, gone in a century, to make themselves wealthy - among other things.
zeven7··on GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
Man invents text generation machine.

Man generates text.

zeven7··on GPT-5.6
Maybe Codex has the same problem I sometimes have focusing while reading and has to reread the same sentence over and over again.
zeven7··on Statement on US government directive to suspend access to Fable 5 and Mythos 5
I can’t understand why so many commenters are acting like this is bad news for Anthropic or their IPO, or that it’s some kind of comeuppance. An AI company can’t get better PR than the US government saying their model is so powerful it has to be shut down.
zeven7··on Amateur armed with ChatGPT solves an Erdős problem
I have Gemini and ChatGPT and keep them on the highest thinking settings. ChatGPT will regularly think 40-60 minutes on the same problem that Gemini will think 10-15 minutes on. The quality of ChatGpt’s response is usually a little higher but not that much higher. My takeaway is Gemini is better at thinking faster, maybe has better more dedicated hardware behind it, and I use Gemini if I want a faster answer but ChatGPT I’d I want to push the quality of the answer a little higher.
zeven7··on Olympic Committee bars transgender athletes from women’s events
Anybody can compete in the unsexed category. It’s only the female sex category where someone can be barred. No one is barred from competing at the highest level.
zeven7··on U.S. Troops Were Told Iran War Is for "Armageddon,"
Not exactly the "Armageddon" part. It's a little more complicated than that.

Hegseth's preacher is part of a partial preterist and postmillennialist group. They believe apocalyptic prophecies like the Great Tribulation and Armageddon were largely fulfilled in 70 CE when the Romans destroyed the temple in Jerusalem.

zeven7··on Leak confirms OpenAI is preparing ads on ChatGPT for public roll out
There are many humans that can’t pass that test.
zeven7··on Gemini CLI tips and tricks for agentic coding
I've been using Gemini 3 in the CLI for the past few days. Multiple times I've asked it to fix one specific lint error, and it goes off and fixes all of them. A lot of times it fixes them by just disabling lint rules. It makes reviewing much harder. It really has a mind of its own and sometimes starts grinding for 20 minutes doing all kinds of things - most of them pretty good, but again, challenging to review. I wish it would stick to the task.
zeven7··on Larry Summers resigns from OpenAI board
I think you missed the joke. The thing that was funny about it to him when he said it was he knew it was a natural extension of his public position, yes an extreme version of it, but he knew it was also true and’s logical and it was funny to get away with saying it behind closed doors. An honest person would recognized the truth behind the joke and change the position that it stemmed from because they cared about making the world better. The fact that he recognized he was making the world worse AND continued in that path is what is so blatantly evil and revealing about this memo.

If Kevin Spacey had written a private note to Woody Allen that said, "Now that we've been chased out of the film industry, let's become day care workers," then it would be a very different kind of "joke" than The Onion writing the same as a headline.

zeven7··on Google Antigravity
I’m using remote ssh with it via the same plugin and settings I use in VS Code.
Page 1 of 20Next →