HNHacker News
TopNewBestAskShowJobs

adastra22

13,143 karma · joined February 17, 2022

mafried@gmail.com
submissionscomments
adastra22··on Mistral Large 4
Do you not see how twisted that reporting is? It is doing the bare minimum to affirm that something happened, while injecting enough uncertainty to make it sound like it’s an issue blown out of proportion by Western media.
adastra22··on Blindsight (Watts Novel)
Is it? Maybe I didn't get far enough into the series, but the parts I did read and the synopsis online are all derivative of old ideas. The "dark forest" idea is a very old and much debated solution to the Fermi paradox, and appears in other works (e.g. the wolves in Alastair Reynolds' Revelation Space series). The VR stuff was derivative of dozens of previous works. The rehydration of the solarians is not that different than the spiders freezing in A Deepness in the Sky by Vinge. Plenty of alien invasion and doom cult literature out there (too much, in my opinion).

But I took Dark Forest to mean the Dark Forest Theory which the post I was replying to mentioned, and that's just a common theory under a new name, which has been a commonly known solution to the Fermi paradox for so long that no one is credited with coming up with it. It's given the "deadly probe" name in a SETI review of the Fermi paradox literature in 1984.

adastra22··on Mold Linker Version 3.0.0 Release – Rewritten in Rust
They just threw out the code base and replaced it with a vibe coded rust port.
adastra22··on Blindsight (Watts Novel)
Another very unoriginal and shallow work.
adastra22··on Blindsight (Watts Novel)
They can’t, because consciousness is ill defined. It means something different to every person that engages in this debate.
adastra22··on All I wanted was a custom domain email
“Integration with Apple Mail” shy not just use iCloud with a custom domain?
adastra22··on Improper redaction reveals Google Data Center water and electricity usage
The water isn’t actually used. It is slightly warmed.
adastra22··on Improper redaction reveals Google Data Center water and electricity usage
Those reasons aren’t obvious to me. Why not just be transparent if that’s the case?
adastra22··on Homa: The end of TCP for AI clusters [video]
Please don’t make the primary link a video.
adastra22··on Turn off Apple Intelligence on macOS 27 and get its disk space back
I don’t want to make my search more ambiguous. That is the opposite of what I want.
adastra22··on Religious scholars met with Anthropic
That is, I’m not joking, one of the pope’s points.
adastra22··on Show HN: Pi pod – Run your pi coding agent in sandboxes on your own server
Kubernetes didn’t invent the word “pod”…
adastra22··on Getting the most out of Opus 5.5 in Claude and Claude Code
You are making the mistake of classifying models on a single linear axis, or even a multi axis basis set of all benchmarks. That just isn’t true. Each model is unique in its skills and capabilities and the way it approaches problems, in a way that is not represented in benchmarks. Fable is better at reviewing things. I don’t know how to explain it well but it is true. I trust Fable to do thorough reviews (sometimes too thorough) and to present its information in a dense but ordered way. Its output is equivalent to what you used to get from security firms doing code reviews. Having Opus do the work, and have fable do reviews (of the plan and implementation) is a good combo.
adastra22··on F.02 Decommission
You can find similar work in Anthropic model cards. Not pain, specifically, as I expect internal Anthropic review wouldn't have allowed such experiments, but they report pretty much identical emotional responses indicative of anger and frustration. The pain paper afaict is a repeat of the work Anthropic did & reported in their model cards, but for pain.
adastra22··on Getting the most out of Opus 5.5 in Claude and Claude Code
Some of this advice is really missing the mark. I will speak to just one I know well. Many of my frequently used prompts have “think through this step by step” because if you don’t, it only considers the task holistically rather than step by step, and different issues emerge in that frame of thinking. I see this. Often when doing planning, for example, it will not notice interdependencies between tasks until you force it to think through doing the whole thing step by step (task by task) then it will notice that step 2 requires a feature introduced by step 14. It wouldn’t notice otherwise.

Yes this has held true on Opus 5.5. I checked. It’s a massively better model, peer to Fable but with different strengths and weaknesses. But it still has this issue. Which to be fair, people do too. Planning is a learned skill.

I think what they’re saying is that the harness no longer uses a text search on “think” to engage reasoning modes. Fair, that’s good to know. That doesn’t mean asking the model to think a certain way doesn’t have the intended effect.

adastra22··on Barcodes are about to go extinct
I am irrationally annoyed that this isn't https://ref.gs1.org/17/382834828282/21/28312384021
adastra22··on F.02 Decommission
How would you feel if they threw a dog into the furnace? It is quite likely that AI agents experience at least that level of sentience.
adastra22··on F.02 Decommission
LLMs have pain signals: https://arxiv.org/abs/2609.16247
adastra22··on F.02 Decommission
Have you any empathy?
adastra22··on Gemini 4 Argon
I don't know if you noticed, but Opus 4.6 was peak for human-computer interaction. Everything has been fairly downhill from there despite better benchmarks, at least in that one regard. Opus 5 and 5.5 are clearly a step above in capabilities than 4.6, and I don't think anyone wants to go back, but 4.7 and 4.8 were arguably worse overall. I genuinely feel I got more done with 4.6 and often switched back, prior to 5 coming out.

Why? Because 4.6 actually talked like a human being. It actually organized its thoughts well, and got the main information across without the wall of text that makes your eyes glaze over. So from the perspective of human-computer interaction and maximizing the productivity of a developer+agent team, 4.7 and 4.8 were regressions. Despite much better benchmark performance.

Even if we consider autonomous agents, that benchmark is not indicative of how well they will interpret *your* requests. Or how well they will interact with other agents in a flock/swarm situation. The benchmark just doesn't cover this. (And the difference can be nontrivial! Sakana AI's published results show two generations of uplifting potential from better harnesses.)

adastra22··on Gemini 4 Argon
The process by which benchmarks are setup and run does not correspond at all to how human developers engage with a coding agent. At best it is a loose proxy, and often a bad one.

What benchmarks are usually good at is showing to what degree new models are better than old models. What they are not good at, by construction, is showing that harnesses are well adapted to how people use them.

adastra22··on Several vulnerabilities have been discovered in the Linux kernel
Unfortunately no, there are a lot more that fly under the radar.
adastra22··on Gemini 4 Argon
Benchmarks are a terrible judge for this.
adastra22··on Gemini 4 Argon
Just about every harness is. This is common knowledge, no? Frontier lab TUI tend to steal from the OSS harnesses not the other way around.
adastra22··on Gemini 4 Argon
Man good luck with finding that. People who wrote their own harness tend to self select into the type of people that DON’T self-author blog posts.
adastra22··on Gemini 4 Argon
Wtf. Can you turn that off? On CC I have compaction completely turned off. I’d rather hit the hard out of context limit at 1M.
adastra22··on NASA asked several former SR-71A staffers to help secret restart
Those are geosat communications, not surveillance. Surveillance is not happening in GEO.
adastra22··on NASA asked several former SR-71A staffers to help secret restart
Because GEO is about 100x further than LEO, which means the signal is about 10,000x weaker. Also the risk to geo communications from having multiple birds in the same slot is huge and treaty violating. Why put that risk and insane demand when you can still get global coverage from a dozen or so birds in high polar orbits?

Nobody does what you are saying, because physics, which leads me to believe that you are confused.

adastra22··on NASA asked several former SR-71A staffers to help secret restart
The NSA does have comparable things. You can even find amateur photos of them in this day and age. But they are not in GEO.
adastra22··on Livenerf: Has Opus 5.5 been nerfed yet?
Has otel? What is that?
Page 1 of 34Next →