HNHacker News
TopNewBestAskShowJobs

petesergeant

5,488 karma · joined October 11, 2019

vivid.art0944@fastmail.com
submissionscomments
petesergeant··on GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence
> largely 2-horse

The absolute frontier is largely 2-horse, but the rest of the pack is very close behind, which I'm grateful for. Grok, Facebook, and the Chinese vendors are producing excellent models.

petesergeant··on Livenerf: Has Opus 5.5 been nerfed yet?
Banks and health insurance are much more consumer friendly outside of the US, usually because of regulation. Turns out you can just tell banks “make transfers cheap and essentially instant” and they’ll do it, rather the bullshit they have in the US.
petesergeant··on Dots: Always-on agents
It is always interesting to get another perspective, and I’ve also found agents to be very good at sysadmin work, including Coolify! But again, I very tightly control what agent has access to what.

Maybe you’re lucky, or maybe the examples I’ve seen (and experienced) of agent overreach are particularly unlucky. I guess at this point it’s about personal comfort level, and mine doesn’t support that type of unfettered access yet.

petesergeant··on ChatGPT Pro 500
Literally just downgraded to the $20 a month one, as they're not grandfathering in the old plans. I don't think this is going to work the way they thought.
petesergeant··on Dots: Always-on agents
I spend literally all my work day, and a good bit of my personal time, talking to agents, getting them to do things on my behalf. Almost always pretty tightly sandboxed. I just don't understand how people using these things haven't had catastrophic failures yet.

I minted what I thought was a minimal-permission Github token for a single action, and the agent I gave it to discovered it had more permissions than I thought, and made use of those permissions. Who is trusting these things with write access to their lives?

petesergeant··on AI companies in race to demonstrate their model most threatening to humanity
This is just special pleading written nicely.
petesergeant··on 500k facial scans at UK stations yield no arrests, 1 false positive
> We know the technology works, but using it in supermarkets to deter shoplifters and prevent attacks on staff is very different from deploying it to catch people on a public transport network. Success in a shop means no such people coming in. Success for the police trying to catch people means people being caught. In that respect the trial doesn’t appear to have been very fruitful.

ahahaha. "My magical face detection software worked because (according to our detection software) none of the undesirables came in!"

petesergeant··on Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
> It sounds like Jev is not a generative model

Why does it sound like that to you?

petesergeant··on AI companies in race to demonstrate their model most threatening to humanity
Cool. What did the guy who invented nuclear weapons do?
petesergeant··on AI companies in race to demonstrate their model most threatening to humanity
As I understand it, you're arguing that all sufficiently dangerous systems are well enough hardened or air-gapped, and that people don't exist who wish to cause harm using these systems? Did I understand right?
petesergeant··on AI companies in race to demonstrate their model most threatening to humanity
"These models are harmless and it's all just marketing" is the HN equivalent of Covid-truthing. I was responding to someone on Bluesky recently who claimed the METR report said the whole coordination thing was hallucinated by agents reading the logs. Hadn't read it themselves of course, and were taking tiny fragments out of context to support it. These people seem to either believe that there's not really been any malicious agent access, or that all systems that can wreak havoc on us are perfectly air-gapped. I really can't wrap my head around it.
petesergeant··on Maybe don't let Muse run your Facebook Marketplace account
I think if you are equating “performing heart surgery” with “a modest drop in already successful inventory management”, we may not have a shared basis in reality from which to converse.
petesergeant··on Maybe don't let Muse run your Facebook Marketplace account
I’m guessing you’re not a programmer. If you were, you’d have seen models go from “kinda helpful for programming” to “usable as a daily driver” about 9 months ago, for example.

“Models aren’t improving incredibly fast” seems a very odd point to be making.

petesergeant··on Maybe don't let Muse run your Facebook Marketplace account
The AI is not ready for this _yet_, but it will be, and FB wanting getting ahead of the game here is potentially good business. It’s all in the public perception of utility vs fuck-up, and it’s far too early to say Zuck got that wrong, and indicators are he got that right.
petesergeant··on Meta Blocks President Lula's Facebook Page, Campaign Ads 2 Weeks from Election
I mean by all means ban Meta everywhere, but this is a Reddit post linking to a substack post that’s a republish of a heavily politicized new source, and fails to mention the page was restored the same day, and also fails the very basic journalist standard of approaching Meta for comment, so I think we can do a little better here?

This would be better, but I guess less anger-tormenting and also you’ll need to translate it yourself: https://agenciabrasil.ebc.com.br/politica/noticia/2026-09/me...

petesergeant··on OpenAI agents tried to bruteforce a UN website's API fields
I'm glad we've moved past "this is all just marketing, there's no security risk!" phase
petesergeant··on Turning GLM-5.3-Flash into a Jev-like decision model
It is obvious. I was going to build my own little toy doing just that ten days ago, and then found four pre-existing projects, so wrote those up instead: https://sgnt.ai/p/jev/
petesergeant··on Anthropic: The Situation Report
> Why are we there when the locals hate the aid workers and actively try to fight them?

Because public health is about protecting everyone. You sound like you don't really care about the philosophical altruistic view, but the cynical reasoning works just as well here: if you can contain it in-situ, in Africa, it stops it spreading to the countries paying for the aid. Instead of gratitude from the direct recipients for how your taxes are being spent, you can instead bask in the warm glow of not catching ebola yourself.

petesergeant··on The Mafia may be keeping fentanyl out of Italy
The rest of the economics behind fentanyl are great for drug producers vs heroin though, so it probably doesn't matter?
petesergeant··on Anthropic: The Situation Report
In before the inevitable, endless cynicism that always floods this kind of thread. This seems like a good thing, and I enjoyed reading about it.
petesergeant··on The Year of Internal Tools
> Numerous Grill Me sessions, using Matt Pocock's wonderful Grill Me skill, which interrogates a plan until the weak parts fall out.

If you haven’t yet drunk the /grilling Kool Aid, give it a go

petesergeant··on Early rogue AI agent activity and attempts to hack found on urlquery.net
What do you see as the negligence angle for the gun makers here?
petesergeant··on Jev in 25 Lines of Python
That's the approach that daseinlabs/open-jev takes, in contrast to the above, which is what TheoLeeCJ/openjev and ekzhang/openjev-sglang do

https://sgnt.ai/p/jev/

petesergeant··on People hooked on vapes try a new way to quit: cigarettes
Last time I managed to get addicted, the gum was a lifeline. I found it pretty easy to switch to, to be honest, and just told myself it was OK if I was on it forever. After about 4 months of as much gum as I wanted, whenever, I started working down through the dosages, eventually making little Franken-gums by using a pill-cutter and regular gum to split the dosage further.

I didn't find the process especially miserable.

petesergeant··on GPT-6 Sol and Luna
I've found Astra to be horrible at making orchestration decisions. I will be trying to use Sol for both. Fable is very good at it though. Worst part of my week is when I hit my Fable usage limit and have to switch to Astra.
petesergeant··on Claude Opus 5.5
Slowing down only makes any sense if you can coordinate a slow-down for everyone.
petesergeant··on Claude Opus 5.5
didn't they say Opus 5 was Fable-level too tho? Let's see, I'm at the point where I don't think benchmarks really tell us very much any more. I'd love it to be as strong as Fable, but I'm skeptical about how that will look in practice.
petesergeant··on AI Has No Wisdom and Neither Will You
Same reason you’d hire engineers rather than expecting the CTO to do all the programming?

It doesn’t take much effort to setup cross-agent reviews and automatic reviews for slop and accretion, while directing design decision questions back to the human to consider. I have had a considerable increase in throughput of code that I designed and made the important decisions about, and that I’m pleased with the quality of, although as always in these discussions, someone will be a long shortly to tell me that that implies I must be a terrible engineer.

petesergeant··on AI Has No Wisdom and Neither Will You
Also AI is fine at creating maintainable software, you just have to nag it to and not accept its first attempt at it, and subject it to peer review. This is plenty similar to human developers.
petesergeant··on Grok 4.7
Grok and Zai have both been excellent as adjunct code-reviews, on their cheapest plans, for me. Fable plans, Opus writes, Codex as primary reviewer, but Grok and Zai usually find something worth fixing that the others have missed. Both are well worth whatever the $20 or so I'm paying for them
Page 1 of 34Next →