HNHacker News
TopNewBestAskShowJobs

jgilias

4,284 karma · joined May 29, 2020

submissionscomments
jgilias··on 'Neanderthals Among Us' review
Lol. Bride kidnapping is still a thing. In my culture it’s not a thing anymore, but it’s still being acted out as part of wedding customs. There’s also a separate word for “people who go after the people who stole the girl”.

I find it more likely the admixture is from girls our ancestors went to take back who were already pregnant.

jgilias··on Sites in ChatGPT
I think there’s a bunch of startups that were basically this. This is the problem with AI startups, you’re just waiting until the labs will come out with what you were building and make you obsolete.

I don’t think it’s all bad though. The labs aren’t going to release anything that makes sense if you’re trying to be vendor independent. So, gateways, vendor agnostic sandboxes, harnesses, etc. That’s where business opportunities may be.

jgilias··on Dots: Always-on agents
OpenClaw but from OpenAI.
jgilias··on GLM-5.3 and the spread of advanced cyber capabilities
Oh boy! First time I got excited of the possibility of using Windows.
jgilias··on Toyota is taking the Corolla electric
This has been tried a bunch of times.
jgilias··on Toyota is taking the Corolla electric
Quality matters though. A 10 years warranty as long as you do the prescribed maintenance with hassle free ownership is a major reason why Toyota’s clientele sticks.

It’s not impossible to unseat them. But it’s also far from given.

jgilias··on Show HN: Jevper – the Jev interface on top of any OpenAI-compatible model
But… why???
jgilias··on Typesafe-computer-use drives a Mac toward a goal for 1/50th of a cent per step
This probably throws a spanner in the wheels there:

> Every piece of reasoning the frontier model does for free has to be rebuilt here as deterministic state.

EDIT: Not to shit on this though. I totally believe that some smart mixture of LLM-reasoning + Jev-style + determinism is going to be pretty amazing.

jgilias··on Pion, an agent designed to run any company autonomously
What harnesses do you use? Ours is basically Claude in a box. There’s some complexity because of that, but the advantage is that it’s very flexible and people who have a bunch of Claude-shaped skills can just basically give those to an agent.

I’m thinking to take a deeper look at Pi. I’m really liking that project.

jgilias··on Pion, an agent designed to run any company autonomously
At $DAYJOB we do something very similar shape-wise. I wonder - it sounds like you have a dedicated agent comms plane? In our case we found that the easiest and most straightforward was to just use our default company chat app directly for this. Because most of the context that the agent workers need to do work is there, but also, it’s just much easier for teams to conceptualise an agent colleague if it just hangs out in their channels.

What do you do here, and how’s it going?

jgilias··on Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher
Maybe it does. Need to run evals to see if it does or doesn’t.

Point was - everything in the context window affects the output. Including “silly” things like “it is known AI can do this”. And that has nothing to do with superstition, as the poster above me seemed to imply.

jgilias··on Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher
Not OP. That’s not implied at all. The fancy autocomplete produces statistically likely continuations to the source text (the context window). For a problem that’s hard for humans one likely continuation is: “this is hard, can’t do”, even though there’s enough in the training corpus of the LLM to actually do it.

So, it follows that adding “pep talk” into the context window reduces the statistical probability of “no, can’t do” coming out as the answer you get.

These things are neither humans, nor deterministic software.

jgilias··on Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher
If it’s 3 days it’s something like 15 minutes, if it’s 3 weeks, that takes a couple hours lol. Seems like there’s some sanity to the estimates after all when you think about it, it’s just the scale it gets wrong due to estimating human time.
jgilias··on We must pace the frontier
A contrarian take - the models aren’t advancing anymore at a pace where each new model would represent a huge capability jump, all being incremental improvements, so the doomsday marketing strategy being invoked since GPT-2 isn’t as effective anymore. Then “pacing” would be a convenient scapegoat to point fingers to when people point out how the new model isn’t _really_ that much better.

“Of course it’s not, we’re pacing!”

jgilias··on Bill Gates tries to install MovieMaker (2003)
I chuckle that this is in Windows XP days. You know, the pinnacle of Windows usability in retrospect.
jgilias··on Smartphone makers don't bother to comply with EU repairability requirements
It does matter, yes. Because it won’t lead to the outcome you want until someone actually gets a painful fine.
jgilias··on Global warming will exceed 1.5-degree limit, UN says
In medicine there’s this well known axiom that if an intervention regime would technically work, but is something that the patients tend to not follow through, then the intervention regime doesn’t work.
jgilias··on GPT-6 Astra
I’m so tired of every model release being touted as AGI or similar. Since <checks notes> GPT-2.
jgilias··on Claude Fable 5.1 and Claude Mythos 5.1
Cool. I’ve realized though that I don’t really need better models anymore. SOTA is good, I just want them faster/cheaper now.
jgilias··on Nvidia agrees to acquire Hugging Face for $13B
Volkswagen has a couple too. Like, actual laboratories with people in white coats working in them. You don’t call Volkswagen “a lab” because of that.
jgilias··on Nvidia agrees to acquire Hugging Face for $13B
Ok, so, what research department did Nvidia grow out of?
jgilias··on Nvidia agrees to acquire Hugging Face for $13B
_why_ have people started to call companies “labs”?!
jgilias··on CEO fired developers to make room for AI. Developers create open source AI CEO
Would be cool to get this thing try and run the vending machine experiment.
jgilias··on The Tariff Cost: analysis of the costs to Americans from new tariffs on Canada
That’s not surprising. He’s, like, really dense.
jgilias··on My Friend Aaron
Hooked me with the first couple sentences, and I just had to keep going, as I just wanted to know what happens next.

Thanks, a good read!

jgilias··on Coding expertise is going to collapse from AI reliance
You have it mixed up. This guy is “the new generation of product people”. The old generation had to be technical enough to be able to grok the systems that they ‘producted’ over. It’s exactly the ‘new generation’ that LLM themselves right into Dunning-Kruger.
jgilias··on Coding expertise is going to collapse from AI reliance
For the completion to work, the source text needs to be ‘good’. That’s a basic kernel of how the thing works. Even with a perfect oracle autocomplete if the source text is ‘bullshit’, the output is too.

Or, slightly changing this. The source text needs to speak the correct vocabulary and language to produce a good completion. See the chat where Terry Tao is doing maths with an LLM. There’s _no way in hell_ I could get to his output because I just have no idea, and can’t speak the language.

Same with any field.

jgilias··on Coding expertise is going to collapse from AI reliance
We have a product guy on the team who was in a deeply not technical role before AI who is trying to do the “hey Claude, read this Jira ticket, implement” thing.

It doesn’t work for the vast majority of tickets he attempts because he doesn’t have the necessary understanding to even start thinking about if the solution that the autocomplete generates is even remotely workable. And that’s with fancy dev loops and whatnot.

The spacer between the keyboard and the chair still matters in my experience.

jgilias··on In Australia, a home battery boom has helped cut wholesale power prices
My neighbour has a battery installation because where I live they subsidise both - the panels and the battery. In the event of a power outage he can keep running his house for ~3 days if being somewhat frugal with consumption.

And the battery thing even looks kind of neat.

jgilias··on How Claude marks AI-generated content
The more they fiddle with the autocomplete system, the more they move away from the autocomplete faithfully producing the completion I need. The more it makes sense to move to an open weights model not served by them.
Page 1 of 30Next →