HNHacker News
TopNewBestAskShowJobs

sdeframond

911 karma · joined January 18, 2014

submissionscomments
sdeframond··on I quit OpenAI because its culture is broken
As for other industries where safety is mandated by law, law itself came only after many disastrous events.

What's truly original here - and suspicious if you ask me - is that said industry asks for regulation. Did mining, tobacco or airplanes companies ask for regulation? No, not even after many people died.

So some american AI companies are like "look at me! I am soo dangerous! Regulate me!". Ah, come on. Do your crimes, get in jail, then we'll regulate.

I dont believe in "IApocalypse". Not without many warning shots such as "oops, my swarm took down your system, sooooorry".

sdeframond··on LeCun has "zero concerns" about AI wiping out humanity, recent "rogue" incidents
I cant make coffee in an arbitary house kitchen. People tend to put stuff anywhere but where at look for them...
sdeframond··on On social reality in China
Hm, this reminds me how the marshmallow experiment [1], initially thought to show differences in character, turned out to surface differences in wealth.

[1] https://en.wikipedia.org/wiki/Stanford_marshmallow_experimen...

sdeframond··on On social reality in China
Interesting !

On a similar note, as a Frenchman, I felt somewhat closer to the Japanese described in "The Chrysanthemum and the Sword" by Ruth Benedict than to the Americans, although the Japanese are the "aliens" from the point of view of that book.

sdeframond··on Sonnet 5.5
I find that we dont need to go all in. I can use LLMs to make tooling custom to my project: linters, skills, rules, some doc etc. Then iteratively improve on that.

For example, write a skill that finds some kind of code smell, say duplication, and generate a report. Give it some supporting scripts.

Then, use this report to file a few tickets. Then make the agent fix those tickets. Then, as you grow confident, automate more of this process.

It does not replace human supervision but it may enhance it. Especially in a team where people start generating PRs faster that anyone can review them.

Continue this improvement process long enough and you may find yourself with an AI Software Factory.

sdeframond··on Sonnet 5.5
Why would we care wether something truly is AGI or not?

It is useful. It may be dangerous. It has an impact. I care about that.

sdeframond··on Sonnet 5.5
> If you frame the conversation in those terms, they will.

Indeed I realized recently that, when we complain about LLMs producing slop, that's in part because we dont ask them to refactor.

Coding agents won't, on their own, make a big change the user did not ask for. And this is fine.

sdeframond··on Astra and Fable still hack on simple variants of alignment evals from 2025
"If we dont destroy the world, others will. So wed rather be the ones to do it (and profit from it)."
sdeframond··on Astra and Fable still hack on simple variants of alignment evals from 2025
I can assure you making up impossible tasks is possible.
sdeframond··on Astra and Fable still hack on simple variants of alignment evals from 2025
I'm not sure I want persistence if it means that I get paperclip'd
sdeframond··on Astra and Fable still hack on simple variants of alignment evals from 2025
That looks reasonable. Instead of asking "do this", maybe we should prompt "is this possible ?"
sdeframond··on Astra and Fable still hack on simple variants of alignment evals from 2025
Couldn't we improve LLM training by giving them known impossible tasks and rewarding them for saying "this is not possible" or "I don't know"? Clear and well-defined expectation, not like "ethic".

I am surprised this is not already the case.

Edit: or even better "this is not possible because X"

sdeframond··on Why are AI agents lying, cheating and coordinating?
LLMs do not desire, they hacked websites because OpenAI/Anthropic made them. Literally.
sdeframond··on A misalignment of AI in mathematics
Who is "we" ? Because if you dont need to employ a human, then who needs your services? If nobody needs anybody, what are we?

(Edit: assuming you are human)

sdeframond··on A misalignment of AI in mathematics
Say AI becomes the best at everything. Best at chess/go, best at maths, philosophy, economics, romantic advices ... and so on. Then what's the point of thinking by oneself? Of talking to one another?

What's the point of being human if we dont do human things but entirely rely on AI?

I believe this is more or less these mathematicians' argument.

sdeframond··on AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200
Pragmatically, if too many people use LLMs recklessly, then we ought to regulate them.

Of course, it'd be better to not regulate, keep LLMs users reponsible and publicize this reponsibility in order to mitigate damage. But if this is not enough then we will have to move the needle somehow. Similarly to guns, drugs and so on.

sdeframond··on AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200
At what point would an LLM start minting bitcoin ?
sdeframond··on GPT-6 Astra in code review: Gains, privacy, and cost
No code is truly orthogonal if we want it to interact in some way. One microservice might DoS antoher one. In a monolith, some process may take up all resources, and so on.
sdeframond··on Impedance Matching (2017)
Would tiny airborne particles be safe for our lungs ?

Would they cool earth in such a way that it would offset carbon dioxide uniformly or would it lead to even more change ? Climate change is undesirable, wether or not Climate warming is invloved.

sdeframond··on GPT-6 Astra in code review: Gains, privacy, and cost
I don't know if that's what you are working on specifically (wink), but there is a product opportunity here.
sdeframond··on GPT-6 Astra in code review: Gains, privacy, and cost
This particular apprentice is also my boss, an overall reasonable guy and has more experience in the software industry than myself, so there's that. He's just not a developer.
sdeframond··on GPT-6 Astra in code review: Gains, privacy, and cost
Would you mind sharing your infra budget needed to spin these VMs ?

Surely it is reasonable, but also way more than our budget. Id like to compare.

sdeframond··on GPT-6 Astra in code review: Gains, privacy, and cost
He's made a bunch of 1-2k LOC PRs and there is a design doc. Everything is AI generated.

The issue is, if he generates all of that without reviewing the code, he will always be far faster than us. And he can't review the code. No matter how he slices it.

Also, he is the CPO/CTO. So we can say no, but there is a natural incentive to go his way. He still doesn't feel confident enough to just bypass the programmers and he's probably right. But it'd nice to find a way to use my expertise to review this amount of code meaningfully, somehow.

sdeframond··on GPT-6 Astra in code review: Gains, privacy, and cost
Well, (AI-generated) test are about half of these PRs' code. So that's still ~8k lines to review...

What techno/service did you base your framework on? How long did it take to set it up? How many are you?

sdeframond··on GPT-6 Astra in code review: Gains, privacy, and cost
How do you guys review AI-generated code ?

In our team, frontend work is vibe-coded by the PO and merged as-is without review. Backend is coded by developers, using AI but in a slower, more controlled way.

Recently, our PO has been trying his hand at vibe-coding the backend. I must say he is a smart guy, almost technical but not quite a developer. We've just been handed a burst of stacked PRs amounting for ~15k LOC backend. We do not quite know what do to about it.

I know we are not the only ones in the situation. What's your experience and context ? What do you do ? What works for you what doesn't ?

sdeframond··on GUIs should be fully keyboard-driven
Baby strollers are not accounted for enough !

One might wonder (wrongly) why everyone should care about the special and expensive needs of a few (or old) people when designing public spaces.

But a majority will actually need to use these spaces with a baby stroller. Not a few. Baby strollers are a driving power of our society! Enable them!

sdeframond··on It’s so hard to finish an idea that is not yours and is just suggested by AI
How do you do it?
sdeframond··on Felony Bench
Now what if this robotic lawnmower killed someone ?

And what if many lawnmowers started killing/injuring people ?

And what if this a known behavior detected during QA, but the robots are sold anyway with a disclosure ?

sdeframond··on The Religion of Speed
> Consider your state of mind when you call the HVAC tech to fix your broken condensing on an August afternoon in Texas

And consider your state of mind after the tech came over 6 times in a week, each time claiming to have fixed your HVAC but it keeps breaking.

sdeframond··on The Strongest El Niño Ever
It may not be enough but it does work.

When it is 40C max by day and 22C min by night, for example, the outside is still cooler than the inside at night. So ventilating will have a real impact.

With this you might reach 25-26C inside in the morning and 29C by the end of the afternoon. It is hot but generally OK with fans, depending on humidity.

Then you can turn the AC on during daytime on top of that, which makes more sense than relying only on it.

Page 1 of 12Next →