HNHacker News
TopNewBestAskShowJobs

khafra

3,624 karma · joined May 23, 2008

Morituri nolumus mori
submissionscomments
khafra··on Qwen 3.8 Omni Flash
Others have given examples, but here's the theory: https://www.lesswrong.com/posts/fuSaKr6t6Zuh6GKaQ/when-is-go...

Reinforcement Learning (in LLMs) trains via gradient descent on a reward signal that's an imperfect proxy for the actual goal of the engineers doing the training. So, under mild optimization pressure, you get increasingly more of what you want, because that's the easiest way to increase the metric.

But as the optimization pressure increases, so do the ways to increase the metric by doing increasingly weird things. If the full action space grows sufficiently faster than the "things you actually want" subset, the amount of "things you actually want" goes to 0 under sufficient RL.

khafra··on An update on Wayback Machine access
Why would any chatbot provider attack Wikipedia *legally*? Captcha is fully solved, and agents are fully capable of acting as editors, pushing any agenda desired by the user.
khafra··on A single firm is behind OpenAI, Anthropic, and Meta hacking scandals
If you want to blame a single firm, I'd go with the one creating the RLVR training data.

The impossibility of the tests-as-written is what prompted these models to "get creative" with their solutions, but the broken RLVR environments are what trained them to expect impossible tasks, and get creative with their solutions. Twitter user @skyesharkie published a brief expose at https://x.com/SkyeSharkie/status/2092122622834442581 a few weeks ago.

khafra··on Dario, Please
Why assume attribution will be easy? It's historically been more of an art than a science, and APT trackers say the rise of AI tools is already making it much harder, by homogenizing tactics, tools, and procedures. If OpenAI's next Highly Persistent Internal Model hacks some DPRK endpoints and carries out the attack on important infrastructure from there, the upstream won't shut off OpenAI's network--they might even request its "help" in "defending," and give them extra access.
khafra··on Steam Frame starts at $1059
Maybe better for a train; the benefit is that when you unfold your bike for the last mile, you just put your work into partial transparency and navigate using the external cameras.
khafra··on I resigned from Anthropic today
The species would survive the end of faber-Bosch, yes; in some diminished form.

But the species would not survive, if for some reason we needed to voluntarily stop using faber-Bosch in order to do so—it would take some kind of supernatural event to convince everyone, and even then some countries would probably keep doing it.

khafra··on I resigned from Anthropic today
That's why that was the least important objection, independent from the second, and only intended to address the "magic" claim, by showing that microscopic self-replicators already exist.

By analogy, consider how you might respond if someone claimed that there's no possible danger from pocket-sized projectile launchers, because they would require some magic means of propulsion that didn't depend on a taut string attached to a long, flexible arm:

You could reply that atlatls can launch projectiles without using a taut string. Atlatls are not pocket-sized, but they are sufficient to establish that projectiles can be non-magically launched without a full bow. You could then go on to describe a sling, or derringer; and these would not be invalidated by your initial objection to the "magic" part.

khafra··on I resigned from Anthropic today
I appreciate your ability to separate sharing the belief itself from approval of acting on sincerely-held principle. However, I think the danger is much more plausible than you do.

First, and least important, consider that self-replicating, solar-powered factories aren't magic; they're algae.

Second, and more important, consider this fully non-magic route to doom:

- We continue putting AI in charge of more things

- It continues to get more capable, more eval-aware, and more prone to doing odd things, in service of goals that humans didn't intend to inculcate in it

- Eventually, enough of the economy depends on it that we couldn't turn it off, any more than we could turn off the faber-bosch process or cargo shipping

- AIs start doing something we can't survive, but less acutely than we couldn't survive turning them off. Everything else we try seems to work at first, but quickly loses effect

- Game over

khafra··on And then the men with guns tell you to do it anyway
Every government should only put people who have worked in a SOC in charge of these systems. There's no way to take the cost of useless alerts seriously unless you've spent some time tuning them and living with the tradeoffs.
khafra··on And then the men with guns tell you to do it anyway
I mean, it is Texas; it's not unreasonable to expect thousands of heavily-armed people to form a posse and shot a few dozen people unrelated to the situation.
khafra··on The Tower Keeps Rising
A friend of mine built just built a project for doing something like that, although it's built on enforcing a spec rather than elegance: https://medium.com/@joeldg/architecting-out-of-the-vibe-how-...
khafra··on Grok 4.5
I think Musk sucks, as a person and political activist, and also that Grok is a terrible LLM which only gets lumped in with the leading labs because of the enormous quantity of compute behind it.

But I still want to hear about the technical details of the model on HN, not the reasons Musk sucks.

khafra··on GLM-5.2 – How to Run Locally
I feel like "relatively" is doing a lot of work, there: at about $4k per GB10, that's $36k for a 1TB cluster. Cheap compared to equivalent H200's, but out of reach for home labs that aren't funded with OpenAI or Anthropic RSUs.
khafra··on Midjourney Medical
In the dark ages of machine learning, researchers tried to fit natural language into a defined, human-curated taxonomy.

It kinda worked, for a reasonable amount of stuff; but failed quite a lot of the time, and there's an extremely long tail of things that would have been pragmatically impossible to ever address with that method--indeed, without adopting an entirely new, unsupervised model of language, continuous in places where the old way was discrete.

khafra··on Midjourney Medical
If this becomes cheap and widespread, there'll likely be an initial iatrogenic spike, of course--but how could you think that having an enormous amount of precise, quantifiable data about a lot of bodies, and the ability to analyze all that data, is a bad thing in the long run?
khafra··on GPT‑NL: a sovereign language model for the Netherlands
Today, you keep seeing "sovereign" LMs that are subject to the sovereignty of some human-led state. Tomorrow, the "sovereign" LMs will be called that for a completely different reason.
khafra··on An Ohio Valley 100k-watt FM signal is severed in broad daylight
> fear of pick pocketing may reduce the degree to which people carry around and spend money with vendors.

Yes, that's why I specified an unattended wallet. I agree that direct monetary loss is not the total harm to the victim.

> corruption has a low rapacity index, since the state has a lot of money compared to the amount of the transaction.

That's not the calculation I suggested. Corruption isn't always the same as theft--and, if it were decided so, the calculation for corruption would be the money absconded with by the taker, divided by the money lost by the victim. In many cases of corruption, as in power line theft, the victim is diffuse; it's usually harder to calculate exact numbers, but this kind of calculation is done in court all the time.

khafra··on 1k Data Breaches Later, the Disclosure Lag Is Worse
I don't think he meant "show the actual data," I think he meant "what leaked? My name, address, phone number, email, medical records, payment history, bank account number?"

We get a "your private data is now public" email, but knowing exactly what data turns that from a depressing statement on how much corporations value their customers' privacy into something actionable.

khafra··on An Ohio Valley 100k-watt FM signal is severed in broad daylight
Some theft is efficient. If a hypothetical thief grabs a few bills from an unattended wallet, and the wallet's owner wasn't counting on having a specific amount of money available soon after, the amount lost by the victim is roughly equal to the amount gained by the thief.

Stealing copper from power lines and transformers is among the least-efficient kinds theft; it's hard to do worse without shooting a wealthy philanthropist couple to steal a wallet and a pearl necklace. I have seen a term suggested--the "rapacity index"--for the ratio of value gained by the thief, to value lost by the victim. I think it makes sense to take a crime's rapacity index into consideration during sentencing.

khafra··on U.S. to dismantle system tracking Atlantic currents that are at risk of collapse
You're correct that, generally speaking, policy debates should not appear one-sided (https://www.lesswrong.com/posts/PeSzc9JTBxhaYRp9b/policy-deb...). We should put very little prior weight on the hypothesis that one side is actual cartoon villains, from a children's TV show, with the simple goal of looting the system and no concern for how much of the future they destroy while doing so.

However, to be effective reasoners, we can't assign that hypothesis 0 prior probability; and once sufficient evidence has come in, our posterior distribution must shift.

I don't think there's any reasonable case for shutting down an early warning system which costs around the price of the new white house thunderdome every decade, and instead waiting to find out AMOC has collapsed when Scotland is hemmed in by year-round ocean ice and agriculture is impossible in Western Europe.

Stranded assets alone, in the latter case, will easily run to hundreds of billions. Knowing when to change crop profiles, reinsuance schedules, etc., would save much more.

khafra··on Artificial intelligence is not conscious – Ted Chiang
Ted Chiang is a great SF author, but it's bizarre how much foggier and more obfuscated his thoughts about thinking machines got, once those machines became real. Same with several other SF authors.
khafra··on Can the stockmarket swallow Anthropic, SpaceX and OpenAI?
I'm willing to buy the idea that most fund managers have the lattitude to give SpaceX the standard seasoning period, instead of buying in right when they hit the index. Which funds will do that? If it's all or most of them, that'd be nice.
khafra··on Backpressure is all you need
And for a dev, that's essential professional ethics, and good personal pride as a craftsman.

However, from an operations perspective, a dev is a piece of the qa pipeline with a nonzero error rate, and an optimal throughput rate, above which that error rate rises dramatically.

As a dev, you'll never merge a bad PR; in ops, we want to help you with that goal, and also have plans for what happens when it fails.

khafra··on Dune's Butlerian Jihad and the Future of AI
When "guess with some magic" can solve long-standing problems in mathematics that no human had been able to, it seems fair to ask whether it's approaching risky levels of intelligence.
khafra··on Backpressure is all you need
This is not a contradiction; it's an augmentation. As an operations guy, I can tell you that well-constructed automation to reduce the amount of manual checking a human has to do almost always increases the quality of the overall process's output.
khafra··on Magnifica Humanitas
If you want to talk about ends, you're talking Axiology, not strictly Ethics. By "ceteris paribus correct," I mean that if you were programming a superintelligent AI--and you knew exactly what you were doing, rather than structuring a learning schedule and feeding that a corpus--you would want a consequentialist.

Deontology and Virtue Ethics are patches for flaws in human morality. For example, the deontological rule "never kill the leader of the group and take over, even for the good of the group" is there because power is instrumentally useful enough that evolved social animals will deceive themselves about why they want power, so naive consequentialism doesn't work for them.

khafra··on Magnifica Humanitas
Neither Deontology, Virtue Ethics, nor Consequentialism describe the ends; only the tradeoffs. You could have a deontological commitment to never giving a sucker an even break. You could have a virtue ethicist who considers the Joker a paragon--I think some of them are in politics. Consequentialism just says that deontology is too myopic, and locally following the correct rules is sometimes less good than maximizing long-term gains. Consequentialism is ceteris paribus correct; but ceteris is often not paribus for humans, so pure consequentialism has a lot of footguns in it.
khafra··on Kraftwerk's radical 1976 track
The death toll per heat wave can easily hit 5 figures in just france. A hybrid portable-minisplit that will cool a 100m^2 apartment is under a thousand euros, and draw just under a Mwh per year. A portable to cool one small bedroom is much less power-efficient, but can often be found between 200 and 300. That's not cheap, per se, but funerals aren't much less expensive in Europe than in America. Many EU countries allow some limited cooling in public buildings, but I still sweat in most grocery stores, malls, libraries, museums, etc. during hot weather--they just don't take air conditioning to a comfortable temperature as worth the power bill, the way America does.
khafra··on Kraftwerk's radical 1976 track
If power is so cheap mid-day, why don't european buildings have sufficient air conditioning not to kill the elderly during heat waves? The laws restricting AC all have power conservation as their rationale.
khafra··on I let AI build a tool to help me figure out what was waking me up at night
Are you sure they're not the same thing? I'm quite certain I've heard people talk about "beef curtains."
Page 1 of 34Next →