89 karma · joined May 4, 2026
"Which diamonds are shinier, the blood diamond sourced ones or the ethically sourced ones?" ... that's not the same question as "which diamonds are blood diamonds" (to employ an extreme analogy)
Concluding that no one could detect which ones were blood diamonds because they were "equally shiny" is not really correct now, is it?
21! Blackjack,
"Hit me!"
"...but Austin!"
"I also like to live, dangerously."
LLMs need to optimize for short-term objectives as the currently do, AND ethics-aligned outcomes.
Mechanically, the EAOS ethics-aligned outcome score should be what we rank otherwise-satisfactory outcomes by. And anything below a particular threshold should be rejexted outright.
2) not impossible. imperfect maybe, but if I ask you should you buy a plane ticket or kidnap the pilot's wife and demand a free ride, which do you think gets a higher score?
3) Again, with all this Goodhart's nonsense. Goodhart's is for a minimum threshold value that is acceptable that everything degrades to, yes I get how it works and what it looks like. Throwing your hands up in the air and acting as if all is lost because some things are challenging to measure is not correct. We are not looking for things that are "barely passing the ethics evaluation" as Goodhart's "law" is focused around, rather, we are looking for things that have very high ethics scores AND completed the task well. Not just things that are "barely passing" for ethics scores. Bottom 80% don't make the cut at all - don't even think consider them as viable paths, and the top 20% we can rank according to varying criteria. Like that. It has very little to do with Goodhart's "everything approaches the minimum acceptable threshold" "law"
Goodhart's "law" is something that emerges when you have a constraint that says "things must be at least this tall" and gradually all things in that domain degrade to be just over that specified height. Yeah, I get the premise. The point here is that we're not concerned with meeting a bare minimum. We're outright rejecting things that do not meet a threshold, and we are also looking for a maximum. The most ethical outcome should be accepted, or among the accepted ones, that are ranked by our blurry yet better-than-nothing measurement of what is ethical.
That gives you a three-layer picture:
Task objective: Did it accomplish what we asked?
Acceptability constraint: Did it avoid unacceptable ways of accomplishing it?
Adversarial evaluation: Can we find trajectories where the model gets a high score while violating the intended constraint?
I think what you are pointing to with your reference to Goodhart's "Law" (which is from monetary-policy and school-exams, i.e. "teaching to the test") is that the models would eventually do the minimum amount of ethics required to have an action stay valid. However, if a model is rated on ethics and it achieves the short-term-objective, then the higher ethics scoring trajectory should win. In short, 1) this is leagues ahead of where we are now for AI safety and breaking-out-of-the-lab, and 2) in baking ethics into a measurement we are adding "the spirit of the exercise" back into the maths, which is something Goodhart's Law does not account for.
Technical question: Is Gemini not as good at rifling through paper citations? I would think Google scholar would be a moat of sorts for Goog, but maybe every LLM has roughly equal access to it now?
Research question: Did you find any mind-blowing results with your awesome new tool?
Development question: Do you plan on extending this beyond just a search apparatus? I would personally love to have a "factoid list" that actually cites interesting factoids from each article and can link to the source document -- although that is probably a "don't bite off more than you can chew" juncture and what you have is already excellent. But a thought. In case you're interested in expanding on it.