707 karma · joined February 9, 2021
As a sanity check, I’d highly recommend trying out your own line of reasoning in cases which you think you might feel differently about. For example, imagine it's instead a certain authoritarian state saying "obey the Party or get blacklisted as a national security threat."
It might feel good in the moment, like "yeah, Anthropic, don't tell the military what to do," because you agree with this particular call. But the government doesn't give that power back, and you might not be so happy about the next way it gets used.
A mutually agreed upon "you may not use our product for X or Y purpose" in a contract clearly doesn't pose any such risk.
Makes sense for GP locally, I guess, but globally we’re all worse off when nobody is motivated to break free from the path of least resistance.
I say just let them duke it out. After a decade of regulatory capture and enshittification, it’s nice to see some actual competition again.
While in fairness the remaining flags seem like they would need more hard coded symbols, it’s still interesting as a procedural plausible flag generator.
Not perfectly neutrally, as you said. But by your definition the only "unbiased" model is one whose output distribution perfectly matches the training distribution, i.e. one that memorized it. All LLMs have some amount of "bias” on literally every possible input.
The tribe names are no different. In the paper they run the same game again, and the bias is different every time. There's no innate preference between them trained into the model, just noise that's revealed due to a lack of any other signal. In a real situation with actually relevant information about the candidate in context, that noise is drowned out.
The more interesting thing to look for would be a bias that's strong enough to persist across different contexts. For example, is "banananow" consistently followed by positive tokens more than "pearian" across a diverse set of realistic prompts, by enough that someone could actually exploit it? The paper shows that’s explicitly not the case for made up tribe names.
What it does show, from what I can gather, is that bias can form inside a feedback loop. The model gets a success or failure result after each hire, and if a hire from one tribe happens to fail early on, the model steers that tribe away from that job for the rest of the game, even though every candidate had the same odds.
And given to the lack of training data on such scenarios, surely the activations are mostly random noise?
It seems much more interesting to look for biases that appear robustly across different realistic scenarios that would actually be influenced by the training data
Doubt you even read your own “645 line policy kernel.” Can you explain what that is and how it works in YOUR OWN words? If you paste Claude at us again we’re gonna know.
Nobody feels entitled to a visa. It’s aspirational, like applying to a job. People do it because they want a better life, knowing full well they could be denied, or that it may well be temporary.
But if you believe immigrants could take something from you, like your job or housing, then it makes sense you’d project this idea of entitlement onto them. As if by applying they’re implicitly laying claim to something owed to you, a citizen.
In reality the economy is not a fixed pie and nothing is being taken from you. Immigrants take jobs, but they also spend money, pay taxes, start businesses, and need houses built, all of which leads to more jobs. It’s a positive feedback loop that benefits everyone, and a big part of how America’s economy grew into what it is today.
Or just vague gesturing about how immigrants aren’t entitled to anything? (but you presumably are!)
What about the sweeping halt of all legal immigrant visa processing?
You get a jmp to some arm64 instructions of your choosing, with a bunch of CPU features locked out.
They don’t provide any source code or documentation (which they easily could) so if want to do literally anything with your “open” device you must first reverse engineer the entire MacOS driver stack like the Asahi project did. Great fun, but it’s a criminal waste of incredibly talented engineer hours.
On top of that, the reason Asahi doesn’t work on M4 and later is that Apple has intentionally modified the ARM core to prevent stock MacOS from running once the hardware is “unlocked,” which makes it way more difficult to reverse engineer.
Bastards.
It’s true we usually don’t care about all the details, which is why abstractions exist. An abstraction hides details by fixing them, and lets you specify the rest precisely.
An LLM isn’t an abstraction in this sense any more than asking a coworker to do something is. The details are not fixed, but decided for you. If you can’t understand or modify what was decided yourself, the only interface left is going back and forth in natural language.
So no, the abstractions won’t become unnecessary. The details don’t just go away. Plenty of people really don’t care about them, and for them the LLM is fine, but only because it’s gluing together the millions of LoC of existing libraries and frameworks where the details have already been fixed, overwhelmingly by humans.
There will always be a need for human professionals who aftually understand all that obscure stuff, and per TFA, it’s looking like it won’t be the ones who went all in on AI.
And if you really want to record all your work and project context somewhere, just let people ... look at that directly? Or summarize and query it with their own LLM if they want to?