249 karma · joined November 21, 2017
It's well worth looking at https://progress.openai.com/, here's a snippet:
> human: Are you actually conscious under anesthesia?
> GPT-1 (2018): i did n't . " you 're awake .
> GPT-3 (2021): There is no single answer to this question since anesthesia can be administered [...]
This is true, but sampling also plays a fairly large role. The model will produce probabilities for the next token, temperature will modify these probabilities somewhat, but different sampling techniques (top-K, top-P, beam search, others) will also change these probabilities.
> I wasn't under the impression that it was to give the user a feeling of "realism", but rather that it produced better results with a slightly random prediction.
My understanding is that it's a bit of both. If the AI responded exactly the same way to every "hi can you help me" prompt, I think users' would call it more robotic. I also think that slightly varying the token prediction helps prevent repetitive text
This example from software doesn't meaningfully hold for neural networks. It's a bit like trying to watch an individual COVID virus duplicate and then attempting to predict the pandemic. It's incredibly complicated and we haven't yet built the tools to help us understand
Yes! loads! (: I want to be able to say statements like "this model will never ask the user to kill themselves" and be confident, but I can't do that today, and we don't know how. Note that we do know how to prove similar statements for regular software.
Common misconception, MoEs do have different "experts", but the model learns when to send input to different experts, and the model does not cleanly send coding tasks to the coding agent, physics tasks to the physics agent, etc. It's quite messy, and not nearly as intepretable as we'd want it to be.
[1]: https://www.reddit.com/r/slatestarcodex/comments/1o6n5ne/why...
How do you define "perfect" data and training? I'd argue that if you trained a small NN to play tic-tac-toe perfectly, it'd quickly memorise all the possible scenarios, and since the world state is small, you could exhaustively prove that it's correct for every possible input. So at the very least, there's a counter example showing that with perfect data and training, models will not get stuff wrong.
But NNs are fundamentally continuous, I don't think it even makes sense to "count" bugs. You can have a list of prompts to which the model gives unwanted output, but it's a completely different ball game compared to regular software.
This seems like a pointless definition of "act"? someone else could use the AI for actions which affect me, in which case I'm very much worried about those actions being dangerous, regardless of precisely how you're defining the word "act".
> when they can literally be implemented with a spreadsheet
The financial system that led to 2008 basically was one big spreadsheet, and yet it would have been correct to be worried about it. "Malicious" maybe is a bit evocative, I'll grant you that, but if I'm about to be eaten by a lion, I'm less concerned about not mistakenly athropomorphizing the lion, and more about ensuring I don't get eaten. It _doesn't matter_ whether the AI has agency or is just a big spreadsheet or wants to do us harm or is just sitting there. If it can do harm, it's dangerous.
I wonder if it's unheard of in junior devs because they're all saints, or because they're not talented enough to get away with it?
Also kinda crazy that all the "native" voice assistants are still terrible, despite the tech having been around for years by now.
I now think it's more accurate to think that someone is an expert relative to someone else, and only for a specific field. But that'll have to be another essay (:
PS: I love your writing, thank you so much for putting it out there (:
Thanks for the vote of confidence (: I'm kicking myself for not figuring out a mailing list before this essay went viral, but I'll cross-post the essay on my substack (https://beyarkay.substack.com/) when it comes out, so you can sign up there to get an email.
> I wonder if we agree on expert aesthetics or not. You write:
So I'm coining "expert aesthetics" as a relatively unused phrase that I can put my own connotations onto. There'll be more in the essay (; but at a high level, I've observed that, as someone becomes an expert in a field, their sense for what's "beautiful" in that field changes, and _generally_ it starts to focus on things that are technically challenging. That is, experts (IME) tend to find technically difficult things _aesthetically_ beautiful, even though novices might not care one bit about the technical skill required.
Examples might help: Wine connoisseurs preferring wine from specific regions or made using specific techniques, while casual drinkers just want something that tastes good. Fashion designers preferring something that's different from last year and riffs off of the current styles, while the general public just want the same old same old. Painters taking delight in still lifes that perfectly capture the reflection of light through a wine glass, while most people just want a pretty sunset or portrait for their wall.
This is all still in flux, but that's the gist of what I'm calling "expert aesthetics".
> novice drives, expert advises
I've not heard this explicitly recommended, but it's so clearly the best way to do things if learning is the goal.