2,643 karma · joined August 22, 2013
What’s the point if you aren’t even going to try to prove a theorem? Or, heck, even really test a hypothesis?
If you include all the information the LLM uses to produce the next token as part of the state, then of course the LLM is a Markov chain.
So would be any other process for sampling continuations of a text, with finite memory.
And, my point is that the inputs it is often fed are not in the convex hull of the inputs in the training data.
When the input space is very high dimensional, this is a common outcome.
I’m not denying that the outputs are causally downstream from the training data. Of course it is.
I’m saying that the inference time inputs aren’t in the convex hull of the training time inputs. This isn’t about saying that the output isn’t because of the training data. Of course it is.
But when you have very high dimensional input space, then even with many inputs in the training data, it is still common for inference time inputs to not be in the convex hull of the train time inputs.
This has nothing to do with the complexities of how the models work after the initial embedding of the tokens as vectors. It’s just about the inputs that appear during training, and the inputs that appear at inference time.
> But an LLM can not infer a concept to which it has no information channel.
Of course! And nothing I said implies otherwise. Really, the point I’m making doesn’t even depend on what the model outputs!
If I took a best fit line from 1 parameter to a 1D output, and then provided that linear model an output that was outside the range of inputs the best fit line was obtained from, that would not be interpolation, it would be extrapolation.
It is similar here, except instead of the input being outside the convex hull due to being further away, it is outside the convex hull due to, like, the shape of the convex hull of training inputs just doesn’t include the point in question.
By "If you need new language" do you mean like, coining new words?
I don't see what would prevent them from doing this? LLMs can process text that includes newly coined terms, and respond to that text in ways that use those newly coined words in accordance with the descriptions of the meanings given for those new words in the prompt. They can also make up new words+definitions when asked to do so. Now, whether they can, without being told to do so, recognize that it would be useful to coin a new word for something, and then start using it, I don't know of any instances of this, but based on the previous two things, I don't see a reason to expect this to be fundamentally beyond what they can do?
I don't know what it would mean for a concept to be "independent of the existing language they are trained on". If there are ideas that can't be expressed in terms of the semantic primes all ideas we can express can be expressed in terms of, then I guess such an idea would be independent of our language, but I think that's a much stricter condition than what you mean (and I'm not sure if there even are any good ideas that can't be indirectly expressed in terms of semantic primes -- I kind of suspect not, unless they are like, ideas that are too big to fit in a human mind anyway).
Of course, the outputs these models produce is causally downstream from the data they are trained on, and the distribution they produce over text is largely based on the distribution over text in the training data, but altered in a number of ways (for example, to make them implement the character of the "assistant" persona).
If it can only interpolate in a literal sense, that means that it only produces good outputs on convex combinations of inputs that appear in the training set. That's what interpolation means. But, if you take the embedding vectors of sentences/prompts, and then take the convex hull of these, it is not typical for new sentences not in the training set to have its embedding vectors be in the convex hull of these.
And I don’t think it’s a good metaphor.
Taking it instead as a metaphorical claim may be more valid, but in that case it doesn’t depend on our understanding of how LLMs work.
But if you actually try to take a convex hull of, some encoding of sentences as vectors? It isn’t true. The outputs are not in the convex hull of the training data.
I guess it’s supposed to be a metaphor and not literal, but in that case it’s confusing. Especially seeing as there are contexts in machine learning where literal interpolation vs literal extrapolation, is relevant. So, please, find a better way to say it than saying that “it can only interpolate”?
Like, “take a random sequence of bits and interpret it as Unicode” is at one end of a scale, and “take a random sequence of words in a language” is just a tad away from it, and the scale continues in that direction for quite a while.
No, it doesn’t.
Many of them are (unfortunately) moral relativists. However, that doesn’t mean their goals are to make the models match their personal moral standards.
While there is a lot of disagreement about what is right and wrong, there is also a lot of widespread agreement.
If we could guarantee that on every moral issue on which there is currently widespread agreement (… and which there would continue to be widespread agreement if everyone thought faster with larger working memories and spent time thinking about moral philosophy) that any future powerful AI models would comport with the common view on that issue, then alignment would be considered solved (well, assuming the way this is achieved isn’t be causing people’s moral views to change).
Do companies try to restrict models in more ways than this? Sure, like you gave the example of about Taiwan. And also other things that would get the companies bad press.
> Language models process signs (representamens) but are blind to when meaning forks — when the same word means different things to different communities.
But, haven’t interpretability results shown that these models internally represent several meanings of the same word, differently? In that case, why would they not already do the same for how words are used differently in different communities?
Like, if one theory says that a hunk of metal actually is made of many microscopic grains of various sizes and orientations, where the sizes and orientations of these grains has an effect on the behavior of the metal, you don’t count the “the sizes and orientations of these grains” as free parameters, do you?
If you accept that every natural number has a successor which is a natural number, and no two natural numbers have the same successor, and that there’s no loops (e.g. by saying that there’s a total order on natural numbers and that any natural number is less than its successor), then there can’t be a finite collection which is all the natural numbers.
You could say “there’s no collection which has all the natural numbers”, which, ok, how do you want to talk about things true of all natural numbers then?
Formulating descriptions of physics without the axiom of infinity (or, without something to play the role of the real numbers) is super icky. You, in practice, can’t do any significant mathematical physics in an ultrafinitistic approach.
Or, something like that?
I think the two logics can emulate one another? Or, at the very least, can describe what the other concludes. I know intuitionistic logic can have classical logic embedded in it through some sort of “put double negation on everything”. I think if you add some sort of modal operator to classical logic you could probably emulate intuitionistic logic in a similar way?
Rather, in classical logic, if you can show that a statement being false would imply a contradiction, you can conclude that the statement is true.
In intuitionistic logic, you would only conclude that the statement is not false.
And, I’m not sure identifying “true” with “provable” in intuitionistic logic is entirely right either?
In intuitionistic logic, you only have a proof if you have a constructive proof.
But, like, that doesn’t mean that if you don’t have a constructive proof, that the statement is therefore not true?
If a statement is independent of your axioms when using classical logic, it is also independent of your axioms when using intuitionistic logic, as intuitionistic logic has a subset of the allowed inference rules.
If a statement is independent, then there is no proof of it, and there is no proof of its negation. If a proposition being true was the same thing as there being a proof of it, then a proposition that is independent would be not true, and its negation would also be not true. So, it would be both not true and not false, and these together yield a contradiction.
Intuitionistic logic only lets you conclude that a proposition is true if you have a constructive/intuitionistic proof of it. It doesn’t say that a proposition for which there is no proof, is therefore not true.
As a core example of this, in intuitionistic logic, one doesn’t have the LEM, but, one certainly doesn’t have that the LEM is false. In fact, one has that the LEM isn’t false.
I don’t know what this claim is supposed to mean.
If it isn’t supposed to have a precise technical meaning, why is it using the word “interpolate”?
Like, if someone mistakes a manikin or scarecrow for an innocent person, and takes action in an attempt to harm that imagined person (e.g. they try to mug the imagined person), they’ve still done something wrong, even though the person they intended to wrong never actually existed.
I guess maybe it kind of depends how strongly and deeply one feels as if the manikin/scarecrow/chatbot is a person? If one is playing make believe using scarecrow, role playing as a mugger, but only as a game, then that’s probably fine I guess. Like, I don’t want to say that it is immoral to play an evil character in a D&D campaign; I don’t think that’s true.
But if one is messing with some ants, and one conceives of oneself as “torturing some ants”, I think one is fairly likely doing something wrong even though I don’t think the ants have a well-being, and there’s nothing wrong with killing a bunch of ants. And I think this is still true even if one has the belief “ants don’t actually have a well-being” at the same time as one conceives of what one is doing as “torturing some ants”.
Also, I feel like “nothing wrong if it does happen” regarding shooting someone, is the wrong perspective. If shooting someone is necessary, then it is necessary, but that doesn’t mean nothing went wrong. Anytime someone gets shot is a time something has gone wrong.