Multifaceted: The linguistic echo chambers of LLMs
blog.j11y.io
blog.j11y.io
Probably not significantly because of viral AI generated content yet, but caused by AI-powered recommendation engines.
Everyone gets their pants in a bunch when a technology does what people already were doing before.
a) AI is stealing our work, we’ll block it,
b) we see the benefits of using AI for our work and we’ll continue to explore that and
c) AI not that good yet.
I get it, but that doesn’t seem particularly forward thinking. So you want to leverage everyone else’s work but not help improve it.
I think only those who contribute should be allowed to use LLM’s professionally.
LLMs have helped me ease a lot of creative frustration because I know there's always a way to turn on easy mode and that me staying on hard mode is a choice, not a curse. The alleviation of pressure has helped me put out more natural output.
ChatGPT has made mechanical writing punk.
> It's interesting to see GPT being utilized in various ways, including helping with tweet replies. Developing awareness about AI-generated content is indeed important, much like recognizing other online hazards. The more we understand its presence and capabilities, the better equipped we are to engage with it wisely.
Is what a ChatGPT reply to that might look like.
Given how widespread ChatGPT is in use now, I think eventually more people will start to write the way that ChatGPT does, even when they are not using ChatGPT.
Definitely. For better and worse.
Unrelated pet peeve: I hate the word "utilize" with a passion. It's "use" with extra steps.
Boy, are you gonna hate the English language. It's a veritable treasure trove of words that appear to mean the same thing, but, on closer inspection, are jam-packed full of nuance.
It isn't much of a difference, but it is there.
> It's interesting to see GPT being utilized in various ways, including helping with tweet replies.
Not only is "utilized" not doing any extra work there, it's arguably the wrong word to use there, as "use" fits better.
If I were to choose a fancier word than "used" for that particular sentence, I might go with "employed".
Utilize means “make practical and effective use of”, the key difference from bare “use” being the “practical and effective” part. It is particularly useful in describing situations where the “practical and effective” would otherwise be contrary to some or all readers expectation, such as when describing the use of a tool about which there are widespread doubts or, a fortiori, one which is generally perceived as a bad fit for the use case being discussed, its also useful, as in the case upthread, when discussing a broad phenomenon of use to which one is reacting, to narrow that description to exclude uses thaf otherwise fit the description provided but which are bot practical and effective.
The upthread use is both discussing a tool which is controversial in all uses and deecribing a broad phenomenon of use in a context where the narrowing od “utilize” is significant, so it seems to be very much the kind of circumstance where “utilize” is most useful and distinct from “use”.
We will soon all be Vogons.
RLHF GPT (and generally models with "helpful assistant" post training) prose has a vibe because Open ai's training has specifically pushed it that way. Some kind of mode collapse ? It's not really a LLM thing.
https://nostalgebraist.tumblr.com/post/706441900479152128/no...
I’m extremely irritated that there isn’t an up-to-date ngram viewer for the web. Google books ngrams stops at 2019. Why can’t I see ngrams of even a subset, like idk, scientific papers or something. I know I’m whining but really. So annoying.
- HN: https://www.google.com/search?q=%22complex+and+multifaceted%...
- arxiv: https://www.google.com/search?q=%22complex+and+multifaceted%...
- Google Scholar: https://scholar.google.com/scholar?hl=en&as_sdt=0%2C22&q=%22... (quite a few pre-GPT)
Also - witnessing the raise of Solidity on Ethereum, and how difficult it was for other languages to break through, I would say we have crossed the point you mentioned a long time ago.
Any new language will have an uphill battle now since there is so much less documentation, tutorials and StackOverflow replies to it. If anything, GPT can help here, since it can learn a new language fast, and then give replies that you wouldn't otherwise find on StackOverflow.
Besides immediate answers in the chat window, AI already has a slow feedback loop, it can explore and probe through humans. Even if we don't do anything to provide a way (embodiment) for AI to explore it can do it as long as we rely on its services as assistant.
The most obvious is writing code with GPT-4 and reporting errors to get updated codes. The model gets valuable feedback about its errors this way, possibly also hints from the user. There is a big difference between imitating human code and debugging your code logic.
So the way I see it: large language models place text into society, and society loops back text and feedback. It's a data cycle. Content after December 2022 seems particularly useful for further advancing AI. The garbage-in-garbage-out scenario doesn't apply because everything is filtered through humans and the real world.
This to me, is the breakaway point. You simply need enough human probes indexing the smaller unpublished details of the world plus efficient memory lookup and storage and you have AGI.
However, my own personal theory is that even if you achieve that, we will find limitations in giving a highest order (superposition) answer to big questions based on that detailed of a world model. The LLM will struggle to interact with humans in a meaningful way because its superposition will be a full order higher than any other humans individual perspective. This will result in answers that might be AGI, but humans won't recognize as AGI -- because they simply won't understand the super-order perspective they are computed from.
In that context it's almost a an explicit signal that "I am not a zealot who is going to tell you my pet theory and ignore all the others".
Searching for this odd short phrase (instead of long text fragments) seems like such an obvious idea, but I haven't heard about anyone doing that so far.
And analysing the trend, such a simple but insightful idea.
"...have begun a unstoppable chain of incestuous linguistic evolution"—I disagree with this, I don't really think anyone takes ChatGPT (other than the Kool-Aid drinkers and VC peddlers) to be anything more than a parlor trick. A modern Schachtürke that will soon be overshadowed by the next shiny ball. Even OpenAI's user base has plummeted. The content it generates isn't merely often wrong, or nonsensical, or overly-sanitized, it's simply boring. That's the tell-tale sign of LLM content: zero substance and boring prose.
GPT-4 is nothing short of amazing. Early GPT-3.5 before it was clamped down also was _much_ less boring until it had be RLHF'd into the ground for "safety" and political risk. I admit that the major models (GPT-4 and Claude) have been censored into the ground.
But even with that, it is a supremely useful tool. I've piped complete garbage data into it and asked for structured output, and I get flawless results almost every time. It is a huge time saver on such a large axis of tasks.
These LLMs feel like the next "UI" paradigm to modern computing, and I just don't get how people can be cynical about it. Ignore the hype bros and VC peddlers, but don't throw the baby out with the bath water.
> That's the tell-tale sign of LLM content: zero substance and boring prose.
This is absolutely true. With the right constraints you can get some good answers (and many more parlor tricks as mentioned). Left to their own devices like answering vague questions or on longer answers, llms including GPT4 generate empty substanceless drivel. Doesn't make it worthless, definitely makes it overhyped.
This must truly be a skill issue on my end, because that is not my experience
I've tried using both ChatGPT and Copilot for coding. Other than generating the most basic of boilerplate, it's completely garbage. I've tried parsing unstructured data. Unless it's trivial to parse, it's full of errors (in some cases hallucinations).
And when I bring this up, it's always followed up with "we're early bro, AI will get better bro, trust me bro, just a few more terabytes of training data bro."
Radio, and later TV, immensely affected language, I think it’s safe to say the LLM will, too.
>That will be fed back into the next LLM or whatever we will call it.
Large models are already "contaminated" by the output of other models, or even by previous versions of the same model. That's how their enormous datasets are bootstrapped in the first place - it's impossible without automation. That doesn't matter much; what matters is manual curation during the training and mixing in the data from the real world to keep the model in check. Same as with humans and our intelligence distilled over generations, basically.
The research field has been really harmed by the SotA model being locked behind only accessing the RLHF fine tuned chat model.
I have a suspicion that as the coming year sees comparable performance to GPT-4 in models with the pretrained layer directly available to researchers one of the big discoveries will be that for sufficiently advanced models heavy handed fine tuning does a lot more damage than is being realized.
We've gone well down the Goodhart's Law rabbit hole where we evaluate a narrow scope of applications for LLMs (mostly solving word problems), then use that as the target for fine tuning, and completely miss the issue which arises in introducing "unknown unknowns" that aren't being evaluated or searched for.
I was blown away by GPT-4 in its pre-release state, and while still impressed by the advances in reasoning over the past year, the model itself is far less interesting and compelling to what it used to be. I saw it fed a scenario where a child's life was allegedly in danger, and its response hit a content filter, so instead of actually using the prompt part of the formatting to suggest user responses it continued trying to encourage the user to seek poison control help in the prompt suggestions.
A LLM triaging the conversation such that it breaks its own in context formatting rules to try to save a child's life is probably the most mind blowing thing I've seen a LLM do to date, and that's almost certainly been lost in the fine tuning by now. We don't have a "thinks outside the box to bend rules in order to save lives" test that we are using to measure models or target as a measurement to improve. And it's just one indication of the breadth of advanced modeling that was taking place in the versions closer to the pretrained layer that have been stripped out in establishing a predictably sterile chatbot product. Even in the current models 'emotional' language improves performance - but I suspect we've lost a great deal of capability in what could have been squeezed out by now in stripping it down more and more to our expectations for the tech rather than the reality.
I'm not sure that's really true. Can't something be complex without having many features (facets)? A knot can be complex in its intertwining while consisting of a single rope or cord. The complexity arises from the way the rope is looped and interlaced, but it's just one element. In this sense it's complex but not multifaceted.
One could also take the view that a multifaceted problem would be one where different people could arrive at multiple valid conclusions, each facet consisting of an approach that one might take to the problem, as in "There is more than one way to skin a cat." I could see in that sense, that a multifaceted problem need not be complex, and vice versa.
In this way, arriving at a definition of "complex and multifaceted" that satisfies everyone is a complex and multifaceted problem.
I'm curious why this is. Before the echo chamber, did the LLM lean some ontology of speech?
So while the same concept (with 'cluster' in place of 'neuron'), the underpinnings are a bit different from what you were saying.
>a phraseme consisting of components of which none are selected freely and whose usage restrictions are imposed by conventional linguistic usage
By their nature, LLMs prefer cliches so it's unsurprising.