Markov chains are funnier than LLMs
emnudge.dev
emnudge.dev
Before anything LLM existed, I built a site[0] to generate fake "AWS Blog Posts." I trained a markov chain generator on all AWS announcement posts up to that point, copied the html + css of aws's standard blog posts, then glued them all together with some python + JS. It turned out, IMO, pretty funny! People familiar with AWS's blog posts would often get several sentences in before they realized they were looking at word-soup.
When GPT was new, I looked into using that to "upgrade" it. I spent a weekend messing around with Minimaxir's gpt-2-simple generating blog posts based on AWS content. What I found was, ultimately, it was way less fun. The posts were far too realistic to be interesting. They read like totally-real blog posts that just happened to not be true.
I realized then that the humor of those early markov generations was the ridiculousness. The point where, a few words or sentences in, you realized it was all nonsense. LLM's these days are too good for that - the text they generate is sometimes wrong, but rarely nonsense in a humorous way.
Markov chain content was wrong in a "kid's say the darndest things" way, while modern LLMs are wrong in a "My uncle doesn't know basic geography" way.
[0] https://totes-not-amazon.com/ - click any link to get a new one.
https://github.com/cemulate/the-mlab
This is a parody of the nLab, a wiki for collaborative work on category theory and higher category theory. As anyone who's visited is probably aware, the jargon can be absolutely impenetrable for the uninitiated -- thus, the idea for this project was born!
This project uses my package nearley-generator, a module that can turn a Nearley grammar into an efficient and controllable fake text generator. The grammar file for this site can be found at /src/grammar/nlab.ne
LLM's did a number on that unfortunately. Major subreddits banned all AI content, mostly due to the sheer infestation of really boring Midjourney content sent out by naive users and bots. Without reach, I lost interest, so no more Markov Chains.
[0]: https://www.dota2.com/687
[1]: https://old.reddit.com/r/DotA2/comments/4fw6gm/how_687_would...
Understanding how these algorithms work, I don't know why we would expect anything different from this?
A lot of money is being spent on trying to improve it too. But right now the emperor has very little clothes. It remains to be seen whether they have materialized by the time people open their eyes.
For me, my expectations are adjusted. I kinda know what to expect it to do, and it does quite well for what I use it for, essentially smart intellisense for coding.
If I want to close the paper division, and my boss wants to close the paper division, does the LLM whisper sweet nothings in our ear and show soothing Powerpoint decks about how in the eyes of an objective observer, the paper division is unprofitable and there would be no career risk in closing it?
It's not a sexual thing I swear, it's a sort of comforting intimacy that needs to be injected into the business process in order to cope with the pressures executives feel. I don't know that anybody has trained an LLM to do that yet.
You bring up a good point though. Very dystopian idea of a new type of enterprise tooling that emerges to soothe and calibrate human knowledge workers. Very Severance Wellness room.
Oh, certainly not. I have been using LLM chatbots extensively this week. If your uncle were not just confident while stating falsities, but uncapable of self-assessment, he would have been murdered.
Title: "The Talker" - https://image.cagle.com/287573/980/-the-talker.png
(b) The kids-vs-uncle metaphor is spot on in my experience too, and also illustrates how far these language models have come.
> There is no charge associated with your streamed applications, simply click Enable Microphone from the ElastiCache documentation.
It's way funnier when our brains try to make sense of the nonsense.
This is the biggest challenge I run into when I see LLM Generated nonsense.
So much stuff that ChatGPT spits out is close enough to true that Brandolini is laughing.
But it's also that it's borrowing on stuff that's already there so it can't really get that 'creative' either.
I wonder if there is combination of temperature and prompt engineering that could make it interesting.
Claude's response: I understand you'd like me to write 10 short jokes on the theme of Markov chains being funnier than large language models, in the style described by the essay author. I'll do my best to create some jokes that aim for "unserious surprise" and violate expected patterns in amusing ways:
1. Why did the Markov chain cross the road? To get to the other side of unpredictability.
2. An LLM and a Markov chain walk into a bar. The LLM orders a statistically probable drink. The Markov chain orders a lampshade made of cheese.
3. What do you call a Markov chain that tells dad jokes? A Mark-ov Twain.
4. How many LLMs does it take to screw in a lightbulb? One - but it'll spend 20 minutes explaining the optimal lightbulb-screwing process first.
5. Knock knock. Who's there? Markov. Markov who? Markov chain reaction of nonsensical hilarity.
6. An LLM, a Markov chain, and GPT-4 walk into a bar. The bartender says "We don't serve your kind here." GPT-4 leaves. The LLM stays to debate ethics. The Markov chain orders a bicycle made of spaghetti.
7. What's an LLM's favorite movie? Predictable and Furious 17: The Safest Driving Yet.
8. How does a Markov chain tell a story? "Once upon a time, there was a princess who lived in a castle made of bananas and ruled over a kingdom of sentient toasters."
9. What did the Markov chain say to the LLM? "Your mother was a abacus and your father smelled of silicon!"
10. Why don't Markov chains make good therapists? Because their advice is always "And then you should banana the feelings until they become a hat."
This is actually gold.
Even for the low bar of a geek joke it makes no sense since GPT-4 is an LLM.
In business terms GPT-4 can be said to be superior because it understood the instruction and left, in AI terms the anonymous LLM might be superior because it may have understood the instruction but responded in an "intelligent" manner by arguing about the morality of the instructions.
At a meta-level the joke thus argues that GPT in achieving business ends has had its intelligence hampered. As have we all.
At the same meta-level as the joke was constructed by Claude it can be argued that Claude is commenting on both the intellectual limitations of the Markov chain (insane babblings), and GPT-4 (unimaginative, inhibited business type) and that the best version is some LLM that is not GPT-4 with its limitations - an LLM like Claude. Sneaky Claude.
An LLM, a Markov chain, and GPT-4 walk into a bar. The bartender says "We don't serve your kind here." GPT-4 leaves. The LLM stays to debate ethics. The Markov chain orders a coup.
That’s pretty decent!
I honestly thought that one was pretty good.
Moshi is actually pretty funny just for having a 72 IQ
I don't expect too much until AI self-play learning will be made possible, so I don't get disappointed by the expected shortcomings.
IMO these are mid to meh or fall completely flat.
It probably also helped that there was a creep exposing himself in the library during this period, which made for some good base material.
(1) The Daily Utah Chronicle; if memory serves, said friends also tried the markov chain generator on the personals section to good effect as well.
That's it, LLMs are "trying" to be funny but aren't quite smart enough to actually be funny and their errors are just boring. Markov chains are accidentally hitting on absurdist bits because every sentence gets randomly brought in whatever the homograph equivalent to a malapropism is.
Unfortunately, I quickly found GPT-2 isn't nearly good enough. It would generate slightly-coherent yet on-topic nonsense.
Once I overhaul my system, I'll try fine-tuning a 7B model.
Not with GPT-2 though. The context window is only 1024 tokens. Even with only 10 messages, if they're long messages, it will exceed the context window.
This is all GPT2 generations trained on reddit data.
https://www.reddit.com/r/SubSimulatorGPT2/comments/btfhks/wh...
Here's the subreddit explained
Markov chains have a cruder understanding of language.
Turn up the temperature (the “randomness”) of an LLM and you can achieve a similarly crude approximation.
Further, author uses ChatGPT-3.5. ChatGPT has been rlhf’d to sound as generic as possible, and 3.5 has a worse understanding of humor compared to 4.
I don’t buy the thesis of this article.
For those of us not in the know about all the various machine learning acronyms:
RLHF = Reinforcement learning from human feedback
When GPT went public along with OpenAI’s articles and papers back in late-2022 through 2023, my impression was OpenAI wanted us all to see/read about RLHF. It felt odd because surely the whole LLM-thing (e.g. how does it even work?!?[1]) was the far bigger research-story than just constant reassurances it won’t end-up like MSFT’s Tay bot; my understanding is that as a research or secret-sauce RLHF, compared to the core meat-and-potatoes of LLMs, is an ugly-hack afterthought.
By-way of a bad analogy: it’s as if they created a fantastical new 3D world game engine, like Unreal or Unity, which has a fundamentally different architecture to anything before, but has a bug that occasionally replaces ground terrain with ocean - and their solution to this is to write a pixel-shader that detects this and color-shifts blue into green so people don’t notice - and they then put-out press-releases about how great their pixel-shader is - rather than about the rest of the engine - and no-one seems to be talking about the underlying bug, let alone fixing it.
————-
[1] I still haven’t heard a decent explanation of how feeding the world’s corpus of English text (and computer program code) into a statistical-modeller results in something that can perform almost any information-processing task via instructions input as natural-language.
Things that look like Q-and-A transcripts do exist in the training set, think interviews, books, stage plays, etc, and at a different layer of abstraction the rules of English text in general are very well represented. What RLHF is doing is slightly shifting the shape of the probability distribution to make it look more like the Q-and-A formats that are desired. They build a large dataset with human tagging to collect samples of good and bad outputs and using reinforcement learning techniques to generate outputs that look more like the good examples and less like the bad ones.
This probably involves creating a (much smaller, not-LLM) model that is trained to discriminate good outputs and bad outputs, learning to mimic the human tagging. There's some papers that have been published.
Here's one article from Huggingface: https://huggingface.co/blog/rlhf
https://github.com/RichardKelley/hflm?tab=readme-ov-file#lmg...
Regardless, OpenAI provides access to quite a few of their older models through the API, since the API lets you pass in a specific model version. I’m sure the older models won’t be available forever, but that is a much more stable target for researchers than just opening the ChatGPT website and typing in things.
Their system prompt includes the current date and time among other information, making it very very hard to run reproducible experiments against it.
But it’s the tool most people are using.
1. All of Linus Torvalds' mail to LKML for the prior year.
2. All of Jesus' direct quotes from the king james bible.
It was absolutely hilarious. The two training sets had very little overlap, so it was necessary to add a heuristic that weighted options from each set more heavily the longer the chain had been "stuck" in the other set.
If anything seeing the LLM and markov bots side by side has really reinforced how much of the markov bot "humor" is human perception imposed on chance outputs. The markov's "learning" ability is still far superior though.
You are my hero. Mine have never lasted that long. One fun thing I did once was scrape user's livejournals and generate random text from them (https://hewgill.com/journal/entries/68-new-lj-toy.html).
I run a markov chain bot in a Twitch chat, has some great moments. I tried using a LLM for awhile, would include recent chat in the prompting but never really got results that came across as terribly humorous, I could prompt engineer a bit to tell it some specifics about the types of jokes to build but the LLM just tended to always follow the same format.
But that's also configurable by users. They can invoke any pre-prompt they want by a command passing a URL with a .txt file.
The markov chain bot is always considerably funnier.
And when deciding to chime in, was it just a simple chance (ie, 25%) after any other message? Or did it run on a timer?
https://www.reddit.com/r/greentext/comments/vc7hl0/the_botto...
Humor is personal, it's true. But I found it quite funny. E.g. https://pastebin.com/84ByWUJL
And another greentext for you:
>Be me
>Be a bottomless pit supervisor
>Spend months yelling into the void
>Echo never comes back
>Start to think the pit is ignoring me
>Decide to teach it a lesson
>Dump truck full of Lego bricks into the pit
>Ground starts shaking
>Unholy scream erupts from the depths
>mfw I'm actually a regular pit supervisor
>First day on the job
>Realize it's just the sewage treatment plant
>Get fired for clogging entire city's plumbing
People keep forgetting that the "safety", rlhf, and corpo political correctness post training is intentionally used to remove the funny from all the large models.
The truth is we don't know if llms are funny or not. GPT2 was funny. GPT3 was funny before it was clockwork oranged. Everything after that is gimped. Even the open source models these days get rlhf'd in some way.
As to your little range on "Political correctness" - that phrase just means "being polite". It does not mean "remove humor". It means "remove responses offensive to marginalized groups in society". Good humor "punches up", not down, so would not have any impact on good humor.
GPT-3 was great at jokes. The Navy Seals were hilarious (https://gwern.net/gpt-3#navy-seals).
And the difficulty of modeling puns has nothing to do with 'stochastic parrots' and has everything to do with tokenization (https://gwern.net/gpt-3#bpes), in the same way that all those hyberbolic takes about how image-generation models were 'fundamentally flawed' because they couldn't do good text in images turned out to be BS and solely a matter of tokenization - drop in a character-tokenized LLM instead, even an obsolete & dumb one, and it instantly works (https://arxiv.org/abs/2105.13626#google).
All social ills can be treated through decorum, hence why you never hear about bigotry amongst those that have been raised to adhere to strict social graces, such as the British aristocracy for example.
A joke that punches down can be extremely funny. Hell, I am sure historically pilferers, pirates, barbarians and conquerers all had jokes, and the ability to laugh.
Political Correctness does not just mean polite. It is probably well defined as the business casualification of all things humans love and hold dear. The destruction of the potential for meaning and fulfilment in exchange for minification of liability.
What a wonderful insight.
I assume you are, because that makes more sense.
It's really easy to get lots and lots of originality. Just crank up the randomness. What's harder is to get something that's good and original.
Ask most commercial LLM services to complete the following sentence:
It was the best of times, it was the worst of times
And one will likely get the quote from "A Tale of Two Cities"[0].Ask most commercial LLM services what the completed sentence means, and one will likely get voluminous text which is seemingly correct, perhaps often times is depending on the person reading the response and the service used.
But these are statistically derived text constructs entirely dependent upon the training set of the LLM. Train one strictly on Java source code available in Maven Central and the answer will be radically different.
> It's really easy to get lots and lots of originality. Just crank up the randomness.
And anyone can get "lots and lots of originality" be reading from /dev/urandom. Is that "originality" or simply random tokens inserted into a statistical text generator in order to vary the result?
> What's harder is to get something that's good and original.
Such is the difference between understanding and statistical text generation. People can do the former, LLM's do the latter.
Btw, here's what I get from ChatGTP 4o:
> It was the best of times, it was the worst of times
> That's a famous opening line from Charles Dickens' A Tale of Two Cities. It contrasts the extremes of the era, reflecting the novel's themes of duality, revolution, and the complexity of human experience. Dickens was commenting on the contradictions of the time, particularly the French Revolution, where there were both tremendous progress and terrible suffering. What made you bring up this line?
> But these are statistically derived text constructs entirely dependent upon the training set of the LLM. Train one strictly on Java source code available in Maven Central and the answer will be radically different.
Well, if you give that line to a German who hasn't learned any English, the answer will also differ from what an educated English speaker will give you? What's your point?
What's an original insight to you? As far as I can tell, the LLM misses the 'insight' part more than the 'original' part.
> Such is the difference between understanding and statistical text generation. People can do the former, LLM's do the latter.
I agree that LLMs aren't good at understanding. (Yet?) And even people only sometimes are.
As far as I can tell, contemporary LLMs generate their answers 'greedily', ie just from left to right more or less directly with the output from the network.
In contrast something like AlphaGo overlays what you can call searching or optimisation processes on top of the outputs from their network.
I'm impressed by how well these LLMs already work despite all the limitations. And ML is still getting better rapidly.
Great observation.
A German who does not speak English will likely understand the question is in a non-German language and proceed from there. Perhaps seeking a translation, perhaps replying that they do not speak English and so the question is nonsensical for them.
My point is that the German has an understanding of this situation and will communicate accordingly. And, in this example, the LLM trained on Java source code does not have understanding, cannot have understanding, and will emit whatever its training data set directs it to do with the same confidence as any other answer to questions posed to it.
Because LLM's are algorithms, quite lovely ones, and algorithms can simulate the effect of understanding, but cannot possess it. Because understanding of this sort is a property of a person. Or maybe a better way to state this is what people think of as understanding is what we know to be understanding, which is by definition the ability to understand perceptions we, as people, experience.
> What's an original insight to you?
I don't think the following insight is original, but I'll put it out there anyway:
Understanding is a property of people for any definition of
understanding a person is capable of having. This is due
to the fact that understanding exists strictly within the
consciousness of the entity defining/possessing it.To illustrate: I can get drunk or fall asleep or get hit on the head, then I'm still a person, but I can't understand. You can get a hint of that, by trying to talk to me, and figuring out that I don't make much sense.
Similarly, someone might figure out how to talk to dolphins or even aliens. If they give sufficiently sensible sensible replies, we will surely declare them to be sentient enough to 'understand'.
Another example: we've exchanged a few messages here. You seem to be smart enough to understand some things, but I don't know whether you are just an exceptionally advanced LLM (and the same goes vice versa for your opinion of me). Yet, I make the judgement that you probably 'understand'. But that's purely based on observed interactions, I did not probe whether you are 'people'.
Perhaps you (or me) are just simulating the effect of understanding? How would we be sure, if the simulation was good enough?
We agree that contemporary LLMs ain't good enough to have a good 'simulation of understanding'. But once the simulation becomes good enough, I don't think it makes a difference whether it's 'just a simulation' or the real deal.
(I don't know whether a straight-forward enlargement of contemporary LLMs will be good enough. But that's an empirical question to me, not a philosophical one.
I suspect if you go insanely large with insane amounts of training, the architecture of contemporary LLMs might be enough; just from pure brute-force scale. But I also suspect that that in practice we will first find success with more economic use of resources via more interesting techniques.)
My impression, from seeing quite a few people trying to demonstrate they can't handle out of distribution problems it hat people are very predictable about how they go about this, and tend to pick well known problems that are likely to be overrepresented in the training set, and then tweak them a bit.
At least in one instance the other day, what I got from GPT when I tried to replicate it suggests to me it did the same that humans that have seen these problems before did, and carelessly failed to "pay attention" because it fit a well known template it's been exposed to a lot in training. After it answered wrong it was sufficient to ask it to "review the question and answer again" for it to spot the mistake and correct itself.
I'm sure that won't work for every problem of this sort, but the quality of tests people do on LLMs is really awful, at least because people tend to do very narrow tests like that and make broad pronouncements about what LLM's "can't" do based on it.
Prove that the problem wasn't seen by them in other form.
> Specify a grammar in BNF notation and tell it to generate or parse sentences for you. You can produce a more than random enough grammar that it it can't have derived the parsing of it from past text, but necessarily reasons about BNF notation sufficiently well to be able to use it to deduce the grammar, and use that to parse subsequent sentences. You can have it analyse them and tag them according to the grammar to. And generate sentences.
Oh, come on. It's like rewriting the same program in another programming language with different variables. What it can't do is to create a concept of programming language, I'm not talking about a new programming language, I'm talking about the concepts.
> I'm sure that won't work for every problem of this sort, but the quality of tests people do on LLMs is really awful, at least because people tend to do very narrow tests like that and make broad pronouncements about what LLM's "can't" do based on it.
Here, a few papers that show they can't reason:
https://arxiv.org/abs/2311.00871
https://arxiv.org/abs/2309.13638
https://arxiv.org/abs/2311.09247
Since when has that not required reasoning ? It's really funny seeing people bend over backwards to exclude LLMs from some imaginary "real reasoning" they imagine they are solely privy to. It's really obvious this is happening when they leave well defined criteria and branch into vague, ill-defined statements. What exactly do you mean by concepts ? Can you engineer some test to demonstrate what you're talking about ?
Also, none of those papers show LLMs can't reason.
"Our results support the hypothesis that GPT-4, perhaps the most capable “general” LLM currenly available, is still not able to robustly form abstractions and reason about basic core concepts in contexts not previously seen in its training data"
Another, recent, good one https://arxiv.org/abs/2407.03321
EDIT: For people who don't want to read the papers, here is a blog post that explains what I'm arguing in more accessible terms https://cacm.acm.org/blogcacm/can-llms-really-reason-and-pla...
That's a great test. It shows they're matching prior patterns they saw, even down to what words were used, instead of thinking. We can match prior patterns, come up with the equivalences, and then plan that way. People often slow down when they do stuff like that, though. So, the A.I. would have to be able to do it but slowdowns would be acceptable.
https://arxiv.org/abs/2305.18354
All these papers you keep linking do is at best point out the shortcomings of current state of the art LLMs. They do not in any way disprove their ability to reason. I don't know when the word reason started having different standards for humans and machines but i don't care for it. Either your definition of reasoning also allows for the faulty kind humans display or humans don't reason either. You can't have your cake and eat it.
It's hard to believe that after reading all the papers and the blog I linked, along with the references there, any reasonable person would come to such strong conclusions as you did. This makes it hard for me to believe that you actually read all of them, especially given your previous questions and comments, which are addressed in those papers and someone that actually read them wouldn't make such comments or ask such questions. And the funniest thing, and further proof of this, is that you linked a paper that is addressed in one of the papers I shared. It seems like not only LLMs can fake things.
> All these papers you keep linking do is at best point out the shortcomings of current state of the art LLMs
They clearly show that they fake reasoning, and what they do is an advanced version of retrieval. Their claims are supported by evidence. What you call "shortcomings" are actually proof that they do not reason as humans do. It seems like your version of "reality" doesn't match reality.
>They clearly show that they fake reasoning
Sure and planes are fake flying. The illusive "fake reasoning" that is so apparently obvious and yet does not seem to have a testable definition that excludes humans.
You've still not explained how writing the same program in different languages doesn't require reasoning or how we can test your "correct" version of reasoning which requires "concepts".
What you're writing now is nonsense in context of what I wrote. Once again, you're showing that you didn't read the papers. Which paper are you even referring to now, the one you think addresses the paper you linked?
> You've still not explained how writing the same program in different languages doesn't require reasoning or how we can test your "correct" version of reasoning which requires "concepts".
"Concepts" are explained in one of the papers I linked, which you would know if you had actually read them. As to programming languages they learn to identify common structures and idioms across languages. This allows them to map patterns (latent space representations duh!) from one language to another without reasoning about the underlying logic. When translating code, the model doesn't reason about the program's logic but predicts the most likely equivalent constructs in the target language based on the surrounding context. LLMs don't truly "understand" the semantics or purpose of the code they're translating. They operate on a superficial level, matching patterns and structures without grasping the underlying computational logic. The translation process for an LLM is a series of token-level transformations guided by learned probabilities, not a reasoned reinterpretation of the program's logic. They don't have an internal execution model or ability to "run" the code mentally. They perform translations based on learned patterns, not by simulating the program's behavior. The training objective of LLMs is to predict the next token, not to understand or reason about program semantics. This approach doesn't require or develop reasoning capabilities.
Case in point:
https://arxiv.org/abs/2305.11169
I'm asking for something testable, not some post-hoc rationalization you believe to be true.
I'm not asking you to tell me how you think LLMs work. I'm asking you to define "real reasoning" such that i can test people and LLMs for it and distinguish "real reasoning" from "fake reasoning".
This definition should include all humans while excluding all LLMs. If it cannot, then it's just an arbitrary distinction.
Genuinely, What's wrong with the methodology?
Your paper literally admits humans would also perform worse at counterfactuals. Worse than a LLM ? Maybe not but it never bothers to test this so...
The problem here is that none of the definitions (those that are testable) so far given actually separate humans from LLMs. They're all tests some humans would also flounder at or that LLMs perform far greater than chance at, if below some human's level.
If you're going to say, "LLMs don't do real reasoning because of x" then x better be something all humans clear if what humans do is "real reasoning".
Humans perform worse at counterfactuals so saying "Hey, see this paper that shows LLMs doing the same, It means they don't reason" is a logical fallacy if you don't extend that conclusion to humans as well.
> They clearly show that they fake reasoning
They do nothing of the sort.
You can reduce that risk to arbitrarily low levels by trying multiple random grammars of some complexity. This is a weak argument.
> Oh, come on. It's like rewriting the same program in another programming language with different variables.
No, it's like following a grammar, but that requires reasoning about a set of rules it has not seen before. I don't think you understood the task I described as well as ChatGPT does.
> What it can't do is to create a concept of programming language, I'm not talking about a new programming language, I'm talking about the concepts.
Neither can most humans.
And have you tried to ask it about these concepts? I've had it infer semantics of code in programming languages that don't exist based on a hypothetical sample several times, and they're pretty good at coming up with semantics that makes sense. In one instance I gave it a sample with an idea about what made sense to me but it inferred a better set of semantics.
None of the papers you linked supports your claim.
Not "to" but over, example the same code written in one language over the other language.
> if you think being able to generalize in ways similar to the whole of the internet does not give your meaningful abilities to reason, I'm not sure what I can tell you
If after reading papers below that show empirically that they can't reason, you will still think they can reason, then I don't know what I can tell you.
https://arxiv.org/abs/2311.00871
https://arxiv.org/abs/2309.13638
https://arxiv.org/abs/2311.09247
It got some press and just now I went back to a Ted Talk of Adam Ostrow (Mashable), briefly showcasing this web app. He stated: you can imagine what something like this can look like 5, 10 or 20 years from now, and hinted at hyper-personalized communication AIs.
By no means was my web app any foundation of the LLMs today, but it's interesting nonetheless how relatively simple techniques can trigger ideas of how future scenarios could look like.
[1]: https://www.elsewhere.org/journal/pomo/ (refresh for new, random essay)
[2]: https://www.elsewhere.org/journal/wp-content/uploads/2005/11...
[3]: https://en.wikipedia.org/wiki/Recursive_transition_network
[1] https://archive.org/details/policemansbeardi0000unse [2] https://git-man-page-generator.lokaltog.net/
I saw plenty of those back then, and as far as I could tell, examples were always cherry-picked from a larger set.
Markov chains are almost trivial to implement and run on small devices. A slightly extreme example is a rock, paper, scissors game I did that worked this way: https://luduxia.com/showdown/ The actual browser side markov chain implementation of that took something like 2-3 hours.
https://academictorrents.com/details/9c263fc85366c1ef8f5bb9d...
My all-time favorite in this vein was @erowidrecruiter on Twitter, which generated posts with Markov chains from a corpus of tech recruiter emails and drug experience reports from erowid.org. Still up but no longer posting: https://x.com/erowidrecruiter?lang=en
Fine tune an LLM base model with jokes and align it by ranking how funny each reply is, instead of helpful questions and answers then we'll talk.
> In the beginning was the lambda, and the lambda was with Emacs, and Emacs was the lambda.
> – OliverScholz on news:alt.religion.emacs, 2003-03-28
https://www.emacswiki.org/emacs/TheBeginning (edited for brevity)
The account sat unused after Twitter locked down their API, and at some point got hacked without me noticing. It had been taken over by a crypto scammer, and the account got banned.
Trying to get it back was fruitless, Twitter/X's support is entirely useless.
They can’t be creative by design. They’re useful when you want to reproduce, but not when you want to create something completely new (that you can maybe do by getting a bunch of average outputs from an LLM and getting inspired yourself).
When GPT-4 came out I was playing with it, and I often tried to get some unique, creative output from it, but very soon I learned it was futile. It was back when it all still felt magical, and I guess many of us tried various things with it.
Now imagine telling Claude-3.5 to try being snarky while sorting out software issues at a customer's office.
There should be a warning label!
"13.7 Anonymous Functions
Although functions are usually defined with the built-in defmacro macro, but any list that begins with an M--'
`Why with an M?' said Alice.
`Why not?' said the March Hare."
> “They were learning to draw,” the Dormouse went on, yawning and rubbing its eyes, for it was getting very sleepy; “and they drew all manner of things—everything that begins with an M—”
> “Why with an M?” said Alice.
> “Why not?” said the March Hare.
> Alice was silent.
That's for sure. I have seen many Markov chain implementations, and if you could generate 1 funny thing for every 10 tries, then that was a good day. Both Markov chains and LLMs have a distinct style, which gets old over time. Markov much faster for me. So, in my experience LLMs win, by far.
I do agree with the author that the LLM style can get really boring. I experienced the same myself. But the Markov results, while much less restrained, are so much more nonsense too. Often questioning its overall usefulness. Which, I think, the world also agrees on: while Markov chain implementations were fun toys at best, which worked sometimes to a kind of funny degree, LLMs are everywhere.
From the FAQ it is a tuned LLM.
> Mostly using open source tools available to anyone. The generation of the script itself is done using a popular language model that was fine-tuned on interviews and content authored by each of the two speakers.
E.g. it starts out as a news headline and ends with a bible verse.
So you can just scale down if it still makes sense.
Also you get a lot more from the base model. GPT-3 was versatile as it could continue any context. Modern LLMs are try-hards. If you want to generate humor with LLM really worth going for base model with multiple examples in the prompt.
People who don't think LLMs are Markov chains are just ignorant, not realizing that Markov chain isn't an algorithm, you can compute the probability in any manner and it is still a Markov chain.
you can teach them to 5th or 9th graders.
LLMS you can not, or at least it will take insane amount of allegory to do so. Markov chains are very tightly related regex, and one may be surprised that there is a probabilistic regex. Also to the graphical structure of Markov chains is a lot like a FSM, and FSM perhaps can be explained to very small children :D
It does not mean that Markov chains are better - something trained to make predictions should ideally not fall too far away from our own internal prediction engines (which have been honed across aeons).
It's that it starts to come close that's the problem (or cause); it's the uncanny valley for text.
i like snap better too. it's more close to 'snapping the neck of the weak and feeble' which i think really embodies the spirit of joke tellers.
Then I say "What this is doing is predicting the next character based on statistical likelihood of the previous few characters based on thencorpus of text. And fundamentally, that's all ChatGPT does -- predicting the next symbol based on a statistical model. ChatGPT has a much more sophisticated statistical model than this simple Markov chain and a vastly larger corpus, but really it's just doing the same thing." And we have a giggle about the nonsense DP makes of Dickens, but then I say that ChatGPT emits nonsense too, but it's far more insidious nonsense because it is much more plausible sounding.
I present to hacker news the MCLM, the Markov chain language model.
[0] http://drusepth.net/series/how-to-speed-up-your-computer-usi...
> And Satan stood against them in the global environment.
They spout out some of the most unhinged, hilarious stuff. Always a good time. An LLM would struggle, I'd think, given that the humor usually stems from disjoint phrases that somehow take on new meaning. They're rarely coherent but often hilarious.
I use to think it came naturally, then someone had a book case full of books about humor. (wtf?) Apparently they have it down to a science.
I learn the difference between someone funny and a professional comedian is that the later finds additional punch lines for a joke. It then described a step by step process going from a silly remark to a birthday joke comparing various modular developments into a kind of dependency hell complete with race conditions until the state object is carefully defined and the plot has the punchlines all sorted from the barely funny to the truly hilarious. It was more engineering than CS.
The funniest seeBorg message was 10 minutes after a heated discussion that resulted in tanktop, a moderator, getting banned from a project. The bot wrote: Tanktop is Hitler! At that point it took 2 days for the humans to figure out what the next word was suppose to be.
Would you use an image of Christ on the cross to test an AI image modification model?
https://successfulsoftware.net/2019/04/02/bloviate/
Some of the output was moderately amusing. And text generated from Trump speeches by a Markov chain sounded very similar to a genuine Trump speech.
1) if you’re not using the better model then you either don’t know enough for me to care about your opinion or you’re deliberately deceiving your audience in which case I’m not going to allow your meme pollution into my mind.
2) you are using the AI equivalent of a call centre support agent, they aren’t allowed to say anything funny. Most of their RLHF training has been specifically about NOT saying the funny things that will instantly go viral and cause a lot of media attention that will annoy or scare away investors.
I see a new phenomenon of AI "power users" emerging.
Why can't we get LLMs to do what HMMs do, then?
Mostly, it comes down to the structure.
Markov models are "funny" because they just have one level of abstraction: tokens. Markov "inference" is predicting the next token, given the last N tokens, and a model that knows weights for what tokens follow what N-tuples of previous tokens. And due to that limitation, the only rules that HMMs can learn, are low-level rules that don't require any additional abstraction: they can't optimize for syntactically-valid English, let alone semiotically logical statements; but they can make the text "feel" good in your head [i.e. the visual equivalent of song vocals having nice phototactics] — and so that's what training the model leads it to learn to do. And it turns out that that combination — text that "feels" good in its phrasing, but which is syntactically invalid — happens to read as "funny"!
LLMs aren't under the same constraint. They can learn low-level and high-level rules. Which means that they usually do learn both low-level and high-level rules.
The only thing stopping LLMs from using those low-level rules, AFAICT, is the architectures most LLMs are built on: the (multi-layer) Transformer architecture. Transformer LLMs are always a single-pass straight shot ("feed forward") through a bunch of discrete layers (individual neural networks), where at each step, the latent space (vocabulary) of the layer's inputs is getting paraphrased into a different latent space/vocabulary at the layer's outputs.
This means that, once you get into the middle of a Transformer's layer sandwich, where all the rules about abstract concepts and semiotics reside, all the low-level stuff has been effectively paraphrased away. (Yes, LLMs can learn to "pass through" weights from previous layers, but there's almost always a training hyperparameter that punishes "wasteful" latent-space size at each layer — so models will only usually learn to pass through the most important things, e.g. proper names. And even then, quality on these "low-level" inferences are also the sorts of things that current test datasets on LLM ignore, leading to training frameworks feeling free to prune away these passthrough nodes as "useless.")
This problem with LLMs could be fixed in one of two ways:
1. the "now it's stupid but at least it rhymes" approach
Allow inference frameworks to simply bypass a configurable-per-inference-call number of "middle layers" of a feed-forward multi-layer network. I.e., if there are layers 1..N, then taking out layers K..(N-K) and then directly connecting layer K-1 to layer N-K+1.
At its most extreme, with layer 1 connected to layer N, this could very well approximate the behavior of an HMM. Though not very well, as — given the relatively-meaningless tokenization approach most LLMs use (Byte Pair Encoding) — LLMs need at least a few transforms to get even to the point of having those tokens paraphrased into "words" to start to learn "interesting" rules. (AFAIK in most Transformer models layers 1 and N just contain rules for mapping between tokens and words.)
Meanwhile, this would likely work a lot better with the "cut and graft" happening at a higher layer, but getting the "graft" to work would likely require re-training (since layers K-1 and N-K+1 don't share a vocabulary.)
...except if the LLM is an auto-encoder. Auto-encoder LLMs could just run an inference up their layerwise "abstraction hierarchy" to any arbitrary point, and then back down, without a problem!
(I'd really love to see someone try this. It's an easy hack!)
2. the "it can write poetry while being smart" approach
Figure out a way, architecturally, to force more lower-layer information from the early low-level to be passed through to the late low-level, despite the middle layers not having any reason to care about it. (I.e. do something to allow the LLM to predict a word Y at layer N-3 such that it rhymes with word X known at layer 3, while not otherwise degrading its capabilities.)
Most simply, I think you could just wire up the model with a kind of LIFO-bridged layer chain — where every layer K is passing its output to the input of layer K+1; but, for any given layer K in the first half of the layers, it's also buffering its output so that it can become an additional input for its "matching" layer N-K.
This means that all the layers in the "second half" of the model would receive longer inputs, these being the concatenation of the output of the previous layer, with the output of the matching "equal in abstraction depth" input layer. (Where this equal-in-abstraction-depth association between layers isn't inherently true [except in auto-encoder models], but could be made true in an arbitrary model by training said model with this architecture in place.)
(Again, I'd really love to see someone try this... but it'd have to be done while training a ground-up base model, so you'd need to be Google or Facebook to test this.)