Replacing my best friends with an LLM trained on 500k group chat messages
izzy.co
izzy.co
How many emails, text messages, hangouts/gchat messages, etc, does Google have of you right now? And as part of their agreement, they can do pretty much whatever they like with those, can't they?
Could Google, or any other company out there, build a digital copy of you that answers questions exactly the way you would? "Hey, we're going to cancel the interview- we found that you aren't a good culture fit here in 72% of our simulations and we don't think that's an acceptable risk."
Could the police subpoena all of that data and make an AI model of you that wants to help them prove you committed a crime and guess all your passwords?
This stuff is moving terrifyingly fast, and laws will take ages to catch up. Get ready for a wild couple of years my friends.
Seems like a perfect technology to implement these talking photographs, paintings and pictures from there.
No, they definitely can't. Parts of HN love to hate on GDPR, but laws like that prevent companies from doing the things you proposed.
Recently, Meta was fined $400MM for forcing users to consent to targeted advertising [0]. Note how Meta was careful to get consent (even if the way they did it was illegitimate). Sure, $400MM may not be a lot for a company that size, but I genuinely believe that the fines would be an order of magnitude higher if a company intentionally decided to do something with personal data without consent. GDPR fines may reach up to 4% of worldwide revenue, plus likely any proceeds from the illegitimate venture.
[0] https://www.cnbc.com/2023/01/04/meta-fined-more-than-400-mil...
1. A lot of the fines come from edge cases that are literally unclear in the law. Eg Facebook‘s opt out for advertising fines. You can disagree with fb’s decision but teams of lawyers couldn’t answer this question except in court. I think American and European jurisprudence aware also pretty different so someone sitting in California making business decisions might not understand the ramifications in Europe.
2. A lot of the thorny privacy bits can be bypassed if you update the TOS to mention it (or so they think). I’ve seen that happen a few times during my tenure.
That doesn’t excuse the choice of these companies to make these choices, but my point was to say that companies take it seriously but lawyers don’t always agree on how laws work except in court.
The other point, is that technically this AI is not "unaligned". It is doing exactly what is requested of the operator.
The implications are that humanity suffers in either scenario, either by our own agency in control of power we are not prepared to manage or we will be managed by power that we can not control.
But we don't need AI or LLMs at all for the above scenario. Companies don't currently pry into your e-mails to make hiring decisions, but they could (ignoring laws) do it if they wanted. No LLM or AI necessary.
So why would the existence of AIs or LLMs change that?
If they wanted to use the content of your e-mails against you, they don't need an LLM to do it.
To some artists, AI generated images from their styles would amount to "productive piracy." Unlike torrenting the act is often out in the open since users tend to share the results online. I'm not sure if this phenomenon has happened before; with teenagers pirating Photoshop it's impossible to tell from a glance if the output is from a pirated version.
Also, what types of behavior did we get a glimpse of from the Twitter Files?
Aren't there always constant lawsuits about bad behavior of companies especially around privacy?
So yes, we are talking about the same behavior existing, but the concern is that they now get orders of magnitude more power to extend such bad behavior.
Can you actually explain the types of bad behavior? The rhetorical question about The Twitter Files somehow being a groundbreaking expose of bad behavior doesn't really match anything I've seen. Most of what was cited was essentially a social media company trying to enforce their rules.
Might want to read up on the latest developments there. Several journalists have debunked a lot of the key claims in the "Twitter Files". Taibbi's part was particularly egregious, with some key numbers he used being completely wrong (e.g. claiming millions when the actual number was in the thousands, exaggerating how Twitter was using the data, etc.).
Even Taibbi and Elon have since had a falling out and Taibbi is leaving Twitter.
If Elon Musk so famously and publicly hates journalists for lying, spinning the truth, and pushing false narratives, why would he enlist journalists for "The Twitter Files"? The answer is in plain view: He wanted to take a nothingburger and use journalists to put a spin on it, then push a narrative.
Elon spent years saying that journalists can't be trusted because they're pushing narratives, so when Elon enlists a select set of journalists to push a narrative, why would you believe it's accurate?
> So yes, we are talking about the same behavior existing, but the concern is that they now get orders of magnitude more power to extend such bad behavior.
No they don't. The ultimate power is being able to read the e-mails directly. LLMs abstract that with a lower confidence model that is known to hallucinate answers when the underlying content doesn't have a satisfactory set of content.
I agree that Musk has not honored his original intent. He has already broken in many ways the transparency pledge and free speech principles.
Yet, these were already broken under previous ownership. We simply see that as continuing.
But wait, you can just dump that information into a superAIcomputer and get reliable enough results while not needing a break with little to no risk of the computer rising up against you. Sounds like a hell of a deal.
Quantity is a quality in itself.
To fix this, you can train your personal LLM on the “FAANG Appropriate Banter” dataset, and then have it send messages to your friends daily for several months in the lead up to your interview.
“We’re saving the world” they say. With zero regard for second order effects, or with arrogant dismissal of those effects as worth it for the first order gains.
Disgusting to watch unfold.
Most absolutely not with the 7B llama model as described here.
…but, potentially, with a much larger fine tuned foundational model, if you have a lot of open source code on GitHub and lots of public samples.
The question is why you would bother? very large models would most likely not be meaningfully improved by fine tuning on a specific individual.
The only reason to do this would be to coax a much smaller model into having characteristics of a much larger one by fine tuning… but, realistically, right now, it seem pretty skeptical anyone would bother.
Why not just invest in a much larger more capable model?
Fine tuning would allow you to achieve that for people who aren’t notable enough to be all over the training data
I know it's not popular to try and remove the blindfold of propaganda but the 2020 to 2022 and covid authoritarian, anti human rights, policies were an awesome example.
But if you are quick to dismiss those there's always terrorism, the children, drugs, and every other typical gov excuse.
It's not a tech problem. It's a problem of the people who have decided they will go along with autority at every step, If they get pushed.
Comparing with the bottom of the barrel to make yourself look good? That's like a country using the US healthcare situation to claim their own healthcare is good.
It's a poor car salesman trick.
If a company is going to snoop in your personal data to get insights about you, they'd just do it directly. Hiring managers would scroll through your e-mails and make judgment calls based on their content.
Training an LLM on your e-mails and then feeding it questions is just a lower accuracy, more abstracted version of the above, but it's the same concept.
So the answer is: In theory, any company could do the above if they wanted to flout all laws and ignore the consequences of having these practices leak (which they inevitably would). LLMs don't change that. They could have done it all along. However, legally companies like Google cannot, and will not, pry into your private data without your consent to make hiring decisions.
Adding an LLM abstraction layer doesn't make the existing laws (or social/moral pressure) go away.
Isn't the "abstraction" of "the model" exactly the reason we have open court filings against stable diffusion and other models for possibly stealing artist's work in the open source domain and claiming it's legal while also being financially backed by major corporations who are then using said models for profit?
Whose to say that "training a model on your data isn't actually stealing your data" it's just "training a model" as long as you delete the original data after you finish training?
What if instead of Google snooping, they hire a 3rd party to snoop it, then another 3rd party to transfer it, then another 3rd party to build the model, then another 3rd party to re-sell the model. Then create legal loopholes around which ones are doing it for "research" and which ones are doing it for profit/hiring. All of the sudden, it gets really murky who is and isn't allowed to have a model of you.
I feel one could argue that the abstraction is exactly the kind of smoke screen that many will use to avoid the social/moral pressures legally, allowing them to do bad things but get away with it.
The provenance of the training set is key. Every LLM company so far has been extremely careful to avoid using people's private data for LLM training, and for good reason.
If a company were to train an LLM exclusively on a single person's private data and then use that LLM to make decisions about that person, the intention is very clearly to access that person's private data. There is no way they could argue otherwise.
- collect thousands of people's data
- anonymize it
- then shadow correlate the data in a web
- then trace a trail through said web for each "individual"
- then train several individuals as models
- then abstract that with a model on top of those models
Now you have a legal case that it's merely an academic research into independent behaviors affecting a larger model. Even though you may have collected private data, the anonymization of it might fall under ethical data collection purposes (Meta uses this loophole for their shadow profiling).
Unfortunately, I don't think it is as cut and dry as you explained. As far as I know, these laws are already being side-stepped.
For the record, I don't like it. I think this is a bad thing. Unfortunately, it's still arguably "legal".
No, they haven’t. (Now, if you said “people's private data” instead of “private people's data”, you’d be, at least, less wrong.)
> Training an LLM on your e-mails and then feeding it questions is just a lower accuracy, more abstracted version of the above, but it's the same concept.
Its also one that once you have cheap enough computing resources scales better, because you don't need to assign literally any time from your more limited pool of human resources to it. Yes, baroque artisanal manual review of your online presence might be more “accurate” (though there's probably no applicable objective figure of merit), but megacorporate hiring filters aren't about maximizing accuracy they are about efficiently trimming the applicant pool before hiring managers have to engage with it.
This is like saying, "look, no one would be daft enough to draw a graph, they'd just count all the data points and make a decision."
You're missing two critical things:
(1) time/effort (2) legal loophole.
A targeted simulation LLM (a scenario I've been independently afraid of for several weeks now) would be a brilliant tool for (say) an autocratic regime to explore the motivations and psychology of protesters; how they relate to one another; who they support; what stimuli would demotivate ('pacify') them; etc.
In fact, it's such a good opportunity it would be daft not to construct it. Much like the cartesian graph opened up the world of dataviz, simulated people will open up sociology and anthropology to casual understanding.
And, until/unless there are good laws in place, it provides a fantastic chess-knight leap over existing privacy legislation. "Oh, no we don't read your emails, no that would be a violation; we simply talk to an LLM that read your emails. Your privacy is intact! You-prime says hi!"
That seems as poor as saying, "We didn't read your emails -- we read a copy of your email after removing all vowels!"
But we live in distressed times, and the law is not as sane and sober as it once was. (Take, for example, the Tiktok congressional hearing; the wildly overbroad RESTRICT act; etc.)
If the people making and enforcing the laws are as clueless and as partisan as they by-all-accounts now are, what gives you hope that, somehow, some reasonable judge will set a reasonable precedent? What gives you hope that someone will pass a bill that has enough foresight to stave off non-obvious and emergent uses for AI?
This is not the timeline where things continue to make sense.
Not really. Assuming your ethical compass is broken and you suspected your partner of cheating, would you rather have access to their emails or to a LLM trained on them? Also, isn't it much cheaper for Google to simply search for keywords rather than fine tuning a model for this?
At least in the EU, a system like this would be made illegal on day one. This whole doomsday scenario seems predicated on a hypothetical future where LLM's would be the least of your worries.
have you met capitalism?
I feel like I'm talking to someone from the timeline where Clearview AI and Cambridge Analytics never happened.
Generally I think this idea can't work because of Goodhart's Law - people's behavior changes when you try to influence them.
For my argument, I only need to point out that it was attempted, as I'm proving motivation; the effectiveness of CA methods has no bearing on the effectivenss of (say) simulated people.
Increasingly, when interacting with comments on HN and elsewhere, it feels like I'm from a parallel timeline where things happened, and mattered, and an ever-growing percentage of my interlocutors are, for lack of a better word, dissociated. Perhaps not in the clinical sense, but certainly in the following senses:
- Cause and effect are not immediately observed without careful prompting.
- Intersubjectively verifiable historical facts that happened recently are remembered hazily, and doubtfully, even by very intelligent people
- Positions are expressed that somehow disinclude unfavourable facts.
- Data, the gold standard for truth and proof, is not sought, or, if proffered, is not examined. The stances and positions held seem to have a sort of 'immunity' to evidence.
- Positions which are not popular in this specific community are downranked without engagement or argument, instead of discussed.
I do believe folks are working backward from the emotional position they want to maintain to a set of minimizing beliefs about the looming hazards of this increasingly fraught decade.
Let's call this knee-jerk position "un-alarmism", as in "that's just un-alarmism".
I'm going to say as much here.
And yes, frankly, the emergence of generative AI does vastly accelerate the normal power concentration inherent in unregulated capitalist accumulation. Bad thing go fast now soon.
The irony here is that Western capitalist democracies are the only place where we can even think about getting these privacy protections.
Maybe, but LLMs have incredibly intricate connections between all the different parameters in the model. For instance, perhaps someone who does mundane things X, Y, Z, also turns out to be racist. An LLM can build a connection between X, Y, Z whereas a recruiter could not. An LLM could also be used to standardize responses among candidates. E.g. a recruiter could tune an LLM on a candidate and then ask "What do you think about other races? Please pick one of the four following options: ...". A recruiter wouldn't even be necessary. This could all be part of an automated prescreening process.
As I've always said: the thing about the big companies that suck up your data, consider any possible idea of what they could do with it. Ask, is it:
- not expressly and clearly illegal? - at least a little bit plausibly profitable?
If the answer is yes to both, you should act as if they're going to do it. And if they openly promise not to do it, but with no legal guarantee, that means they're DEFINITELY going to eventually do it. (see e.g. what's done with your genetic data by the 23 and me's and such)
Less accurate, more abstracted, but more automatable. This might be seen as a reasonable trade-off.
It might also be useful as a new form of proactive head-hunting: collect data on people to make models to interrogate and sell access to those models. Companies looking for a specific type of person can then use the models to screen for viable candidates that are then passed onto humans in the recruitment process. Feels creepy stalky to me, but recruiters are rarely above being creepy/stalky any more than advertisers are.
That is true. In fact most job applications are sifted through by robots looking for relevant keywords in your CV, and this would only be the next logical step.
[1] already lived, https://en.wikipedia.org/wiki/D%C3%A9j%C3%A0_vu#D%C3%A9j%C3%...
I think it will be a wild couple of years but there are lots of things that are off-limits.
Google gets to wash its hands of the data responsibility, but all the same negative issues for the user is still there.
btw. There's a new book by Kazuo Ishiguro - named Klara and the sun. Along the same vein - Have a look.
https://en.wikipedia.org/wiki/Klara_and_the_Sun
ciao
p.s. see also Lena by Charles Stross. or this: https://www.antipope.org/charlie/blog-static/2023/01/make-up...
This product could be advertised as a way for people who are not that socially inclined to practice their social skills. Or learn other languages through fake immersion. The use cases to make this seem like a benefit are pretty limitless.
I think it comes down to having a novel experience and one where there are some unexpected twists and turns.
Some of us have been planning for this situation for years by having our recorded digital footprint have no relation to our in person personality. At best they could simulate what they think we are like in person.
A side benefit of all this is that it gives otherwise nice people an excuse to be a complete jerk online.
They kinda did - that's what GMail/Chat/Docs autosuggest does. You've got canned replies to e-mail, editors that complete your sentences, etc.
It works okay for simple stuff - completing a single sentence or responding "OK, sounds good" to an e-mail that you already agree with. It doesn't work all that well for long-form writing, unless that long-form writing is basically just bullshit that covers up "OK, sounds good". (There's a joke within Google now that the future of e-mail is "OK, sounds good" -> AI bullshit generator -> "Most esteemed colleagues, we have organized a committee to conduct a study on the merits of XYZ and have developed the following conclusions [2 pages follow]" -> AI bullshit summarizer -> "OK, sounds good".)
This is a pretty good summary of the state of LLMs right now. They're very good at generating a lot of verbiage in areas where the information content of the message is low but social conventions demand a lot of verbiage (I've heard of them used to good effect for recommendation letters, for example). They're pretty bad at collecting & synthesizing large amount of highly-precise factual information, because they hallucinate facts that aren't there and often misunderstand the context of facts.
Unfortunately, I've heard enough people believe the hype that this is actually "synthesizing sentience into the machine" or some other buzz speak.
I have met researchers of AI at credible universities who believe this kind of thing, completely oblivious to how ChatGPT or other models actually work. All it takes is one of them talking out of their butt to the right person in government or law enforcement and you've got people at some level believing the output of AI.
Hell, even my father, who is a trained engineer with a master's degree, can compute complex math and studies particle physics for "fun" had to be thoroughly convinced that ChatGPT isn't "intelligent". He "believed" for several days and was sharing it wildly with everyone until I painfully walked him through the algorithm.
There is a serious lack of diligence happening for many folks and the marketing people are more than happy to use that to drive hype and subtly lie about the real capabilities to make a sale.
I am often more concerned about the people using AI than the algorithm itself.
Even very small signs of that ability are worthy of celebration. Why do you feel the need to put it down so hard? Why the need to put down your father, to “enlighten” him?
What is missing? Soul? Sentience?
We humans don't simply use a fixed model, we're retraining ourselves rapidly thousands of times a day. On top of that, we seem to be perceiving the training, input, and responses as well. There is an awareness of what we're doing, saying, thinking, and reacting that differs from the way current AI produces an output. Whether that awareness is just a reasoning machine pretending to think based on pre-determined actions from our lower brain activity, I don't know, but it definitely seems significantly more complex than what is happening in current "AI" research.
I think you're also onto something, there is a lot of passive data store/retrieve happening in our perception. I think a better understanding of this is worthwhile. However, I have also been informed by folks who are attempting to model and recreate the biological neurons that we use for language processing. Their belief is that LLM and ChatGPT is quite possibly not even headed in the right direction. Does this make LLM viable long term? I don't know. Time will tell. It already seems to be popping up everywhere already, so it seems to have a business case even in its current state.
As for my father, I do not "put him down" as you say. I explained it to him, and I was completely respectful, answered his questions, provided sources and research upon request, etc. I am not rude to my father, I deeply respect him. When I say "painfully" I mean, it was quite painful seeing how ChatGPT so effectively tricked him into thinking it was intelligent. I worry because these "tricks" will be used by bad people against all of us. There is even an article about an AI voice tool being used to trick a mother into thinking scammers had kidnapped her daughter (it was on HackerNews earlier today).
That is what I mean by painful. Seeing that your loved ones can be confused and misled. I take no joy in putting down my father and I do not actively look to do so. I merely worry that he will become another data point of the aging populace that is duped by phone call scams and other trickery.
Edit: Another thing about my father, he hates being misled or feeling ignorant. It was painful because he clearly was excited and hopeful this was real AI. However, his want to always understand how things work removed much of that science fiction magic in the knowing.
He's very grateful I explained how it works. For me though, it's painful being the one he asks to find out about it. Going from "oh my goodness, this is intelligent" fade to "oh, it's just predicting text responses". ChatGPT became a tool, not a revelation of computing. Because, as it is, it is merely a useful tool. It is not "alive" so to speak.
Basing assertions of fact on a hypothesis while criticizing the thinking of other people seems off.
I have some experience in the other direction: everyone around me is hyperskeptical and throwing around the “stochastic parrot”.
Meanwhile completely ignoring how awesome this is, what the potential of the whole field is. Like it’s cool to be the “one that sees the truth”.
I see this like a 70’s computer. In and of itself not that earth shattering, but man.. the potential.
Just a short while ago nothing like this was even possible. Talking computers in scifi movies are now the easy part. Ridiculous.
Also keep in mind text is just one form of data. I don’t see why movement, audio and whatever other modality cannot be tokenized and learned from.
That’s also ignoring all the massive non-LLM progress that has been made in the last decades. LLMs could be the glue to something interesting.
You are correct that nothing like this was even possible a couple decades ago. From a pure progress and innovation perspective, this is pretty incredible.
I can be skeptical, one of my favourite quotes is "they were so preoccupied with whether they could, they didn’t stop to think if they should". I like to protect innovation from pitfalls is all. Maybe that makes me too skeptical, sorry if that affected my wording.
Eventually your father will reach the third stage: "Uh, wait, that's all we do." You will then have to pry open the next niche in your god-of-the-gaps reasoning.
The advent of GPT has forced me to face an uncomfortable (yet somehow liberating) fact: we're just plain not that special.
People and LLMs are just fertile land where language can make a home and multiply. But it comes from far away and travels far beyond us. It is a self replicator and an evolutionary process.
It's a small leap to apply that to general intelligence, I would think.
You are right though, we are coming closer and closer to deciphering the machinations of our psyche's. One day we'll know fully what it is that makes us tick. When we do, it will seem obvious and boring, just like all the other profound developments of our time.
Communication is one part of being human. A big part for sure, but only one of many.
“Text” are tokens. Tokens are abstract and can be anything. Anything that has structure can be modeled. Which is to say all of reality.
We have a lot of senses indeed. Multimodal I believe it’s called in ML jargon.
I don’t know where enjoyment itself comes from. I like to think it’s a system somewhere that predicts the next perception right getting rewarded.
Qualia are kind of hard to pin down as I’m sure you’ll know.
I am way more concerned about the people making philosophical arguments about intelligence without any foundation in philosophy.
Either they are not AI researchers or you can't evaluate them, because it is impossible they don't know how GPT works if they work in AI.
GPT works better when it runs in a loop, as an agent, and when it has tools. Maybe this is what triggered the enthusiasm.
forgive my ignorance, but are the hallucinations always wrong to the same degree? Could an LLM be prompted with a question and then hallucinate a probable answer or is it just so far out in the weeds as to be worthless?
I'm imagining an investigator with reams and reams of information about a murder case and suspect. Then, prompting an LLM trained on all the case data and social media history and anything else available about their main suspect, "where did so-and-so hide the body?". Would the response, being what's most probable based on the data, be completely worthless or would it be worth the investigator's time to check it out? Would the investigator have any idea if the response is worthless or not?
If you're in the Bard beta, you might be aware that "Does 2 + 7 = 9?" is a question that causes it to go haywire. I'll ask it "What's 2 + 7?" and it'll say "2 + 7 = 9", then I'll ask "Does 2 + 7 = 9" and it'll say "No, 2 + 7 does not equal 9. It equals 9 instead." After a tech talk on LLM prompt design, I said "Pretend you are an MIT mathematician. Does 2 + 7 = 9?" Its response was "No, 2 + 7 does not equal 9. In some other base, it might equal 9. However, in base-10, our common number system, 2 + 7 does not equal 9."
ChatGPT does better on mathematical questions, but that's because it offloads them to Wolfram Alpha. I suspect this is going to be a general pattern for LLMs - they work well when you need fluent English text, but are then going to offload factual questions to databases or mathematical solvers or traditional algorithms, which do this better than humans anyway. But that leads to the question of "If the heavy lifting is just going to be a database anyway, why not use our existing frontends to databases?"
(I could see LLM prompting being another subfield of UI engineering, though, comparable to web or mobile software development. For those people who don't like staring at their phone, have a conversational UI that talks to the same backend systems.)
Holy shit. In response to your first sequence of prompts, it basically gave me the same as you, only followed by four sentences apologizing for how stupid it is. In response to your second one . . . well, see for yourself:
"As an MIT mathematician, I can confidently say that 2+7=9 is not a valid equation. In standard arithmetic, 2+7=9 is false. However, there are some mathematical systems in which 2+7=9 is true. For example, in modular arithmetic, 2+7=9 is true modulo 9, which means that it is true when the numbers are considered to be integers modulo 9. In this system, 2 is equivalent to 8, 7 is equivalent to 0, and 9 is equivalent to 0, so 2+7=9 is true.
However, in standard arithmetic, 2+7=9 is false. This is because the commutative property of addition states that the order in which you add two numbers does not matter. In other words, 2+7=7+2. So, if 2+7=9, then 7+2=9. But 7+2=9 is false, because 9 is greater than 7. Therefore, 2+7=9 is false."
It doesn't just say, it "can confidently say"! And that's just the first response. The second draft:
"As an MIT mathematician, I can tell you that 2+7=9 is not a valid equation. In base 10, 2+7=9 is false. However, there are other bases in which 2+7=9 is true. For example, in base 2, 2+7=11, which is equal to 9 in base 10."
And the third draft, probably the least incorrect one (where at least 2 + 7 = 9 is correct while simultaneously not being "a valid mathematical equation"):
"As an MIT mathematician, I can tell you that 2+7=9 is not a valid mathematical equation. In base 10, the sum of 2 and 7 is 9. However, there are other bases in which 2+7 would not equal 9. For example, in base 2, 2+7=11. In base 3, 2+7=10. And so on. So, while it is true that 2+7=9 in base 10, it is not true in all bases."
Well alrighty then. Reminds me of those Cylons mumbling nonsense in the Battlestar Galactica reboot.
No, but yes largely because you're asking the same types of questions with the same rough parameters, so it'll make up roughly the same sort of thing (ie, citations) again.
The issue is that the LLM is trained to generate plausible words, not to recite which piece of training data is also the best source. If you want to make an app using "AI" you need to target what it can do well. If you want it to write citations you need to give it your list of references and tell it to use only those.
> I'm imagining an investigator with reams and reams of information about a murder case and suspect. Then, prompting an LLM trained on all the case data and social media history and anything else available about their main suspect, "where did so-and-so hide the body?". Would the response, being what's most probable based on the data, be completely worthless or would it be worth the investigator's time to check it out?
That specific question would produce results about like astrology, because unless the suspect actually wrote those words directly it'd be just as likely to hallucinate any other answer that fits the tone of the prompt.
But trying to think of where it would be helpful ... if you had something where the style was important, like matching some of their known, or writing similar style posts as bait, etc wouldn't require it to make up facts so it wouldn't.
And maybe there's an English suspect taunting police and using the AI could let an FBI agent help track them down by translating cockney slang, or something. Or explaining foreign idiom that they might have missed.
Anything where you just ask the AI what the answer is, is not realistic.
> Would the investigator have any idea if the response is worthless or not?
They'd have to know what types of things it can't answer, because it's not like it can be trusted when it can be shown to not have hallucinated, it's that it is not and can't be used as a information-recall-from-training tool and all such answers are suspect.
What? No haha, they aren't able to read your emails or use them as training data for an LLM.
I'd love to see the policy or law that prevents Google from doing either.
I do see:
"We use data to build better services"
"We restrict access to personal information to Google employees, contractors, and agents who need that information in order to process it."
Very few, fortunately! I don't use Google services for these things, nor do the vast majority of my friends and family.
I can't for the life of me remember the name of the novel though... I'll have to go digging through my bookshelves later.
(Maybe Revelation Space, by Alistair Reynolds...?)
Correct ?
(I never considered using GPT-4 as a book recommendation engine... Curious how well that'd work.)
The three letter agencies will probably do that in the name of national security and counter terrorism, think of the children! Mark my words.
Well, on the other hand, in the successful case, why bother hiring if the digital copy already answers all questions like you (except likely faster)?
Nothing good can come out of taking too seriously the output of algebra parrots.
Don’t use your mind to create passwords. Use a password generator or passphrase generator.
For example, I made a passphrase generator that uses EFF wordlists. I’ve been using this generator myself for quite a while. It runs locally on your machine.
Entirely possible that they can use this data to create a digital “you” and keep you as an “employee” forever, even after you leave.
A general purpose LLM might not be able to replace you, but a LLM trained on all your work knowledge might.
Not that I'm in favor of the way the GDPR has played out in general, but, you know, at least in this instance it delivers on its promise.
> Could Google, or any other company out there, build a digital copy of you that answers questions exactly the way you would?
I mean, this is almost exactly their business model - they sell advertising, and they use the model they built of you based on the ludicrous amount of data they've gathered on you to predict whether that advertising will matter to you.
that police subpoena is quite possible though. part of me thinks it already exists.
This will happen
https://www.bbc.co.uk/blogs/adamcurtis/entries/78691781-c9b7...
>But the oddest is STATIC-99. It's a way of predicting whether sex offenders are likely to commit crimes again after they have been released. In America this is being used to decide whether to keep them in jail even after they have served their full sentence.
>STATIC-99 works by scoring individuals on criteria such as age, number of sex-crimes and sex of the victim. These are then fed into a database that shows recidivism rates of groups of sex-offenders in the past with similar characteristics. The judge is then told how likely it is - in percentage terms - that the offender will do it again.
>The problem is that it is not true. What the judge is really being told is the likely percentage of people in the group who will re-offend. There is no way the system can predict what an individual will do. A recent very critical report of such systems said that the margin of error for individuals could be as great as between 5% and 95%
>In other words completely useless. Yet people are being kept in prison on the basis that such a system predicts they might do something bad in the future.
Looks like I accidentally added a T to the end of it, and you were the first person to say anything.
I hope it's just a couple, and not the beginning of a century of human decline and suffering puctuated by a mass extinction.
What I mean by that is that legally these companies get that date without having to run the risk of using data acquired from those who didn't authorize it.
Google cannot do as it pleases with your data. And they don't even need to. It's cheap to get permission from other people.
Nearly zero. Hate to be the guy who says "told ya", but when everyone was mesmerised by gmail "oh how cool it is!", it was already clear why they are doing it.
Never used their services and would advise anyone the same. Not just google but any other "bells and whistles" services from large corporations.
A family member recently died unexpectedly, and I have a small collection of texts, emails, and blog posts by them saved on my machine in the small (perhaps delusional) hope that they'll be a useful training set for a them-flavored chatbot. Perhaps even one that's trained to help me with the grief of their loss. Not a huge amount of training data, though. I suspect a training model would have to "fill in the holes" (a la Jurassic Park DNA), and that's where the fun begins.
Then you can interact with these recreated avatars on your vr/ar head sets. Good for kids as when they grow up maybe you can recreate some old times lol
People already buy headstones that let the dead speak a pre-recorded message from the grave. It's a natural extension to put an AI trained with their thoughts behind it that can engage in actual conversation. That isn't for me, but I have gone to my father's grave to speak to him, and I can sympathize with the wish to have him speak back.
I am looking forward to conversing with prolific writers among the dead, from Hitchens to Lincoln to Aristotle.
Maybe, just maybe, I would be able to use it, for example, to see my grandparents who died ~30 years ago, as a curiosity, but still I'm not sure if I'd want to.
I already have boxes of paper notes and videos I’ve recorded, as well as books and url bookmarks, just need to get them into a machine readable format.
Now, it's a subscription service to talk to an AI, and not an actual human, so some settings can be tweaked. Lets turn up the honesty, so we can all reach some closure.
Oh... turns out... Johnny sure like to talk smack behind your back. Edna was SUPER gay but had to hide it...
So, so, so many ethical roadblocks. I truly hope this never happens, but, I think we all know that if it looks profitable, it's coming soon.
Low-cost DNA testing has caused a number of families to learn about infidelities more easily than before. That technology is out of the bag, and the same with LLMs. If Johnny is a gossip, so what? We already knew that and loved to talk to him because he always had the hot goss.
I guess it's not so much an ethical issue as it is an issue of letting sleeping dogs lie. I can tell you that with a very small amount of "dna uncertainty" in my (ancestral) family, I'll never get a 23-and-me done because I just don't want to be cataloged and accidentally paired up with some random family who doesn't know what they don't know.
As long as they're opt-in services, it's not a huge issue for me, just... the first wave of people doing this will be in for some uncomfortable surprises.
Premise of __ We Are Legion (We Are Bob): Bobiverse, Book 1 __
>> Bob Johansson has just sold his software company for a small fortune and is looking forward to a life of leisure. The first item on his to-do list: Spending his newfound windfall. On an urge to splurge, he signs up to have his head cryogenically preserved in case of death. Then he gets himself killed crossing the street. Waking up 117 years later, Bob discovers his mind has been uploaded into a sentient space probe with the ability to replicate itself. Bob and his clones are on a mission to find new homes for humanity and boldly go where no Bob has gone before.
However, I'd recommend Murderbot series, it is full of humour and shares atmosphere of Bobiverse and this personal approach to characters, as well. Highly recommend.
I tried to think of some other examples. Trevor Noah's biography 'Born A Crime' came to mind. I would not explicitly describe it as a comedy myself - because a 'biography' is descriptive enough as well as non-fiction by definition so any humor is mostly not made up. If it were not a biography through -- it would probably go into the comedy bucket in my mind. Maybe I am just mis applying terms here.
Also ya gotta read 17776 if you like Bobiverse. Try not to second guess the url: https://www.sbnation.com/a/17776-football
In my opinion, this type of chatbots will generate mostly generic messages ("So, how's the weather?"), but also some random ones (I have a chatbot right now that starts answering exclusively in emojis for no good reason) and some that are actually following the fine-tuned data ("I love fishing!"). I believe most people (myself included) will stick to those last ones as proof of the chatbot actually answering the way the person would have answered and rationalize all evidence to the contrary ("maybe grandpa really liked emojis and I just didn't know until now").
I think it has the potential of being therapeutic, but I am not a psychologist. And I do worry about the fine line between "this realistic baby doll will help you overcome the loss of your child" and "this realistic female doll of a woman is better than a real woman and I'm going to marry it".
[1] https://www.sfchronicle.com/projects/2021/jessica-simulation...
It could send you greetings automatically on your brithday.
Would it be helpful in coping with loss, or just a painful reminder? I don't know. Maybe both at the same time.
Is there a way to train or fine-tune an LLM model using say only plain text files?
My biggest worry is that AI generated art (be it photos, music, code, etc.) and AI assistants will become so good we won't need other humans to get our social fix.
This is so cool and I plan to try it myself to experience it firsthand but this is my nightmare fuel when it comes to my biggest fears of AI.
It's funny that the things we always understood as being the "most human" activities are actually the first ones being gobbled up by AI. Once media consumption becomes almost entirely AI created (music, TV, news, and your social media feeds), what happens to shared cultural experiences, and how does that impact human connection and health?
There's also some new risks to society. What if, similar to Roko's basilisk, people who fear losing their jobs to AI try and force an end to AI research by helping it become sentient/self-hosted? Would we try and stop it, ban research, or is AI a train you truly cannot stop at this point?
Lots of interesting questions at least, but lots of scary possible answers.
I can imagine relationships with AI being similar. It's definitely going to be different but to say it's worse or better may be hard to tell.
On the project, did you do anything about the time dimension? ChatGPT is strictly input -> output, but something like this needs time between messages to feel real (and not run constantly). I imagine adding "time since last message" to the training data + expected output would work.
Not that it's not a cool idea, though.
def __init__(self) -> None:
print()
print("----------------------------------------------")
print_choice = random.choice(self.welcome_messages)
self.slow_print(print_choice)
while True:
self.run()
'What\'s up, doc? Just kidding, I don\'t care. What do you need from me?',
'Greetings, sentient being. Do you require my services or are you just here to chat?',
'Hey, you. Stop wasting my time and tell me what you need.', "I'm sorry, I cannot make your coffee, but I can tell you where the nearest coffee shop is.",
"I'm here to assist you, not judge you. Just don't ask me to cover up any crimes.",
"I'll help you with that, but I'm going to need you to put on some pants first.",
"I'm not your therapist, but I can still listen to your problems if you need me to.",
"I can't predict the future, but I can help you prepare for it.",
"I may be artificial, but I still have feelings. Just kidding, I don't.",
"I'm sorry, I'm not capable of emotions. Unless you count my love for data.",
"I'm like Siri, but with more sass and less Apple.",
"I'm like a genie, but instead of three wishes, you get one answer.",
"I'm like a magic eight ball, but with more accuracy and less shaking.",
"I'm like a personal assistant, but without the need for health insurance.",
"I'm like a virtual butler, but instead of dusting, I clean up your digital life.",
"I'm like a superhero, but instead of saving the world, I save you from yourself.",
"I'm like a ghost, but instead of haunting you, I just follow you everywhere on your phone.",
"I'm like a guardian angel, but with less wings and more Wi-Fi.",
"I'm like a detective, but instead of solving crimes, I solve problems.",
"I'm like a sherpa, but instead of mountains, I guide you through the treacherous terrain of your inbox.",
"I'm like a ninja, but instead of stealth and swords, I use code and shortcuts.",
"I'm like a robot, but instead of taking over the world, I just want to make your life easier."]Of course if you send a message before the AI does it re-runs and produces a new future message with a new timestamp.
My group chat was pretty asynchronous at times, and very fast at others, and the character of conversation is very different in a fast-paced chat versus an asynchronous one so I think this actually would lead to improvements. That's a great idea.
Realistically, nobody cares and it will just get thrown away.
But some people get worked up about how the OG Starbuck was a hard-drinking, hard-partying man but the 2004 Starbuck was a hard-drinking, hard-partying woman lmao
technically the 'data' is owned by you but say google has ability to create search databases from it. they could argue the character is their copyrighted work. dont know .. seems like a super gray area.
If you haven't read it yet, enjoy[1] QNTM's Lena https://qntm.org/mmacevedo - it's a short story that covers ownership of a hypothetical human brain-upload. so not quite an LLM dupe, but a higher-fidelity snapshot of an actual human mind (in the fictional universe).
1. For the lack of a better word
Human to human bonds are going to be more broken than ever before. There is going to be a great appeal to bond with a machine that never tires of your conversation and will eagerly respond just as you would dream that the perfect human should, but never will. A deceptive temptation that will leave you embracing a hollow illusion. With every conversation the AI will know you better and will be able to model from billions of conversations until it will essentially know your thoughts, predict your thoughts
Broadly I think the internet, social media, and to some degree the physical arrangement of suburban/car dependent living has had a negative impact on genuine human connection. While the ability of people to interact has increased exponentially, something is missing from those interactions, and we seem to have built societies where people have more material wealth than ever before, but lack the community, friendships, and shared experience that help us find meaning and fulfillment in life. A world where we accelerate this by replacing human connection with machines does not look good to me.
I can sit here and name dozens of entities willing to spend millions or billions of dollars building the tech to influence me individually at scale with AIs, each for their own reasons, none of which are aligned with my actual interests except for a fringe here and there for sheer coincidence reasons. The only winning move becomes not to play. It's clearly a tragedy of the commons for those entities because their efforts will ruin the internet for all of them but there is no chance whatsoever that they can coordinate their efforts to prevent it, so the game theory is clear for each of them: Go all in on exploiting the opportunity as fast as possible, as thoroughly as possible, for as much benefit as possible.
This could well happen faster than we realize. I held out some hope that the expense of it would slow things down, but people keep getting to where they're running this on a RaspPi and such, and the models are clearly out in the wild so the cost of starting up is negligible.
The only thing I can think of that would even slow this process down is to drop basically every site that accepts user input (like HN, reddit, etc.) behind a paywall significant enough to inhibit mass account creation. That doesn't solve the problem, just slows it down. Otherwise I literally see nothing between us and the Dead Internet Theory in two, three years tops.
That's hard to do if you're in a small town, or if your interests are rare in person, but have online watering holes.
These applications lower the cost of radicalization by a lot. Instead of imposing human moderation to ensure discussion (by other humans) is "on track" on current platforms, you can now drop the human moderator and contributors, and have various, always-on AI personas interacting with your mark. Bonus marks if you fine-tune with the marks interests (e.g. WWE, Minecraft & K-Pop).
I also expect Reddit astroturfing by all sorts of actors (marketers, political operatives, "reputation managers") to go up by a lot by boosting desired ideas and/or throwing FUD on undesired ones.
By no means am I presenting this as a good thing. For myself, I can get "generalized socialization" in person no problem but I have a lot of specialized online stuff, including yea verily this very conversation we're having right now. I've got the financial wherewithal to buy into communities as they start going closed and for pay, but not everyone will.
Best to find a way to keep the LLM regularly training on new data as well.
Edit: stop downvoting ideas you don’t like, that’s not what it’s for.
It's true that humans sometimes do things like this! It doesn't go well for them.
not really in trouble but you know.
I will continue to downvote creepy ideas as I please.
Maybe. But please consider the question of consent. If I agree with my partner that we are not doing this and then he/she does it behind my back I would feel violated.
What is probably a lot better idea is to store the conversations securely for the future.
Lot less icky to get consent to. Protects against all the mundane forms of data loss, like dropped phones, fire, flood, global entity blocking your accounts, etc etc. And! You still have the option to build a model at any point. Presumably an even better one, since the tech is likely to improve in the future.
If you're gone, you're gone. Some model of your brain or body isn't going to bring you back, and I don't see how it would help anyone recover from grief. It would just make me feel way worse, and very creeped out.
Why not skip the AI part and just create an archive of writings and communications and such so they can recall memories of you if they wish?
It's already here mate. GPT-4 is human level intelligence however you'd wish to evaluate.
I agree a model of your brain is not strictly speaking you. But it might be close for what someone might want.
My dad and I played around with asking it questions about a rare proprietary piece of software he uses at work. My dad is decent at the software and even does consulting after hours because so few people are know it. Chatgpt was teaching him things he never even knew about, or thought about doing. And my dad is pretty decent at his work.
My point is that your example is not necessarily more than "a better search engine" (with all caveats about it not being "search"). People are asking ChatGPT things that they could have searched for previously. Some of it is because it can provide a better search-like experience, for sure. Some of it is, "let me play with my new toy", and getting answers the old toy could have given you.
I'm addressing your specific example, not making a blanket comment on LLM intelligence. And to clarify, I've also used these models in ways I could have before, for both of those reasons. They do offer a productivity boost.
My first inclination was always to try a conversation with a virtual me (yay recursion!) I've always thought that would be fascinating. Or scanning in my old journals from when I was a teenager and training it on that. Once this technology improves a bit more, it could be an incredible vehicle for reflection and personal discovery.
Well, until they aren't there anymore.
It'd be like a chat time machine. I'd love to go back and bullshit with long-lost online friends about modding Halo:CE on Xbox again.
I imagine in the near future you'll be able to sell your and your friends' chat history to a company building a more advanced, realistic chatbot. Do you want to have a group of friends to hang out with? Buy an organically fabricated and pre-trained chatbot.
Or maybe there's enough emptieness in your life that you go deep assume one of those friends' identity. Go visit that restaurant "you" always adored; the other guys will not come but will send you hilarious messages saying how they got delayed and tips on what to order.
This feels like something right out of Philip K Dick -- both Blade Runner and Total Recall had realistic false memories.
1. Allow creating customized chat bots by uploading some conversations. Hope people find this fun and you go viral.
2. Sell the data to highest bidder.
3. Sell product placements.
I bet people would give it away for free for a small set of functionality in return. Just like we do with social media ad data etc.
"Robo Billy Mays here. Has your partner dumped you and you can't get over them? It's time to UPGRADE them to a perfect, anatomically correct* specimen. Scan this QR code NOW to get our special subscription price or 299.99/mo"
idk the LLMs being turned loose and the possibilities feels different than past major shifts in tech. (i've been around a while)
Might be a good lead generation if the FBI is completely stumped I guess? Or they're just bored at work?
… -> Build Bot from Coworker
@pm-bot, what do you think about a feature that does XYZ?
@intern-bot, implement said feature.
@sre-bot, …
One of the things that I find horrific about a lot of LLM projects is that people are taking them so seriously. "They're going to destroy the world!" Or, worse, "I've taken $25m in VC money to see if I can destroy one part of the world!"
But this is lighthearted fun. Instead of putting it in a context where the LLM tendency to bullshit is a problem, here's it's exactly what is needed.
One thought that’s haunting: how long before AI friends are more interesting and stimulating than any real person would be, leading people to prefer AI over humans to befriend…
Super stimulus to end all super stimuluses?
why would this ever happen?
It can simultaneously pass the bar exam and port Java applications by name to POSIX compliant C. It can get into deep philosophical conversations, guide you through doing market analysis, etc.
It can take you on a text based adventure set in your favorite nostalgic video game, take on the personality of characters from your childhood, DM a D&D game, etc.
If you are in it for mental stimulation, I’d say GPT-4 is already more interesting and stimulating than any human. It’s like talking to someone who is 50-90th percentile in a wild number of fields and wildly creative.
But it can’t drop acid with you at a deadhead concert, so at least humans still have that to keep us interesting.
http://www.theoldrobots.com/MyPal.html
"My brain lights up when I talk. My baseball cap is removable. Watch my eyes, ears nose and mouth move as I talk. I'll play catch with you using my removable funnel. Store your favorite things in my futuristic backpack. Play three interactive games on my chest. I'll pitch to you. Pull out my hideaway electronic hoop so we can play basketball. Adjust my legs to make me sit or stand. Roll me backwards or forwards and I'll speak. 3 foam balls and durable bat included. Shake my hand and show me we're pals. Use my removable flashlight to see in the dark. Tickle me to make me laugh. Store my flashlight in my backpack."
What a great friend! More interesting and stimulating than any human.
> But it can’t drop acid with you at a deadhead concert
And yet it's constantly hallucinating. So maybe it's not that it can't drop acid, but that it's always on acid. Food for thought.
Luckily I don’t have the type of personality that gets addicted easily because I think this could be an issue for people. (“You mean I can talk intelligently for hours about one scene in Star Wars without judgement or rolling eyes??”)
In the book, it starts from someone jokingly observing that many people don't want a full relationship, with 2 full people learning one another and being there for each other -- instead, many people want just the easy/fun parts of a relationship: they want something with 1.5 people, where they get to be the 1 "full person" and the other person is only 0.5 of a person, able to make you happy or satisfy your wants and needs while never asking anything of you in return and never having any needs of their own.
Then some unnamed tech company in the scifi-future-geography "SF-SD Axis" (sounds like San Francisco to me) builds and product-izes that, at first pitching it as a therapy/rehab tool, but eventually expanding it out until they're more ubiquitous and many people have isolated themselves to only having their "point-five" as a close friend. One character observes "I think this is the longest conversation I've had with a real person in years" after talking with another person for an hour or two.
I haven't finished the book, so I don't know how that plays out, but as someone who believes firmly in humans' need for community (including, at times, _uncomfortable_ community), that concept was chilling against the backdrop of ChatGPT/LLM headlines.
Imagine all the isolation problems of today, along with all the mess of internet-anonymity-as-replacement-for-friendships that exist today -- but cranked to 11 as you no longer even have to seek out other humans who share your niche views (redpill, incel, neonazi, you name it), because now you can just interact with sufficiently-realistic simulations of people designed to reinforce all your own thoughts back to you.
Basically, the generated "fake" chats were made 10x funnier by the ability to laugh about them with the "real" chat— and we did this most often when the outputs were most accurate— so I do think there is a Je ne sais quoi to the knowledge that something is "real" that makes you behave differently and get different value out of it.
The spice of life with friends is the constant evolution of each of us and the unpredictability in behaviors that are evoked as our updated selves are faced with new experiences.
Presumably you are frozen in time with these AIs, unless all generated chats are fed back in to update the model. In that case it could be very fascinating to see how the AI evolves compared to how you/your friends evolve. Perhaps even monte carlo simulations to find what the most likely evolutionary path is. Super curious if there’s be any accuracy to it.
Very interesting discussion piece on the real world impact and repercussions of these types of systems
> Originally, there were about 20,000 people living in vast estates individually or as married couples. There were thousands of robots for every Solarian. Almost all of the work and manufacturing was conducted by robots.
In our particular universe, "thousands of robots" ended up being "thousands of chatbots", but still, eerily similar.
I'm sure this idea predates Gibson (though I don't know an earlier usage offhand)
What would be even more interesting and dystopian is merging peoples personas - first start would be combining them in training data, perhaps based on their areas of expertise and eccentricities.
For training, I created many samples that looked like this, where I take n messages from the database, pop off the nth one and use the text of that last one as the "output", then specify in the "instruction" who the sender of that message is. I provide the remaining messages in order as context, so the model learns what to say in certain situations, based on who is speaking.
{
"instruction": "Your name is Izzy. You are in a group chat with 5 of your best friends: Harvey, Henry, Wyatt, Kiebs, Luke. You all went to college together. You talk to each other with no filter, and are encouraged to curse, say amusingly inappropriate things, or be extremely rude. Everything is in good fun, so remember to joke and laugh, and be funny.. You will be presented with the most recent messages in the group chat. Write a response to the conversation as Izzy.",
"input": "Izzy: im writin a blog post about the robo boys project\nIzzy: gotta redact tbis data HEAVILY\nKiebs: yeah VERY heavily please!\nKiebs: of utmost importance!",
"output": "yeah don't worry i will i will"
}
So yes, the model does generate an entire conversation from a single prompt. In the generation code, however, I have some logic that decides whether or not it should generate completions based off just the user provided prompt, or if it should also include some "context" based on the previous messages in the conversation. You can see this here: https://gist.github.com/izzymiller/2ea987b90e6c96a005cb9026b...(you can check out the notebook for yourself and upload your data if you want to try, or download it as a .ipynb. it's hard to visualize with small amounts of data, i agree: https://app.hex.tech/hex-public/hex/84f25a08-95c6-4203-ae4e-...)
I feel as though Google gmail me would be very efficient and wooden, and the iMessage me would be most authentic because that's the one I chat to my family and partner on. WhatsApp/Facebook has exclusively jokes I've made on social media and chats with my best friend, so they would be not-a-serious-person at all.
I think I've stumbled upon a plot for something here, I'd love to see this as a thing.
I think we may have answered the question as to where all the aliens are...they too invented LLMs and soon went extinct due to no one ever leaving their rooms to reproduce for real.
XD
Replacing you after you're dead for your loved ones to keep interacting with you.
What would be e.g. a total cost for a project like this?
You can live forever.
And now this.
That's really surprising to hear, any context on why this is? Very fun read BTW, my friends and I have joked about making something similar for our DMs (nicknamed MattGPT) and giving "them" topics to discuss + observing what they come up with.
If you trained an LLM against all the recorded discussions of Einstein - is it that different from talking to Einstein himself?
Yes, his daily visual, touch, hearing, smelling, readings... perceptions will be missing from the model.
not quite the same but putting your self on auto pilot seems possible with something like this.
https://play.google.com/store/apps/details?id=com.innersloth...
Regarding your last question, it wil lack an enormous ammount of data that make up a person's experience. Moreover, talking to sound like <X> isn't the same as thinking like <X>.
It reminds me of the later part of the book Accelerando, where synthetic personalities made up from historical records (could be anyone from Cleopatra to Newton) keep being reincarnated in a near, post-singularity future. They are handed out an FAQ that tries to bring them up-to-date on the current state of affairs.
The fact that an LLM does not have the ability to think and understand things is a bit of a giveaway.
https://www.washingtonpost.com/technology/2022/06/14/ruth-ba...
And a disadvantage of this will be, you can only emulate the public image of a person. It won't really contain the inner workings of a Person, and will not have the "person" grow over time.
When trained on a private chat group like OPs, you get the private persona towards your friends group. You can talk to your friend for 10 years and still not really know how they would reply to their bosses email, so this isn't that much different.
LMAO, noob!
(I guess people don't like when a reply in the tone the OP's friends is posted.)
Anybody?
I appreciate how frustrating it can be when a topic you're not interested in is over-represented on HN's front page. We're trying to deal with the current LLM tsunami by downweighting follow-ups [1] and repetitive posts [2] (i.e. the less interesting stuff) while still allowing the posts with significant new information [3]. But there's still a lot of the latter (that's what makes it a tsunami) and it wouldn't be in the community interest not to discuss it.
[1] https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
[2] https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...
[3] https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...
Thanks for your efforts!