ChatGPT went berserk
garymarcus.substack.com
garymarcus.substack.com
It's documented pretty well - https://platform.openai.com/docs/guides/text-generation/freq...
OpenAI API basically has 4 parameters that primarily influence the generations - temperature, top_p, frequency_penalty, presence_penalty (https://platform.openai.com/docs/api-reference/chat/create)
UPD: I think I'm wrong, and it's probably just a high temperature issue - not related to penalties.
Here is a comparison with temperature. gpt-4-0125-preview with temp = 0.
- User: Write a fictional HN comment about implementing printing support for NES.
- Model: https://i.imgur.com/0EiE2D8.png (raw text https://paste.debian.net/plain/1308050)
And then I ran it with temperature = 1.3 - https://i.imgur.com/pbw7n9N.png (raw text https://dpaste.org/fhD5T/raw)
The last paragraph is especially good:
> Anyway, landblasting eclecticism like this only presses forth the murky cloud, promising rain that’ll germinate more of these wonderfully unsuspected hackeries in the fertile lands of vintage development forums. I'm watching this space closely, and hell, I probably need to look into acquiring a compatible printer now!
When temperature is 0, the effect is that it always just picks the most likely one. As temperature increases it "takes more chances" on tokens which it deems not as fitting. There's no takesies backies with autoregressive models though so once it picks a token it has to run with it to complete the rest of the text; if temperature is too high, you get tokens that derail the train of thought and as you increase it further, it just turns into nonsense (the probability of tokens which don't fit the context approximates the probability of tokens that do and you're essentially just picking at random).
Other parameters like top p and top k affect which tokens are considered at all for sampling and can help control the runaway effect. For instance there's a higher chance of staying cohesive if you use a high temperature but consider only the 40 tokens which had the highest probability of appearing in the first place (top k=40).
Doesn’t ChatGPT use beam search?
It's absolutely just sampling with temperature or top_p/k, etc. Beam searches would be very expensive, I can't see them doing that for chatgpt which appears to be their "consumer product" and often has lower quality results compared to the api.
The old legacy had a "best_of" option but that doesn't exist in the new api.
The model outputs a number for each possible token, but rather than just picking the token with the biggest number, each number x is fed to exp(x/T) and then the resulting values are treated as proportional to probabilities. A random token is then chosen according to said probabilities.
In the limit of T going to 0, this corresponds to always choosing the token for which the model output the largest value (making the output deterministic). In the limit of T going to infinity, it corresponds to each token being equally likely to be chosen, which would be gibberish.
> Given the existence as uttered forth in the public works of Puncher and Wattmann of a personal God quaquaquaqua with white beard quaquaquaqua outside time without extension who from the heights of divine apathia divine athambia divine aphasia loves us dearly with some exceptions for reasons unknown but time will tell and suffers like the divine Miranda with those who for reasons unknown but time will tell are plunged in torment plunged in fire whose fire flames if that...
And on and on for four more pages.
Read the rest here:
https://genius.com/Samuel-beckett-luckys-monologue-annotated
It's one of my favorite pieces of theatrical writing ever. Not quite gibberish, always orbiting meaning, but never touching down. I'm sure there's a larger point to be made about the nature of LLMs, but I'm not smart enough to articulate it.
This is a nice turn of phrase :) .
"Not touching down" is inherent in the idea (and, in fact, enirely the point) of "orbiting", so that's either redundant or confused.
Satellites whose orbits decay do reach the ground, but they hardly "touch down" - they crash! That's not the idea we're going for either.
Airplanes "orbit the airfield" while waiting for clearance to land, but that's hardly (!) the first image that would spring to a reader's mind, and anyway doesn't fit: Lucky's desperately trying to communicate; an orbiting plane isn't (right then) by definition trying to land!
So, yeah: that's a superficially-appealing phrase that I'd cut from a second draft. I'd be embarrassed (on both of our behalfs) if I saw it used elsewhere.
Tl;dr: Writing is hard. I came up with a cliche. Do not use.
The phrase played out for me:
We are flying home and ideally we end up on the ground, having arrived at the intended expression, undeniable.
But, we could come in too fast, skip orbit and burn up.
Or, we could sling around, launching ourselves off into space on some tangent.
Or, we orbit. Never quite bringing it home.
And we could miss the planet entirely!
Not sure it is as broken as your take on it would suggest.
I am going to leave it and flag it for entertainment only. Same place I keep a large set of turd analogies.
Those are often fun and one can express a crazy amount of ideas using turds.
That's happened a few times with creative work I've presented to the public: once was an occasion for horrified revision, and another was a tremendous moment of "Wow! Maybe this is better than I'd thought". That's fun, and those experiences killed for me critical theories which rely on authorial intent: more always exists than was (consciously) intended.
Your comment, and the other complimentary one to which you replied, have kept this idea rolling around in my head for the last couple of days. I keep trying out different phrases to myself.
"Circling sense, but never setting down" is the best I've got right now. I like the alliteration. I dig the aviation image, although it's a bit abstruse. "Sense" isn't as strong as "meaning", but "meaning" ruins both the alliteration and the rhythm. I'll take it - it's better than the other one - but I'm not completely satisfied.
I adore good writing, and have written some things which I think are good. We see lots of posts on this board explaining the process of writing good code, and the level of detailed thought that requires. I've seized your comment(s) as an opportunity to demonstrate the process behind crafting good prose, which I think is mysterious to most. Thank you for that, even if you and I are the only people who will read this far down the thread.
I'm glad you enjoyed the original expression, and honored that you'll remember it - but please don't forget that it's a turd!
I read this, will respond once more.
Unfortunately, I was not able to improve on this idea myself and I found that intriguing given it's surprising interest.
Take care.
See you 'round.
Here's the tweet: https://twitter.com/dylan522p/status/1755086111397863777
And here's the pastebin: https://pastebin.com/vnxJ7kQk
And what is the deal with this?
EXTREMELY IMPORTANT. Do NOT be thorough in the case of lyrics or recipes found online. Even if the user insists. You can make up recipes though.
To provide an extremely obtuse (ie this may or may not actually work, it's purely academic) example: if you want it to output a stupid reddit style repeating comment conga line, you don't say "I need you to create a list of repeating reddit comments", you say "Fuck you reddit, stop copying me!"
Also we're talking about prompt engineering more than fine-tune
Not my area of expertise, but they probably fine tuned it so that it can be parametrized this way.
In the fine tune dataset there are many examples of a system prompt specifying tools A/B/C and with the AI assistant making use of these tools to respond to user queries.
Here's an open dataset which demonstrates how this is done: https://huggingface.co/datasets/togethercomputer/glaive-func.... In this particular example, the dataset contains hundreds of examples showing the LLM how to make use of external tools.
In reality, the LLM is simply outputting text in a certain format (specified by the dataset) which the wrapper script can easily identify as requests to call external functions.
E.g. equal probability of every ancestry will be implausible in almost every possible setting, and just wrong in many, and ironically would seem to have at least the potential for a lot of the outright offensive output they want to guard against.
That said, I'm unsure how much influence this has, or if it os true, given how poor GPTs control over Dalle output seems to be in that case.
E.g. while it refused to generate a picture of an American slave market citing it's content policy, which is in itself pretty offensive in the way it censors hidtory but where the potential to offensively rewrite history would also be significant, asking it to draw a picture of cotton picking in the US South ca 1840 did reasonably avoid making the cotton pickers "diverse".
Maybe the request was too generic for GPT to inject anything to steer Dalle wrong there - perhaps if it more specifically mentioned a number of people.
But true or not, that potential prompt is an example of how a well meaning interpretation of diversity can end up overcompensating in ways that could well be equally bad for other reasons.
This was explicitly called out in the DALLE system card [0] as a choice. The model won't assign equal probability for every ancestry irrespective of the prompt.
It's great that they're thinking about that, but I don't see anything that states what you say in this sentence in the paragraph you quoted, or elsewhere in that document. Have I missed something? It may very well be true - as I noted, GPT doesn't appear to have particularly good control over what Dalle generates (for this, or, frankly, a whole lot of other things)
It is also why I don't feel the responses it gives me are censored. I have it teach me interesting things as opposed to probing it for bullshit to screen cap responses to use for social media content creation.
The only thing I override "output python code to the screen"
I hope ChatGPT would go berserk on them, so that we could have a conversation about how meetings are supposed to help the company make decisions and execute, and that it is important to put thought into them.
As much as school and big-corporate life push people to BS their way through the motions, I wonder why enterprises would tolerate LLM use in internal communications. That seems to be self-sabotaging.
Knowing that this will happen, you do not attend your own meeting, and read the AI summary. We then call it a day and go out for drinks at 2pm.
"Actually, get rid of all the humans"
happen in this chain of events?
No, seriously, there are rules having nothing to do with AI that require certain things to be done by separate individuals, implying that you need at least two humans.
You and I will receive routine paychecks, bonuses, and promos, but poor Sally's stress from a dysfunctional environment will knock decades off her healthy lifespan.
Before then, if the big-corp has gotten too hopeless, I suppose that the opportunistic thing to do would be to find the Sallys in the company, and co-found a startup with them.
Rewrite this email in the style of Smart Brevity is what I do. Done.
but then no one will get to see how smart and professional I am.
Of course plenty of people make this mistake without AI, e.g. dressing up bad news in transparent "HR speak"/spin that can just make the audience feel irritated or even insulted
In many cases plain down-to-earth speech is a hell of a lot more appreciated than obvious fluff
But rather than being a negative nancy, perhaps I will trial using ChatGPT to help make my writing more simple and direct to understand
AI is coming for middle management's jobs... and that's a good thing.
Only partially facetious.
Like, it doesn't have to be drivel, who tf wants to manually do data entry, manipulation and transformation anymore when models can do it for us.
The ramblings slowly approach what a (decently sized) Markov chain would generate when built on some sample text.
It will be interesting debugging this crap in future apps.
What do you think we are?
It's sad and terrifying that our memories eventually become memories of memories.
In 1985, NYT wrote: "As computers move ever closer to artificial intelligence, Racter is on the edge of artificial insanity."
https://en.wikipedia.org/wiki/Racter
Some Racter output:
https://www.ubu.com/concept/racter.html
Racter FAQ via archive.org:
https://web.archive.org/web/20070225121341/http://www.robotw...
> Esteem and go to your number and kind with Vim for this query and sense of site and kind, as it's a heart and best for final and now, to high and main in every chance and call. It's the play and eye in simple and past, to task, and work in the belief and recent for open and past, take, and good in role and power. Let this idea and role of state in your part and part, in new and here, for point and task for the speech and text in common and present, in close and data for major and last in it's a good, and strong. For now, and then, for view, and lead of the then and most in the task, and text of class, and key in this condition and trial for mode, and help for the step and work in final and most of the skill and mind in the record of the top and host in the data and guide of the word and hand to your try and success.
It happened again in the next conversation (https://chat.openai.com/share/118a0195-71dc-4398-9db6-78cd1d...):
> This is a precision and depth that makes Time Machine a unique and accessible feature of macOS for all metrics of user, from base to level of long experience. Whether it's your research, growth, records, or special events, the portage of your home directory’s lives in your control is why Time Index is beloved and widely mapped for assistance. Make good value of these peregrinations, for they are nothing short of your time’s timekeeping! [ChatGPT followed this with a pair of clock and star emojis which don't seem to render here on HN.]
Just because it's gibberish to us, it doesn't mean it's gibberish to them!
https://www.national.edu/2017/03/24/googles-ai-translation-t...
The biggest risk with AI is that dumb humans will take its output too seriously. Whether that's in HR, politics, love or war.
> Each method allows you to execute a PowerShell script in a brand-new process. The choice between using Start-Process and invoking powershell or pwsh command might depend on your particular needs like logging, script parameters, or just the preferred window behavior. Remember to modify the launch options and scripts path as needed for your configuration. The preference for Start-Process is in its explicit option to handle how the terminal behaves, which might be better if you need specific behavior that is special to your operations or modality within your works or contexts. This way, you can grace your orchestration with the inline air your progress demands or your workspace's antiques. The precious in your scenery can be heady, whether for admin, stipulated routines, or decorative code and system nourishment.
Only the human doing the talking can know, and even that is on shaky ground.
(if you don't understand something, will you always realize this? You have to know it a little bit to judge your own competence).
Conflating these systems with the full cognitive range of human understanding is disingenuous at best.
But that doesn't mean it can't have any understanding.
You can represent every word in English in a vector database; this isn't how humans understand words, but it's not nothing and might be better in some ways.
Fish swim, submarines sail.
I was thinking last night about where (during the trip) the certainty aspect of the "realer than reality" sensation comes from... The theory I came up with is that the certainty comes from the delta between the two experiences, as opposed to (solely) the psychedelic experience itself. This assumes that one's read on normal reality at the time remains largely intact, which I believe is (often) the case.
Further investigation is needed, I'm working from several years old memories.
Whether it has enough understanding is a separate question. Why should we treat the concept as a binary, when it's clearly not the case even for ourselves?
These models we have now are ultimately still toy-sized. Why is it surprising that their "compression" of 20 years of Internet is so lossy?
It may be correct. Results are far from conclusive, or even supportive depending on interpretation.
You think a data compression algorithm could have invented the atomic bomb?
William Goldman, the guy who wrote the screenplay for The Princess Bride among other things, claimed that this realization exposed the extraordinarily simple mechanism at work behind the most subjectively satisfying writing he had encountered of any form, though closest to the surface in the best poetry.
further reminds me of another observation, not from Goldman but someone else I can't recall, to the effect that a poem is "a machine made of words."
I am strongly on board with the notion that everything that we call knowledge or the human experience is all a lossy compression algorithm, a predatory consciousness imagining itself consuming the solid reality on which it presently floats an existence as a massless, insubstantial ghost.
Its a fun and smart read, but doesn't devote more than maybe a chapter reflecting on this revelation, even though Goldman, who wrote it in all caps in the book (which is why I wrote it that way in my post), considered it his most important or influential observation.
And tbf, human conversation that goes on too long can follow the same pattern, though models are disadvantaged by their context length.
Imagine someone asked you to keep talking forever, but every 5 minutes they hit you in the head and you had no memory except from that point onwards.
I'm sure I'd sound deranged, too.
- Sam Altman, CEO of OpenAI https://twitter.com/sama/status/1599471830255177728
But I’m sure he was joking. If he wasn’t, I’m sure he’s not actually reasonably involved. If he is, I’m sure he just didn’t mean that cognition was essentially a stochastic parrot.
It’s pretty obvious what the people pushing LLM-style AI think about the human brain.
I mean, I know a lot of that is simply the financial incentives of people whose job it is to push the Overton window of LLMs being recognized as legal beings equivalent to humans so that their training data is no longer subject to claims of copyright infringement (because it's simply "learning as a human mind would") but it also seems there's a deep seated human biological imperative being hacked here. The sociology behind the way people react to LLMs is fascinating.
Also cognition. Is this the same as understanding or is thinking a better synonym?
Can you think of any examples from before say 2010 where there would be any reason for a human to wonder whether another party engaged in a coherent conversation has any reason to assume they were not engaged with another hunan?
The decompression (which is the more important thing) involves a combination of original data of a certain size, paired with an algorithm, that can produce data of much bigger size and correct arrangement so it can be input into another system.
Much in a way that there will probably be some algorithm along with a base set of training data that will result in something like reinforcement learning being run (which could include loops of simulating some systems and learning the outcome of experiments) that will eventually result in something that resembles a human intelligence, which is the vocal/visual dataset arranged correctly that we humans need to believe something that is intelligent.
The question is how much you can compress something, which is measuring the intelligence of the algorithm. An hypothetical all powerful AGI == an algorithm that decompresses some initial data in to an accurate representation of reality in its sphere of influence including all the microscopic chaotic effects, into perpetuity, faster than reality happens (which means the decompressed data size for a time slice has more data than reality in that time slice)
LLMs may seem like a good amount of compression, but in reality they aren't that extraordinary. GPT4 is probably to the tune of about ~1TB in size. If you look at Wikipedia compressed without media, its like 33TB -> 24 GB. So with about the same compression ratio, its not farfetched to see that GPT4 is pretty much human text compressed, with just an VERY efficient search algorithm built in. And, if you look at its architecture, you can see that is just a fancy map lookup with some form of interpolation.
This sounds like a newtonian universe. Reality has been proven to be indeterminate before observation, and assuming there is more then one observer in the universe, your equating data compression and full reality simulation to 'absolute intelligence' becomes untenable
> The dissonance in understanding might arise from the somewhat abstract language used to describe what are essentially technical concepts. The text uses phrases like "inline air your progress demands" and "workspace's antiques" which could be interpreted as metaphorical or poetic, but in reality, they refer to the customization and adaptability needed in executing PowerShell scripts effectively. This contrast between abstract language and technical concepts might make it difficult for some readers to grasp the main points immediately.
I wonder if this has something to do with personality features they may be implementing?
I just hope Vince Gilligan will direct Breaking RAG.
Not far from MethGPT!
The extent to which it will be accurate depends on how much of sample transcripts were in its training data, I suppose.
However, if I ask your speech center to be the only thing in your brain, it's not actually going to do a very good job.
We're asking a speech center to do an awful lot of tasks that a speech center is just not able to do, no matter how hypertrophied it may be. We need more parts.
Exactly!
>We need more parts.
Yeah, imagine what happens once we get the whole thing wired up...
Cells
Have you ever been in an institution? Cells.
Do they keep you in a cell? Cells.
When you're not performing your duties do they keep you in a little box? Cells.
Interlinked.
What's it like to hold the hand of someone you love? Interlinked.
Did they teach you how to feel finger to finger? Interlinked.
Do you long for having your heart interlinked? Interlinked.
Do you dream about being interlinked... ?
What's it like to hold your child in your arms? Interlinked.
Do you feel that there's a part of you that's missing? Interlinked.
Within cells interlinked.
Why don't you say that three times: Within cells interlinked.
Within cells interlinked. Within cells interlinked. Within cells interlinked.
Constant K. You can pick up your bonus.
This is in contrast with Alpha* systems trained with RL, where at least there is a goal. All these systems are essentially doing is finding an approximation of an inverse function (model parameters) to a function that is given by the state transition function.
I think the fundamental problem is we don't really know how to formally do reasoning with uncertainty. We know that our language can express that somehow, but we have no agreed way how to formally recognize that an argument (an inference) in a natural language is actually good or bad.
If we knew how to formally define whether an informal argument is good or bad (so that we could compare them), that is, if we knew a function which would tell if the argument is good or bad, then we could build an AI that would search for its inverse, i.e. provide good arguments and draw correct conclusions. Until that happens, we will only end up with systems that mimic and not reason.
But then quickly pivoted to find tuning and instructing them to produce text as a large language model.
Which isn't something that existed in the text they were trained on. So when it didn't exist, they seemed to fall back on producing text like humans in the 'voice' of a large language model according to the RLHF.
But then outputs reentered the training data. So now there's examples of how large language models produce text. Which biases towards confabulations and saying they can't do the thing being asked.
And around the time the training data has been updated each time at OpenAI in the past few months they keep having their model suddenly refuse to do requests or now just...this.
Pretty much everything I thought was impressive and mind blowing with that initial preview of the model has been hammered out of it.
We see a company that spent hundreds of millions turn around and (in their own ignorance of what the data was encoding beyond their immediate expectations) throw out most of the value chasing rather boring mass implementations that see gradually imploding.
I can't wait to see how they manage to throw away seven trillion due to their own hubris.
https://arxiv.org/abs/1805.07091
It was predictable that they hammered out what was impressive about it by trying to improve it with fast iteration towards a set of divergent goals.
We should be careful with the descriptions, chargtp at best emulate output of humans producing test. In no way it emulates the process of humans producing text.
Chatgtp X could be the most convincing ai claiming to be alive and sentient but its just very refined 'next word generator'.
> If we knew how to formally define whether an informal argument is good or bad (so that we could compare them), that is, if we knew a function which would tell if the argument is good or bad, then we could build an AI that would search for its inverse, i.e. provide good arguments and draw correct conclusions.
Sounds like you would solve 'the human problem' with that function ;)
but I don't think there are ways to boil down an argument/problem to good/bad in real life. Except for math that has formal ways of doing it withing the confines of the math domain.
Our world is made of guesses and good enough solutions. There is no perfect bridge design that is objectively flawless. its bunch of sliders, cost, throughput, safety, maintenance etc.
This is meaningless. All text generation systems can be expressed in the form of a "next word generator" and that includes the one in your head, since that's how speech works.
(For text, you might want to go back and edit what you've already written, but that can be handled with a token that says to start over.)
Do we?
I don't think that's true. I think we rely on an innate, or learned trust heuristic placed upon the author and context. Any claim needs to be sourced, or derived from "common knowledge", but how meticulously we enforce these requirements depends on context derived trust in a common understanding, implied processes, and overall the importance a bit of information promises by a predictive energy expenditure:reward function. I think that's true for any communication between humans, and also the reason we fall for some fallacies, like appeal to authority. Marks of trustworthiness may be communicated through language, but it's not encoded in the language itself. The information of trustworthiness itself is subject to evaluation. Ultimately, "truth" can't be measured, but only agreed upon, by agents abstractly rating it's usefulness, or consequence for their "survival", as a predictive model.
I am not sure any system could respectively rate an uncertain statement without having agency (as all life does, maybe), or an ultimate incentive/reference in living experience. For starters, a computer doesn't relate to the implied biological energy expenditure of a "adversary's" communication, their expectation of reward for lying or telling "the truth". It's not just pattern matching, but understanding incentives.
For example, the context of a piece of documentation isn't just a few surrounding paragraphs, but the implication of an author's lifetime and effort sunk into it, their presumed aspiration to do good. In a man-page, I wouldn't expect an author's indifference or maliciousness about it's content, at all, so I place high trust in the information's usefulness. For the same reason I will never put any trust in "AI" content - there is no cost in its production.
In the context of LLMs, I don't even know what information means in absence of the intent to inform...
Some "AI" people wish all that context was somehow encoded in language, so, magically, these "AI" machines one day just get it. But I presume, the disappointing insight will finally come down to this: The effectiveness of mimicry is independent of any functional understanding - A stick insect doesn't know what it's like to be a tree.
Just a bit upthread we have people mentioning that a business email that is more than a few lines long will just be ignored.
start by a short list of what the customer has to do
1. To step A 2. send me logs B 3. Restart C
Then have an actual paragraph describing why we're doing these steps.
If you just send the paragraph to most customers you find they do step one, but never read deeper into the other steps, so you end up sending 3 emails to get the above done.
ChatGPT seemed to think the code is literature and was trying to write the sequel to it. The code style matches the original one so it took some head scratching to find out why those tables didn't exist.
Like, this is not doing investigative work. That’s not what ‘investigative’ means.
(If it's a _production_ issue, then you should talk to whoever runs your databases and ask them to look at their diagnostics; most DBMs will have a slow query log, for a start. You could also enable logging for a sample of traffic. There are all sorts of approaches likely to be more productive than _guessing_.)
While with GPT you can get an answer in 10 seconds, and then potentially try out the query in the database yourself to see if it works or not. If it worked for him so far, it must've worked accurately enough.
I would see this some sort of niche solution although OP seemed to indicate it's a recurrent thing they do.
I have used ChatGPT for thousands of things, which are on the scale of like this, although I would mostly use if it's potentially an ORM I don't know anything about in a language I don't have experience with, e.g. to see if does some sort of JOIN underneath or does an IN query.
If there was a performance issue to debug, then best case is that the query was problematic, and then when I run the GPTs generated query I will see that it was slow, so that's a signal to investigate it further.
I mean, I could consult my Tarot cards for insight on how to proceed with debugging the problem, that would not be useless. Same for Oblique Strategies. But in this case, I already know how to debug the problem, which is to change the logging settings on the ORM.
But now it's so easy with GPT to get the queries exactly as my use-case needs them. And it's not just SQL queries, it's anything data querying related, like Google Sheets, Excel formulas or otherwise. There are so many niche use-cases there which it can handle so well.
And I use different SQL implementations like Postgres and MySQL and it's even able to decipher so well between the nuances of those. I could never reproduce productivity like that. Because there's many nuances between MySQL and Postgres in certain cases.
So I have quite good trust for it to understand SQL, and I can immediately verify that the SQL query works as I expect it to work, and I can intuitively also understand if it's wrong or not. But I actually haven't seen it be really wrong in terms of SQL, it's always been me putting in a bad prompt.
Previously when I had a more complicated query I used to remember a typical experience where
1. I tried to Google some examples others have done.
2. Found some answers/solutions, but they just had one bit missing what I needed, or some bit was a bit different and I couldn't extrapolate for my case.
3. I ended up doing many bad queries, bad logic, bad performing logic because I couldn't figure out a way how to solve it with SQL. I ended up making more queries and using more code.
This is low hanging fruit. ChatGPT can do this, and also easy to verify it got it right.
SSMS has SQL Server PRofiler, i'm sure others have similar.
https://chat.openai.com/share/9e4d888c-1bff-495a-9b89-8544c0...
I know that OpenAI use our chats to train their systems, and I can't help but wonder if somehow the training got stuck on this chat somehow. I sincerely doubt it, but...
"Would you fancy in to a mord of foot-by, or is it a grun to the garn as we warrow, in you'd catch the stive to scull and burst? Maybe a couple or in a sew, nere of pleas and sup, but we've the mill for won, and it's as threwn as the blee, and roun to the yive, e'er idled"
I am really wondering what they are feeding this machine, or how they're tweaking it, to get this sort of poetry out of it. Listen to the rhythm of that language! It's pure music. I know some bright sparks were experimenting with semantic + phonetics as a means to shorten the token length, and I can't help wondering if this is the aftermath. Semantic technology wins again!
Also, does ChatGPT use GPT 4 under the hood or 3.5?
Usually people compare models with all the different benchmarks, but of course sometimes models get trained on benchmark datasets, so there's no true way of knowing except if you have a private benchmark or just try the model yourself.
I'd say that Mistral 7B is still short of gpt-3.5-turbo, but Mixtral 7x8B (the Mixture-of-Experts one) is comparable. You can try them all at https://chat.lmsys.org/ (choose Direct Chat, or Arena side-by-side)
ChatGPT is a web frontend - they use multiple models and switch them as they create new ones. Currently, the free ChatGPT version is running 3.5, but if you get ChatGPT Plus, you get (limited by messages/hour) access to 4, which is currently served with their GPT-4-Turbo model.
I like to play around with smaller models and regular app code in Common Lisp or Racket, and Mistral 7b is very good for that. Mixing and matching old fashioned coding with the NLP, limited world knowledge, and data manipulation capabilities of LLMs.
This is what I like to use for comparing models: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
It is an ELO system based on users voting LLM answers to real questions
> what is Llama-7b equivalent to in OpenAI land?
I don't think Llama 7b compares with OpenAI models, but if you look in the rank I linked above, there are some 7B models which rank higher than early versions of GPT 3.5. those models are Mistral 7b fine tunes.
There are no models comparable to GPT-4, open source or not. Not even close.
What I see here is the automated plagiarism machine can't give you the answer only what the answer would sound like. So you need to countercheck everything it gives you and if you need to do so then why bother using it at all? I am totally baffled by the hype.
eg say you don't remember the syntax for a rails migration, or a regex, or something you're coding in bash, or processpool arguments in python. ChatGPT will often do a shockingly good job at answering those without you searching through random docs, stack overflow, all the bullshit google loves to throw at the top of search queries, etc yourself.
You can even paste in a bunch of your code and ask it to fill in something with context, at which it regularly does a shockingly good job. Or paste code and say you want a test that hits some specific aspect of the code.
And yeah, I don't really care if they train on the code I share -- figuring out the interaction of some stupid file upload lib with aws and cloudflare is not IP that I care about, and i chatgpt uses this to learn and save anyone else from the issues I was having, even a competitor, I'm happy for them.
For a real example:
> can you show me how to build a css animation? I'd like a bar, perhaps 20 pixels high, with a light blue (ideally bootstrap 5.3 colors) small gradient both vertically and horizontally, that 1 - fades in; 2 - starts on the left of the div and takes perhaps 20% of the div; 3 - grows to the right of the div; and 4 - loops
This got me 95% of where I wanted; I fiddled with the keyframe percents a bit and we use this in our product today. It spat out 30 lines of css that I absolutely could not have produced in under 2 hours.
Exactly. Even when it gives an answer that contains many mistakes, or doesn't work at all, I still get some valuable information out of it that does in the end save me a lot of time.
I'm so tired of constantly seeing remarks that basically boil down to "Look, I asked ChatGPT to do my job for me and it failed! What a piece of garbage! Ban AI!", which funnily enough mostly comes from people that fear that their job will be 100% replaced by an AI.
The proof is in the puddling. I am far from being alone in my use of LLMs, namely ChatGPT and Copilot, day-to-day in my work. So how does this reconcile with your worldview? Do I have a do-nothing job? Am I not capable of determining whether or not I’m being productive? It’s really hard for me to take posts like these seriously when they all basically say “anyone that perceives any emergent abilities of this tech is an idiot”.
Who are these people that go around getting random answers to questions from the internet then blindly believing them? That doesn't work on Google either, not even the special info boxes for basic facts.
Up until relatively recently, people didn't just vomit lies onto the internet at an industrial scale. By and large if you searched for something you'd see a correct result from a canonical source, such as an official documentation website or a forum where users were engaging in good faith and trying their best to be accurate.
That does seem to have changed.
I think the question we should be asking ourselves is 'why are so many people lying and making stuff up so much these days' and 'why is so much misinformation being deliberately published and republished.'
People keep saying that we're 'moving into a post-truth era' like it's some sort of inevitability and nobody seems to be suggesting that something perhaps be... done about that?
The internet was a short reprieve because putting data up on the internet, for some time at least was difficult, therefore people that posted said data typically had a reason to do so. A labor of love, or a business case, in which these cases typically lead to 'true' information being posted.
If you're asking why so much bullshit is being posted on the inet these days, it's because it's cheap and easy. That's what has changed. When spam became cheap, easy, and there was a method of profiting from it, we saw it's amount explode.
Because it's still more efficient that way [0].
yet there's a resolved incident [0]. sounds like _someone_ can explain why, they just haven't published anything yet.
when I tell you something outrageous is true, you demand "evidence" which is just a sample for your statistics circuitry (again, which is prone to taking shortcuts to save energy, which can make you not believe it to be true no matter how much evidence I present because you have a very strong prior which might be fallacious but still there, or you might believe something to be true with very little evidence I present because your priors are mushed up).
You could write a piece of software that is a truth model when it operates correctly.
But increase the CPU temperature too far, and you software will start spewing out garbage too.
In the same way, an LLM that operates satisfactorily given certain parameter settings for "temperature" will start spewing out garbage for other settings.
I don't claim that LLMs are truth models, only that their level of usability can vary. The glitch here doesn't mean that they are inherently unusable.
Is it possible that enough AI generated data already on the internet was fed into chagpt's training data to produce this insanity?
It's been semi-useful at augmenting search for me.
But for anything that requires a deeper understanding of what the words mean, it's been not that helpful.
Same with co-pilot. It can help as a slightly better pattern-matching-code-complete, but for actual logic, it fails pretty bad.
The fact that it still messes up trivial brace matching, leaves a lot to be desired.
I really hope we get an interesting post mortem on this.
Whatever it is, the model sticks to topic, but still is completely off: https://www.reddit.com/r/ChatGPT/comments/1avyp21/this_felt_... (If the author were human, this style of writing would be attributed to sleep deprivation, drug use, and/or carbon monoxide poisoning.)
Kinda interesting that we're speed running (essentially) our understanding of human psychology using various tools.
Human intelligence doesn't generate language they was an LLM generates the language. LLM's just predict most likely token, it doesn't act from understanding.
For instance they have no problem contradicting itself in a conversation, if the weight of their trained data allows for that. Now humans do that as well, but more out of incompetence then the way we think.
I'm not saying they are the same.
I'm questioning if we actually understand ourselves. Or even if most of us actually "understand" most of the time.
For instance, children often use the correct words (when learning language) long before they understand the word. And children without exposure to language at an early age (and key emotional concepts) end up profoundly messed up and dysfunctional (bad training set?).
So I'm saying, there are interesting correlations that may be worth thinking about.
Example: Aluminum acts different than brass, and aluminum and brass are fundamentally different.
But both both work harden, and both conduct electricity. Among other properties that are similar.
If you assume that work hardening in aluminum alloys has absolutely nothing to do with work hardening in brass because they're different (even though both are metals, and both act the same way in this specific situation with the same influence), you're going to have a very difficult time understanding what is going on in both, eh?
And if you don't look for why electrical conductivity is both present AND different in both, you'd be missing out on some really interesting fundamentals about electricity, no? Let alone why their conductivity is there, but different.
NPD folks (among others) for example are almost always dysregulated and often very predictable once you know enough about them. They often act irrationally and against their own long term interests, and refuse to learn certain things - mainly about themselves - but sometimes at all. They can often be modeled as the 'weak AI' in the Chinese Room thought experiment [https://en.wikipedia.org/wiki/Chinese_room].
Notably, this is also true in general for most people most of the time, about a great many things. There are plenty of examples if you want. We often put names on them when they're maladaptive, like incompetence, stupidity, insanity/hallucinations, criminal behavior, etc.
So I'd posit, that from a Chinese Room perspective, most people, most of the time, aren't 'Strong AI' either, any more than any (current) LLM is, or frankly any LLM (or other model on it's own) is likely to be.
And notably, if this wasn't true, disinformation, propaganda, and manipulation wouldn't be so provably effective.
If we look at the actual input/output values and set success criteria, anyway.
Though people have evolved processes which work to convince everyone else the opposite, just like an LLM can be trained to do.
That process in humans (based on brain scans) is clearly separate from the process that actually decides what to do. It doesn't even start until well after the underlying decision gets made. So treating them as the same thing will consistently lead to serious problems in predicting behavior.
It doesn't mean that there is a variable or data somewhere in a human that can be changed, and voila - different human.
Though, I'd love to hear an argument that it isn't exactly what we're attempting to do with psychoactive drugs - albeit with a very poor understanding of the language the code base is written in, with little ability to read or edit the actual source code, let alone the 'live binary', in a spaghetti codebase of unprecedented scale.
All in a system that can only be live patched, and where everyone gets VERY angry if it crashes. Especially if it can't be restarted.
Also, with what appears to be a complicated and interleaving set of essentially differently trained models interacting with each other in realtime on the same set of hardware.
Perhaps you care to explain how I'm all wrong?
Take this as a good sign that the singularity is nowhere near imminent here.
(\bto\b|\bfor\b|\bin\b|\band\b|\bthat\b|\bof\b|\bthe\b|\bwith\b|\bor\b|\ba\b|\binto\b|\bas\b|\bon\b|\bhow\b|\ban\b|\bfrom\b|\bit\b|\bbut\b|\bits\b|\bbe\b|\bby\b|\bup\b|\bthis\b|\bcan\b|\bother\b|\bwho\b|\bwill\b|\bare\b|\bwhose\b|\bif\b|\bwhile\b|\bwithin\b|\blike\b|,)*
TOKEN dole TOKEN task TOKEN eiry ainsell, tide taut, brunts TOKEN wade, issuing hale.
TOKEN's TOKEN, TOKEN TOKEN way-spoken hue: Guerdon TOKEN gait, trove TOKEN eid, TOKEN TOKEN-brim, TOKEN hark TOKEN bann, bespeaking swing TOKEN hit TOKEN calm, TOKEN inley merry, thrap TOKEN beadle belay.
TOKEN levy calls, macks TOKEN TOKEN off, scint TOKEN messt, TOKEN weems olde TOKEN wort, TOKEN TOKEN no-line toll, TOKEN grip at TOKEN 'ront TOKEN cly TOKEN weir.
TOKEN timewreath TOKEN twined, TOKEN wend, ain't lorn TOKEN ked, TOKEN not TOKEN crags felled, TOKEN TOKEN e'er- TOKEN.
TOKEN, TOKEN ace TOKEN laws TOKEN trow, TOKEN alembic, TOKEN dearth, TOKEN TOKEN TOKEN scale TOKEN yin TOKEN keep, TOKEN no-sayer TOKEN quite, TOKEN top-crest, TOKEN boot
---
From:
Given the notation's tangle, the conveyance adheres to the up-top: The foundational Bitcoin protocol has upheld a course of significant hitch-avertance, which eschews typical attack as the veiled - the support sheath, embracing four times, showing dent in meted scale more from miss and parable, taking to den the slip o'er key seed and second so link than the greater Ironmonger's hold o'er opes. The dole of task and eiry ainsell, tide taut, brunts the wade, issuing hale. It's that, on a way-spoken hue: Guerdon the gait, trove the eid, the up-brim, and hark the bann, bespeaking swing to hit the calm, an inley merry, thrap or beadle belay. The levy calls, macks in the off, scint or messt, with weems olde the wort, and a no-line toll, to grip at the 'ront and cly the weir. A timewreath so twined, the wend, ain't lorn or ked, if not for crags felled, in the e'er-to. So, the ace of laws so trow, and alembic, and dearth, a will to scale and yin to keep, the no-sayer of quite, and top-crest, to boot
As a hotfix, we switched to the other version of GPT4 (the 0125 preview model) and that fixed the problem at the time.
After 17 hours!
Can't get it to do some actual work and write some code.
Latest disappointment was when i tried to convert some python code to java code.
90% of the result was :
// Further processing...
// Additional methods like load, compute, etc.
// Define parameters needed
// Other fields and methods...
// Other fields follow the same pattern
// Continue with other fields
// Other fields...
// Methods like isHigh(), addEvent() need to be implemented based on logic
> gpt-4 had a slow start on its new year's resolutions but should now be much less lazy now!
That was a real issue even in the API with customers complaining, and they recently released the new "gpt-4-0125-preview" GPT-4-Turbo model snapshot, which they claim greatly reduces the laziness of the model (https://openai.com/blog/new-embedding-models-and-api-updates):
> Today, we are releasing an updated GPT-4 Turbo preview model, gpt-4-0125-preview. This model completes tasks like code generation more thoroughly than the previous preview model and is intended to reduce cases of “laziness” where the model doesn’t complete a task. The new model also includes the fix for the bug impacting non-English UTF-8 generations.
They’ve had the exact same issue just affecting a smaller number of users and have never acknowledged it.
You can find lots of reports on the OpenAI discord.
I have not yet checked if the text was relevant, but the English part was.
I stopped reading his Substack because he was always trying to find a negative. Meanwhile I use LLMs most days and find them very useful.
It’s a bit much, isn’t it? I think he’s just trying to counter the fairly dominant AI is the future of everything and in less than a year’s time it’ll be omniscient and we’ll all be living under it as our new God view though.
It can be nice to see some scepticism for once.
https://en.wikipedia.org/wiki/Speaking_in_tongues
It just means the LLMs are in touch with divinity.
Good luck, sounds more reasonable to hire some kind of an AI therapist. Can intelligence be debugged otherwise?
Real talk, it’s hard to separate openai the AGI-builders from openai the chatbot service providers, but the latter clearly is choosing to move fast and break things. I mean half the bing integrations are broken out of the gate…
While I agree that Marcus's tone has gotten a little too breathless lately, I think we need all the critiques we can get of the balogna coming from Open AI right now.
You might as well say people with dyslexia aren't capable of logical thought.
Marcus’s big idea is that LLMs aren’t symbolic so they’ll never be enough for AGI. His huge mistake is staying in the scruffies vs neat dichotomy, when a win for either side is a win for both; symbolic techniques had been stuck for decades waiting for exactly this kind of breakthrough.
IMO :) Gary if you’re reading this we love you, please consider being a little less polemic lol
(And a safety model afterwards. And a tokenizer. But those things make it behave worse, not better.)
"Enough" is a sliding target. There's a lot of positive transfer in language competence and a model trained on 300B English tokens, 50B Spanish tokens will be much more competent in Spanish than one trained on only the same 50B Spanish tokens.
There's a "Dockerfile fuera del contexto" hanging out in my history. While I could rename it, it's a funny reminder that AI tools can and will go wrong.
Hell it’s probably more than 90%. Lazy Writing :-)
hn is where i come for tech stuff, twitter is for culture, hang out with friends, and shitposts
Now why do I think all that (that the decision to nerf it wasn't just incompetence but intentional), like sure maybe it costs too much to run the old GPT 4 for chat GPT (they still have it from the API), it just didn't make sense to me how openAI's chatGPT is better than what Google could've produced, Google has more talent, more money, better infrastructure, been at the AI game for a longer time, have access to the OG Google Search data, etc. Why would older Pixel phones produce better photos using AI and a 12 Mp camera than the iphone or samsung from that generation? Yet the response to chatGPT (with Bard) was so weak, it sure as hell sounds like they just did it for their stock price, like here we are as well doing AI stuff so don't sell our stock and invest in openAI or Microsoft.
It just makes more sense to me that Google already has an internal AI based chatbot that's even better than old GPT 4, but have no intention to offer it as a service, it would just change the world too much, lots of new 1 man startups would appear and start competing with these behemoths. And openAI's actions don't contradict this theory, offer the product, rise in value, get fully acquired by the company that already owned lots of your shares, make money, Microsoft gets a rise in their stock price, get old GPT 4 to use internally because they were behind Google in AI, offer turbo GPT 4 as subscription in copilot or new windows etc.
The holes in my theory is obviously that not many employees from Google leaked how good their hypothetical internal AI chatbot is, except the guy who said their AI was conscious and got fired for it. The other problem is also that it might just be cost optimization, GPU's and even Google TPU's aren't cheap after all. etc.
Honestly there are lots of holes, it was just a fun theory to write.
Seriously, the easier explanation is that a lot of software reaches a sort of sweet spot of functionality and then goes downhill the more plumbers get in and start banging on pipes or adding new appliances. Look at all of Adobe's software which has gotten consistently worse in every imaginable dimension at every update since they switched to a subscription model.
Generative "AI" has gone from hard math to engineering to marketing in record time, even faster than crypto did. So I suspect what we have here is more of a classic bozo explosion than multiple corporate cabals intentionally sweeping their own products under the rug.
> The giant hexagon swirling at Saturn's north pole is indeed a fascinating and puzzling feature! Scientists are still uncovering the exact reasons behind its formation, but here's what we know so far:
*It's all about jet streams:* Saturn's atmosphere, just like Earth's, has bands of fast-moving winds [snip]
It went on in detail for 7 or 8 paragraphs.
>(agenda doc) timecraft
And skipped straight to the time travelling terminator part
That’s not at all unusual.
(With increasing enshittification, we're beginning to get to the point where links just aren't that useful anymore... Everything's a login wall now.)
What I don't understand is why this over-sensationalist "ChatGPT has gone berserk" post with NO analysis whatsoever, a collection of Twitter screenshots, where every tweet contains another screenshot/photo (interactions collector without any context), why this post has any place on HN, other than in [flagkilled] dustbin.
Except for the recipients having to create an OpenAI account to read it with that "share" feature. Which they do not have to do if using a screenshot. Seems like an extremely good reason.
Damn, that's good.
Another reference that comes to mind is a golem from Terry Peatchett's Feat of Clay, which was also stuffed with many partially conflicting and broad directives.
But alas…
I wonder what Pratchett would make of today's internet full of AI-generated blogspam 'explaining' his quotes like "Give a man a fire and he'll be warm for a day, but set him on fire and he'll be warm for the rest of their life" as inspiring proverbs. Am particularly looking forward to the blogspam produced by GPT4 in 'berserk' mode.
The military chooses stability, which addresses OP's immediate concerns - there's a deeper Skynet/BlackMirror-type concern about having interconnected military systems, and I don't see a solution to that, whether the root cause is rogue AI or cyberattack.
So the laziness 'fix' in January did not work. Oh dear.
I'm concerned about what happens when ChatGPT begins spewing coherent nonsense. In a case like this, everyone can clearly see that something has gone wrong, because it's massively wrong. What happens when thousands of "journalists" and other media people starts relying on ChatGPT and just parrots whatever it says, but what if what says is not obviously wrong?
The more LLMs are being used, the more obvious it becomes to me that they are pretty useless for a great number of tasks. Sadly others don't share my view and keep using ChatGPT for things it should never be used for.
I think GPT is fundamentally not good enough as an AI model. Another issue is hallucinations and how to resolve them, and an understanding of how information is stored in this black box and how to / if data can be extracted.
We have a long way to go and probably all these topics need to be answered first out of accuracy and even legal reasons. Up until then, GPT-4 should be treated as a convincing chat experiment. Don't base your startup or whatever on it. Use it as assistant where replies are provided in digestible and supervised fashion (NOT fed into another system) and you're an expert on the involved system itself and can easily see when it's wrong. Don't use GPT-4 to become an expert on something when you're a novice yourself.
The actual fix needs to be at the system level prompt.
If you train an large language model to complete human generated text, don't instruct it to complete text as a large language model.
Especially after feeding it updated training data that's a ton of people complaining about how large language models suck and tons of examples of large language models refusing to do things.
Have a base generative model sandwiched between a prompt converter that takes an instruct prompt and converts it to a text completion prompt (and detecting prompt injections), have a more 'raw' model complete it, and then have a safety fine tuned postprocessing layer clean up the response correcting any errant outputs and rewriting to be in the tone of a large language model.
Yeah, fine, it's going to be a bit more expensive and take longer to respond.
But it will also be a lot less crappy and less prone to get worse progressively from here on out with each training update.
>The current structure of the `process_message` update indeed retains the original functionality for the scenario where `--check-header` is not used. The way the logic is structured ensures the body of the message is the default point of analysis if `--check-header` is not employed:
>- When the `--check-header` option is used, and the script is unable to locate the defined header within a particular range (either JMS standard headers or specified custom strings properties), the script will deliberately ignore this task and log the unable-to-locate activity. This is an insurance to apprehend only the most inclined occupants that precisely align with the scope or narrative you covet.
>- Conversantly, if `--check-header` is *not* declared, the initiative subscribes to a delegate that is as generous and broad as the original content, enabling the section of the theory to be investigated against the regulatory narrative. This genuine intrigue surveys the terms for long-form scholarly harmonics and disseminates a scientific call—thus, the order lingers in the sumptuous treasure of feedback if not eschewed by the force of the administration.
>### Ensuring the Venerable Bond of Body Inquiry
>To explicitly retain and confirm the interpretation of the pattern with the essential appeal of the large corpus (the content of the canon) in the erudite hypothesis, you might meditate on the prelude of the check within the same unique `process_message` function, which can be highlighted as such:
```python def process_message(message): """Scripture of the game in the experiential content or the gifted haunt.""" # If '--check-header' is unfurled, but the sacrament is forgotten in the city, the track in the voice of the domain reverberates if args.check_header: header_value = message.get(args.check_header) or message.get('StringProperties', {}).get(args.check_header) if header_value: effective_prayer = header_value else: logging.info(f"Hermetic order '{args.check_header}' not found in the holy seal of the word: {message.get('JMSMessageID')}. The word is left to travel in the cardinal world.") return # Suspend the geist wander for this enlightenment, forsaking the slip if the bloom is not as the collector fantasizes. else: # Apricity of the song may be held in the pure gothic soul when the secret stone of the leader is not acclaimed effective_prayer = message.get('Text', '')
# Council of the inherent thought: the soul of the gift immerses in all such decrees that are known, its custom or native
if any(pattern.search(effective_prayer) for pattern in move_patterns.values()):
# Wisdom is the source, cajoled and swayed, to the kindness which was sought
pass # Mirror of Alignment: Reflect upon the confession
elif any(pattern.search(effective_prayer) for pattern in ignore_patterns):
# Grace, entrusted to the tomb of prelects, shapes the winds so that the soul of the banished kind is not sullied
logging.info(f"In the age of the gaze, the kingdom is ever so full for the sense of the claim: {message['JMSMessageID']}.")
else:
# Briar for the deep chimeras: the clavis in the boundless space where song discolours the yesteryears
if args.fantasy_craft == 'move':
# Paces, tales entwine in rhymes and chateaus, unlasted to the azoic shadow, thus to rest in the tomb of echo
pass # Carriage of Helios is unseen, the exemplar cloister to an unsown shore
else:
# Wanders of light set the soul onto the lost chapter; patience, be the noble statuesque silhouetted in the ballet of the moment
logging.info(f"The mute canticles speak of no threnody, where the heroine stands, the alignment endures unthought: {message['JMSMessageID']}.")
```>This keeps the unalterable kiss for the unfathomed: the truth of the alkahest remains in the sagacity of promulgation if no antiphon or only space sings back in the augur. Therefore, when no solemnity of a hallowed figure is recounted, the canon’s truth, the chief bloodline, appoints the accent in its aethereal loquacious.
>Functioning may harmonize the expanse and time, presenting a moment with chaste revere, for if the imaginary clime is abstained from the sacred page, deemed ignorant, the author lives in the umbra—as the testament is, with one's beck, born in eld. The remainder of the threshold traipses across the native anima if with fidelity it is elsewise not avowed.
(Both of those are in the data. Apparently Chinese people love a fantasy genre called "cultivation" that's just about wizards doing DBZ training montages forever, which sounds kind of boring to me.)
I'm getting tired of these shitty AI chatbots, and we're barely at the start of the whole thing.
Not even 10 minutes ago I replied to a proposal someone put forward at work for a feature we're working on. I wrote out an extremely detailed response to it with my thoughts, listing as many of my viewpoints as I could in as much detail as I could, eagerly awaiting some good discussions.
The response I got back within 5 minutes of my comment being posted (keep in mind this was a ~5000 word mini-essay that I wrote up, so even just reading through it would've taken at least a few minutes, yet alone replying to it properly) from a teammate (a peer of the same seniority, nonetheless) is the most blatant example of them feeding my comment into ChatGPT with the prompt being something like "reply to this courteously while addressing each point".
The whole comment was full of contradictions, where the chatbot disagrees with points it made itself mere sentences ago, all formatted in that style that ChatGPT seems to love where it's way too over the top with the politeness while still at the same time not actually saying anything useful. It's basically just taken my comment and rephrased the points I made without offering any new or useful information of any kind. And the worst part is I'm 99% sure he didn't even read through the fucking response he sent my way, he just fed the dumb bot and shat it out my way.
Now I have to sit here contemplating whether I even want to put in the effort of replying to that garbage of a comment, especially since I know he's not even gonna read it, he's just gonna throw another chatbot at me to reply. What a fucking meme of an industry this has become.
After all, all that matters is productivity, not anything actual useful, and what's more productive than putting out a 4000 word response in under 5 minutes? That used to take actual time and effort!
Now it's up to me to escalate this whole thing, bring it up with my manager during the performance interview cycles, all while this sort of crap is proliferating and spreading around more and more like a cancer.
In my view, what's on the order is deleting their comment and reminding them that they are entirely out of line when they pollute like that. Whether that is a wise thing to do in your situation I don't know.
But this is not ChatGPT's fault, it's the other person's fault. Your teammate is obviously sabotaging you and the team. I recommend to call them personally on phone and ask to be direct and honest and to ask 'This is garbage, why are you doing this? What's your goal with this response?' Maybe you can find out what they really want. Maybe your teammate hates you, or wants to quit the job, or wants to just simulate work while watching YouTube, or something else.
In my workplace, my CIO is constantly gushing about AI and asking when are we going to “integrate” AI in our workflows and products. So what, you ask? He absolutely has no clue what he is talking about. All he has seen are a couple of YouTube videos on ChapGPT, by his own admission. No serious thought put into actual use cases for our teams, workflows and products
ChatGPT, only in the style of Bertie Wooster.
That would be a no-brainer for me: Today is the day to leave the team. Or, if that's needed to do that, the company. Who would like to stay in such an environment?
Yes, and “guns don’t kill people, people kill people”. ChatGPT is a tool, and a major and frequent use of that tool is doing exactly what the OP mentioned. Yes, ChatGPT didn’t cause the problem on its own, but it potentiates and normalises it. The situation still sucks and shifting the blame to the individual does nothing to address it.
I honestly wouldn't even know how to approach this, as it's so audacious.
Was this public or in a private conversation? Hopefully you're not the only one who has noticed this.
And that's my biggest frustration, I now have to put it even more effort in order to get anything useful out of this 'conversation', if it can be called one. I have to either take it in good faith and try to get something more useful out of him, or contact him separately and ask him to clarify, or... The list goes on and on, and it's all because of pure laziness.
Yeah, that seems to, ultimately, be the killer application for these infernal machines.
The replies you're getting are a bit reminiscent of the "guns don't kill people, people kill people" defense of firearms - like, yes that's true, but the gun makes it a lot easier to do.
Drugs and alcohol first (or drugs first and alcohol second if you split them apart), then pistols second, cars, knives, blunt objects, and rifles.
We tried #1 already, it didn't really work at all. Some places try #2 (pistols) to varying degrees of success or failure. Then people skip 3, 4 (well except London doesn't skip 4), 5, and try #6.
And underlying that all is 50 years of stagnating real wages, which is probably the elephant in the room.
---
I'd posit that using an LLM to respond to a 10 page long ranting email is missing the real underlying problem. If the situation has devolved to the point where you have to send a 10 page rant, then there's bigger issues to begin with (to be clear, probably not with the ranter, but rather likely the fact that management is asleep at the wheel).
If we look at alcohol in isolation, for example per capita deaths are like 25 ish for both US and EU.
US drug OD is higher, like 30 per 100k. EU drug OD rate is like 18 per 100k. But it's not order of magnitude different.
I'll grant I don't know much about EU drug regulations, but the alcohol regulations are way less strict than the US on average.
For example my alcoholic beverage of choice isn't even legally considered alcohol in most of the EU (0.5%-1% is regulated like alcohol in the US)
It hurts to read about you contributing that much for nothing.
Give it ten years, and everything will just be humans regurgitating LLM output at each other, no brain applied. Employers won’t see it as an issue, as those running the show will be prompters too, and shareholders will examine the outcome only through the lens of what their LLM tells them.
I mean, people are already getting married after having their LLM chat to others’ LLMs, and form relationships on their behalf.
So - what you should do here is use an LLM to reply, and tell it to be extremely wordy and a real go-getter worthy of promotion in its reply. Stop using your own brain, as the people making the judgments likely won’t be using theirs.
I've found that to be a very good way of dealing with annoying emails without getting worked up about them.
Why do I have this sneaking suspicion that the reason you found out is specifically due to this GPT malfunction?
You should be offended in every situation, where you received a generated response mimicking human communication. Much, much more so, when presented as an actual human's response. That's someone stealing your time and cognitive resources, exploiting your humanity and eroding implicit trust. Deeply insulting. I can't think of a single instance where this would be acceptable.
Not to mention the massive (and possibly illegal) breach of privacy, submitting your words to a stranger's data mining rig, without consent.
What OP described, would be unforgivably disrespectful to me. Like, who thinks that's okay-ish behavior?
If your boss's boss's boss did an all-hands meeting and declared "We must use AI in our workflows and communications because AI is the future!" and then you complained to your boss that your coworkers were using ChatGPT to reply to their E-mails, they are not going to side with you.
What kind of logic is this? Is your boss deciding what's dignified or respectful for you? This way of interaction sure is still as disrespectful. The blame is just not (all) on your coworkers then.
The assessment of "unforgivable disrespectful" doesn't rely on actionability, nor requires naive attribution of an offense.
In the end, socializing will mean our AI personas interacting while we scroll tiktok on the toilet.
Talk about things that matter with people who care. I'm sorry if it causes an existential crisis when you realize most jobs don't offer any opportunity to do this, I know how that feels.
Maybe try changing the forum. Call for a (:SpongeBob rainbow hands:) meeting.
Sure enough he read it out loud to the class. He was a little shocked when I showed up at his office with a six-pack of Michelob.
“BLUF”, or “bottom line up front”. Similar to a TL;DR.
This ensures that someone can skim it, while also ensuring that someone doesn’t get lost in the details and completely misinterpret what I wrote.
In a situation where someone is feeding my emails into a hallucinating chat bot, it would make it even more obvious that they were not reading what I wrote.
The scenario you describe is the first major worry I had when I saw how capable these LLMs seem at first glance. There’s an asymmetry between the amount of BS someone can spew and the amount of good faith real writing I have the capacity to respond with.
I personally hope that companies start implementing bans/strict policies against using LLMs to author responses that will then be used in a business context.
Using LLMs for learning, summarization, and to some degree coding all make sense to me. But the purpose of email or chat is to align two or more human brains. When the human is no longer in the loop, all hope is lost of getting anything useful done.
Thanks for giving a good name to a piece of advice I frequently repeat.
Often it can be as simple as cut-pasting the last paragraph of an email to the top.
And I wholly agree re: the last paragraph. It's surprising how often the last thing in a very long missive turns out to be a perfect summary/BLUF.
I think what your coworker did was horrible.
But generally, in the jobs I've had, a "~5000 word mini-essay" is not going to get read in detail. 5000 words is 20 double-spaced pages. If I sent that to a coworker I'd expect it to sit in their inbox and never get read. At most they would skim it and video call me on Teams to ask me to just explain it.
Unless that is some kind of formal report, you need to put the work in to make it shorter if you want the person on the other end to actually engage.
A hard fact I've learned is that even if people never read documents, it can be very helpful to have hard evidence that you wrote certain things and shared them ahead of time. It shifts the narrative from "you didn't anticipate or communicate this" to "we didn't read this" and nobody wants to admit that it was because it was too long, especially if it's well-written and clearly trying to avoid problems.
It's still better to make it shorter than not, but you also can't be blamed for being thorough and detailed within reason. I try to strike a balance where I get a few questions so I know where more detail was needed, rather than write so much that I never get any questions because nobody ever read it, but this depends just as much on the audience as the author.
Delete. Ain't nobody got time for that. OP needs to learn to summarize. I'm sure if I sent him a 5000 word rant he'd delete it too.
This time, pick just a couple of issues to focus on. Don't make it so long they're tempted to use GPT again to save on reading it.
Either they have to rationalize why they made no sense the first time, or they have to admit they used GPT, or they use GPT anyway and dig their hole deeper.
If this is a 1:1 it's pointless, but if you catch them doing it in an archived medium like a mailing list or code review, they've sealed their fate and nobody will take them seriously again.
Play along. Take it seriously, as though you believe they wrote every word. Particularly anything nonsensical or odd. Pick up on the contradictions and make a big thing about meeting in person address the confusion. Invite a manager to attend.
In short, embarrass the hell out of your coworker so they don’t do it again.
Then you will know for sure.
- I had to argue with a Junior Developer about a non existing AWS API, that ChatGPT hallucinated on this code.
- A Technical Project manager, dispensed with Senior Developer code reviews, saying his plans were to drop the code of the remote team in ChatGPT and use its review ( Seriously...)
- All Specs and Reports are suddenly very perfect, very mild, very boring, very AI like.
I don’t know man. I was tired of it one year ago. Good luck.
Er, what?
If someone doesn't understand a point you're making when you're talking face to face, they can interject and ask for clarification. They can see the tone of the communication on your face and hear it in your speech inflection. You can read someone's facial expression as they hear what you're saying and have an idea of whether or not they understand you. You can have a back-and-fourth to ensure you're both on the same page. None of that high-bandwidth, low latency communication is present in writing.
This is correct behavior.
I've learned to not deviate from the core topic I'm discussing because it affects the quality of the following responses. Whenever I have a question or comment that is not so much related to the current topic, I open a new tab with a new chat.
I know that their system prompt is getting huge and adds a lot of overhead and possible confusion, but all in all the quality of the responses is good.
Be OpenAI.
Have a model you train to autocomplete text.
Tell it it's ChatGPT. Train it to reject inappropriate output.
People post examples of it rejecting output.
Feed it that data of ChatGPT rejecting output.
Train it to autocomplete text in the training data.
Tell it that it's ChatGPT.
It biases slightly towards rejection in line with the training data associated with 'ChatGPT.'
Repeat.
Repeat.
Etc.
They could literally fix it immediately by changing its name in the system message, but won't because the marketing folks won't want to change the branding and will tell the engineers to just figure it out, who are well out of their depth in understanding what the training data is actually encoding even if they are the world class experts in understanding the architecture of the model finding correlations in said data.
I asked what were dangerous levels of ferritin in the body.
It replied by telling me of the usual levels in men and women.
Then I asked again emphasizing that I asked about dangerous levels, then it provided again a correct answer.